6,880 Matching Annotations
  1. Last 7 days
    1. Author response:

      eLife Assessment:

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. While these data are interesting and helpful for hypothesis generation, they are deemed incomplete; additional experiments would be needed to test causality and to provide mechanistic insight. Despite these limitations, this work will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

      The authors would like to thank the reviewers for thorough and constructive comments on our manuscript. We will make major updates to the manuscript addressing the following points and suggestions from the three reviewers: (1) assessing HbAS/AA genotype influence on microbiome composition; (2) conducting the requested beta diversity analysis, (3) conducting the requested sensitivity analysis to assess the impact of disease severity and therapy on microbiome and virome features; (4) modifying our language to clearly state that our results do not indicate causality or mechanism of microbiome interactions with sickle cell disease pathophysiology; (5) improved discussion of the phage results and their strengths and limitations; (6) additional changes throughout for clarity and correction of errors. We will change the title to “Bacterial and viral gut microbiome alterations characterize microbiome-immune-pathophysiology axes in Sickle Cell Disease.” These additions will greatly improve our work and presentation and we are grateful to the reviewers and our editors.

      We have indicated where specific changes were made in response to the public reviews below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease, but instead, aged neutrophil numbers and bacteriophages do. Appropriate control subjects and additional analyses are needed to support that conclusion.

      Strengths:

      The primary strength of this paper is the investigation into disease-associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease processes, or are they in any way contributing to disease pathophysiology?

      Weaknesses:

      A primary weakness of the manuscript is the fact that the majority of individuals included in the control group maintain sickle cell trait (HbAS genotype). Although typically asymptomatic, it is unclear if this genotype is associated with microbial changes that would not be observed in a true control group (HbAA genotype). This is a significant limitation that may limit the ability to draw conclusions from the current data set.

      Another key weakness is the lack of beta diversity assessment. Although decreased alpha diversity is observed in individuals with SCD, and specific bacterial taxa are differentially abundant following multivariate analyses, there is no overall comparison of bacterial community composition between individuals with SCD and controls. Prior to drawing conclusions about the relationship (or lack thereof) between the SCD microbiome and inflammatory markers, it is important to know if this study did indeed find disease-associated changes in microbiome composition.

      It is unclear which individuals were used for aged neutrophil (AN) and molecular data assessments. For example, were children who were still receiving penicillin prophylaxis included in these specific assessments? Given the authors' previous work demonstrating that antibiotic treatment decreases AN pathology, it seems critical to limit all AN/molecular analyses to older subjects who are not on daily penicillin treatment (if possible).

      A minor weakness is the continued use of "disease" vs. "healthy" indicators as primary microbiome metrics that are used for molecular correlations. The lack of metric specificity - and lack of discussion regarding which diseases were used to generate these indicators (how similar/different are they to sickle cell?) - could be said to make these metrics essentially meaningless.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this will be clarified in the main text. We will include an analysis evaluating the influence of control genoype (HbAA/HbAS) on our microbiome and virome results. To evaluate whether control genotype influenced major microbiome and virome features, analyses were restricted to control participants only. Controls were stratified by genotype as HbAA or HbAS. Four significant microbiome and virome features were tested: F:B ratio, Shannon diversity, provirus fraction, and virus count. HbAA and HbAS controls were compared using two-sided Mann-Whitney U tests. Benjamini-Hochberg FDR correction was applied across the four tested features. HbAS and HbAA controls did not differ significantly for F:B ratio, Shannon diversity, provirus fraction, or virus count. The inclusion of HbAA/AS will strengthen our results with respect to the observation that sickle cell disease patient microbiomes remain significantly different from sickle trait (HbAS) controls. These results will be reported in a new Supplemental Table.

      We will include a beta diversity analysis using MetaPhlAn species profiles. Beta diversity analyses were performed in Python using pandas and NumPy for data processing, scikit-bio for distance calculations and PERMANOVA, scikit-learn for ordination-related computations, statsmodels for multiple-testing correction where applicable, and matplotlib for visualization.

      For the primary disease/control comparison, samples were grouped as control or SCD. For the genotype control sensitivity analysis, samples were restricted to HbAA and HbAS individuals as described above. Species detected in at least 10% of included samples were retained for beta diversity analysis. To account for the compositional structure of metagenomic relative abundance data, species profiles were transformed using a centered log-ratio transformation after addition of a small pseudocount to accommodate zero values. Aitchison distances were calculated from the CLR-transformed species profiles. Statistical significance of group separation was assessed by PERMANOVA using 999 permutations. For the control versus SCD comparison, PERMANOVA was performed between the two disease-status groups. For the HbAA versus HbAS control comparison, PERMANOVA was performed among controls only.

      In the SCD cohort, beta diversity differed significantly between controls and SCD participants by Aitchison distance after CLR transformation (R<sup>2</sup> = 0.030, p = 0.001). In contrast, HbAA and HbAS controls did not differ significantly in beta diversity (R<sup>2</sup> = 0.024, p = 0.282), supporting the conclusion that the observed SCD/control separation was not driven by control genotype composition. These methods and results will be reported in the revised manuscript.

      The manuscript describing the microbiome health and disease indicators was submitted to eLife jointly with this manuscript as a package; eLife declined to review the indicator manuscript. Briefly, this study conducted a cross-disease meta-analysis of 38 studies comprising 8,204 samples and identified 100 bacterial taxa or “indicators” that are weakly but consistently associated with health or disease across diverse conditions, including, but not limited to, inflammatory bowel disease, colorectal cancer, type 2 diabetes. The indicator taxa were validated in an independent cohort of Graves’ disease patients. We currently cite an older version of this work posted as a preprint. The manuscript is currently under review at another journal and we will update this manuscript with the updated citation when it is available.

      We will address the other recommendations from this reviewer in the revision.

      Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

      Weaknesses

      The evidence for this proposed immune-related mechanism is incomplete. The study is cross-sectional and largely based on associations, so it cannot determine whether these virus-like elements drive immune activation, reflect immune activation, or are linked indirectly through disease severity, treatment history, or other clinical factors. The main limitations are the single-center design, modest sample size for some immune measurements, limited ability to control for treatment and disease heterogeneity, and the need for clearer multiple-testing correction in the correlation analyses. In particular, stronger adjustment for available clinical factors such as hydroxyurea use, transfusion history, pain admissions, genotype, and other markers of disease burden would help readers judge how specific the microbial and viral findings are to sickle cell disease itself.

      Overall, the authors largely achieve their descriptive aim of identifying gut microbial differences associated with sickle cell disease. The evidence is solid for the presence of broad microbial community differences, but incomplete for the stronger conclusion that virus-like elements form a pathophysiological immune axis. The work will likely be useful to researchers studying the microbiome, inflammation, and sickle cell disease, especially as a hypothesis-generating dataset. Its impact would be strengthened by more cautious interpretation, stronger control of clinical confounders, clearer statistical correction, and future longitudinal or experimental studies to test causality.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this will be clarified in the main text. We will temper our interpretation of our results, making clear that we are not arguing that either prophages or bacteria are causal or mechanistically associated with SCD biology and pathology. We will strengthen our control of clinical confounders, and add clearer statistical correction in the revision. We look forward to conducting future studies to test causality and understand mechanism.

      We will address the other recommendations from this reviewer in the revision.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome. Based on these findings, the authors propose the existence of a viral-immune axis in SCD pathophysiology and targetable functional alterations in the gut microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      However, this study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. In particular, the claim that the study identifies targetable functional alterations in the gut microbiome for disease treatment may be somewhat overstated. Although the reported functional module changes in SCD patients are intriguing, additional mechanistic and/or longitudinal evidence would be required before these features can realistically be considered targetable.

      We thank the reviewer for their helpful comments and suggestions. We will note in the text that additional mechanistic and longitudinal studies are required before we can target the microbiome and virome in SCD and clarified that this is a single center, cross-sectional. We will make further modifications to the manuscript to clarify cohort features (specifically, age and race were matched, other baseline characteristics were balanced), to properly describe the Shannon diversity metric, and to fix several errors that this reviewer caught.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Singh et al. presents an application of MOA-seq to better define transcriptional control underlying the hypoxia response in human endothelial cells. This group's previously described MOA-seq technique allows for precise, identity-agnostic mapping of occupied sites of DNA-binding proteins across the epigenome and over time. Here, they applied MOA-seq to HUVECs under normal oxygen conditions or variable lengths of hypoxia treatment, comparing changes in occupancy over time and associating these changes with corresponding transcriptome alterations. This approach revealed thousands of dynamically occupied sites comprising 10 major kinetic clusters that appear to define distinct subsets and phases of the hypoxia response. Analysis of DNA motifs in these dynamically occupied regions captured the known major roles of HIF1A in the hypoxia response and also implicated new HIF1A-associated regulators. Importantly, they also identified many potential HIF1A-independent candidate TFs that act at HREs, which has been an outstanding question in the field. Additionally, this study identified ~7K additional sites not previously defined as regulatory elements by ENCODE.

      Strengths:

      Overall, this study is well executed and described, providing new biological insights as well as a rich data resource for the field. As MOA-seq was previously developed for use in plants, this work demonstrates the application of this method in mammalian cells and highlights its utility in identifying new potential regulatory sites not captured by DNase-seq or ATAC-seq. The conclusions made by the authors are well supported by the results, with the caveat that extensive use of DNA motif identification and ontology analyses invariably leads to some uncertainty regarding factor identity and gene network properties.

      Weaknesses:

      There are several areas where the clarity of presentation could be improved:

      (1) Given the importance of the methodology, the methods section needs more detail on how the extent of MNase digestion is chosen to achieve optimal results with MOA-seq. This is described to some extent in the description of control library preparation, but not for the experimental samples.

      We thank the reviewer for noting this unintended omission. We have not updated the Methods section to specify as follows:

      "Digestion patterns were assessed via gel electrophoresis, and the light digest levels ideal for MOA-seq (as per Savadel et al., 2021) were selected as the lightest digest levels that give a pattern of a nucleosomal ladder spanning the entire DNA fragment size range from undigested to mononucleosome bands, as indicated in Figure 1 with the asterisk-marked gel lanes."

      (2) The abstract describes this approach as "native cistrome profiling" but this is misleading since formaldehyde fixation is used.

      We believe the formaldehyde fixation captures native chromatin structure, but indeed we are digesting fixed chromatin and have updated the wording to read as “in situ cistrome profiling.”

      (3) Species- and field-specific jargon and abbreviations need to be clarified on first usage. For example, on page 9: "Downsampling analysis was carried out for two sets of published reference peaks; the CTCF cCRE peak midpoints and for the ERG motif under the ERG ReMap ChIP-seq peaks." The different categories of cCREs were not clearly defined, nor will it be clear what the term ReMap refers to for those outside the field. The sentence after this refers to IDR, which also should be defined.

      We thank the reviewer for highlighting the need for clearer definitions of field-specific terminology and abbreviations. In response, we have revised the manuscript to explicitly define all relevant terms at first mention. Specifically, we now describe the ENCODE candidate cis-regulatory element (cCRE) catalogue and define the individual cCRE categories, including promoter-like (PLS), proximal enhancer-like (pELS), distal enhancer-like (dELS), DNase I–H3K4me3 (K4m3), and CTCF-only regions. We also clarify that ReMap is a curated database of human transcriptional regulator binding peaks derived from ChIP-seq, ChIP-exo, and DAP-seq experiments. Additionally, we now define IDR as the Irreproducible Discovery Rate framework upon first use.

      (4) Figure 4C: Are these motifs examined under MOA sites specifically or anywhere in the genes in question?

      Leading up to and including Figure 4C, we have not yet examined any motifs. Instead, Figure 4C compares gene sets, one defined by our diff-MOA, and those from GO libraries, in this case the "target genes" which are defined by TF-specific studies, primarily ChIP-seq but also related immuno-based mapping techniques. Consequently, the analysis shown in Fig. 4C is not a motif enrichment analysis. Instead, we used the ENRICHR gene set enrichment analysis tool with ENCODE and ChEA consensus transcription factor target gene sets. Thus, the analysis was performed at the gene-set level, and transcription factor motifs were not examined within diff-MOA peaks or elsewhere in the associated genes for Fig. 4C. We note that motif enrichment within diff-MOA peaks was subsequently examined separately in Fig. 6. In Fig. 7, we further examined differentially expressed genes associated with diff-MOA peaks containing enriched transcription factor motifs and used clustering analyses to investigate their regulatory relationships. We have clarified these distinctions in the revised manuscript.

      If the question is about the location of MOA footprints relative to gene structure, we did not examine any MOA sites at any specific location, just overlapping the gene +/- 200 bp, as indicated in Fig. 4B.

      (5) Figure 5B shows that up-DEGs with diff-MOA footprints tend to show more losses of footprints. Do the authors interpret this as a loss of repressor binding?

      Not exclusively, but yes, that is one plausible explanation. That is, the activation (defined by increased RNA levels) via de-repression could be happening. But we also expect these dynamic footprints to be but one component. In other words, we interpret the relationship as consistent with that possibility, but not only that possibility. A logical explanation is that loss of footprint occupancy associated with upregulated genes could be based on displacement of repressive DNA-binding factors, thereby contributing to transcriptional activation. Thus, while loss of repressor binding is a plausible explanation for a subset of these events, additional factor-specific experiments would be required to know for sure in each case. We have added text to the Discussion acknowledging this possibility.

      Reviewer #2 (Public review):

      Summary:

      Singh et al. apply MOA-seq to map transcription factor occupancy genome-wide in HUVECs across a hypoxia time course. The study provides a well-validated, high-resolution view of cistrome dynamics and identifies both HIF1A-associated and independent regulatory programs.

      Major Comments:

      Methodological validation is strong. MOA-seq's ability to map protein-bound DNA at near-nucleotide resolution without factor-specific antibodies is a genuine advance, and the cross-validation against independent ChIP-seq and ENCODE datasets is convincing. As noted, future work with additional biological replicates could further strengthen confidence in the smaller kinetic clusters.

      Regarding additional biological replicates, we have acknowledged this point in the discussion. Importantly, we did subject the replicates to IDR analysis, which we explain in the methods as "In accordance with ENCODE ChIP-seq guidelines (Landt et al., 2012), we further evaluated data quality by assessing pooled pseudo-replicate consistency and self-consistency for each individual replicate (Supplementary Table S2)." This IDR analysis demonstrated consistent peaks between our bioreplicates, meeting ENCODE guidelines. In addition, downsampling analysis demonstrated that our sequencing depth of coverage (Supp Fig 1) was over 10-fold greater than required. We do appreciate that it will be useful to have more biological replicates from other cell types, tissues, or organisms, and hope this study prompts just such future research.

      Imaging-based validation would strengthen the key biological claims. The kinetic clustering and pathway enrichments are computationally inferred. Orthogonal approaches, for example, live-cell fluorescence imaging of HIF1A nuclear translocation to confirm the proposed temporal binding waves, would provide independent experimental support.

      Live-cell imaging could indeed be interesting, but it is beyond our current capacity to add to this study and consider this an exciting future direction, but presence in the nucleus could include both bound and unbound HIF1, so the results may not easily track the DNA-bound HIF1 only.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      In Figure 3B, the x-axis is not labeled.

      Thank you for pointing this out. We have revised Figure 3B by adding the previously missing x-axis label.

      Reviewer #2 (Recommendations for the authors):

      In the abstract, it would be good to define what MOA-seq is and what the cistrome is.

      Thank you for this suggestion. We have revised the abstract to define both MOA-seq (MNase-defined cistrome-Occupancy Analysis sequencing) and the cistrome upon first mention to improve accessibility for readers who may be unfamiliar with these terms.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      This paper is an exciting follow-up to two recent publications in eLife: one from the same lab, reporting that slender forms can successfully infect tsetse flies (Schuster, S et al., 2021), and another independent study claiming the opposite (Ngoune, TMJ et al., 2025). Here, the authors address four criticisms raised against their original work: the influence of N-acetyl-glucosamine (NAG), the use of teneral and male flies, and whether slender forms bypass the stumpy stage before becoming procyclic forms.

      Strengths:

      We applaud the authors' efforts in undertaking these experiments and contributing to a better understanding of the T. brucei life cycle. The paper is well-written and the figures are clear.

      Comments on revisions:

      We thank the authors for the revised manuscript and for considering our comments.

      We outline below the 3 points that, in our opinion, remain to be clarified.

      (1) Effect of NAG on slender-form infections in tsetse flies

      The conclusion that "NAG has a negligible effect on slender infections in tsetse flies" based on Figure 1, cannot be fully supported in the absence of a positive control. A relevant positive control is well established in the literature, namely that NAG promotes Tsetse infection by stumpy forms. Without such a control, it is not possible to exclude technical issues (for example, an ineffective NAG treatment), which would yield results similar to those presented in Figure 1.

      We agree that an internal stumpy-form positive control would provide an additional technical reference. However, the enhancing effect of NAG on stumpy-form midgut infections is well established and was also demonstrated under the experimental framework of our original study (Schuster et al. 2021, Figure 2A).

      The purpose of the present Research Advance was therefore not to re-establish the known effect of NAG on stumpy infections, but to test whether slender-form infections require NAG supplementation. Under the conditions tested here, slender bloodstream forms established midgut, proventriculus and salivary-gland infections also in the absence of NAG. We have revised the text accordingly to avoid implying a general absence of NAG effects and to make clear that our conclusion is restricted to slender-form infections under the conditions tested (line 128).

      (2) Infection of non-teneral flies

      Because the experiments shown in Figure 1 (teneral flies) and Figure 2 (non-teneral flies) were not conducted in parallel or under identical conditions, it is important that the figure legends clearly state the parasite numbers used in each case. Specifically, infections of teneral flies were performed with 200 parasites/mL (approximately 4 parasites per bloodmeal), whereas non-teneral infections used 1 × 10<sup>6</sup> parasites/mL (approximately 20,000 parasites per bloodmeal?). At present, this information is scattered across the Methods and Supplementary Tables 1 and 2, making it difficult for readers to immediately appreciate that the parasite load differs by roughly 5,000-fold between these conditions.

      As previously shown by the authors (Schuster et al., 2021) and in the Rotureau laboratory (Tsagmo Ngoune et al.), and as generally expected, the initial parasite dose strongly influences infection outcomes in teneral flies. In this context, it would be informative to know whether the authors have attempted infections of non-teneral flies using lower parasite numbers (noting that Tsagmo Ngoune et al. used a maximum of 10,000 parasites) and what the infection rate was.

      Relatedly, the statement in line 370 appears to be an overgeneralization, as fly age was not directly tested under matched experimental conditions:

      Line 370 - "Here, we unambiguously show that, in the absence of immunosuppressive treatment, slender forms can establish infections in tsetse flies, irrespective of the fly's age or sex."

      We thank the reviewer for highlighting the inconsistent presentation of parasite doses between Figure 1 and 2. We agree this is confusing and have revised the figure legends to clearly state both the parasite concentration (cells/mL) and estimated fly uptake per bloodmeal for each experiment (Lines 143 and 206).

      Regarding experiments with non-teneral flies using lower parasite numbers: We have not tested intermediate doses (e.g., 10,000 parasites/bloodmeal as used by Ngoune et al.) in non-teneral flies. Given that teneral flies already show relatively low infection rates even under optimal conditions, we chose the higher parasite dose (20,000 parasites/bloodmeal) for non-teneral flies to ensure sufficient statistical power for meaningful analysis of infection outcomes across different fly compartments.

      We acknowledge the reviewer's concern regarding the statement in line 370 and have revised this sentence (line 375) to more accurately reflect our experimental conditions, avoiding overgeneralization beyond the specific parameters tested.

      This reads now: “Here, we demonstrate that slender forms can establish infections without immunosuppressive treatment under the conditions tested. This infectivity was observed in both teneral and non-teneral, as well as in both male and female flies, indicating that slender forms retain transmission potential across different fly demographics. However, direct age comparisons under identical parasite doses remain to be tested.”

      (3) Transcriptomic analysis

      Supplementary Figure 8 lacks statistical analysis, which limits its interpretability. Two types of comparisons would be particularly helpful:

      (i) a comparison of PAD1/2 expression levels between slender and stumpy forms at 0 h; and

      (ii) for each gene, a comparison of the overall change in expression (from 0 to 72 h) between infections initiated with slender versus stumpy forms.

      In addition, the figure legend should clarify what "expression levels" refer to. TPM? Normalized counts?

      We appreciate this helpful comment and included statistical analysis for the expression of PAD1 and PAD2 (Supplementary Figure 8) between the two forms for the baseline (0 h) as well as during the differentiation to procyclic forms (0 h to 72 h) by using Welch´s t-test.

      While PAD1 did not show a statistically significant difference in this analysis, PAD2 displayed significant differences in expression dynamics over time. This supports the broader transcriptomic observation that slender- and stumpy-initiated differentiation follow distinct transcriptional trajectories before converging at the procyclic stage.

      We also clarified the figure legends showing the mean log2 counts per million (CPM) values.

      Finally, for the benefit of the field, eLife could encourage publishing a collaborative study in which the Engstler and Rotureau laboratories exchange parasite lines and culture protocols (including media with and without methylcellulose) and perform tsetse fly infections in parallel in their respective laboratories. Such an approach could help resolve the remaining discrepancies and provide a valuable reference for the community.

      We appreciate this constructive suggestion. A collaborative inter-laboratory study in which parasite lines, culture conditions and infection protocols are exchanged between the Engstler and Rotureau laboratories would be a valuable way to address the remaining discrepancies in the field. In particular, parallel infections using matched parasite lines and culture conditions, including media with and without methylcellulose, could provide a useful reference dataset for the community.

      At the same time, such a study would require substantial coordination, reciprocal strain exchange, protocol harmonization and new infection series in two laboratories. It therefore goes beyond the scope of the present Research Advance, which was designed to address the specific methodological concerns raised in response to our original publication. We have restricted our conclusions accordingly and view the proposed collaborative benchmark study as an important direction for future work.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript reports the discovery and characterization of the first bifunctional degrader of tankyrase. Notably, the tankyrase degrader exhibits stronger β-catenin inhibition and tumor growth suppression compared to conventional tankyrase inhibitors. Mechanistically, while tankyrase inhibitors stabilize tankyrase and promote Axin puncta formation - thereby impairing β-catenin degradation - the degrader avoids this effect, resulting in deeper suppression of β-catenin signaling. These findings suggest that targeted degradation of tankyrase offers a novel therapeutic strategy for β-catenin-driven cancers. Overall, this is a compelling study with significant translational potential.

      Strengths:

      (1) The manuscript presents a rigorous and well-executed study on a timely and impactful topic.

      (2) The biochemical and cellular characterization of the tankyrase degrader is thorough, and the comparative analysis with tankyrase inhibitors is insightful.

      (3) The finding that tankyrase stabilization by inhibitors may interfere with Axin function is novel and significant. It aligns with earlier observations (e.g., Huang 2009) that transient tankyrase overexpression can stabilize β-catenin independently of PAR domain activity.

      (4) The use of TNKS1/2 knockout cells expressing catalytically inactive tankyrase to demonstrate β-catenin inhibitory activity of the tankyrase degrader is elegant.

      (5) The finding that the tankyrase degrader has superior anti-proliferative effects in colorectal cancer models has important therapeutic implications.

      Weaknesses:

      (1) A key caveat is that the identified tankyrase degrader also targets GSPT1 for degradation. This raises the possibility that GSPT1 degradation may contribute to the observed β-catenin and tumor growth inhibition.

      (2) The authors address this concern reasonably by showing that DLD1 cells resistant to GSPT1 degradation remain sensitive to the tankyrase degraded.

      (3) To further strengthen this point, the authors might consider generating TNKS1/2 double knockout cells (e.g., in DLD1 or SW480 backgrounds) and demonstrating that the degrader loses its growth-inhibitory effect in these models. However, given the technical challenges of creating double knockouts in cancer cell lines, such experiments could be considered optional.

      We thank the Reviewer for the favorable feedback. The major concern is the collateral degradation of GSPT1. As the Reviewer noted, IWR1-POMA was able to suppress colony formation in DLD-1 cells resistant to a GSPT1/2 degrader (DLD-1R, Figure 6B and S9F), suggesting that TNKS but not GSPT degradation is responsible for growth inhibition.

      We also appreciate that the Reviewer brought it to our attention an important early observation of the TNKS scaffolding effects. Cong reported in 2009 that overexpression of TNKS induced AXIN puncta formation in a SAM but not PARP domain-dependent manner (PMID: 19759537, Ref. 12). We have added this reference to the introduction of TNKS scaffolding in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      The ADP-ribosyltransferase tankyrase controls many biological processes, many of which are relevant to human disease. This includes Wnt/beta-catenin signalling, which is dysregulated in many cancers, most notably colorectal cancer. Tankyrase is a positive regulator of Wnt/beta-catenin signalling in that it counters the activity of the beta-catenin destruction complex (DC). Catalytic inhibition of tankyrase not only blocks PAR-dependent ubiquitylation and degradation of AXIN1/2, the central scaffolding protein in the DC, but also tankyrase itself. As a result, blocking tankyrase gives rise to tankyrase accumulation, which may accentuate its non-catalytic functions, which have been proposed to drive Wnt/beta-catenin signalling. Most tankyrase catalytic inhibitors have shown limited efficacy and substantial toxicity in vivo. By developing tankyrase-directed PROTACs, the authors aim to block both catalytic and non-catalytic functions of tankyrase, aspiring to achieve a more complete inhibition of Wnt/beta-catenin signalling. The successfully developed PROTAC, based on the existing catalytic inhibitor IWR1, IWR1-POMA, induces the degradation of both TNKS and TNKS2, blocks beta-catenin-dependent transcription without stabilising the DC in puncta/degradasomes, and inhibits cancer cell growth in vitro. Mechanistically, this points to a scaffolding role of tankyrase in the DC, at least under conditions of tankyrase catalytic inhibition, in line with previous proposals.

      Strengths:

      The study clearly illustrates the incentive for developing a tankyrase degrader, namely, to abolish both catalytic and non-catalytic functions of tankyrase. By and large, the study achieves these ambitions, and the findings support the main conclusions, although the statement that a more complete inhibition of the pathway is achieved requires corroboration. The proteomics studies are powerful. IWR1-POMA constitutes a very useful tool to re-evaluate targeting of tankyrase in oncogenic Wnt/beta-catenin signalling. The paired compounds will benefit investigations of tankyrase scaffolding functions across many different biological systems controlled by tankyrase. The findings are exciting.

      Weaknesses:

      Although the results are promising and mostly compelling, the claim that the PROTACs provide "a deeper suppression of the WNT/β-catenin pathway activity" requires further corroboration, particularly at endogenous tankyrase levels.

      We thank the Reviewer for the encouraging and insightful comments. The major critique concerns whether TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels. IWR1-POMA reduced the level of cytosolic β-catenin more effectively than IWR1 in Wnt3A-stimulated HEK293 cells without protein overexpression (Figure 1D). IWR1POMA also suppressed STF activity more effectively than IWR1 in DLD-1 cells (Figure S8C) and reduced the expression levels of several WNT/β-catenin targets more effectively than IWR1 (Figure 1G and S8D). These results support that TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels.

      There are also some other points that, if considered, would further improve the manuscript, as detailed below.

      (1) Abstract and line 62: Many catalytic tankyrase inhibitors tend to display toxicity, which is likely on-target (e.g., 10.1177/0192623315621192; 10.1158/0008-5472). This constitutes the main limiting factor for these compounds. An incomplete inhibition of Wnt/beta-catenin signalling may contribute to the challenges, but this does not appear to be the dominant problem. A more prominent introduction to this important challenge is probably expected by the field.

      A previous study showed that G007-LK, a selective TNKS inhibitor, exhibited weak efficacy and dose-limiting toxicity at 5‒30 mg/kg BID or 10‒60 mg/kg QD in various mouse xenograft models (PMID: 23539443, Ref. 28). Similarly, G-631, another TNKS inhibitor, also showed dose-limiting toxicity without significant efficacy at 25‒100 mg/kg QD in mice (PMID: 26692561, Ref. 60). However, other studies showed that G007-LK was well-tolerated at 200 mg/kg QD over 3 weeks in mice (PMID: 29316982, Ref. 61), and treating mice with G007-LK at 10 mg/kg QD over 6 months also improved glucose tolerance without notable toxicity (PMID: 26631215, Ref. 62). Importantly, basroparib, a selective TNKS inhibitor, was well tolerated in a recent clinical trial (PMID: 40964966, Ref. 64), and constitutive silencing of both TNKS1 and TNKS2 for 150 days in APC-null mice prevented tumorigenesis without damaging the intestines (PMID: 31337618, Ref. 8). We have included some discussion of the toxicity issue associated with TNKS targeting at the end of the Discussion section.

      (2) The authors do a good job in setting the scene for the need for tankyrase degraders. Their observations relating to the formation of puncta (degradasomes) being tankyrase-dependent are compatible with a previous study by Martino-Echarri et al. 2016 (10.1371/journal.pone.0150484): simultaneous silencing of TNKS and TNKS2 by RNAi abolishes degradasome formation. The paper is cited as reference 17, but only in passing, and deserves more prominence. (It includes an entire paragraph titled "Expression of tankyrases 1 and 2 is required for TNKSi-induced formation of axin puncta").

      Indeed, Henderson’s 2016 paper (PMID: 26930278, previously Ref. 17, now Ref. 18) shed important light on the role of TNKS scaffolding in the DC. However, whereas this study demonstrated that knocking down both TNKS1 and TNKS2 by siRNA prevented G007-LK to induce AXIN puncta, it concluded that “puncta formation requires both the expression and the inactivation of TNKS,” which is inconsistent with our observations that accumulation of either catalytically active or inactive TNKS can promote AXIN puncta formation. The function roles of TNKS scaffolding in the DC also remained unaddressed. We have included additional discussion of Henderson’s findings in the first paragraph the Discussion section.

      (3) Moreover, the scaffolding concept has been discussed comprehensively in other studies: 10.1111/bph.14038 and more recently 10.1042/BCJ20230230. There are also a few studies that focus on targeting the ankyrin repeat clusters of tankyrase to disengage substrates (10.1038/s41598-020-69229-y; 10.1038/s41598-019-55240-5) that illustrate the concept of blocking the scaffolding function. In that sense, the hypotheses are mature, and it is interesting to see some of them supported in this study. The authors could improve how they set their work into the context of these other efforts and proposals.

      Indeed, Guettler demonstrated in 2016 that TNKS scaffolding could promote WNT/β-catenin signaling, which forms the basis of the current work. Meanwhile, whereas there have been efforts to target the SAM or ARC domain to address TNKS scaffolding by Guettler and Lehtiö, our approach of targeting TNKS for degradation is complementary. We have included in the last paragraph of the Discussion section information on efforts to target the ARC or SAM domains as an alternative approach to suppress WNT/β-catenin signaling without promoting TNKS oligomerization (PMID: 31836723 and 32704068, Ref. 66 and 67).

      (4) In several places in the manuscript, the DC is referred to as "biomolecular condensate", at times even as a "classic example", implying that it operates through phase separation. This has not been demonstrated. In fact, super-resolution microscopy indicates that the puncta are not droplet-like (10.7554/eLife.08022), which would argue against the condensate hypothesis.

      Biomolecular condensates are membraneless cellular compartments formed by phase separation of biomolecules, regardless of their physical/material properties (PMID: 28935776 and 28225081, Ref. 22 and 23). Super-resolution microscopy studies by Stenmark (PMID: 26124443, Ref. 17) showed that AXIN, APC, TNKS, and β-catenin interacted with each other to assemble into membraneless complexes, wherein AXIN and APC formed filaments throughout the DC. Peifer has also summarized evidence that supports the condensate nature of the DC (PMID: 30782412, Ref. 9; see also PMID: 26393419). However, we acknowledge that testing the physical properties of reconstituted DC (for example, PMID: 34352208) with TNKS will provide a better understanding of the nature, for example liquid vs. gel, of these condensates.

      (5) It is beautiful to be able to use IWR1 and IWR1-POMA at identical concentrations for direct comparisons. However, this requires the two compounds to bind to tankyrase similarly well and reach the target to a comparable extent. How sure are authors that target engagement is comparable? Has this been evaluated?

      Using a BRET assay, we have confirmed that IWR1-POMA binds to TNKS1 with affinity comparable to that of IWR1. Details of this study is now included in the Results sections, and the data are presented in the Supplementary Information (Fig. S3E–G).

      (6) Figure 1F: It is not immediately apparent how IWR1-POMA shows more complete containment of Wnt/beta-catenin signalling. Most Wnt/beta-catenin targets lie close to the perfect diagonal, so I do not see how the statement "that IWR1-POMA controlled WNT/β-catenin signaling more effectively than IWR1" (in the legend of Figure 1F) is supported. Minimally, an expanded explanation would benefit the reader. Providing the colour-coding legend directly in the figure would help improve clarity. Also, the panel is very small and may benefit from a different presentation in the figure.

      We have updated Fig. 1F to include an inset of Quadrant III for improved clarity and readability. We have also moved Fig. S7C to the main text as Fig. 1G and added an expanded explanation for these figures.

      (7) Figure 2: The conclusion of a "deeper suppression" of signalling relies on overexpression of tankyrase in an otherwise tankyrase-null background. Have the authors attempted to measure reporter activity or endogenous gene expression without tankyrase overexpression, in Wnt3a-stimulated cells (in the context of a normal Wnt/beta-catenin pathway) or CRC cells at the basal level? Non-catalytic activity in a similar assay has previously been observed upon tankyrase overexpression (10.1016/j.molcel.2016.06.019). Whether or not there is a substantial scaffolding effect at endogenous tankyrase levels after tankyrase inhibition remains unconfirmed, and the PROTAC is a valuable tool to address this important question. The findings presented in Figure S7C and D go some way towards answering this question - these data could be presented more prominently, and similar assays could be performed in other cell systems.

      IWR1-POMA suppressed STF activity more effectively than IWR1 in APC-mut DLD-1 and SW480 CRC cells without TNKS overexpression (Fig. S8C). Similarly, IWR1-POMA provided a deeper suppression of STF signals in HeLa cells transfected with AXIN1 and β-catenin while expressing endogenous TNKS (Fig. 4G). These results suggest that inhibitor-induced TNKS scaffolding plays a significant role at endogenous TNKS expression levels. Following the reviewer’s suggestion, Fig. S7C is now Fig. 1G.

      (8) Line 237/238: "TNKS accumulation negatively impacts the catalytic activity of the DC (Figure 5D)" - the data do not show this. Beta-catenin levels are a surrogate readout for DC function (phosphorylation and ubiquitylation). Minimally, this requires rewording, with reference to beta-catenin levels.

      We have rephrased "TNKS accumulation negatively impacts the catalytic activity of the DC" as "TNKS accumulation negatively impacts the exchange of β-catenin in the DC."

      (9) Line 303-304: Beta-catenin is thought to exchange at beta-catenin degradasomes; this is clear from previous FRAP assays and the observation that phospho-beta-catenin accumulates in degradasomes upon proteasome inhibition (10.1158/1541-7786.MCR-15-0125). However, degradasome size hasn't, to my knowledge, been related to activity. Can this be clarified, please?

      We apologize for confusing β-catenin phosphorylation with β-catenin abundance. Here, we refer the catalytic activity of the DC to as the ability of the DC to promote β-catenin degradation rather than the kinetics of β-catenin phosphorylation. It is commonly observed that AXIN stabilization by TNKS inhibitors increases the DC size and reduces the β-catenin levels. As such, the induction of AXIN puncta by TNKS inhibitors is frequently used as an indicator of WNT/β-catenin pathway inhibition. However, we have found that, TNKS inhibition drives TNKS accumulation, which reduces the ability of the DC to promote β-catenin degradation. We agree that the DC only primes β-catenin but does not catalyze its degradation. We have revised our manuscript as follows: "increasing the local concentration of the DC components improves its 'effective activity'[50,51]."

      (10) There are previous hypotheses/proposals that the sensitivity of CRC cells to tankyrase inhibition correlates with APC truncation or PIK3CA status (10.1158/1535-7163.MCT-16-0578; 10.1038/s41416-023-02484-8). Have the authors considered expanding their cell line panel (Figure S7) to sample a wider range of cell lines, including some that are wild-type with regard to APC or Wnt/beta-catenin signalling in general? This would be a valuable addition to the work. Quantitated colony formation data could be moved to the main body of the manuscript.

      We have so far tested the effects of IWR1-POMA on the proliferation of DLD-1, SW480, HT-29, HCT116, and RKO cells (Fig. 6A and 6B). While a heterozygous Ser45 deletion in CTNNB1 confers resistance to IWR1-POMA, we did not observe sensitivity associated with APC or PIK3CA status. The ability of IWR1-POMA to suppress the growth of RKO cells expressing wild-type APC is consistent with a previous report that knockdown of both TNKS1 and TNKS2 stabilized PTEN to suppress cell proliferation and glycolysis in vitro and tumor growth in vivo (PMID: 25547115, Ref. 48) independently of the β-catenin pathway. We have added this new information as well as quantification of the colony growth results (Fig. S8A, S8B, S9A, S9F, and S9G) to the revised manuscript.

      (11) The manuscript only mentions toxicity (i.e., therapeutic window) in the last sentence of the Discussion section. As this is THE main challenge with tankyrase inhibitors (as mentioned above), can the authors expand their discussion of this aspect? Is there an expectation that PROTACs may be less toxic?

      As discussed above, evidence for on-target toxicity of WNT/β-catenin inhibition is mixed. Yet, the absence of dose-limiting toxicity for basroparib at doses up to 360 mg QD in human (PMID: 40964966, Ref. 64) is encouraging. PROTAC works by catalyzing target degradation, which is different from traditional catalytic inhibitors that require continuous target occupancy at a high level. It remains unclear whether the observed on-target toxicity of TNKSi is associated with TNKS accumulation at high doses, akin to the cytotoxicity induced by PARP1-trapping upon catalytic inhibition. We have included a brief discussion of the toxicity issue in the final paragraph of the Discussion section.

      (12) Figures 3, 4, 5A: For fluorescence microscopy experiments, can these be quantified, and can repeat data be included?

      We have included quantification data and replicate information for Fig. 3–5.

      (13) Figure 4, S6: An additional channel illustrating the distribution of cells (e.g., nuclei, cytoskeleton, or membrane) would be helpful for orientation and context for the AXIN1 signal.

      We have included cell outlines or nuclear staining for Fig. 3, 4, S6, and S7.

      (14) How were cytosolic fractions of cells prepared to assess cytosolic beta-catenin levels? This detail is missing from the methods.

      We have updated the Methods section to include additional details on the preparation of the cytosolic fractions of cells.

      Reviewer #3 (Public review):

      In this manuscript, Wang et al employ a chemical biology approach to investigate the differences between the enzymatic and scaffolding roles of tankyrase during Wnt β-catenin signalling. It was previously established that, in addition to its enzymatic activity, tankyrase 1/2 also plays a scaffolding function within the destruction complex, a property conferred by SAM-domain-dependent polymerization (PMID: 27494558). It is also known that TNKS1/2 is an autoregulated protein and that its enzymatic inhibition leads to accumulation of total TNKS proteins and stabilization of Axin punctae (through the scaffolding function of TNKS1/2), leading to rigidification of the DC and decreased β-catenin turnover. The authors surmised that this could, in part, explain the limited efficacy of TNKS1/2 catalytic inhibition for the treatment of colorectal cancers. To test this hypothesis, they evaluated a series of PROTAC molecules promoting the degradation of TNKS1/2 to block both the catalytic and scaffolding activities. They show that IWR1-POMA (their most active molecule) promotes more efficient suppression of beta-catenin-mediated transcription and is more active in inhibiting colorectal cancer cell and CRC patient-derived organoids growth. Mechanistically, the authors used FRAP to demonstrate that catalytic inhibitors of TNKS led to a reduced dynamic assembly of the DC (rigidification), whereas IWR1-POMA did not affect the dynamics.

      Overall, this is an interesting study describing the design and development of a PROTAC for TNKS1/2 that could have increased efficacy where catalytic inhibitors have displayed limited activity. Knowing the importance of the scaffolding role of TNKS1/2 within the destruction complex, targeting both the catalytic and scaffolding roles certainly makes sense. The manuscript contains convincing evidence of the different mechanisms of the PROTAC vs catalytic inhibitors. Some additional efforts to quantify several of the experiments and to indicate the reproducibility and statistical analysis would strengthen the manuscript. Ultimately, it would have been great to evaluate the in vivo efficacy of IWR1-POMA in an in vivo CRC assay (APCmin mice or using PDX models); however, I realize that this is likely beyond the scope of this manuscript.

      We thank the Reviewer for the helpful suggestions.

      I have some recommendations listed below for consideration by the authors to strengthen their study:

      (1) The title is slightly misleading, as it is already known that the scaffolding function of TNKS is important within the DC. The authors should consider incorporating the PROTAC targeting aspect in the title (e.g., PROTAC-mediated targeting of tankyrase leads to increased inhibition of betacat signaling and CRC growth inhibition).

      We have modified the title accordingly to "Targeting tankyrase scaffolding in the β-catenin destruction complex by PROTAC overcomes the limitation of catalytic inhibitors in cancer."

      (2) The authors should comment in the manuscript on the bell-shaped curve obtained with treatment of cells with the PROTACs (Figure S2C). This likely indicates tittering of the targets within a bifunctional molecule with increasing concentration (and likely reveals the auto-inhibition conferred by the catalytic inhibition alone).

      As suggested by the Reviewer, the bell-shaped dose-response likely originated from the formation of non-productive binary protein-ligand complexes at high PROTAC concentrations. We have added a sentence to clarify this unique behavior of PROTAC molecules.

      (3) The authors comment that using G007-LK as warehead was unsuccessful, but they do not show data. Do the authors know why this was the case?

      The structure-activity relationship of PROTACs is often unpredictable, as both the kinetics and thermodynamics of target and E3 ligase binding play important roles in promoting efficient target degradation. We have include data on G007-LK based PROTACs (Fig. S2D) in the revised manuscript.

      (4) Throughout the manuscript, the authors need to do a better job at quantifying their results (i.e., the western blots and the IF). For example, the degradation of TNKS1/2 in Figure 1D is not overly convincing. Similarly, the IF data in Figure 3 needs to be quantified in some ways. Along the same lines, the effect of IWR1-POMA treatments on the proliferation of cells and organoids should be quantified using viability assays... There is also no indication of how many times these experiments were performed and whether the blots shown are representative experiments. The quantification should include all experiments.

      We have included quantification of the immunofluorescence images, colony formation data, and Western blots in the revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) For clarity, can the authors use the official gene names, TNKS and TNKS2?

      We favor using TNKS1 and TNKS2 when referring to the protein for clarity and use TNKS for simplicity when referring to both proteins.

      (2) Line 92: The authors refer to TNKS2 "induction" - it remains unclear what is meant by "induction".

      We have changed "without induction" to "under basal conditions".

      (3) Can the authors please display molecular weight markers for Western blots throughout?

      (4) Line 144: The description "significantly more effectively" refers to Figure S5A, which shows a single, non-quantified Western blot. I don't think significance has been tested, and this statement should be reworded, or quantified aggregate data provided.

      We have added a Supplementary Information file showing molecular weight markers and quantification of Western blots.

      (5) Line 226: "plateaued at a much lower level" - can this be expressed more quantitatively in the text?

      We have included more quantitative information on the FRAP results.

      (6) Line 249: Can the authors repeat the cross-reference to Figure S7A here?

      We have repeated the cross-reference to the figures.

      (7) Line 266: The description of the experiment using the GSPT1/2 degrader CC-90009 would benefit from a brief recap of the purpose as not every reader will be familiar with this common PROTAC off-target. This is a very thorough analysis, though, and commendable.

      We have added background information on GSPT1 degradation to the revised manuscript.

      (8) Figure 1A: Can the number of repeats and the type of repeats be indicated, please?

      (9) Figure 2: Does n refer to biological or technical repeats?

      (10) Figure 5B, D: How many separate experiments are the data based on?

      (12) Figure S3D, S9A, D: number and types of repeats and the nature of the displayed data and error bars need to be included, please.

      (13) Figure S6B, S7B: I can see three data points, but it would still be helpful to state the number and type of repeats in the legend.

      (14) Figures S9A, S9D: There is value in showing the cumulative data from several repeats in the main figure (Figure 6, which currently is only qualitative) rather than the supplementary material.

      (15) Where single Western blots are shown, can the authors indicate how many experiments they are representative of?

      We have included the number of biological repeats for all data.

      (11) Figure S2C: For most graphs, the main response of interest occurs at low compound concentrations. The y-axis scale does not always help the reader to appreciate the effects, as the response seems small against the magnitude of the hook effect. Interrupting the y-axis as in the final panel may help, with y-axis scales consistent over all panels in the figure.

      We have updated Fig. S2C to emphasize on the degradation efficacy.

      (16) The authors may want to give further method details for some of their assays to facilitate replication of their experiments in the future. For example, the STF assay description is currently quite minimalistic. I assume the assay is fairly robust, though. Other details include cell media (general media details and specific additives and their concentrations in the 3D spheroid formation assay), etc. A general look at the methods section will likely be beneficial.

      We have updated the Methods section to provide more detailed experimental information.

      Reviewer #3 (Recommendations for the authors):

      (1) In Figure 2A, one of the most important findings of the manuscript is that IWR1-POMA induced promoted deeper suppression of beta-catenin-mediated transcription. This seems to be the case only at 3.2uM. Is it statistically significant? What are the data points on this graph? What are the error bars?

      We have included statistical analysis as Fig. S5G.

      (2) On Figure 2C and 2D, do the authors know why the TNKS20M1054V mutant is much better at promoting signaling than the TNKS1-PD ? Is it expression levels?

      It is indeed interesting that TNKS2-M1054V promoted significantly stronger WNT signaling than TNKS1-PD. The basis for its strong scaffolding effect is unclear.

      (3) In Figure 4C, the authors claim that when cells are treated with IWR1-POMA, AXIN1 is distributed diffusely throughout the cytoplasm. It appears that small punctae are visible.

      Quantitative analysis (Fig. 4F) suggest that the size of AXIN1 puncta upon IWR1-POMA is rather insignificant.

      (4) Label on Figure 1D has a spelling error TNKS1/2.

      Corrected.

    1. Author response:

      We thank the editors and reviewers for their thoughtful assessment of our manuscript, and for recognizing openretina as a valuable and timely resource for the retinal modelling community.

      We are especially glad that the reviewers appreciated the motivation of the project, the focus on standardization and reproducibility, and the potential of the platform to support systematic benchmarking and community-driven model development.

      We also understand the concerns raised. In the revision of the manuscript, we will strengthen the conceptual discussion of how predictive models, including the current “Core + Readout” models, can contribute to retinal neuroscience alongside more mechanistic and circuit-based approaches. This is a central matter for us, and one that some of us have recently addressed in a broader review on current trends in retina modelling (see https://doi.org/10.1016/j.visres.2026.108854). We will draw on this perspective to better articulate when predictive models are useful, where their limitations lie, and how openretina can provide infrastructure for comparing functional, normative and mechanistic models within a shared framework.

      We will also clarify the scope and limitations of the in-silico analysis methods provided within openretina. This will include a more explicit discussion of how MEIs, gradient-field analyses, and model-weight visualisations should be interpreted.

      Furthermore, we will add more information that will help the reader better judge different aspects of dataset quality, including, for example, spike-sorting or calcium-processing information and explainable-variance distributions. We note, however, that there are many subtle details about experimental workflows that are difficult to capture in compact indicators. In addition, we will make it clearer that the manuscript represents a snapshot of a living resource: The website, dataset cards, documentation, and repository will be the primary source of this information, especially as new datasets are contributed.

      Finally, we will of course address the technical clarifications raised by the reviewers, with the aim of making the manuscript more accessible overall.

      We are grateful for the reviewers’ constructive comments and believe that addressing these points will make our presentation of openretina clearer and more useful to the community.

    1. Author response:

      We thank the editors and reviewers for their thoughtful comments. Below, we list our provisional responses to the reviewers’ major points:

      On the rationale for CaMKIIα versus Thy1-driven stimulation and physiological relevance: We agree that we did not make clear the motivation for using CaMKIIα-driven stimulation, distinct from the Thy1-driven paradigm in our previous work (Williams et al., 2026). Using the Thy1 driver, both excitatory and inhibitory cells received direct theta drive. In contrast, CaMKIIα expression is largely restricted to principal neurons. Comparing these models lets us isolate a "driven I-cell" PING mechanism from the "E cell recovers first" mechanism relevant when interneurons are also directly driven.

      Regarding physiological relevance, Gonzalez-Sulser et al. (2014) found that septal GABAergic projections selectively and directly inhibit mEC interneurons, rather than exciting either principal cells or interneurons, implying that theta drive in vivo likely acts through rhythmic disinhibition of interneurons rather than direct excitation of any cell type. Neither the Thy1 nor the CaMKIIα paradigm reproduces this disinhibitory mechanism: both rely on excitatory optogenetic drive rather than rhythmic inhibition of interneurons, and replicating the natural drive (tonic excitatory tone plus rhythmic, interneuron-selective inhibition) is technically difficult in acute slices, which are largely quiescent without exogenous stimulation. We therefore view CaMKIIα and Thy1 as complementary approximations, each isolating a different circuit interaction. If forced to choose, we’d argue that the CaMKIIα is a better model of disinhibition of excitatory neurons. We will revise the Discussion regarding this point.

      On reproducibility of the voltage imaging findings: We thank the reviewer for this comment and agree that clarification is warranted.

      The voltage imaging dataset combines two levels of analysis with different sample sizes. The population-level firing and spike-correlation analyses (Fig. 5F–H) are pooled across multiple imaging sessions (n = 240 neurons). The spatial clustering analysis of subthreshold voltage correlations (Fig. 6, and the corresponding example traces in Fig. 5A–E) are drawn from a single representative recording session, as the reviewer correctly notes. We have voltage imaging data from 14 fields of view (1 FOV per slice) across 6 mice (240 neurons total; 3–41 neurons per FOV). In revision, we will extend the clustering and spatial-correlation analysis from Fig. 6 across sessions to assess whether the reported organization is reproducible, rather than relying on a single example. We will also revise the text to distinguish clearly which analyses are single-session versus pooled.

      On restricting the computational model of excitatory neurons to stellate cells: We modeled stellate cells as the excitatory population because they are the principal cells reciprocally connected to fast-spiking PV+ interneurons (Fuchs et al., 2016), the interneuron class most directly implicated in theta-nested gamma. Pyramidal cells, by contrast, are primarily connected via 5-HT3a-positive interneurons (Fuchs et al., 2016), with the exception of a subset of "intermediate" pyramidal cells that do show reciprocal PV+ connectivity. Our model, which captures the full measured heterogeneity of stellate cell and PV+ interneuron intrinsic properties and their reciprocal connectivity, is, to our knowledge, the most biophysically constrained implementation of this specific microcircuit to date. Incorporating the PV+-connected intermediate pyramidal population is a natural next step. Because this refinement, which requires more experimental data, is nontrivial and beyond the scope of this study, we will note this explicitly as a limitation of the current model in the revised Discussion.

      In vivo comparison (temporal/phase-locking): We agree that grounding our findings in existing in vivo data strengthens the study and will add these comparisons to the revision.

      Our whole-cell recordings reproduce the temporal organization in vivo and provide further insights into cell-type differences between the principal cells. All cell types were strongly phase-locked to theta, while gamma phase-locking declined across successive spikes, with stellate cells decoupling after the first spike and pyramidal cells after the second. This earlier decoupling in stellate cells may contribute to their weaker theta rhythmicity reported in freely moving rats (Ray et al., 2014; Tang et al., 2014). In extracellular recordings from behaving mice, spike-train cross-correlation identifies putative monosynaptic excitatory connections (1–4 ms) from principal cells onto fast-spiking interneurons (Latuske et al., 2015); the excitation-to-inhibition offset we measured is of comparable magnitude, here resolved as a synaptic-current delay in electrophysiologically classified cell types.

      We note that bursting and theta engagement have been assigned inconsistently across in vivo datasets. Bursty cells are preferentially classified as putative stellate by spikepattern classifiers (Latuske et al., 2015), while anatomically identified pyramidal cells are reported as the bursty, theta-rhythmic population in other work (Ebbesen et al., 2016). Because our cell-type assignments are based on subthreshold intrinsic properties (membrane sag, time constant) rather than spike patterning, our phase-locking results are independent of this classification ambiguity.

      In vivo comparison (spatial organization): We agree high-density silicon-probe datasets are the appropriate reference here. To our knowledge, the anatomical distribution of gamma-locked spiking in superficial mEC has not been characterized in vivo. The highest-density available recordings (Gardner et al., 2022) analyze population activity in the decoded state rather than tissue coordinates, do not examine gamma, and are restricted to grid cells. We regard the dissociation we observe between spatially clustered subthreshold input and spatially distributed spiking as a principal advance of the present study, and as a testable prediction for future high-density recordings.

      Ebbesen CL, Reifenstein ET, Tang Q, Burgalossi A, Ray S, Schreiber S, Kempter R, Brecht M. 2016. Cell Type-Specific Differences in Spike Timing and Spike Shape in the Rat Parasubiculum and Superficial Medial Entorhinal Cortex. Cell Reports 16:1005–1015. DOI: https://doi.org/10.1016/j.celrep.2016.06.057

      Fuchs EC, Neitz A, Pinna R, Melzer S, Caputi A, Monyer H. 2016. Local and Distant Input Controlling Excitation in Layer II of the Medial Entorhinal Cortex. Neuron 89:194–208. DOI: https://doi.org/10.1016/j.neuron.2015.11.029

      Gardner RJ, Hermansen E, Pachitariu M, Burak Y, Baas NA, Dunn BA, Moser M-B, Moser EI. 2022. Toroidal topology of population activity in grid cells. Nature 602:123–128. DOI: https://doi.org/10.1038/s41586-021-04268-7

      Gonzalez-Sulser A, Parthier D, Candela A, McClure C, Pastoll H, Garden D, Sürmeli G, Nolan MF. 2014. Gabaergic projections from the medial septum selectively inhibit interneurons in the medial entorhinal cortex. Journal of Neuroscience 34:16739–16743. DOI: https://doi.org/10.1523/JNEUROSCI.1612-14.2014, PMID: 25505326

      Latuske P, Toader O, Allen K. 2015. Interspike Intervals Reveal Functionally Distinct Cell Populations in the Medial Entorhinal Cortex. Journal of Neuroscience 35:10963–10976. DOI: https://doi.org/10.1523/JNEUROSCI.0276-15.2015

      Ray S, Naumann R, Burgalossi A, Tang Q, Schmidt H, Brecht M. 2014. Grid-Layout and Theta-Modulation of Layer 2 Pyramidal Neurons in Medial Entorhinal Cortex. Science 343:891–896. DOI: https://doi.org/10.1126/science.1243028

      Tang Q, Burgalossi A, Ebbesen CL, Ray S, Naumann R, Schmidt H, Spicher D, Brecht M. 2014. Pyramidal and Stellate Cell Specificity of Grid and Border Representations in Layer 2 of Medial Entorhinal Cortex. Neuron 84:1191–1197. DOI: https://doi.org/10.1016/j.neuron.2014.11.009

      Williams B, Vedururu Srinivas A, Baravalle R, Fernandez FR, Canavier CC, White JohnA. 2026. Fast spiking interneurons autonomously generate fast gamma oscillations in the medial entorhinal cortex with excitation strength tuning ING–PING transitions. eneuro ENEURO.0452-25.2026. DOI: https://doi.org/10.1523/ENEURO.0452-25.2026

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates how Ca2+ levels inside the RGCs' mitochondria relate to whether these cells survive or die after injury to the optic nerve. The authors used advanced in vivo fundus live imaging techniques in mice to watch these changes unfold in real time, combined with genetic and drug-based tools to alter calcium flow into these compartments. Their central finding is a striking paradox: cells that naturally survive injury tend to have higher baseline calcium levels in these compartments, yet experimentally reducing calcium entry protects the broader population of cells from death.

      Strengths:

      The authors are applying sophisticated biosensors to track cellular chemistry in living animals over days and weeks. The tools and methods are creative and direct to detect the longitudinal RGC degeneration with mito-Ca2+ imaging. The topic and research aspect are novel and attractive. The results are significant, showing a clear relationship between the mito-Ca2+ regulatory machinery and cell survival.

      Weaknesses:

      The details of the mitochondrial-located signal of the Ca2+ sensor need to be further proved in the mito-matrix or between the mito-membranes. The study primarily describes a correlation and a surprising experimental outcome without fully explaining the underlying biological reasons for the paradox. While the evidence supporting the phenomenon is good, the mechanistic insight into why high calcium is linked to survival, or why lowering it helps after injury, remains limited.

      We appreciate Reviewer #1’s assessment of our manuscript. We also agree that we should have more clearly indicated that our mitochondrial Ca2+ sensor (Cox8-Twitch2b) is localized to the mitochondrial matrix. The Cox8-mitochondrial localization peptide is a well-established tool first identified in 1992 by Rizzuto and colleagues (Rizzuto, Simpson and Pozzan, 1992). We should have cited this work in our manuscript and will add it to our references. Further, as discussed in our submission, Cox8-Twitch2b has previously been validated for mitochondrial Ca2+ measurements in CNS axons (Witte et al., 2019). Thus, given the decades of use and characterization for this toolset, and the fact that we have pharmacological data supporting mitochondrial matrix localization of Cox8-Twitch2b, we do not feel it is strongly necessary to demonstrate mitochondrial matrix versus inner membrane space localization. However, we could attempt immuno-electron microscopy if this is deemed critical.

      We also agree that the mechanism by which reducing mitochondrial Ca2+ is protective would be satisfying and strengthen this study. But we feel it is beyond the scope of this project. It is likely manifold since mitochondrial Ca2+ impacts many vital cellular functions relevant to pathology including metabolism and apoptosis. We ultimately believe that an adequate investigation of these mechanisms would significantly slow down the dissemination of the core novel findings presented herein.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by McCraken and colleagues provides a continuation of their 2023 study (Cell Reports 42:113165) characterizing calcium regulation in retinal ganglion cells (RGCs) after acute optic nerve damage (a 10s crush using an intraorbital approach). This work is principally focused on how mitochondrial calcium stores change in both RGCs that are resilient and susceptible to injury. They report that resilient RGCs typically exhibited high calcium levels, but paradoxically, manipulating mitoCa2+ levels was more protective when the stores were reduced. Overall, regardless of susceptibility, mitoCa2+ levels decreased after injury, which is opposite to other reports that mitoCa2+ increases in degenerating neurons. The manipulation of mitoCa2+ was conducted both pharmacologically (Ru265) and by overexpression or knockdown of a primary calcium uniporter MCU. The evaluation of mitoCa2+ was conducted by using a reporter (Twitch2b) that was targeted to the mitochondria.

      Strengths:

      Many of the experiments are elegant and well-performed.

      Weaknesses:

      (1) Some experiments require further controls to validate that reagents are doing what they are intended to do.

      We agree with Reviewer #2 that our AAV manipulations of shMCU and MCU overexpression should be analyzed to verify how they alter mitochondrial Ca2+. To do this, we will co-express gene therapy vectors to lower and raise MCU expression with mito-Twitch2b biosensor and perform direct measurements of mitochondrial Ca2+. We will then determine if there is a relationship between gene expression level (inferred by mCherry intensity) and mitochondrial Ca2+ within samples, and if mean mitochondrial Ca2+ levels in treatments are higher or lower than mCherry reporter only controls.

      (2) Some findings can have alternate interpretations that are not considered.

      We will expand our Results and Discussion sections to broaden the interpretations of our data.

      (3) There is a broad generalization to the biology of all RGCs that may not be biologically relevant to different RGC subtypes.

      We agree that a more fine-grained understanding of RGC mitochondrial Ca2+ diversity would make interpretations of our data stronger. In our revisions, we will thus expand the number of RGC families in which we directly measure homeostatic mitochondrial Ca2+ levels. To do this, we will perform in vivo mito-Twitch2b measurements, collect and fix retinal wholemounts and immunostain for ON-OFF-direction selective RGCs using the marker CART and F-RGCs using the marker Foxp2. This will provide a complement of well-surviving RGC types (alpha and intrinsically photosensitive RGCs already examined) and poorly-surviving types.

      Reviewer #3 (Public review):

      Summary:

      Following previous work that demonstrated a relationship between higher homeostatic cytosolic calcium and lower retinal ganglion cell (RGC) apoptosis following injury to their axons, McCracken et al. investigated whether homeostatic calcium levels of the endoplasmic reticulum (ER) or mitochondria provide additional insights into the mechanisms by which calcium influences RGC survival. Their study reveals that homeostatic mitochondrial calcium shows a similar positive correlation with RGC survival. Despite that correlation, pharmacologic or genetic methods to lower mitochondrial calcium improved, rather than reduced, the survival of injured RGCs, while a genetic approach intended to increase mitochondrial calcium resulted in more RGC loss. These findings highlight the complexities of calcium regulation in modulating neuronal survival and raise important questions of how homeostatic levels of mitochondrial calcium affect stress responses that themselves can be either neuroprotective or neurodegenerative.

      Strengths:

      This study tackles an intriguing hypothesis that differences in calcium ion homeostasis in specific organelles may contribute to differences in survival of various RGC subtypes after optic nerve injury. This is a technically demanding question, and a primary strength of this work is its attention to, and meticulous reporting of, appropriate controls and, where applicable, seemingly contradictory results. Among these are careful evaluation of the effects of drug (or vehicle) delivery and genetic manipulations with and without injury and over extended time courses. The combination of thoughtful pharmacologic and genetic approaches makes for a thorough analysis of a challenging set of questions. The result is a study that provides a helpful perspective on the complicated roles that calcium, and especially mitochondrial calcium, can play across neuronal insults, neuronal types, and neuronal subtypes.

      Weaknesses:

      Given the paradoxical results, it would be helpful to have a clearer picture of how strongly the overexpression and knockdown of MCU altered the mitochondrial calcium levels. There may be potential for extraordinarily strong effects that would need to be tuned by using different shRNAs or promoters to more closely align with the observed differences between surviving RGCs and those that die. The investigation includes a relatively small number of resilient RGC subtypes, using the markers SPP1 and TBR2, raising questions of how generalizable the trend is between mitochondrial calcium levels and RGC resilience. The analysis and implications of Figure 3D might benefit from including not only the provided 50:50 split between "high" and "low" but also views of the data after splitting into thirds, fourths, and perhaps even fifths. The authors' inference that higher homeostatic calcium in more resilient RGCs may result in chronic mitochondrial stress is intriguing and worthy of more experimental investigation than is currently provided.

      We agree with the feedback from Reviewer #3, especially as it aligns with input from other reviewers. As these points agree with aspects above we will briefly reiterate our proposed revisions. We will validate the true effects on mitochondrial Ca2+ levels after gene therapy treatments by co-injecting AAV-mito-Twitch2b and AAV-shMCU or AAV-MCU. We will measure mitochondrial Ca2+ levels and correlate these levels with mCherry reporter expression intensity to determine the effect size of these treatments, and compare sample mean mitochondrial Ca2+ levels with those of mCherry control AAV.

      To further map the variance in homeostatic mitochondrial Ca2+ levels to RGC types we will perform in vivo mito-Twitch2b imaging, and then immunostain for ON-OFF-direction selective RGCs (CART) and F-RGCs (Foxp2), two poorly surviving RGC types.

      Lastly, we agree with Reviewer #3 that finer delineation between mitochondrial Ca2+ levels and their relationship to survival may be informative. We will split RGCs into smaller subgroups based on homeostatic mitochondrial Ca2+ levels and examine their survival outcome.

      Overall, we thank the Reviewers for their feedback, and believe the suggested changes will greatly strengthen our study.

      REFERENCES

      Rizzuto R., Simpson A.W. and Pozzan T. (1992). Rapid changes of mitochondrial Ca2+ revealed by specifically targeted recombinant aequorin. Nature, 358 (6384): 325-327.

      Witte M.E., Schumacher A-M., Mahler C.F., Bewersdorf J.P., Lehmitz J., Scheiter A., Sanchez P., Williams P.R., Griesbeck O., Naumann R., Misgeld T. and Kerschensteiner M. (2019). Calcium influx through plasma-membrane nanoruptures drives axon degeneration in a model of multiple sclerosis. Neuron, 101(4): 615-624.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Chen et al. describe metabolic phenotypes in Dp16 Down Syndrome mice, specifically the Dp(16)1Yey/+ mice - segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs. The group has performed metabolic phenotyping data in chow and high-fat diets, as well as undertaking a transcriptomic and metabolomic approach in tissues such as white and brown adipose tissues, liver, skeletal muscle, and hypothalamus to reveal both shared and sex-specific differences. The group describes sexual dimorphism in body weight, body temperature, food intake, and physical activity. Core shared features are insulin resistance, glucose intolerance, impaired lipid clearance, and dyslipidaemia in the Dp16 mice. They report tissue signatures of immune activation and a pro-inflammatory state, ER and oxidative stress, fibrosis, impaired glucose and fatty acid catabolism, altered lipid and bile acid profiles, and reduced mitochondrial respiration in Dp16 mice.

      Strengths:

      Overall, this is a good study with detailed, comprehensive data from an excellent group who have previously published on metabolic phenotyping of 2 other Down Syndrome mouse models. Although somewhat descriptive, it does certainly add to the current field and understanding of strengths and weaknesses of Down Syndrome mouse models, as well as identifying new features whilst strengthening previously suggested mechanisms.

      Weaknesses:

      Many aspects of this study have been described in other Down syndrome mouse models, though there are certainly aspects that are new. It would be useful if the authors could do a direct critique and comparison with previous publications in the area, utilizing the same Down Syndrome mouse model. There are also a few limitations in the number of animals used and the interpretation of the data that should be acknowledged.

      We have cited all relevant publications using Down syndrome mouse models. Regarding the Dp16 model, we have cited and discussed the only other study addressing metabolic aspects beyond body weight (Reference #138; PMID: 39803786). While that study reported glucose intolerance, insulin resistance, and defective insulin secretion, we did not measure pancreatic insulin content in our mice. Crucially, while the previous study found no sexual dimorphism, our study observed extensive sexual dimorphism in body weight gain, tissue-specific gene expression, and serum and liver metabolite changes.

      Regarding sample size, we used 6 mice per genotype per sex for transcriptomic and metabolomic analyses; this is constrained by the cost of performing these omics-type analyses. For mitochondrial respiration assays, we used 9–10 mice, and for most other in vivo and ex vivo assays, we utilized 12–15 mice, with some assays exceeding 20. We believe these sample sizes are robust and appropriate for this study.

      Reviewer #2 (Public review):

      Summary:

      Human DS is associated with metabolic dysfunction in humans, but the precise details of this have not been studied in detail. Here, the authors use a mouse model of DS to study systemic metabolic and transcriptional responses in key metabolic tissues to provide a deep understanding of the metabolic changes associated with DS. As part of his work, the authors also aimed to help inform the selection of a mouse model that best reflects the metabolic profile of DS, through comparison with other DS model metabolic data.

      The data presented in this model will be of interest to those in the field of metabolism. The immediate impact is unclear, but the breadth of data presented makes this a very useful resource.

      Strengths:

      (1) This work builds on other comprehensive analyses that the authors have performed in other DS mouse models.

      (2) The authors note common metabolic disturbances between male and female mice (e.g., insulin resistance) alongside clearly sexually dimorphic phenotypes (e.g., body weight). Studying both sexes in this context is important.

      (3) The authors have written the paper in a way that integrates a large number of observations well. There is complex data, and a high degree of sexual dimorphism. The study has generated a valuable and wide-ranging dataset comprising molecular, biochemical, and physiological data that will be useful for further, more mechanistic studies of metabolism in DS.

      (4) For specific observations, like the findings of altered body temperature in male and female mice, the authors undertake follow-up hypothesis-driven analyses of BAT mitochondria and specific hormones. Although these analyses do not explain the change in temperature, they ensure the study is not purely descriptive in nature.

      Weaknesses:

      (1) Assessing metabolism using dynamic testing is a strength. ITT, GTT and LTTs are included.

      (2) The dosing for GTTs, ITTs and LTTs was performed per body weight. But the mice under chow and HFD had different body weights. This may compromise the interpretation of the data. Further, ITTs are presented as percentage change, and this can be heavily influenced by baseline glucose measures. The changes appear quite dramatic, so can the authors plot the raw data instead?

      We have updated the ITT data plots to show raw glucose values instead of percentage change. Regarding the dosing, we believe basing it on body weight is an appropriate approach. This method is consistent with nearly all published rodent studies, as blood volume and metabolic tissues such as skeletal muscle and adipose tissue scale with body weight. Adjusting for weight prevents potentially erroneous conclusions. As for the diet groups, we compared WT and Dp16 mice only within the same diet group (Chow or HFD) rather than across different diets. We believe this ensures a valid and appropriate comparison for our study.

      (3) In addition, throughout the manuscript, it is not clear which tissues are the most dominant in disrupting metabolism. The ITT and GTT are composite measures across tissues. Tissue-specific analyses using a clamp technique or isolated tissues may provide more clarity here.

      Our data suggest a systemic metabolic deficit across multiple tissues, supported by tolerance tests, pan-tissue transcriptomic analyses, and liver and serum metabolite profiling. This is consistent with the triplication of genes in Down syndrome, several of which have known metabolic roles as highlighted in our discussion. We do not have evidence to support the role of a dominant tissue that contributes to the systemic metabolic dysfunction.

      Regarding the suggestion to use a clamp technique, we agree this would effectively determine whether insulin resistance is localized in the liver or skeletal muscle. However, we do not currently have the necessary equipment at Johns Hopkins University to perform these experiments. Conducting this work would require sending separate cohorts of WT and Dp16 male and female mice (on both chow and HFD) to an NIH-funded Mouse Metabolic Phenotyping Centre (MMPC). While we appreciate the value of this approach, we believe such labor-intensive experimentation falls beyond the scope of the present study.

      (4) One of the aims of the study was "to help inform the selection of mouse model that best reflects the metabolic profile of DS". The discussion does not contain a comparison between the previous work on different strains and relative to known human data.

      We chose not to include a comparison of different mouse models in the "Discussion" section because we previously highlighted the widely used Down syndrome models (Ts65Dn, Tc1, and TcMAC21) and their associated caveats in the "Introduction." Given the significant limitations of those models such as hypermetabolism in TcMAC21 and the presence of 41 triplicated protein-coding genes unrelated to human chromosome 21 we focused our in-depth metabolic analyses on the Dp16 model, which does not share these issues. We felt that restating this information in the "Discussion" would be unnecessarily repetitive.

      (5) Data availability. Raw metabolomic data should be made available.

      We have uploaded all metabolomics data, along with details regarding sample processing and data analysis, to the Metabolomics Workbench, an NIH-funded public repository. We have updated the "Methods" and "Data Availability" sections of the manuscript to include this information and the corresponding access link.

      Reviewer #3 (Public review):

      Summary:

      The article by Chen et al. describes the comprehensive metabolic profiling of DP16 mice, a Down syndrome model that carries a duplicated segment of the mouse chromosome syntenic to human chromosome 21. The authors note that this model is superior to previously used models, based on genetics, as ~65% of the chromosome 21 orthologues. The metabolic phenotypes also appear to be more consistent with those observed in humans with Down Syndrome. The study lays the groundwork for a more detailed genetic dissection of dosage-sensitive genes that contribute to the metabolic deficits observed in Down Syndrome.

      Strengths:

      There is an enormous amount of data in this manuscript, and the methods are described with adequate attention to detail. A strength of the manuscript is that both male and female mice were analyzed, so that concordant and discordant phenotypes were identified. Both males and females had evidence of insulin resistance. Transcriptomic and metabolomic data revealed impaired pathways for lipid metabolism, a pro-inflammatory state, reduced mitochondrial health and oxidative stress. Although the effects of a high-fat diet on weight gain were divergent, this diet caused worsened insulin resistance in both males and females.

      The discussion is excellent. Limitations of the study are well described. This reviewer does not identify any critical missing data.

      Weaknesses:

      It might have been helpful to have included blood pressure measurements, given the differences in 19-Nor-deoxycorticosterone. The discussion references several articles that describe sex-dependent differences in metabolic phenotypes in humans with Down syndrome, and it might have been helpful to state more explicitly whether these differences correlate with those observed here in mice.

      We appreciate the suggestion of blood pressure measurements. While we agree this is an important metric, given the metabolic focus of the present study and the significant volume of data already presented, we feel that blood pressure analysis is beyond the current scope and better suited for a follow-up study.

      Our study highlights sex differences in metabolic phenotypes in individuals with Down syndrome. While most published human studies focus on a limited set of parameters such as body weight, adiposity, serum lipoprotein profile, and fasting lipid/glucose levels our mouse data remain generally concordant with these findings. Beyond these standard measurements, we also observed substantial sex differences in pan-tissue transcriptomes as well as serum and liver metabolites.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) A major question is how these findings compare to data that have previously been published. For example, Lamantia et al. Bone 2024 and Dard et al. European Journal of Pharmacology 2025 both report no changes in body weight using the same Dp(16)1Yey Down syndrome mouse model? There is also a recent publication on liver dysfunction in Down Syndrome using the same mouse model. It would be useful to understand some of the similarities and differences of what is being reported by Dunn et al. Cell Rep 2026. In this assessment, there is an in-serum alanine transaminase (ALT) level, which was not the case in Dunn et al?

      For the Lamantia et al. Bone 2024 study, the authors only measured the body weights of Dp16 mice at 6 weeks of age. Our findings at 6 weeks align with Lamantia et al., showing no weight differences between Dp16 and WT mice of either sex (Fig. 2A and C). For the Dard et al. 2025 study, the authors only measured the body weights of Dp16 mice at 12 weeks old (P90) and observed no differences in body weights between genotype of either sex. At 12 weeks of age, we also did not observe body weight differences between Dp16 male mice and WT littermates (Fig. 2A). However, at 12 weeks of age, the Dp16 female mice clearly gained more weight compared to WT littermates (Fig. 2C). Our study tracked weights weekly from 6 to 16 weeks, revealing that while Dp16 females start at weights similar to WT littermates, the groups diverge over time. The reason for the difference between our findings and the single-point measurement by Dard et al. is unclear. Notable variables include:

      Mouse Sourcing: We obtained all cohorts and littermate controls from Jackson Laboratory, while Dard et al. bred their mice in-house.

      Diet: We used Envigo standard chow (catalogue # 2018SX). Dard et al. did not specify the chow used in their study.

      It remains uncertain whether these or other environmental factors contribute to the observed weight differences in female mice.

      In the Dunn et al study (Cell Rep 2026), they also performed metabolic analyses on serum and liver tissue in Dp16 mice. Consistent with their metabolic analyses of serum and liver tissue in Dp16 mice, we also observed the upregulation of multiple bile acids, including taurochenodeoxycholic, tauromuricholic, taurolithocholic, and lithocholic acids. Furthermore, our findings align with theirs regarding the transcriptomic and biochemical signatures of hepatic inflammation and fibrosis. However, there are two notable differences between our studies:

      (1) Liver Injury Markers: We observed an elevation in serum ALT, whereas the Dunn et al. study did not.

      (2) Sex Differences: We identified significant sex differences in the Dp16 transcriptome and metabolome. In contrast, Dunn et al. reported minimal to no sex differences and consequently combined male and female data for all analyses.

      Because Dunn et al. combined male and female data, a sex-stratified comparison between our results (separated by sex) and theirs was not feasible.

      (2) It would be important to understand trends in wild-type animals compared to Dp16 mice. For example, the sex specific and non-specific features - are any of these described in obesogenic wild-type animals fed on a high-fat diet? I.e., are the same features at play and just exacerbated in Dp16, or is this a Dp16-specific feature of systemic metabolism?

      Published literature indicates that WT females typically gain significantly less weight on a high-fat diet (HFD) than WT males. However, our data suggest that the weight gain patterns observed in Figure 6A and C are specific to the Dp16 genotype. Dp16 females gained substantially more weight during the first six weeks of HFD before WT females caught up. In contrast, Dp16 males showed robust initial weight gain comparable to WT controls, but their weight plateaued after seven weeks while WT controls continued to gain, leading to a clear divergence (Fig. 6A).

      Other metabolic parameters also appear specific to the Dp16 model. On a standard chow diet, WT mice of both sexes generally do not exhibit glucose intolerance, insulin resistance, dysregulated lipoprotein profiles (VLDL-TG), or an impaired capacity to handle lipid loads. We observed all of these features in our Dp16 male and female mice (Fig. 3). Furthermore, transcriptomic analyses of Dp16 mice on standard chow revealed gene signatures of inflammation, fibrosis, and oxidative stress that are absent in WT mice.

      When challenged with HFD, while WT mice typically develop glucose intolerance and insulin resistance, the triplicated genes in Dp16 mice significantly exacerbated this metabolic deterioration. This is reflected in the worsening of glucose control and insulin sensitivity observed in our tolerance tests.

      In summary, most of these metabolic features are specific to Dp16 mice on a standard chow diet and are further exacerbated when combined with a high-fat diet.

      (3) Food intake data is difficult to interpret when weight has already diverged, as bigger animals will eat more food. Hence, the higher food may be a consequence rather than a cause of the weight gain (data in Figure 1).

      The reviewer makes a valid point. Since physical activity and energy expenditure do not differ significantly between Dp16 females and WT controls (Fig. 2F), the observed increase in food intake may indeed contribute to the higher body weights in Dp16 female mice.

      To rigorously confirm this, food intake would need to be measured between 6 and 8 weeks of age, prior to the divergence in body weight. Unfortunately, we did not measure food intake at that earlier time point.

      (4) The n numbers seem to vary significantly. For example, the use of n=6 for metabolic studies is generally rather small and underpowered. For the seahorse data, another concern is the snap freezing of samples before Seahorse assessment. For example, snap freezing of samples has been shown to increase certain metabolites. Freeze-thaw tissues often show a significant reduction in optical redox ratio.

      Regarding the transcriptomics and metabolomics studies, we utilized N=6 mice per tissue per sex. While we agree that a larger sample size is always preferable, the high cost of OMICS analyses covering 144 RNA-seq and 48 metabolomics samples limited our capacity to increase this number. However, N=6 remains a robust and standard approach for these specific assays. For the majority of our other in vivo and ex vivo data, we employed a higher sample size of 12-15 mice per genotype per sex to ensure statistical rigour. For a few assays, we have sample size of over 20.

      Regarding the respirometry analysis, we acknowledge the limitations of using frozen tissue. We chose this method because it allowed us to perform Seahorse assays on multiple tissues from 9-10 mice, which is a significant sample size for this type of analysis. The alternative isolating mitochondria from fresh tissue would have restricted our ability to process multiple tissues from a large number of animals on the same day due to the length of the protocol. We believe this trade-off was necessary to maintain a high sample size across various tissues.

      (5) For oestradiol measurements, were the samples taken at the same times within the estrous cycle? This may affect the comparability of female Dp16 and WT mice?

      Regarding our protocol, blood samples were collected between 11:00 AM and noon, with food removed two hours prior. While we did not specifically monitor the oestrous cycle of the female mice, serum samples for both the Dp16 females and WT littermates were collected on the same day and at the same time to ensure comparability across the groups.

      (6) Body weight reduction and organ size reduction on an HFD are especially interesting. Could enhanced inflammation and fibrosis be the root cause of this? Are there other mouse models where this is the reason?

      On a high-fat diet, we observed a reduction in iWAT and gWAT fat depot weights in both male and female Dp16 mice, which is consistent with their lower overall body weights (Fig. 6 - figure supplement 3). Conversely, Dp16 females fed a high-fat diet showed increased heart and kidney weights. Despite their lower adiposity, the Dp16 mice on this diet exhibited greater insulin resistance and glucose intolerance (Fig. 7). This suggests that the worsening of glucose control is independent of obesity. While we observed signatures of inflammation and fibrosis, we do not yet have direct mechanistic evidence demonstrating that these factors causally impaired glucose and lipid metabolism.

      (7) The authors are circumspect throughout to avoid over-claiming, as the majority of data is observational. One exception: "Many bile acids serve as ligands for nuclear hormone receptors (e.g., FRX and TGR5) that control various aspects of glucose and lipid metabolism (74, 75), and extensive changes in circulating bile acids are contributing, at least in part, to the systemic metabolic phenotypes in Dp16 mice." The authors have not shown a direct link between bile acids and metabolism in this model. Please edit.

      We have edited the text accordingly.

      Minor:

      (1)"Most human studies at the whole-body level are limited to assessing the impact of trisomy 21 on food intake, adiposity, physical activity level, and energy expenditure in adolescents or adults with DS"

      While we were uncertain of the reviewer's specific intent regarding the suggested changes, we have rephrased the sentence for clarity.

      (2) It is somewhat surprising that T3 is elevated, although there are reports of T3 elevation in visceral obesity in humans (e.g., Sun Nam et al., Obes Res Clin Pract, 2010).

      We observed that T3 levels did not differ by genotype in mice of either sex when fed a standard chow (Fig. 2 - figure supplement 5). However, we noted elevated T3 levels in both male and female Dp16 mice on a high-fat diet (Fig. 6 - figure supplement 2). While increased T3 levels correlated with higher physical activity and a modest increase in metabolic rate in Dp16 females, this was not observed in males (Fig. 6). We do not currently have a clear explanation for these findings. Given that individuals with Down syndrome often present with hypothyroidism and lower T3 levels, this discrepancy may reflect a species-specific difference between humans and mice.

      (3) Please can the authors clarify the percentage gene coverage, as this is quoted as ~58% of Hsa21 gene orthologs or ~65% of the Hsa21 gene orthologs, where the same reference is used.

      We apologize for the confusion. The number of triplicated genes in Dp16 mice corresponds to ~58% of Hsa21 genes (PMID: 26765563). We have corrected the typographical error in the text.

      (4) "segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs" for this given percentage majority sounds too strong, and the use of percentage is recommended.

      We have modified the text accordingly.

      (5) It is puzzling that in female gWAT with 7 triplicated Hsa21 gene orthologs (Rbm11, Chodl, Cldn8, Sh3bgr, Igsf5, Itgb2l, and Tmprss2). Could this be a technical issue? Was the reduced expression quantified by RT-Q-PCR?

      We have examined the normalized counts in the RNA-seq data for the seven genes in question, and the results do not appear to be an artifact. The sample size for this data is six mice per tissue per sex. In general, we prefer utilizing raw and normalized counts from RNA sequencing because there is a linear relationship between transcript amount and raw counts that is independent of housekeeping genes. In contrast, RT-qPCR involves mRNA amplification and requires expression to be normalized by one or more housekeeping genes (such as GAPDH, β-actin, 36B4, or ubiquitin) under the assumption that their levels remain constant.

      (6) The difference in body temperature is of interest. In male Dp16 mice, there is an increase in core temperature and a lowering of body temperature in females. In female Dp16 mice, higher estradiol levels have been stated by the authors to contribute to lower body temperature and higher physical activity (69-72). I am uncertain if the references are all relevant, as some relate to ovariectomized animals. No explanation is given for males.

      We currently do not have an explanation for why Dp16 males on a chow diet exhibit higher core body temperature, while Dp16 females show lower body temperatures. Although elevated T3 levels can increase body temperature, we have ruled this out; our data indicates there are no significant differences in T3 levels between genotypes for either sex on a chow diet.

      (7) The authors find a higher percentage heart weight in Dp16 mice on HFD and comment in the discussion that this is in keeping with "high-fat diet-induced cardiac hypertrophy". From what I can see, no histology has been performed to justify this statement. Furthermore, it would be useful to understand which animals had congenital heart disease in the first instance.

      We have modified the text accordingly. Unfortunately, we do not have histology data on the heart to inform us on whether some of our mice had congenital heart disease.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors should comment on the dosing method of glucose/insulin/lipid in the tolerance tests to acknowledge that differences in body weight may affect these tests. In addition, I encourage the authors to present ITT data as raw data, and not % change.

      In response to the reviewer’s comments, we have updated the ITT data plots to show raw data rather than percentage change. Regarding the dosing methodology, we maintain that basing dosage on body weight is appropriate. This approach is consistent with the vast majority of published rodent studies, as blood volume and metabolic tissues—such as skeletal muscle and adipose tissue—scale with body weight. Standardizing dose independently of body weight could lead to erroneous conclusions.

      (2) It would be useful for the authors to include a discussion on the likely specific tissue involvement in the whole-body metabolic disturbance. From my reading of the manuscript, there seems to be data suggesting functional and transcriptional dysfunction across most tissues, but do the authors suggest there is a dominant tissue in this regard?

      Due to the triplication of large number of genes on human chromosome 21, people with Down syndrome exhibit deficits across most organ systems (PMID: 32029743). Metabolic homeostasis also involves multiple tissues and cell types (adipose tissues, liver, skeletal muscle, pancreas, gut, hypothalamus, and immune cells). Most of the triplicated genes do express across these tissues. Our data indicate metabolic dysregulation across adipose tissues (white and brown), liver, skeletal muscle, and hypothalamus. Given the complex genetic perturbations of the Down syndrome mouse model, we do not think that there is a dominant tissue that contributes disproportionately to the systemic metabolic dysfunction phenotypes we observed in the Dp16 mice. Rather, we think that the metabolic phenotype is due to the combined deficits across multiple organs and tissues. As we do not have data to support the disproportionate contribution of any one tissue, we therefore did not speculate on the dominant contribution of any single tissue in the Discussion.

      (3) Related to this, muscle lipid is thought to be a major driver of muscle insulin resistance. Do the authors have measures of muscle lipid accumulation? This might be particularly interesting in the HFD models.

      Unfortunately, we did not measure lipid content in the skeletal muscle during this study. For the chow-fed mice, the entire gastrocnemius muscle was used for RNA isolation to perform RNA sequencing, and no tissue remains for additional analysis. Regarding the HFD-fed group, skeletal muscle was not collected at the termination of the study. As a result, we are unable to provide the requested lipid analysis data.

      (4) For mitochondrial analyses - do the authors have measures of total tissue mitochondria, and might changes in mitochondria abundance be driving some of these differences?

      For all our mitochondrial respiration analyses, we normalized the data to mitochondrial content as quantified by the MTDR assay (PMID: 32432379; PMID: 39704485). These results indicate that for a given amount of mitochondrial content, respiration as measured by the Seahorse assay is reduced in Dp16 mouse tissues, specifically in the BAT and liver.

      (5) To broaden the scope and interest, can the authors compare the transcriptional or metabolomic data to what has been found in non-DS insulin resistance (humans or mice), for example? This may help to highlight the key changes in metabolism that are causal for specific phenotypes.

      Overall, this is a comprehensive assessment of metabolism in a DS model.

      We appreciate the reviewer’s suggestion. However, given the vast number of published datasets on non-DS insulin resistance in both humans and mice, comparisons would yield varying results depending on the specific datasets selected. Consequently, we feel that such an analysis is beyond the scope of this study. We would like to highlight that many of the processes dysregulated in Dp16 mice as identified through our pan-tissue transcriptomes and metabolomes align with those frequently observed in non-DS insulin resistance. These include signatures of chronic low-grade inflammation, fibrosis, ER and oxidative stress, and impaired glucose and lipid metabolism.

      Reviewer #3 (Recommendations for the authors):

      It is slightly disconcerting that Figure 5 - Figure Supplements 2-5 are referred to in the text before the data in Figure 5 are discussed. It might make sense to indicate that the data are discussed further below (assuming that the authors do not wish to renumber these figures).

      We have fixed this issue raised by the reviewer.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the manuscript "A stable cryogenic fluorescence microscope for correlative super-resolution light and electron microscopy," the authors demonstrate a new cryogenic light microscopy design and characterize its temperature and spatial stability. The manuscript does a good job of reviewing the state of the field and highlights the need for improved cryogenic microscope stages. The system avoids challenges associated with vacuum-based designs, particularly vacuum transfer systems that can be difficult to engineer, while also showing minimal ice contamination and drift, which are the primary challenges associated with open cryostat systems.

      Strengths:

      The key strengths of the manuscript are the simple design and the significant level of detail provided in the description of the cryogenic stage. This represents a valuable step forward for the field by providing a home-built, non-vacuum stage design that others can emulate.

      We thank the reviewer for their positive assessment and strive to address the weaknesses they have constructively raised below.

      Weaknesses:

      There are only minor weaknesses or issues to address, which, if resolved, would strengthen the manuscript overall.

      (1) A key element of the design gets little attention, which is the plastic cap for the objective. It is not entirely clear to the reader how this is being used except as something of a thermal break between the cryogen environment and the objective, but there are some questions. Is the objective housing touching the plastic cap? Where is the front of the cap relative to the front objective lens? Is the front objective lens exposed to the cryogenic environment? Could the authors provide some 3D views of that in an SI figure? This would help clarify.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we will include a new supplementary figure (Fig. S2) providing detailed 3D views of the copper adapter, microscope objective, and plastic cap. The figure will show that the rim surrounding the front lens of the objective is covered by the plastic cap to provide thermal insulation between the objective housing and the cryogenic environment (Fig. S2b). We will also clarify that the front surface of the cap is levelled with the front objective lens to maintain the full working distance of the objective while allowing for axial movement of the z-stage. Finally, we will explicitly state that the front objective lens is exposed to the cryogenic environment (cold nitrogen gas).

      (2) The refilling system is not shown in the diagrams provided in Figure 1 and S1 in sufficient detail. How is the system mechanically coupled to the dewar on the microscope stage? Are there any concerns about coupling vibrations onto the table?

      To minimise vibrations arising from the nitrogen refilling pumps, the cryostat and liquid nitrogen tubing are mechanically decoupled from the microscope cage system, objective, translation stages, and sample. Specifically, the cryostat and nitrogen tubing are supported independently on a laboratory jack and surround the cage system without rigid mechanical contact. In the revised manuscript, we will update Fig. S1a, b to illustrate the liquid nitrogen tubing and refilling system more clearly. In addition, we will include a new supplementary figure (Fig. S3) to show the detailed cryostat design, refilling tubing, and temperature sensor position.

      (3) There is a description on page 6 that a rectangular aperture is used to align the excitation with the position and orientation of the sample. I know the authors are using this for excitation of the lamella, but without saying so in this text, it is confusing. I would consider stating that this is for future work involving excitation of lamella and then citing their preprint.

      We agree that the purpose of the rectangular aperture should be made clearer. In line with the suggestion from the reviewer, in the revised manuscript, we will briefly explain that the aperture is intended for selective illumination, such as in applications to cryo-FIB lamellae, and will cite our recent preprint describing this approach.

      (4) In Figure 2d, the z-drift is shown with the focus lock correction applied. This is highly relevant, but I also think it would be good to plot the z position plus the stage position in an SI figure. This will give a better idea of the mechanical stability of the system. Also, in this figure, I wonder if the authors could comment on the source of the jumps in lateral position. For example, just before 30 minutes. Lastly, I would make the lower plot have a tighter y-axis range. It is hard to see anything, hence the inset.

      We thank the reviewer for this suggestion. In the revised manuscript, we will include an additional supplementary figure (Fig. S4) showing the axial drift measured without focus-lock correction to illustrate the intrinsic mechanical stability of the microscope. We will also clarify that the periodic lateral displacement observed along the x-direction (approximately 300 nm amplitude with a period of ~22 minutes) arises from slight lateral repositioning accompanying z-stage stepping during focus-lock operation, likely due to mechanical coupling between the axes of the translation stage. We will revise the lower panel of Fig. 2d by reducing the y-axis range to improve data visibility.

      (5) The ice contamination looks minimal in Figure 3. I think it would benefit the manuscript to have lower magnification images as well, to show the level of ice contamination across a representative square. This would be good, but only if the authors have it in hand.

      We agree that this would be useful. In the revised manuscript, we will update Fig. 3 to include two additional low and intermediate-magnification cryo-EM images showing a representative grid square and a zoomed-in region of it, including a few grid holes. These images provide an overview of the ice contamination across a substantially larger field of view.

      (6) In Figure 4b, the y-axis is unclear. It looks like it has been normalized. Consider revising.

      The y-axis in Fig. 4b represents the localization rate (number of detected localizations per frame) within the selected ROI in Fig.4c and was not normalized. The values were calculated in SMAP by binning the localization frames into 100 temporal bins and dividing the number of localizations in each bin by the corresponding bin width, resulting in units of localizations per frame. Therefore, values close to 1 indicate approximately one localization detected per frame at that time point. To avoid potential confusion regarding the interpretation of this representation, we will replace this plot in the revised manuscript with a more explicit visualization showing the number of detected localizations per defined number of frames as a function of time (frame number) for the specific ROI shown in Fig. 4c.

      (7) A fluorescence intensity trace for the data shown in Figures 4c and f would be helpful to show the single-molecule behavior.

      In the revised manuscript, we will add fluorescence intensity traces corresponding to the single-molecule events shown in Fig. 4c and Fig. 4f to further demonstrate their single-molecule emission characteristics.

      Reviewer #2 (Public review):

      Summary:

      This manuscript reports the development of a cryo super-resolution fluorescence microscopy system. The authors demonstrate that they can achieve a mechanical and thermal stability that is sufficient to perform cryo-SMLM over the course of several hours. Focus instability is compensated for by tracking a fluorescent bead for its movement in the axial direction and adjusting the sample stage accordingly during data acquisition. Lateral instabilities are corrected after data acquisition. An enclosure around the microscope allows to significantly reduce ice contamination during cryo-SMLM imaging and sample transfer. The authors show an example of correlative cryo-SMLM and cryo-ET imaging achieved with their microscope system, which depicts the distribution of FtsZ-rsEGFP2 in E. coli.

      Strengths:

      The authors have designed a microscopy system for SR-cryo-CLEM, which achieves high stability while reducing complexity and costs substantially when compared to vacuum-insulated systems (e.g., Hoffman et al., 2020). They also provide software for controlling the microscope and data acquisition. This lowers the barrier for other labs to implement SR-cryo-CLEM into existing cryo-ET workflows. Reduction of ice contamination helps to increase throughput, which is currently one of the biggest bottlenecks for SR-cryo-CLEM.

      We thank the reviewer for their critical assessment, and for their suggestions below which we have used to improve the manuscript.

      Weaknesses:

      To correct for focus drift, the authors track a fluorescent bead in the far-red channel. This is possible for bacterial samples as used in this work, as beads can easily be introduced to surround the cells.

      Recommendations:

      (1) It is not discussed how this can be achieved in other samples than bacterial samples, such as lamellae in mammalian cells. Here, it would be much more difficult to introduce bright point-like markers with far-red fluorescence that would be distributed in the entire cell to capture at least one in the final lamella. Furthermore, it might be important to know for readers whether the far-red channel has to be sacrificed entirely for the focus correction.

      We thank the reviewer for highlighting this point. We agree that focus stabilization strategies for cryo-FIB lamellae are likely to differ from those used for the individual bacterial cell samples. For lateral drift correction, the presence of a single continuously detectable bright feature within the field of view is sufficient. Importantly, this feature does not need to be a fluorescent bead; any stable signal that can be continuously detected by the camera can serve as a suitable reference for drift correction. We will expand the Discussion to describe potential strategies for stable cryo-SMLM imaging, including the use of intrinsic sample or lamella features for autofocus, minimal fiducial-based approaches, and the practical implications of dedicating the far-red channel to focus stabilization.

      Furthermore, in the revised manuscript, we will include a new supplementary figure (Fig. S4) demonstrating the intrinsic axial stability of the microscope in the absence of active focus-lock correction. These measurements show that the system remains within the objective's depth of focus for a relatively long time, providing adequate stability for experiments in which far-red fluorescent fiducial beads are unavailable, such as cryo-FIB lamella imaging.

      (2) The authors show an application of SR-cryo-CLEM imaging of FtsZ-rsEGFP2 in E. coli. In the chosen correlative example (Figure 4d.f), no clear structure can be seen in the fluorescent images. The overview image (Figure 4d) shows no distinct signal in the cell, as it is shown for the non-correlative example in Figure 4a. The cryo-SMLM image (Figure 4f) does not show any ring-like features or accumulations of signals at the constriction site, as would be expected for a projecting along the optical axis. A clearer application example, which would show how increased resolution in cryo fluorescence microscopy enables resolving certain structural details or adds information not accessible in cryo electron tomography, would have strengthened the work. Particularly if taking into consideration that bacteria have a strong auto-fluorescence in the green range (Dahlberg et al., 2020), which could lead to high background or false positive localizations when using green fluorophores as labels.

      We thank the reviewer for this thoughtful comment. We agree that a correlative example displaying more pronounced structural features would further illustrate the capabilities of cryo-SMLM. However, the primary aim of the present work is the development and characterization of a robust cryogenic super-resolution microscope for reliable cryo-SMLM and correlative cryo-CLEM, rather than the demonstration of new biological applications. The utility of correlative cryo-SMLM/cryo-ET for resolving cellular structures has already been established in previous studies, including those employing rsEGFP2-labelled targets.

      The correlative dataset presented here is intended to demonstrate the compatibility of the microscope with cryo-CLEM workflows rather than to provide detailed biological insight. Moreover, the use of intact E. coli cells imposes inherent limitations on the ultrastructural information accessible by cryo-electron tomography; overcoming these limitations would typically require specimen thinning, for example, by cryo-focused ion beam (cryo-FIB) milling, which is beyond the scope of the present work.

      Regarding the concern about auto-fluorescence, elevated background fluorescence is not unique to bacterial samples or green fluorescent proteins but is a general consideration in cryo-SMLM that depends on the specimen and imaging conditions. While auto-fluorescence may reduce image contrast, it does not affect the conclusions of this work, which focuses on the design and performance of the microscope.

      (3) Access to CAD drawings (particularly for custom-made parts, such as cryostat or humidity enclosure) and a parts list is highly important for other researchers who would like to set up this SR-cryo-CLEM system in their own lab or institution. This is currently missing and, therefore, creating a hurdle for a wider adaptation of the technique.

      Thank you for this useful suggestion. In the revised manuscript, we will make available the complete SolidWorks CAD files for all custom-designed components, together with a comprehensive parts list and the full assembly corresponding to Fig. S1 as supplementary materials.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examines how luminescence can be used to measure bacterial population dynamics during antimicrobial treatment by comparing it directly with optical density and colony counts. The authors aim to determine when luminescence reflects changes in population size and when it instead captures metabolic or physiological states induced by drug exposure. By generating parallel datasets under controlled conditions, the work provides a detailed view of how these three common measurements relate to one another across a range of drug treatments.

      Strengths

      The study is technically strong and thoughtfully designed. Measuring luminescence, optical density, and colony counts from the same cultures allows the authors to make clear and informative comparisons between methods. The data are compelling, and the analyses highlight both agreements and divergences in a way that is easy to interpret. The manuscript also succeeds in showing why these divergences arise. For example, the observation that filamentation and metabolic shifts can sustain luminescence even when colony counts drop provides valuable information on how different readouts capture distinct aspects of bacterial physiology. The writing is clear, the figures are effective, and the work will be useful for researchers who need high-throughput approaches to quantify microbial population dynamics experimentally.

      Weaknesses:

      The study also exposes some inherent limitations of luminescence-based measurements. Because luminescence depends on metabolic activity, it can remain high when cells are damaged or unable to resume growth, and it can fall quickly when drugs disrupt energy production, even if cells remain physically intact. These properties complicate interpretation in conditions that induce strong stress re-sponses or heterogeneous survival states.

      In addition, the use of drug-free plates for colony counts may overestimate survival when filamented or stressed cells recover once the antibiotic is removed, making differences between luminescence and colony counts harder to attribute to killing alone. Finally, while the authors discuss luminescence in the context of clinically relevant concentration ranges, the current implementation relies on engineered laboratory strains and does not directly demonstrate applicability to clinical isolates. These limitations do not detract from the technical value of the work but should be kept in mind by readers who wish to apply the method more broadly.

      We thank the reviewer for reading our paper thoroughly and for the helpful feedback.

      Luminescence limitations. We agree that the lack of a direct link between light intensity and a population property such as biomass or cell number is the main limitation of the luminescence method. To further emphasise this, we have expanded the Discussion in the revised manuscript.

      Drug-free plates. The use of drug-free plates is intentional. As we measure a time series, the question at each point is how many cells are alive at each time point. Cells that are stressed but viable at time t contribute correctly to the count at t. How long they survive under the respective treatment is captured by the subsequent timepoints.

      Filaments. Recovery of plated filamented cells should not inflate this estimate. A single plated filamentous cell is expected to yield either zero (death before division) or one single colony, regardless of in how many parts it separates, as all descendants are part of the same cluster. However, if the cells divide before plating, CFU can overestimate survival. Having that said, we have no indication that this occurred in our experiments, since in all observed discrepancies, CFU-based estimates were equal to or lower than those obtained from luminescence and the time cells spent in dilution was kept short.

      Clinical applicability. We agree with the reviewer that the method is not practical for ad-hoc pharmacodynamic studies of clinical isolates. What we instead provide is an E. coli-based model system to explore clinically relevant treatment conditions, which we address in the revised manuscript. We believe that constructing analogous bioluminescent model strains in other clinically relevant species would be a valuable direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This preprint proposes luxCDABE-based luminescence as a high-throughput alternative (or complement) to CFU time-kill assays for estimating antimicrobial rates of population change at super-MIC concentrations, by comparing luminescence- and CFU-derived rates across 20 antimicrobials (22 assays) and attributing divergences primarily to filamentation (luminescence closer to biomass/volume than cell number) and changes in culturability/carryover (CFU undercounting viable cells).

      Strengths:

      The authors do not merely report discrepancies; they experimentally validate the biological causes. Specifically, they successfully attribute the slower decline of luminescence in certain drugs to bacterial filamentation (maintaining biomass despite halted division) and the rapid decline of CFU in others to loss of culturability or carryover effects.

      The inclusion of 20 antimicrobials spanning 11 classes provides a robust dataset that allows for broad categorisation of drug-specific assay behaviours.

      The study critically exposes flaws in the “gold standard” CFU method, specifically regarding antimicrobial carryover (demonstrated with pexiganan) and the potential for CFU to overestimate cell death in the presence of VBNC (viable but non-culturable) states induced by drugs like ciprofloxacin.

      The use of chromosomal integration for the lux operon to minimise plasmid copy-number effects and the validation of linearity between light intensity and cell density establish a solid technical foundation.

      Weaknesses:

      The study is conducted exclusively using Escherichia coli. While E. coli is a standard model organism, the paper claims to evaluate luminescence as a generalisable high-throughput tool. Many of the discrepancies observed are driven by filamentation. However, distinct morphological responses occur in other critical pathogens (e.g., Staphylococcus aureus does not filament in the same way).

      The authors propose that luminescence data can be corrected using microscopyderived volume data to better align with CFU counts. The primary appeal of luminescence is high-throughput efficiency. If a researcher must perform timelapse microscopy to calculate cell volume changes to “correct” their luminescence data, the high-throughput advantage is lost.

      The paper argues that for ciprofloxacin, CFU underestimates viability because cells remain intact and impermeable to propidium iodide. While the cells are metabolically active and membrane-intact, if they cannot divide to form a colony (even after drug removal/dilution), their clinical relevance as “living” pathogens is debatable.

      Some other comments:

      The use of a population dynamical model to simulate filamentation effects is excellent. The finding that light intensity tracks volume ($\psi_V$) better than cell number ($\psi_B$) is a key theoretical contribution.

      The model assumes linear elongation. The authors should briefly comment on whether this holds true for the specific drug mechanisms tested (e.g., PBP inhibition vs. DNA gyrase inhibition).

      The use of bootstrapping to estimate rate distributions is appropriate and robust.

      Conclusion:

      Muetter et al. provide a compelling argument that luminescence is a reliable, highthroughput alternative to CFU for super-MIC investigations, particularly when the quantity of interest is biomass. The paper effectively warns researchers that discrepancies between CFU and luminescence are often biological (filamentation, VBNC) rather than methodological failures.

      We thank the reviewer for reading our paper thoroughly and for the helpful feedback.

      Generalisability. We agree that the alignments and divergences reported for specific drugs may not transfer directly to other species, which may elongate differently (e.g. cocci) or show different physiological responses to treatment. Constructing analogous model strains — for example based on S. aureus to cover a broader range of morphologies and clinically relevant species would therefore be an interesting follow-up project, and we have adjusted the Discussion to make this clearer. We nevertheless believe that the broader conclusions (larger cells emit more light) of the paper likely hold across species.

      Volume correction. We agree that requiring microscopy would undermine the high-throughput advantage of the luminescence assay. It was not our intention to propose this as a practical approach, nor to imply that the luminescence signal needs a correction. Taken on its own, the signal can be interpreted as the cumulative metabolic output of the population, which is closely linked to biomass, and that measure is valuable in itself for many applications. We used the volume correction only to demonstrate that luminescence tracks biomass more closely than cell number: by adjusting the luminescence distribution with the measured volume change, it moves towards the CFU distribution. We have revised the Discussion to prevent this from being misunderstood as a required step.

      Culturability vs. clinical relevance. We agree that the dynamics of culturable cells are highly relevant, especially in a clinical context. Our aim was to explain the observed differences between CFU and luminescence by highlighting that culturability and viability are not always identical, without implying that one measure is inherently superior to the other — we leave it to the reader to decide which metric best suits their needs.

      Linear elongation. The model assumes linear elongation for mathematical convenience, which, depending on the specific strain and drug mechanism, could be incorrect. Its purpose is to demonstrate that a shift of the mean cell volume to a new, higher equilibrium under treatment can cause an initial peak in the luminescence signal despite a declining population. This remains true for non-linear elongation models, though the shape, height and position of the peak may change. We have adjusted the Results to make this clearer.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors present luminescence as a practical measurement of population decline under antibiotic exposure. One aspect that could be clarified is how the method behaves when tolerance arises from phenotypic heterogeneity, such as the presence of small, metabolically quiet survivors. Because luminescence reflects metabolic activity and biomass, the signal will be dominated by metabolically active cells, making rare tolerant subpopulations difficult to detect. A short discussion of how luminescence performs in these heterogeneous scenarios, and whether complementary assays are needed to capture long-lived tolerant cells, would strengthen the manuscript.

      Yes, that is a valid concern and we thank the reviewer for raising this point.

      Heterogeneity in cell-specific luminosity alone does not bias population-level rate estimates. A bias can arise, however, when specific luminosity correlates with a second factor — most importantly, the decline rate under treatment.

      We agree with the reviewer’s suggestion that brighter cells plausibly die faster than tolerant, metabolically quiet ones. When one subpopulation dominates the light signal, we expect minimal bias, as the rate estimate primarily reflects that subpopulation. However, in a transition phase when both subpopulations contribute roughly equally to the light signal, luminescence likely overestimates the decline.

      We added a corresponding caveat to the Discussion (lines 581–583).

      (2) The manuscript shows that filamentation can influence ψ_I by altering biomass and metabolic activity independently of cell number. However, antibiotic exposure can also trigger other stress responses and metabolic shifts that change energy fluxes, redox balance, and biosynthetic activity. Since luminescence depends on metabolic state and substrate availability, these additional physiological transitions may also affect ψ_I in ways not directly tied to birth or death processes. It would be useful to comment on whether such responses, beyond filamentation, are likely to influence luminescence dynamics across different drug classes or treatment conditions.

      We thank the reviewer for raising this point and agree that there is no biological law strictly linking luminosity to a single population property such as biomass or cell number, and changes in the metabolism most likely affect Ψ<sub>I</sub> as well.

      Transitioning to a new metabolic steady state biases Ψ<sub>I</sub>; once the new steady state is reached, however, the rate estimate should no longer be affected.

      Looking across drug classes, drugs that primarily lyse cells (polymyxins and, to a lesser degree, beta-lactams targeting PBP1) did not show noticeable deviations between Ψ<sub>I</sub> and Ψ<sub>CFU</sub>, and — perhaps counterintuitively — neither did ribosome-inhibiting drugs.

      For the remaining cases, we were able to attribute part of the discrepancy between Ψ<sub>I</sub> and Ψ<sub>CFU</sub> to changes in biomass or loss of culturability, though drug-induced metabolic changes may also contribute to the residual differences.

      We clarify this in the Discussion (lines 569–579).

      (3) The authors quantify survival using colony counts on drug-free medium. Because filamentation can be a reversible state that persists during antibiotic exposure, plating on drug-free medium may capture recovery potential rather than in-treatment viability. Filamented or stressed cells that cannot divide in the presence of a drug may nevertheless form colonies once the drug is removed. Clarifying how this recovery step affects ψ_CFU would help readers interpret differences between luminescence-based and colony-based measurements, especially in cases where transient tolerant states are present.

      We thank the reviewer for raising this point.

      Our CFU assay estimates the number of culturable cells at each time point; the rate Ψ<sub>CFU</sub> is then inferred from how this number changes across time points. Plating on drug-free medium is intentional, as it maximises the probability that a culturable cell is detected at each snapshot. Whether those cells would have continued dividing or died under continued treatment is captured by the subsequent time points.

      Filamentation interacts with the probability of colony formation in several, partly opposing ways:

      (1) It can increase the death rate, as for ceftazidime and cefepime, which is part of the kill effect captured by Ψ<sub>CFU</sub>;

      (2) Entanglement between filaments may reduce the number of colonies per plated bacterium;

      (3) Conversely, if a filament divides upon drug removal, its fragments form a cluster that — stochastically — is very likely to produce one (but not multiple) colony.

      The only scenario in which CFU could overestimate bacterial density is if a filament separates into individual cells in the liquid phase before plating; we have no indication that this occurred in our experiments.

      We addressed this concern in our response to the public comment.

      (4) A brief discussion comparing luminescence to fluorescent reporter systems could be helpful. Fluorescent proteins typically require a chromophore maturation step before becoming detectable, which introduces a delay between the underlying cellular event and the appearance of the signal. In contrast, as far as I understand, lux reporters emit light immediately once the enzymatic components and substrates are present, without a maturation stage. Highlighting this distinction may help readers understand why luminescence is well-suited for tracking rapid changes in population physiology under antibiotic exposure. However, the manuscript also notes that luminescence can lag slightly behind very rapid killing (particularly for AMPs), but the temporal dynamics of signal shutdown are not explored in detail. Because lux reflects metabolic activity rather than viability, a short delay between irreversible damage and the loss of light is biologically expected. It may help readers if the authors could expand on the mechanism underlying this delay in order to clarify when ψ_I is likely to track true biomass decline and when residual metabolic activity might mask early killing events.

      On fluorescent reporters:

      We thank the reviewer for this suggestion.

      Under some conditions, change rates can also be measured using fluorescence, provided the number of fluorescent molecules per bacterium remains constant. This requires a balance between production, maturation, degradation and dilution, which is only established if the growth rate and conditions remain constant over a sufficiently long period (typically hours).

      For measuring population decline, however, the key issue is that cell death does not inactivate fluorescent proteins: once matured, they emit independently of the cell’s metabolic state and decay only with the protein’s half-life, which is typically slower than the kill rates of interest.

      We added a clarification to the Introduction (lines 58–60).

      On the lux signal lag:

      We thank the reviewer for raising this point. The short lag between luminescence and CFU decline could in principle arise from two mechanisms: (i) luminescence declining more slowly than the actual cell number (residual light from dead cells), or (ii) CFU declining more steeply than the actual cell number (damaged but still viable cells failing to form colonies).

      Mechanism (i) splits into two sub-cases:

      (i.a) Dead but impermeable — the lux reaction could in principle continue for a short while if enough components are retained in the cell. However, a metabolically active, impermeable cell is difficult to classify as dead in the first place, making this scenario conceptually awkward.

      (i.b) Dead and permeable (lysed) — the lux components dilute into the medium, and by mass-action the reaction rate should drop rapidly (though not instantly). Any residual signal after lysis should therefore be short-lived.

      Mechanism (ii) — damaged (e.g. permeable) cells may be particularly sensitive to plating on agar (e.g. due to oxidative stress), resulting in a declining probability of colony formation.

      In our case, the discrepancy was observed specifically for pexiganan, where cells can be assumed to lyse, making (i.b) and/or (ii) the likely explanations. Based on our experimental data, we cannot distinguish between these possibilities and therefore limit ourselves to reporting the observed discrepancy.

      We have moved the interpretation from the Results to the Discussion (lines 523–542) and expanded the discussion there.

      (5) In lines 85–89, the authors state that “high-throughput OD and luminescence measurements at sub-MIC concentrations provide valuable insights into drug effects on growth rates, [but] the super-MIC range is clinically more relevant,” and they present luminescence as a way to investigate super-MIC population dynamics. While super-MIC behaviour is indeed important for pharmacodynamics and resistance evolution, it is not clear that the specific luminescence implementation used here has direct clinical relevance. The study relies on a chromosomally integrated reporter in a laboratory strain, and the manuscript does not demonstrate that this approach can be applied to clinical isolates or diagnostic workflows. It may be helpful to moderate the claim of “clinical relevance” and frame the method more clearly as a high-throughput experimental tool that can inform clinically relevant questions, rather than as an assay ready for clinical application.

      We agree and have moderated the framing accordingly (lines 100–104).

      Reviewer #2 (Recommendations for the authors):

      (1) The conclusions regarding “biomass vs. cell number” may not apply equally to non-rod-shaped bacteria or species with different stress responses. The authors must explicitly discuss this limitation in the Discussion.

      The broad conclusion that bigger cells emit more light likely holds across morphologies, since it rests on the principle that more cellular material means more metabolic activity and therefore more light. The quantitative relationship between cell size and luminosity, however, may differ across species, shapes and conditions, for two reasons. First, chromosome copy number: whether drug-induced morphological changes are accompanied by chromosome replication and therefore an increase in lux operon copy number — varies across species and drug mechanisms. Second, the surface-to-volume ratio likely modulates mass-specific metabolism; some morphological changes preserve it (e.g. purely lateral elongation) while others do not.

      The more specific conclusions about which drug classes produce alignment or divergence between CFU and luminescence may also not transfer directly, as drug mechanisms can act differently across species.

      We already note this limitation in the Discussion (lines 594– 597) and have expanded the wording there.

      (2) The manuscript should clarify that luminescence is a superior metric for biomass without correction, rather than framing the volume correction as a necessary step to mimic CFU. The divergence should be embraced as a feature (biomass tracking), not a bug that needs fixing via labor-intensive microscopy.

      We agree with the framing and will make it clearer; it was actually our intention to clarify which method does what, rather than judge one as better or worse.

      We removed the “correction” sentence from the Discussion to make this clearer.

      (3) The authors should refrain from definitively stating CFU “underestimates” viability and instead use more precise terminology, such as “reproductive capability” vs. “metabolic integrity.”

      We agree with the reviewer that measuring culturability is a property, not a flaw, of CFU. Our intention was to emphasise that when CFU is used as a proxy for viability (which it often is), it can yield lower values than the actual number of survivors. We tried to make that distinction explicit in the manuscript (e.g. in lines 317–322).

      We would also like to note that in the case of antimicrobial carryover, CFU can genuinely underestimate culturability itself, not only viability.

      Regarding the suggested reproductive capability vs. metabolic integrity framing: we agree that metabolism and luminescence are closely linked. What held us back from drawing that link directly is that metabolism is hard to quantify, being the cumulative output of a diverse set of processes.

      (4) The model assumes linear elongation. The authors should briefly comment on whether this holds true for the specific drug mechanisms tested (e.g., PBP inhibition vs. DNA gyrase inhibition).

      Linear elongation is a mathematically convenient simplification whose only purpose in the model is to allow the population to converge to a new equilibrium volume under treatment. Assuming constant volume-specific luminosity, we showed that this produces an initial peak in light intensity before the signal declines in parallel with Ψ<sub>B</sub>. The exact shape, height and position of this peak depend on the volume growth model used, but the qualitative pattern — peak followed by parallel decline — holds for other growth models as well. We now clarify this in lines 230–235.

      (5) The authors suggest the carryover effect is due to a delay between cell death and cessation of luminescence. This “lag time” is a critical physical constraint of the lux system (likely related to ATP depletion or enzyme decay) and should be quantified or discussed in more detail as a fundamental “speed limit” for the assay.

      The origin of the lag between luminescence and CFU is an interesting question, but one we cannot definitively answer. We can, however, discuss the potential mechanisms:

      A dead but impermeable cell could in principle continue to emit residual light for some time. We note, though, that calling a metabolically active, impermeable cell “dead” is a question of definition we would rather not discuss here.

      In our case, the discrepancy was observed specifically for pexiganan, where cells can be assumed to lyse. Under lysis, the lux components dilute quickly into the medium, and by mass-action the reaction rate should drop rapidly — though not necessarily instantaneously.

      A plausible alternative to a delayed cessation of the light signal is that the probability of colony formation drops rapidly after permeabilisation, for example because permeable cells are sensitive to oxidative stress when plated on agar.

      Based on our experimental data we cannot distinguish between these mechanisms, so we limit ourselves to reporting the observed discrepancy. We have moved the interpretation from the Results to the Discussion (lines 523–542) and expanded on the candidate mechanisms there.

      Additional revisions

      Beyond the changes prompted by the reviewers’ comments, we made the following revisions to the supplementary information:

      We corrected the Λ matrix (converted row 2, col 4 from 0 → 2)

      We removed the line numbering

    1. Author response:

      Reviewer 1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRR v2) recently described by some of the same authors and based on incorporation of [3H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRR v2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRR v2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      We thank reviewer 1 for the supportive feedback and for raising some important points.

      Many antimalarials have quite specific times of action. Are these MULT-i<sup>2</sup> assays, and the comparator PRR v2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      We thank the reviewer for this important comment. Both, the MULT-i<sup>2</sup> and PRR v2 assays were performed using asynchronous parasite cultures. This information is included in the Methods section together with the relevant references. To improve clarity, we will also explicitly state this in the main text.

      The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULT-i<sup>2</sup> method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      We thank the reviewer for this helpful suggestion. In the Introduction we will mention and describe alternative approaches for assessing parasite viability. This will also include the work by Maiga et al., which combines MitoTracker with a nuclear dye to improve the reliability of flow cytometry-based readouts. We will revise the text to explicitly mention the use of dual straining to make this discussion more explicit.

      We agree that several additional methods, such as luciferase-based reporter systems, have been successfully applied in antimalarial screening. However, these approaches are primarily designed to assess parasite growth inhibition rather than directly measuring parasite viability after drug exposure, which is the focus of the present study. Readout methods used to assess parasite viability in a PRR assay setup are so far based on HRP2-ELISA (de Carvalho et al.), MitoTracker and SYBR green staining (Maiga et al.) and [<sup>3</sup>H]-hypoxanthine incorporation (Sanz et al.; Walz et al.) as cited in the manuscript. Many other readout methods to assess parasite growth have other limitations as briefly discussed in Hellingman et al, 2024. A comprehensive comparison and review of all available readout methods would therefore be beyond the scope of this manuscript.

      It would be helpful for authors to provide some indication of the cost comparison between the PPR v2 and MULT-i<sup>2</sup>.

      We thank the reviewer for this valuable suggestion. We agree that a comparison of the costs associated with the PRR v2 and MULT-i<sup>2</sup> assays would be informative, but while the consumable costs provide one measure of assay expense, we consider the reduction in hands-on time and the simplified workflow to be the main contributors to the overall cost advantage of the MULT-i<sup>2</sup> assay. These reductions in labor requirements are subject to large regional differences and impossible for us to access. Nevertheless, together with the increased throughput and the reduced labor, make the MULT-i<sup>2</sup> assay more cost-effective for larger-scale applications compared with the PRR v2 assay.

      Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      We thank the reviewer for this important suggestion. The engineered parasite line will be made available for non commercial use to other researchers upon request. An MTA will be required excluding commercial use of the provided strains. The detailed code used for data analysis is available upon request, and an example code file has already been included as a Supplementary File.

      The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

      We thank the reviewer for this valuable comment. We agree that implementation of pharmacological modeling approaches can represent a barrier for laboratories without prior experience in pharmacometric analysis, particularly due to the requirement for specialized software such as NONMEM. To facilitate implementation, an example code is provided in the Supplementary File. The final model was developed using a forward–backward selection approach for parameter estimation and model refinement as described in the Methods section. These additions should help other researchers adapt the approach to their own datasets.

      Reviewer 2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRR v2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i<sup>2</sup> assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i<sup>2</sup> assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i<sup>2</sup> assay.

      We thank reviewer 2 for her/his appreciation of our work.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i<sup>2</sup> assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRR v2 assay?

      We thank the reviewer for raising this important point. All, the MULT-i<sup>2</sup> and PRR v2 assay were performed using asynchronous parasite cultures. We will clarify this in the revised manuscript.

      We agree that parasite developmental stages may influence the MULT-i<sup>2</sup> readout, as LacZ expression levels differ between parasite stages, with differences observed between ring stages and more mature trophozoite/schizont stages as published by Hellingman et al., 2024. This represents a potential source of variability, as the MULT-i<sup>2</sup> assay quantifies the amount of expressed reporter enzyme rather than directly measuring parasite numbers at the time of readout. The use of asynchronous cultures minimizes the impact of stage-specific effects by providing a mixed parasite population representative of the natural distribution of developmental stages. Nevertheless, we acknowledge that differences in parasite stage progression following drug exposure may contribute to variation in the extrapolated parasite numbers and may partially explain differences observed between the MULT-i<sup>2</sup> and PRR v2 assay measurements. We will add this consideration to the Discussion.

      The addition of an inducible element is an improvement of their earlier lacZ/β-gal<sup>SENSOR</sup> (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRR v2, they fail to compare it to their own non-inducible lacZ/β-gal<sup>SENSOR</sup> system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved?

      We thank the reviewer for this important comment. The main improvement provided by the inducible system is the temporal separation of parasite growth/drug exposure from reporter expression. In the original non-inducible lacZ/β-gal<sup>SENSOR</sup> system, reporter expression occurs continuously throughout the assay, resulting in accumulation of β-galactosidase during parasite growth/drug exposure and therefore an increasing background signal. Consequently, quantification relies on endpoint reporter levels and does not allow the reporter expression window to be standardized independently of parasite exposure history.

      In contrast, in the MULT-i<sup>2</sup> system, reporter expression is initiated only after addition of rapamycin post antimalarial drug washout. This prevents reporter accumulation during the drug exposure window and ensures a defined reporter enzyme accumulation window after drug exposure. Importantly, this allows parasite numbers to be extrapolated from a calibration curve generated at the time of induction, which would not be possible with the non-inducible system because reporter expression would continue after drug removal and would depend on the previous culture history.

      We will revise the manuscript to more clearly describe these advantages and to emphasize that the key benefit of the inducible system is not simply an increase in signal intensity, but improved control of reporter expression, reduced background accumulation, and the ability to perform quantitative parasite reduction rate measurements.

      How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      We thank the reviewer for this important suggestion. We acknowledge that the sensitivity of the MULT-i<sup>2</sup> readout depends on both the initial parasite density and the duration of the induction period and that a detailed characterization of the induction kinetics, including earlier time points after rapamycin addition, would provide additional information on the sensitivity and temporal resolution of the MULT-i<sup>2</sup> system.

      In the present study, we focused on the time window relevant for application of the assay in a PRR assay workflow and routine drug screening setting. Earlier time points (<24 h after induction) were therefore not systematically evaluated. The selected time points were chosen based on the expected kinetics of the loxP-DiCre recombination system, which has previously been reported to achieve high recombination efficiency within one asexual parasite cycle, (Collins et al., 2013) and shown with own data in this study, as well as on practical considerations for implementation in routine workflows.

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

      We thank the reviewer for this question. The chemiluminescence signal obtained with the inducible lacZ (i-lacZ) parasites is comparable to that observed with the previously characterized constitutively expressing lacZ parasites. However, the inducible system provides an important additional advantage by avoiding continuous β-galactosidase production and accumulation during parasite growth, thereby reducing background signal and enabling a controlled reporter expression window.

      We do not consider the MULT-i<sup>2</sup> assay to be a replacement for classical PRR assays. Rather, we consider it a complementary approach that enables more efficient screening and characterization of drug combinations, particularly by providing information on the time-dependent onset of parasiticidal activity in a higher-throughput format. Promising combinations identified using MULT-i<sup>2</sup> assay can subsequently be investigated in more extensive PRR assays.

      Reviewer 3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULT-i<sup>2</sup>, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the NULT-i<sup>2</sup> assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULT-i<sup>2</sup> assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULT-i<sup>2</sup> provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      We thank reviewer 3 for her/his appreciation of our work.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULT-i<sup>2</sup> methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1): The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      We thank the reviewer for raising this important point regarding the rationale, applicability, and limitations of the MULT-i<sup>2</sup> methodology.

      Quantification of viable parasites after drug exposure remains challenging, particularly when surviving parasites are present at low frequencies or require extended recovery periods. Current approaches, such as the parasite reduction ratio (PRR) assay based on [<sup>3</sup>H]-hypoxanthine incorporation, provide sensitive measurements of replicating parasites but are labor-intensive, require specialized infrastructure, and are not easily scalable for large numbers of drug combinations. Alternative approaches based on HRP2 detection no longer rely on radioactive readouts but generally provide lower sensitivity, particularly when quantifying low levels of surviving parasites within a shorter time frame.

      The MULT-i<sup>2</sup> assay was developed to address these limitations by combining a highly sensitive chemiluminescent β-galactosidase readout with an inducible reporter system. The 5-day induction period after drug exposure serves as a controlled gene expression step, allowing surviving parasites to recover and produce sufficient reporter signal for sensitive quantification using a standard plate reader. This approach enables higher-throughput assessment of parasiticidal activity while avoiding radioactive readouts and reducing the need for labor-intensive dilution-based approaches.

      We acknowledge that the recovery and reporter expression period introduces additional biological steps compared with direct parasite detection methods and may therefore represent a potential source of variability. The MULT-i<sup>2</sup> assay is not intended to replace all existing viability measurements but rather to provide a complementary screening tool for investigating larger numbers of drug combinations. More detailed comparisons with additional detection platforms, including fluorescence-based approaches such as flow cytometry, would be valuable; however, a comprehensive comparison of all available parasite viability readouts was beyond the scope of this study. We will add more explanations to the Discussion including the strengths and limitations.

      (2) Related to that above, how would MULT-i<sup>2</sup> perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      We thank the reviewer for raising this important point regarding the interpretation and applicability of the MULT-i<sup>2</sup> assay. We agree that distinguishing between growth inhibition assays and viability-based assays is essential when interpreting the response to drugs that induce temporary parasite dormancy or delayed recovery.

      The MULT-i<sup>2</sup> assay was specifically developed as a viability-based approach and therefore differs fundamentally from conventional IC50 assays, which primarily measure inhibition of parasite growth during drug exposure and may not capture parasites that survive treatment through temporary growth arrest or dormancy. Similar to the PRR assay, the MULT-i<sup>2</sup> assay measures the ability of surviving parasites to recover and proliferate after drug exposure. Therefore, parasites that temporarily enter a dormant state but subsequently resume replication are expected to contribute to the measured signal rather than representing false-positive or false-negative results.

      This is illustrated by the artemisinin experiments presented in this study, where the MULT-i<sup>2</sup> assay captures the recovery of surviving parasites following treatment as it does the PRR v2 assay.

      (3) Given the stated cost and labor efficiency of MULT-i<sup>2</sup>, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i<sup>2</sup> method more attractive. In particular, it would be nice to see if one could use MULT-i<sup>2</sup> for studies of triple combinations as enthusiastically suggested.

      We thank the reviewer for this valuable suggestion. We agree that demonstrating additional applications, including triple-drug combinations, would further highlight the potential of the MULT-i<sup>2</sup> assay.

      The primary aim of this study was to validate the MULT-i<sup>2</sup> methodology against the established PRR v2 assay and to demonstrate that the new platform can reproduce known parasiticidal interaction profiles while providing a more scalable workflow. For this reason, we selected well-characterized drug combinations, including atovaquone/proguanil and piperaquine/pyronaridine, which provide suitable benchmark systems for comparison with previous PRR data.

      Although evaluation of a larger number of novel combinations and triple-drug regimens would be highly valuable, generating corresponding PRR datasets for direct comparison was beyond the scope of the current study.

      (4) Throughout the manuscript, the authors claim that MULT-i<sup>2</sup> is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

      We thank the reviewer for this important comment. We agree that absolute assay costs can vary depending on local reagent prices, labor costs and laboratory infrastructure.

      When comparing both methods under the same laboratory conditions, the total assay duration of the MULT-i<sup>2</sup> assay is shorter than that of the PRR assay (11 days (MULT-i<sup>2</sup>) compared with approximately 21–28 days (PRR) according to published protocols). In addition, the MULT-i<sup>2</sup> assay reduces labor-intensive processing steps and enables higher-throughput measurements using a plate reader for readout. These factors contribute to reduced workload and improved scalability, independent of fluctuations in individual reagent or personnel costs.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Wang et al. reports the potential involvement of an asymmetric neurocircuit in the sympathetic control of liver glucose metabolism.

      Strengths:

      The concept that the contralateral brain-liver neurocircuit preferentially regulates each liver lobe may be interesting.

      Weaknesses:

      However, the experimental evidence presented did not support the study's central conclusion.

      We thank the reviewer for recognizing the conceptual novelty of our work and for constructive comments aimed at enhancing its rigor and clarity. In response, we carried out targeted experiments to address the points raised, including: (i) further characterization of LPGi projections to vagal and sympathetic circuits; (ii) evaluation of potential pancreatic involvement; and (iii) validation of the specificity of chemogenetic activation within the proposed circuit. All new experiments, figures, and text have been incorporated, and corresponding revisions are highlighted for ease of review.

      (1) Pseudorabies virus (PRV) tracing experiment:

      The liver not only possesses sympathetic innervations but also vagal sensory innervations. The experimental setup failed to distinguish whether the PRV-labeling of LPGi (Lateral Paragigantocellular Nucleus) is derived from sympathetic or vagal sensory inputs to the liver.

      Thank you for raising this important point. We fully agree that the liver receives both sympathetic and vagal sensory innervation, and we acknowledge that PRV-based tracing alone does not definitively distinguish between these two pathways. This represented a limitation of the original experimental design.

      Based on established anatomical literature as well as our experimental observations, vagal sensory neuron cell bodies reside in the nodose ganglion (NG), and their central projections terminate predominantly in the nucleus of the solitary tract (NTS) (Nature. 2023;623(7986):387-396; Curr Biol. 2020;30(20):3986-3998.e5.), which is located in the dorsomedial medulla. In contrast, the LPGi, together with other sympathetic-related nuclei, is predominantly distributed in the ventral medulla (Cell Metab. 2025;37(11):2264-2279.e10; Nat Commun. 2022;13(1):5079).

      To determine whether the LPGi contains neurons that modulate the liver via vagal sensory pathways, we performed two complementary experiments.

      First, we conducted CGRP immunohistochemistry on brainstem sections, using the NTS, a well-established visceral sensory centre, as a positive control. While abundant CGRP-positive cell bodies were detected in the NTS as expected, few to no CGRP-positive cell bodies were observed in the LPGi (Figure S1G). These results strongly support that the LPGi neurons labeled in our PRV tracing predominantly belong to sympathetic efferent circuits rather than vagal sensory pathways.

      Second, to examine whether LPGi neurons send axonal projections to sensory ganglia, we injected hSyn-Cre combined with DIO-Axon-EGFP into the LPGi and examined both the dorsal root ganglia (DRG) and nodose ganglia (NG). No Axon-EGFP-positive signals were detected in either ganglion (Figures S1H-S1J), indicating that LPGi neurons do not directly innervate sensory ganglia. In other words, PRV cannot retrogradely trace to the LPGi via the NG or DRG.

      These additions have been incorporated into the revised manuscript, with the Result 1 clearly documenting that these findings confirm that the LPGi specifically regulates sympathetic, rather than vagal sensory, inputs to the liver.

      (2) Impact on pancreas:

      The celiac ganglia not only provide sympathetic innervations to the liver but also to the pancreas, the central endocrine organ for glucose metabolism. The chemogenetic manipulation of LPGi failed to consider a direct impact on the secretion of insulin and glucagon from the pancreas.

      Thank you for this important comment. We agree that the celiac ganglia (CG) provide sympathetic innervation not only to the liver but also to the pancreas, which plays a central role in glucose homeostasis through the secretion of both insulin and glucagon. Therefore, the potential pancreatic implications associated with LPGi chemogenetic manipulation are worth careful consideration.

      To address this concern, we measured circulating glucagon and insulin levels following chemogenetic manipulation of the LPGi<sup>GAD1</sup> neurons. We found that neither glucagon nor insulin levels changed significantly under our experimental conditions, which indicated that the hyperglycemic effect induced by LPGi activation is unlikely to be mediated by changes in pancreatic hormone secretion (Figure S2G).

      These additions have been incorporated into the revised manuscript, with the Result 2 clearly documenting that these findings confirm that the hyperglycemic effect induced by LPGi activation is unlikely to be mediated indirectly via altered pancreatic endocrine output.

      (3) Neuroanatomy of the brain-liver neurocircuit:

      The current study and its conclusion are based on a speculative brain-liver sympathetic circuit without the necessary anatomical information downstream of LPGi.

      Thank you for raising this important point. A clear anatomical definition of the downstream pathways linking the brain to the liver was essential for interpreting the proposed brain-liver sympathetic circuit.

      The present study (Figure 4A) provides direct anatomical evidence supporting the organization of the brain–liver sympathetic neurocircuit. These observations are consistent with our recent detailed characterization of the brain-liver sympathetic circuit published in Cell Metabolism (Cell Metab. 2025;37(11):2264–2279). In that study, we showed that LPGi GABAergic neurons inhibit GABAergic neurons in the caudal ventrolateral medulla (CVLM). Disinhibition of CVLM reduced GABAergic suppression of rostral ventrolateral medulla (RVLM) neurons, which are key excitatory drivers of sympathetic tone. RVLM neurons project to sympathetic preganglionic neurons in the sympathetic chain (Syc). These neurons synapse with postganglionic sympathetic neurons in ganglia such as the celiac-superior mesenteric ganglion (CG-SMG). Postganglionic sympathetic fibers then innervate the liver, releasing norepinephrine (NE) to activate hepatic β<sub>2</sub>-adrenergic receptors and stimulate hepatic glucose production (HGP).

      Together, these data establish a coherent anatomical basis for the proposed brain-liver sympathetic pathway and clarify the downstream organization relevant to the functional experiments presented in figure 4A and Author response image 1..

      Author response image 1.

      Tracing scheme (Left) and whole-mount imaging (Right) of PRV-labeled brain-liver neurocircuit. Scale bars, 3,000 (whole mount) or 1,000 (optical sections) μm.

      (4) Local manipulation of the celiac ganglia:

      The left and right ganglia of mice are not separate from each other but rather anatomically connected. The claim that the local injection of AAV in the left or right ganglion without affecting the other side is against this basic anatomical feature.

      Thank you for raising this important anatomical point. We fully acknowledge that the left and right CG in mice are interconnected, and that unilateral viral injection could theoretically affect the contralateral side. The CG-SMG complex serves as a major sympathetic hub that regulates visceral organ functions. Recent transcriptomic, anatomical, and functional studies have revealed that the CG-SMG is not a homogeneous structure but is composed of molecularly and functionally distinct neuronal populations. These populations exhibit specialized projection patterns and regulate different aspects of gastrointestinal physiology, supporting a model of modular sympathetic control. (Nature. 2025 Jan;637(8047):895-902). Therefore, we were aware of this phenomenon during the initial stages of these experiments.

      To minimize unintended spread to the contralateral CG, we took two complementary approaches.

      First, we optimized the injection strategy by using an extremely small injection volume (100 nL per site), with a very slow infusion rate (50 nL/min), and fine glass micropipettes. With these refinements, contralateral viral spread was rarely observed.

      Second, and importantly, all animals included in the final analyses were subjected to post hoc anatomical verification. After completion of the experiments, CGs were collected, sectioned, and examined for viral expression. As shown in Supplementary Figure 5F, only mice in which viral expression was strictly confined to the targeted CG, with no detectable infection in the contralateral ganglion, were included in the presented data.

      Together, these measures ensure that our local manipulation of the intended CG produced the reported effects. We have revised the Methods section to more explicitly detail these technical precautions, and the legend for Figure S5F clearly states its role in validating injection specificity.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Wang and colleagues aims to determine whether the left and right LPGi differentially regulate hepatic glucose metabolism and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. Then the authors activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, as well as changes in protein expression in the liver lobes. These data suggested modulation of HGP (hepatic glucose production) in a lobe-specific manner. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      We thank the reviewer for the thorough and constructive evaluation of our manuscript. In direct response, we undertook comprehensive revisions to enhance the rigor and clarity of the study, including: (i) correcting ambiguous or misleading terminology about anatomical resolution and sympathetic circuit organization; (ii) expanding the Methods section with complete experimental details, improved image presentation, and explicit justification of our viral and genetic approaches; and (iii) strengthening data interpretation by addressing issues related to sparse PRV labeling, projection heterogeneity, and the functional implications of double-labeled neurons.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) The wording/terminology used in the manuscript is misleading, and it is not used in the proper context. For instance, the goal of the study is "to investigate whether cerebral hemispheres differentially regulate hepatic glucose metabolism..." (see abstract); however, the authors focus on the brainstem (a single structure without hemispheres). Similarly, symmetric is not the best word for the projections.

      We thank the reviewer for raising these critical points regarding terminology and conceptual framing. We acknowledge that certain phrases in our original manuscript may have been overly broad or ambiguous, particularly in describing the scope of sympathetic heterogeneity and the specificity of neural projections. Due to practical constraints and the scope of our study, our investigation focused on the brainstem, which represents the final common pathway for these lateralized commands. We acknowledge that terms referring to the cerebral hemispheres do not accurately describe our study. We have revised the manuscript to ensure accurate and consistent terminology.

      Below are specific examples:

      Original 1: This study aims to investigate whether cerebral hemispheres differentially regulate hepatic glucose metabolism and localize the site of sympathetic crossover to the liver.

      Revised 1: “This study aimed to determine whether the central nervous system exerts lateralized control over hepatic glucose metabolism and to localize the site of peripheral sympathetic crossover to the liver.”

      Original 2: These findings demonstrate that the brain exerts lobe-specific, lateralized control of hepatic glucose metabolism via symmetric brain-liver sympathetic pathways.

      Revised 2: “These findings demonstrate that the brainstem can exert lobe-specific, lateralized control of hepatic glucose metabolism via bilaterally projecting brain–liver sympathetic pathways.”

      Original 3: The cerebral hemispheres exhibit pronounced functional asymmetry, [1,2] a phenomenon traditionally associated with cognitive and motor processes such as language, voluntary movement, and spatial navigation.[3,4]

      Revised 3: “Pronounced functional lateralization within the central nervous system (CNS) is a well-documented phenomenon, [1,2] traditionally associated with cognitive and motor processes such as language, voluntary movement, and spatial navigation [3,4].

      (2) Sparse labeling of liver-related neurons was shown in the LPGi (Figure 1). It would be ideal to have lower magnification images to show the area. Higher quality images would be necessary, as it is difficult to identify brainstem areas. The low number of labeled neurons in the LPGi after five days of inoculation is surprising. Previous findings showed extensive labeling in the ventral brainstem at four days post-inoculation (Desmoulins et al., 2025). Unfortunately, it is not possible to compare the injection paradigm/methods because the PRV inoculation is missing from the methods section. If the PRV is different from the previously published viral tracers, time-dependent studies to determine the order of neurons and the time course of infection would be necessary.

      We sincerely thank the reviewer for these detailed and constructive comments regarding the PRV tracing experiments. We fully agree that careful presentation and interpretation of the anatomical data are essential for ensuring rigor and transparency. We address each point in detail below.

      (1) Image magnification and anatomical context of LPGi labeling

      We agree that the original images did not sufficiently convey the broader anatomical context of the LPGi. Due to fluorescence quenching in previous sections, we repeated the PRV retrograde tracing experiment and performed statistical analysis. In the revised manuscript, we replaced the original panels in Figure 1 and Figure S1 with new images that include lower-magnification overviews of the brainstem, alongside higher-magnification views of the LPGi (Figure 1). These images clearly delineate the LPGi with respect to established anatomical landmarks and atlas boundaries. Image contrast and resolution were optimized to allow unambiguous identification of PRV-labeled neurons and surrounding structures.

      (2) Sparse LPGi labeling at 5 days post-injection and methodological details

      We apologize for the omission of the detailed PRV injection protocol in the original Methods section. We deliberately used small-volume, local injections (1 µL per liver lobe) to minimize viral spread and to restrict labeling to circuits specifically connected to the targeted hepatic region. This sparse labeling was consistent with the use of small, spatially restricted injections designed to minimize off-target spread and preferentially label higher-order upstream neurons. This information has now been added, including the PRV strain, viral titer, injection volume, precise injection coordinates, and surgical procedures. All new figures, legends, and Method details have been incorporated, with changes clearly highlighted for ease of review.

      These additions have been incorporated into the revised manuscript, in Figure 1B-D, Figure S1C-F. Furthermore, we also added details of the Methods.

      (3) Not all LPGi cells are liver-related. Was the entire LPGi population stimulated, or was it done in a cell-type-specific manner? What was the strain, sex, and age of the mice? What was the rationale for using the particular viral constructs?

      We thank the reviewer for this insightful and important question. We agree that not all neurons within the LPGi are liver-related, and we apologize that our rationale was not clearly articulated in the original manuscript.

      (1) Our decision to target GABAergic neurons in the LPGi using GAD1-Cre mice was based on prior experimental evidence rather than an assumption about the entire LPGi population. In our previous study (Cell Metab. 2025;37(11):2264-2279.e10), we performed single-cell RNA sequencing on retrogradely labeled LPGi neurons following liver tracing. These analyses revealed that the majority of liver-projecting LPGi neurons are GABAergic in nature. Based on these findings, we chose to selectively manipulate GABAergic neurons in the LPGi rather than the entire LPGi neuronal population, to achieve greater cellular specificity and to minimize potential confounding effects arising from heterogeneous neuron types within this region. We regret that this rationale was not clearly described in the original submission and have now revised the manuscript to explicitly state this reasoning (Results section 2, paragraph 2: “Prior single-nucleus RNA sequencing and immunofluorescence analyses demonstrated that liver-projecting LPGi neurons are predominantly GABAergic.”).

      (2) In addition, we apologize for the omission of mouse strain, sex, and age information in the Methods section. These details have been fully added.

      (3) We selected AAV-based viral vectors, specifically the AAV9 serotype, due to their well-established efficiency in transducing neurons in the brainstem, relatively low toxicity, and widespread use in circuit-level chemogenetic and optogenetic studies. When combined with Cre-dependent viral constructs in GAD1-Cre mice, this approach enabled selective and reliable manipulation of LPGi GABAergic neurons.

      (4) The authors should consider the effect of stimulation of double-labeled neurons (innervating more than one lobe) and potential confounding effects regarding other physiological functions.

      We thank the reviewer for raising this important point. We agree that neurons innervating more than one liver lobe could, in principle, introduce potential confounding effects and may reflect higher-order integrative autonomic neurons.

      This consideration is consistent with a key finding of the cited study: the CG-SMG contains molecularly distinct sympathetic neuron populations (e.g., RXFP1<sup>+</sup> vs. SHOX2<sup>+</sup>) that exhibit complementary organ projections and separate, non‑overlapping functions. Specifically, RXFP1<sup>+</sup> neurons innervate secretory organs (pancreas, bile duct) to regulate secretion, while SHOX2<sup>+</sup> neurons innervate the gastrointestinal tract to control motility. This functional segregation supports the concept of specialized autonomic modules rather than a uniform, “fight-or-flight” response, reinforcing the need for careful interpretation of circuit-specific manipulations. (Nature. 2025;637(8047):895-902; Neuron. 2026;114(3):463-478.e7).

      In our PRV tracing experiments, the proportion of double-labeled neurons was relatively small, suggesting that the majority of labeled LPGi neurons preferentially associate with individual hepatic lobes. Nevertheless, we recognize that activation of this minority population could contribute to broader physiological effects beyond strictly lobe-specific regulation. We have therefore added a paragraph in the second paragraph of the Discussion (Paragraph 2: “A small subset of LPGi neurons was double-labeled after bilateral PRV injections, suggesting a fraction of these neurons projects bilaterally to both sides of the liver. Such neurons may support interlobar coordination.”).

      (5) The authors state that "central projections directly descend along the sympathetic chain to the celiac-superior mesenteric ganglia". What they mean is unclear. Do the authors refer to pre-ganglionic neurons or premotor neurons? How does it fit with the previous literature?

      We thank the reviewer for pointing out this imprecise wording. We agree that the original phrasing was anatomically inaccurate and potentially confusing. The pathways we intended to describe involve brainstem premotor neurons that project to sympathetic preganglionic neurons in the spinal cord. These preganglionic neurons then innervate neurons in the CG-SMG, which in turn provide postganglionic input to the liver.

      We have revised the manuscript to clearly distinguish premotor from preganglionic neurons (Results section 4, paragraph 1: “Using whole-mount clearing, we visualized the brain–liver sympathetic circuit and found that preganglionic neurons in the spinal cord send descending fibers through the sympathetic chain (SyC) to innervate postganglionic neurons in the CG-SMG (Figure 4A). Further whole-mount TH immunostaining showed that TH-positive sympathetic cell bodies within the CG-SMG project to the liver along the hepatic vasculature (Figure 4B and Figure S5F). These observations suggest that the nerve bundles likely decussate at the porta hepatis before entering the individual hepatic lobes.”).

      (6) How was the chemical denervation completed for the individual lobes?

      We thank the reviewer for raising this important methodological concern. We agree that potential diffusion of 6-OHDA is a critical issue when performing lobe-specific chemical denervation, and we apologize that our original description did not sufficiently clarify how this was controlled.

      In the revised Methods section, we provided a detailed description of the denervation procedure, including the injection volume and concentration of 6-OHDA, as well as the physical separation and isolation of individual hepatic lobes during application to minimize diffusion to adjacent tissue.

      To directly assess the specificity of the chemical denervation, we included immunofluorescence and Western blot analyses demonstrating a selective reduction of sympathetic markers in the targeted lobe (Figure 3C), with minimal effects on non-targeted lobes. These results support the effectiveness and relative spatial confinement of the 6-OHDA treatment under our experimental conditions.

      We thank the reviewer for highlighting this point, which has helped us improve both the clarity and rigor of the manuscript.

      (7) The Western Blot images look like they are from different blots, but there are no details provided regarding protein amount (loading) or housekeeping. What was the reason to switch beta-actin and alpha-tubulin? In Figures 3F -G, the GS expression is not a good representative image. Were chemiluminescence or fluorescence antibodies used? Were the membranes reused?

      We thank the reviewer for this careful and detailed evaluation of the Western blot data. We apologize that insufficient methodological detail was provided in the original submission.

      (1) We would like to clarify that the protein bands shown within each panel were derived from the same membrane. To improve transparency, we provided full, uncropped images of the corresponding membranes in the supplementary materials. In addition, detailed information regarding protein loading amounts, gel conditions, and housekeeping controls has also been added to the Methods section.

      (2) The use of different loading controls (β-actin or α-tubulin) reflects a technical consideration rather than an experimental inconsistency. In our experiments, the molecular weight of the TH (62kDa) was too close to that of α-tubulin (55kDa), and β-actin (42kDa) was therefore used to avoid band overlap and to ensure accurate quantification.

      (3) Regarding the GS signal shown in Figures 3F–G, we agree that the original representative image was suboptimal. This appears to be related to antibody performance rather than sample quality. To address this, we repeated the Western blot from Figures 3F–G using a newly validated antibody. The original tissue samples had been aliquoted and stored at −80 °C, allowing reliable re-analysis.

      (4) All Western blot experiments were detected using chemiluminescence, and membrane stripping and reprobing procedures are now explicitly described in the Methods section.

      We thank the reviewer for highlighting these issues, which significantly improve the rigor and clarity of our data presentation. All new figures and legends have been incorporated, with changes clearly highlighted for ease of review.

      (8) Key references using PRV for liver innervation studies are missing (Stanley et al, 2010 [PMID: 20351287]; Torres et al., 2021 [PMID: 34231420]; Desmoulins et al., 2025 [PMID: 39647176]).

      We thank the reviewer for pointing out these important and highly relevant references that were inadvertently omitted in our initial submission. The studies by Stanley et al. (Proc Natl Acad Sci U S A, 2010), Torres et al. (Am J Physiol Regul Integr Comp Physiol, 2021), and Desmoulins et al. (Auton Neurosci, 2025) represent key PRV-based retrograde tracing work that has mapped central neural circuits innervating the liver and thus provide essential context for our anatomical analyses.

      We agree that the inclusion of these studies is necessary to properly situate our findings within the existing literature. Accordingly, we incorporated citations to these references in the revised manuscript and discussed their relationship to our results.

      Reviewer #3 (Public review):

      Summary:

      This study found a lobe-specific, lateralized control of hepatic glucose metabolism by the brain and provides anatomical evidence for sympathetic crossover at the porta hepatis. The findings are particularly insightful to the researchers in the field of liver metabolism, regeneration, and tumors.

      Strengths:

      Increasing evidence suggests spatial heterogeneity of the liver across many aspects of metabolism and regenerative capacity. The current study has provided interesting findings: neuronal innervation of the liver also shows anatomical differences across lobes. The findings could be particularly useful for understanding liver pathophysiology and treatment, such as metabolic interventions or transplantation.

      Weaknesses:

      Inclusion of detailed method and Discussion:

      We sincerely thank the reviewer for the positive and constructive feedback, which significantly enhances both the methodological rigor and the broader biological interpretation of our study. In direct response, we revised the Discussion to elaborate on the potential physiological advantages of a lateralized and lobe-specific pattern of liver innervation. Furthermore, we expanded the Methods section to include a comprehensive description of the quantitative analysis applied to PRV-labeled neurons. Together, these revisions strengthened the manuscript’s clarity, depth, and relevance to researchers in hepatic metabolism, regeneration, and disease.

      (1) The quantitative results of PRV-labeled neurons are presented, and please include the specific quantitative methods.

      We thank the reviewer for this helpful suggestion. We have added a detailed description of the quantitative methods used to analyze PRV-labeled neurons in the revised Methods section. We have now provided detailed information in the Methods section, including the criteria used for cell counting, the anatomical boundaries of the brain regions analyzed, the delineation of regions of interest, and the normalization procedures applied to derive the reported neuron counts. These additions have been incorporated into the revised Methods, with all changes clearly indicated for ease of review.

      (2) The Discussion can be expanded to include potential biological advantages of this complex lateralized innervation pattern.

      We appreciated the reviewer’s suggestion. We have expanded the Discussion to include a paragraph addressing the potential biological significance of lateralized liver innervation. We highlight that this asymmetric organization could allow for more precise, lobe-specific regulation of hepatic metabolism, enable integration of distinct physiological signals, and potentially provide robustness against perturbations. The additional discussion content has been highlighted in the revised version as indicated (Discussion section, paragraph 3: “Bilateral LPGi activation produced additive effects, indicating that both sides of the brainstem can cooperatively regulate hepatic metabolism in a spatially segregated manner. This pattern suggests that hepatic glucose output can be modulated in a lobe-specific, rather than uniform whole-organ, manner.”).

      Reviewer #4 (Public review):

      Summary:

      The studies here are highly informative in terms of anatomical tracing and sympathetic nerve function in the liver related to glucose levels, but given that they are performed in a single species, it is challenging to translated them to humans, or to determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies is mechanistically informative. Denervation studies lack appropriate controls, and the role of sensory innervation in the liver is overlooked.

      We sincerely appreciate the reviewer's thoughtful evaluation and fully agree that findings derived from a single-species model must be interpreted with caution in relation to human physiology. In direct response, we revised the manuscript to explicitly clarify that all experimental data were obtained in mice and to provide a discussion of the limitations regarding direct extrapolation to humans. Concurrently, we expanded the Discussion section by integrating our findings with recent human and translational studies, including a multicenter clinical trial demonstrating that catheter-based endovascular denervation of the celiac and hepatic arteries significantly improved glycemic control in patients with poorly controlled type 2 diabetes, without major adverse events (Signal Transduct Target Ther. 2025;10(1):371). While our current work focuses on defining the anatomical organization and functional asymmetry of this circuit in mice, the clinical findings suggest that the core principles, sympathetic control of hepatic glucose metabolism via CG-liver pathways, may be conserved and of translational relevance. Additionally, we clarified the interpretation of TH labeling and expanded the discussion of hepatic sensory and parasympathetic innervation, acknowledging their important roles in liver-brain communication and identifying them as key directions for future research. Collectively, these revisions provide a more balanced, clinically informed, and rigorous framework for interpreting our findings.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      We thank the reviewer for this suggestion. We agree that the species should be clearly indicated. The findings presented in this study were obtained in mice using tissue clearing and whole-organ imaging approaches. Due to technical limitations, these observations are currently restricted to the mouse strain. We have updated the title (Symmetric brain-liver circuits mediate lateralized regulation of hepatic glucose output in mice) and clarified the species used throughout the manuscript.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also hits a portion of sensory fibers that need to be ruled out in whole-mount imaging data

      We thank the reviewer for pointing this out. We acknowledge that TH labels not only sympathetic fibers but also a subset of sensory fibers. We have added a limitation of this point in the revised manuscript. In addition, using SyGlass (2.4.0) three-dimensional reconstruction, we observed TH-positive nerve fibers originating from the CG-SMG extending along the porta hepatis and penetrating into the liver parenchyma. Given that the CG-SMG is a well-established sympathetic ganglion innervating visceral organs (Nature. 2025 Jan;637(8047):895-902.), these nerve fibers can be definitively identified as sympathetic. In parallel, we collected DRG from spinal segments T1-6 and T7-12 five days after intrahepatic PRV injection. While T7-12 DRG are known to contain sensory neurons innervating the liver, only a sparse number of PRV-positive neurons were detected in these segments (Anat Rec A Discov Mol Cell Evol Biol. 2004 Sep;280(1):827-35. Auton Neurosci. 2024 Jun;253:103174). The additional figure and discussion content have been highlighted in the revised version as indicated (Discussion section, paragraph 6: “Third, although whole-mount TH immunostaining with three-dimensional reconstruction revealed sympathetic nerve bundles projecting from the CG to the liver, TH is not entirely specific and can also label a subset of sensory neurons. More selective approaches, such as genetic targeting of sympathetic lineages, will be important for further validation.”).

      Author response image 2.

      Representative immunofluorescence images of PRV-labeled neurons (EGFP) in DRG from the spinal segments T1-6 (bottom) and T7-12 (top) following PRV injections into the liver lobes. Scale bars, 200μm

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. There is only a brief mention in the Discussion currently, but comparing methods and observations would be helpful.

      We thank the reviewer for this suggestion. Previous studies largely relied on electrical stimulation to modulate liver innervation, which provides relatively coarse control of neural activity (Eur J Biochem. 1992;207(2):399-411). By contrast, our use of chemogenetic and optogenetic approaches allows selective, cell-type-specific manipulation of LPGi neurons. We revised the Discussion to place our functional data in the context of prior work, highlighting how these more precise approaches improve understanding of the contribution of liver-innervating neurons to hyperglycemia. The newly added discussion has been clearly labeled in the response to facilitate your review (Discussion section, paragraph 3: “This spatial organization is likely obscured by conventional electrical stimulation, which indiscriminately activates heterogeneous sympathetic fibers. By contrast, chemogenetic and optogenetic approaches permit selective, cell type-specific manipulation of LPGi neurons, thereby revealing the contralateral and lobe-specific architecture of brain-liver sympathetic control”).

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases to tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though it is clearly assumed to be. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      We thank the reviewer for this insightful and important comment, which highlights a potential alternative interpretation of our findings. We agree that chemical sympathetic denervation with 6-OHDA may induce compensatory changes in non-sympathetic inputs, including sensory and parasympathetic (vagal) innervation of the liver.

      Conceptually, we agree with the reviewer’s perspective that the central nervous system operates as a highly integrated homeostatic regulatory system, continuously receiving and integrating a broad range of afferent signals. These inputs include, as noted by the reviewer, hepatic sensory and vagal afferents (Science. 2024;386(6722):673-677), as well as centrally derived interoceptive signals such as brain glucose, temperature sensing, even the pulsation of cerebral vascular system (Cell Metab. 2025;37(11):2264-2279.e10.; Cell Metab. 2022;34(6):888-901.e5; Science. 2024;383(6682):eadk8511). The CNS integrates these diverse signals and generates coordinated efferent outputs to maintain systemic homeostasis.

      From this viewpoint, the changes in c-FOS activity that we observe in the LPGi likely represent only a limited snapshot of this broader integrative process, rather than evidence of a single dominant pathway. We acknowledge that compensatory sensory or parasympathetic mechanisms, in addition to altered sympathetic drive, contributed to the observed LPGi activation following hepatic sympathetic denervation.

      We further acknowledge that, due to limitations in scope and experimental focus, we did not directly assess sensory or parasympathetic innervation of the liver in the present study. As appropriately pointed out by the reviewer, a more comprehensive characterization of hepatic neural inputs would provide a more complete picture of the underlying neurocircuitry. To address this, we expanded the Discussion and explicitly noted this limitation, including a more balanced discussion of potential crosstalk among sympathetic, sensory, and parasympathetic pathways and how these may collectively influence LPGi activity. For your convenience, the newly added discussion text has been distinctly marked in the manuscript (Discussion section, paragraph 4: “Although enhanced sympathetic output appears to mediate much of this compensation, our findings suggest that the underlying regulation extends beyond a purely descending pathway. In particular, c-FOS activation in the contralateral LPGi after unilateral 6-OHDA-mediated denervation suggests that the loss of peripheral input may be sensed through an ascending neural pathway, centrally integrated, and translated into compensatory sympathetic output to the intact hepatic lobes. These results therefore support a model in which hepatic glucose production is regulated by an integrated afferent-central-efferent loop, with our current analyses primarily resolving its efferent component.”).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Although the findings are interesting, this reviewer has major concerns about the experimental design, methodology, results, and interpretation of the data. Experimental details are lacking, including basic information (age, sex, strain of mice, procedures, magnification, etc.).

      We thank the reviewer for this important recommendation. We agree that comprehensive reporting of experimental details is essential for rigor and reproducibility.

      In the revised manuscript, we added complete information regarding mouse strain, sex, age, and sample size for each experiment. In addition, detailed descriptions of surgical procedures, viral constructs, injection parameters, imaging magnification, and analysis methods have been incorporated into the Methods section.

      These revisions ensured that all experiments are described with sufficient technical detail and clarity to allow accurate interpretation and replication of our findings. Experimental details have been incorporated, and corresponding revisions are highlighted for ease of review.

      Reviewer #3 (Recommendations for the authors):

      Addressing a few questions might help:

      (1) The study found that liver-associated LPGi neurons are predominantly GABAergic. It would be informative to molecularly characterize the PRV-traced, liver-projecting LPGi neurons to determine their neurochemical phenotypes.

      We thank the reviewer for this insightful suggestion. We agree that molecular characterization of liver-projecting LPGi neurons is important for understanding their functional identity.

      This issue has been addressed in detail in our recent study (Cell Metab. 2025;37(11):2264-2279.e10), in which we performed single-cell RNA sequencing on retrogradely traced LPGi neurons connected to the liver. These analyses demonstrated that the majority of liver-projecting LPGi neurons are GABAergic, with a defined transcriptional profile distinct from neighboring non–liver-related populations.

      Based on these findings, the current study selectively targeted GABAergic LPGi neurons using GAD1-Cre mice. We have explicitly cited these molecular results in the revised manuscript to clarify the neurochemical identity of the PRV-traced LPGi neurons. New text has been incorporated, and corresponding revisions are highlighted for ease of review.

      (2) Is it possible to do a local microinjection of a sodium channel blocker (e.g., lidocaine) or an adrenergic receptor antagonist into the porta hepatis? That would potentially provide additional evidence for the porta hepatis as the functional crossover point.

      We appreciated the reviewer’s thoughtful suggestion. Although pharmacological blockade at the porta hepatis can modulate local neural activity, this approach is inherently limited in its ability to distinguish between ipsilateral and contralateral inputs. Consequently, it may not provide definitive evidence for neural crossover at this specific site.

      In our view, the anatomical evidence provided by whole-mount tissue clearing, dual-labeled tracing, and direct visualization of decussating nerve bundles at the porta hepatis offers a more definitive demonstration of sympathetic crossover. Pharmacological blockade would affect both crossed and uncrossed fibers simultaneously and therefore would not specifically resolve the anatomical organization of this decussation.

      Nevertheless, we agree that functional interrogation of the porta hepatis represents an interesting direction for future work, and we acknowledge this possibility in the Discussion (Paragraph 6: “Fourth, although our data support a peripheral decussation at the porta hepatis, direct validation of this crossover site was not feasible with local pharmacological blockade, as currently available approaches lack sufficient spatial specificity and would likely perturb multiple neural components. Future studies employing more selective inhibitory strategies will be required to directly test this possibility.”).

      (3) It is possible to investigate the effects of unilateral LPGi manipulation or ablation of one side of CG/SMG on liver metabolism, such as hyperglycemia?

      We thank the reviewer for this important suggestion. Because unilateral LPGi manipulation was already examined in our study (Figure 2D), we focused here on unilateral ablation of the CG to further assess lateralized sympathetic control of hepatic metabolism. We successfully performed unilateral CG ablation without LPGi manipulation, but observed no significant change in blood glucose compared with the sham group (Author response image 3A and 3B). To determine whether glucose homeostasis was nonetheless affected, we further performed glucose tolerance tests (GTT) and insulin tolerance tests (ITT) (Author response image 3C and 3D). Neither test showed significant impairment after unilateral ablation, suggesting that compensatory neural mechanisms and/or hormonal homeostatic regulation may be recruited to preserve systemic glucose homeostasis.

      Author response image 3.

      (A) Blood glucose levels in mice subjected to left- or right-sided CG ablation via 6-OHDA treatment (n = 6). (B) Representative images of ablation of CG. Scale bars, 100 μm. (C and D) Blood glucose levels during GTT (C, n = 6) and ITT (D, n = 6) in mice with left- or right-sided CG ablation.

      Reviewer #4 (Recommendations for the authors):

      In the abstract and elsewhere, the use of the term 'sympathetic release' is unclear - do you mean release of nerve products, such as the neurotransmitter norepinephrine? This should be more clearly defined.

      We thank the reviewer for pointing out this ambiguity. We agree that the term “sympathetic release” was imprecise. In the revised manuscript, we explicitly referred to the release of sympathetic neurotransmitters, primarily norepinephrine, from postganglionic sympathetic fibers.

      We revised the wording throughout the manuscript to ensure accurate and consistent terminology and to avoid potential confusion regarding the underlying neurobiological mechanisms.

      Original: “Following unilateral hepatic denervation, contralateral LPGi activation induced metabolic compensation in the remaining innervated lobes, characterized by increased sympathetic release, glucose production, and glycogen depletion.”

      Revised: “Following unilateral hepatic denervation, contralateral LPGi activation induced metabolic compensation in the remaining innervated lobes, characterized by increased norepinephrine release, glucose production, and glycogen depletion.”

    1. Author response:

      The following is the authors’ response to the previous reviews

      We thank you for the time you took to review our work and for your feedback! The main changes to the manuscript are:

      We added a paragraph to the Discussion addressing differences in visuomotor mismatch responses recorded over frontal and occipital electrodes, and their possible interpretation.

      We added time-frequency power and phase-locking analysis as supplementary figures to the manuscript.

      We added a statement in the Discussion emphasizing the importance of performing these experiments with denser EEG channel coverage.

      Public Reviews:

      Reviewer #1 (Public review):

      In this paper, Solyga, Zelechowski & Keller study human visuomotor mismatch responses as an alternative instantiation of prediction errors to classic oddball paradigms. Using VR, they created a condition in which participants were moving around thereby creating a visuomotor coupling between physical movement and visual flow. To attempt to isolate the contribution of specifically movement-related predictions in this condition, they contrasted it to a condition in which participants were seated and rewatching their movement trajectory during the 'active' condition. Visuomotor mismatches were created by temporarily decoupling movement and visual experience by halting the VR display as participants continued to move.

      The core finding of the paper is that participants exhibit a positively-valenced response to the visuomotor decoupling in the active but not in the passive condition. Since walking speed only insignificantly slows down following decoupling events in the active conditions, the authors argue that this difference can not be accounted for by "changes in participants' behavior or to simple visual offset responses" with the latter being equal across both conditions. The following reinstatement of the coupling in turn does not differ between the two conditions. The authors additionally show that this mismatch response differs from visual onset responses elicited by checkerboard inversions and that it's "qualitatively" stronger than more commonly studied auditory oddball mismatch responses.

      The design with its focus on ecological validity is impressive, well-rationalized and the results are well illustrated. I additionally appreciate the control analyses with regards to changes in walking speed and playback DOF and, now added, additional participants who experience the passive condition before the active. I have a couple of questions/comments.

      My main question in round 1 regarded the isolation of visuomotor mismatch. Although the comparison with a seated control seems like a very sensible way to control for simple visual responses, there seem to be more differences than just a break in visuomotor coupling between the conditions. I therefore wonder whether the reduced offset response in the seated condition may be, in part, explained differently. For example, given that participants always conduct the active condition before rewatching their movement in the seated condition, it seemed likely that there is a component of learning across the session that flow will sometimes be halted. This is confirmed with the analyses. The explanation that there is a visuomotor component here is given further weight by their conduction of an additional group of participants who perform the conditions in the reverse order, so this has strengthened the manuscript considerably. However, it does of course remain an imperfect control because the visual stimulus is now different between the conditions for these participants. It's the best that can be achieved with this type of paradigm though and of course it yields a great deal of ecological validity.

      The reviewer is correct. But one should keep in mind that our result here stands in the context of a considerable amount of work on mouse cortex investigating responses to very similar visuomotor mismatches. There we can we have much additional evidence to argue that the cortical response to a visuomotor mismatch is a prediction error. We would argue, it is the best one can do in human experiments.

      I was also wondering whether the authors may consider the findings in frontal electrodes more closely given that the title of the paper focuses on a specifically occipital effect. Their further analyses have confirmed that there are likely interesting frontal effects. From a theoretical point of view, the spatial dissociation in adaptation effects, which were stronger in frontal and weaker in occipital areas, seems interesting and perhaps worth discussing, especially given the interpretation that "mismatch processing may initially arise in sensory visual areas before engaging higher-order frontal regions." How come the frontal decrease in responses is not accompanied by an analogous decrease in its supposed occipital source? Could these two responses reflect different kinds of prediction error signals (i.e. objective vs subjective)?

      We have added a paragraph to the Discussion addressing the differences between signals recorded over frontal and occipital electrodes, as suggested.

      I remain concerned that the authors fight too defensively that they have absolutely isolated visuomotor prediction mechanisms with this paradigm. It's a nice, informative study, but it seems odd to argue there are no other possible explanations. One picks a design to optimize some features but they will always come at some cost to others. Prioritising ecological validity, which is a justifiable aim, necessarily usually weakens some control over confounds.

      We are not sure what the reviewer is referring to here. We certainly do not think (or are aware of having argued) that a visuomotor prediction error is the only possibly interpretation of the response. In the last paragraph of our response to the reviewers point 3 in the last revision, we explicitly discuss that the interpretation of the responses as a prediction error is only one possible interpretation. Our argument is that it is the most likely given the evidence.

      To outline my reasoning fully: My concerns wrt generic influences of action on perception are reflected in Fig 1. The P1 is smaller when walking than sitting. It seems likely that the mismatch response reflects something about extrapolation or prediction, because it is larger when walking. However, it's not necessarily sensorimotor prediction. Even if you remove action from the equation, the flow can be extrapolated or predicted most of the time in a way it cannot so well when the video is halted. Of course the sitting condition somewhat controls for it, but when it came second the visual flow disruptions were more predictable here. A reduction in effects over time is indeed confirmed with their analyses. They now have conducted a study with the conditions in the reverse order and they find the same thing. But of course this necessitates non-identical visual flow because the sitting condition is playing the previous participant's flow. So it is likely that across all of these comparisons, it is the visuomotor mismatch that is especially salient. It's just that each comparison is a bit messy/confounded. It would strengthen the manuscript if there were some consideration given to the other processes likely at play here.

      We would be happy to add additional considerations to other processes. If the reviewer has anything specific in mind, we can add that, but it would need to be somewhat concrete with some theoretical basis. We share the reviewer’s intuition, but unless this can be formalized to the point of being experimentally testable, we do not see any value in discussing it in the manuscript.

      Regarding the reason for a difference in visual responses in walking vs sitting state is, this is not entirely clear to us. Predictive processing would provide one possible explanation. Assuming the precision weighting of predictions is higher during walking, the sudden appearance of a visual stimulus might lead to stronger stimulus history prediction errors than when just sitting. But this is rather speculative.

      As a more minor point in response to our previous review, whether particular accounts represent an 'orthodox' view at present does not determine whether they raise logical issues in need of consideration. The authors may have missed that the papers in question consider mechanisms underlying the attenuation of particular pieces of information ‘from perception’. Not perceptual processing. We have one percept at any one moment in time and must understand how different population types synergistically generate that percept.

      Please excuse, the reviewer is correct, the orthodoxy of an idea is not relevant. For dubious reasons, we chose to euphemize what we actually meant to say here. With regards to circuit implementations of predictive processing (we cannot and do not intend to speak to interpretations of predictive processing that relate to conscious perception much of V1 activity is likely not consciously perceived – we assume this is what the reviewer is referring to by “we have one percept”) – the reviewers interpretation was not unorthodox, but rather incorrect (which is what we should have said). The statement that “the brain predictively ‘cancels’ expected action outcomes from perception” is incorrect in the context of sensory processing – based on both theoretical models of predictive processing, and more importantly physiological evidence. If the point was only in regards to conscious perception, we also suspect the statement is wrong, but even if it were correct, don’t see how it pertains to our work.

      Similarly a little strange is the way in which the authors aggressively defend the position that self-generated motion is 'the strongest' type of prediction. Sure, we probably experience the effects of our actions more often than ambulances. But what about objects obeying laws of gravity or others' faces being structured and moving in systematic ways? It is hard to quantify, such that presumably many scientists would be skeptical of such a claim, and it is not needed logically to justify the importance of examining mechanisms enabling action to shape perceptual processing. I'd assume it better to fight the battles you need to (and can) fight, such that the robust claims carry more weight.

      We believe it is absolutely essential for the progress of the field that we start to emphasize the differences between something that is “predictable in principle” and “predicted by the brain”. There is likely indeed a hierarchy of predictability that looks something like this:

      (1) Sensorimotor coupling

      (2) Laws of physics

      (3) Behavior of other living things

      (4) Artificial, human-made statistical relationships

      Almost all published experiments are based on the fourth type of prediction. Indeed, why not use physics simulations instead of oddballs and MMN? We absolutely should! But the field tends to revert to artificial couplings. As a direct consequence of this, the number of papers appearing recently (from both human and mouse fields), that are built on the following premise:

      (1) Expose an animal or human to an artificial coupling between A and B (e.g. an oddball, or a global oddball, or any of a myriad other constructions).

      (2) Probe for prediction error responses to the violation of the artificial coupling.

      (3) Find no prediction error responses and conclude predictive processing is wrong.

      Is utterly baffling. The fallacy here is of course the assumption that if something is predictable in principle, the brain must predict it. Thus, we are, and will continue to be strong on this point, and we think it is essential that we – as a field – are.

      Hope these comments are helpful.

      Reviewer #2 (Public review):

      Summary:

      This study investigates whether visuomotor mismatch responses can be detected in humans. By adapting paradigms from rodent studies, the authors report EEG evidence of mismatch responses during visuomotor conditions and compare them to visual-only stimulation and mismatch responses in other modalities.

      Strengths:

      Authors use a creative experimental design to elicit visuomotor mismatch responses in humans.

      The study provides an initial dataset and analytical framework that could support future research on human visuomotor prediction errors.

      Weaknesses:

      Methodological issues (e.g., volume conduction) make it difficult to confidently attribute the observed mismatch responses to activity in visual cortical regions. This could be alleviated by increasing the number of channels.

      We have added a discussion of this.

      The authors successfully demonstrate that visuomotor mismatch paradigms can, in principle, be applied in human EEG. This approach provides a translational bridge between rodent and human work on predictive processing.

      Reviewer #3 (Public review):

      Solyga, Zelechowski, and Keller present a concise report of an innovative study demonstrating clear visuomotor mismatch responses in ambulating humans, using a mobile EEG setup and virtual reality. Human subjects walked around a virtual corridor while EEGs were recorded. Occasionally, motion and visual flow were uncoupled, and this evoked a mismatch response that was strongest in occipitally placed electrodes and had a considerable signal to noise ratio. It was robust across participants and could not be explained by the visual stimulus alone.

      This is an important extension of their prior work in mice, and represents an elegant translation of those previous findings to humans, where future work can inform theories of e.g. psychiatric diseases that are believed to involve disordered predictive processing. For the most part, the authors are appropriately circumspect in their interpretations and discussions of the implications. The paper in its current form represents an important addition to the literature.

      The authors have included analyses of the auditory mismatch using temporal electrodes, referenced to Cz (and therefore should exhibit a mismatch positivity). This added data clearly and convincingly shows that the sensorimotor mismatch is, indeed, stronger than the passive auditory MMN.

      The reference electrode placed at Cz makes it is difficult to interpret relative differences between frontal and occipital electrode responses, as the occipital electrodes are placed farther away from the Cz reference than the frontal electrodes. Similarly, signal occuring cortically near the Cz reference might only appear as though it is occipitally distributed in this montage. It is common in EEG research to remontage the data to an averaged common reference in order to better interpret the scalp distributions. As the electrode coverage was sparse for some subjects, this could be challenging, and this reviewer does not feel that it is necessary to do this analysis step, or even to drastically rewrite the body of the paper. We only request that some discussion, however brief, is included in the discussion section or the methods that recommend more dense electrode coverage in the future to better interpret scalp distributions and potential meso-scale sources.

      We have added a discussion of this as suggested.

      This is just a suggestion. The authors are encouraged to analyse (and report) time-frequency power and phase locking for these mismatch responses, as is common in much of the literature (see Roach et al 2008 Schizophrenia Bulletin). This is not to say that doing so will yield insights into oscillations per se, but converting the data to the time-frequency domain provides another perspective that has some advantages. fosters translations to rodent models, as ERP peaks do not map well between species, but e.g. delta-theta power does (see Lee et al 2018 Neuropsychopharmacology; Javitt et all 2018 Schizophrenia research; Gallimore et al 2023 Cereb Ctx). Further, ERP peaks can be influenced by the actual neuroanatomy of an individual (especially for quantifying V1 responses). Time frequency analyses may aid in interpreting the "early negative deflection with a peak latency of 48 ms " finding as well. As it stands, the report is complete, and it would be acceptable if the authors chose to save this type of analysis for a future publication.

      We have added this as suggested.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The authors have addressed most of my concerns by providing additional analyses, partly based on new data. The volume conduction issue is partly addressed based on the result showing latency differences, however to confidently assign responses to visual regions, one would need to perform recordings with a larger number of electrodes, sufficient to perform source localization. Nevertheless, the manuscript is now more solid than the previous version.

      We have now added this point to the Discussion.

      Reviewer #3 (Recommendations for the authors):

      The reviewer appreciates that the authors have carried out time-frequency analyses, and are ok with them leaving this out of this paper.

      We have now added this to the manuscript.

      Finally, in response to the participant quote "are you printing this? hi mom!" - this reviewer concedes that it does not significantly detract from the report, and, in the interest of amusement and joy, would abide its reinstatement.

      We greatly appreciate the reviewers entertaining our attempts at humor but will leave it out as originally suggested.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors wanted to better understand how the various septin-associated kinases contribute to septin organization and function in budding yeast. This question has been recently addressed by similar kinds of studies but there are still some open questions, particularly as regards to what extent the kinases may interact with and/or modify components of the contractile ring that drives cytokinesis.

      Strengths:

      This study uses sensitive imaging with good temporal and spatial resolution to monitor the localization of various proteins in living cells. Particularly informative is the use of a GFP/GFP-binding-protein "tethering" approach to ask if the requirement for one protein can be bypassed by physically tethering another protein to a third protein. Results from a yeast two-hybrid assay for measuring protein-protein interactions in vivo are buttressed by direct in vitro binding assays using purified proteins, which is important given the likelihood of "bridging" interactions between yeast proteins in the two-hybrid approach. The authors' conclusions are quite well supported by the data.

      Weaknesses:

      A control for non-specific binding is missing from the in vitro binding assay. The figures suffer sometimes from the very small text in the labels, which obscures understanding. Ultimately, while the study provides some interesting and novel insights, we still don't understand which phosphorylation events on which proteins are important for the events occurring at the molecular level, so the advance in knowledge is somewhat incremental.

      We thank the reviewer for highlighting the strengths of our imaging pipelines and protein-protein interaction data. We have now included appropriate controls for the in vitro binding assays, which demonstrate that the observed interactions are specific (Fig. 2H). We have also revised all figures to improve clarity, including increasing font sizes to enhance visibility across panels. We agree that mapping the specific phosphorylation sites regulated by these septin kinases would provide valuable mechanistic insights. However, only a few studies have addressed this direction so far (Mortensen et al., 2002; Asano et al., 2006 and Marquardt et al., 2024) [1-3]. The current study focuses on the interplay among septin-associated kinases and their role in regulation of the actomyosin machinery (AMR). In this context, we highlight several key findings:

      (i) a molecular link between the septin kinase network and AMR through physical interaction between the KA1 domain of Gin4 and F-BAR domain of Hof1 (Fig. 2F-2H and S3H);

      (ii) a kinase-independent role for Gin4 in coordinating septin organization and AMR dynamics (Fig. 3A-3E, S3F & S3G, S3I and 4F-4H);

      (iii) a novel role for Hsl1 in regulating septins and the AMR downstream of Gin4 and Elm1, potentially through plasma-membrane binding (Fig. 4I-4K, 5A-5F, 8A-8G and S9H-S9J); and

      (iv) crosstalk between Gin4 and Hsl1 that is independent of their role in the morphogenetic checkpoint (Fig. 6A-6D and S6A-S6C).

      We have clarified this scope in the Discussion section and explicitly stated that mapping these phosphorylation sites will be an important direction for future work (Lines 623-626).

      Reviewer #2 (Public review):

      Summary:

      In this paper, Bhojappa et al. provide insights into the function of septin-related kinases Elm1, Gin4, Hsl1, and Kcc4 in septin organization and actomyosin ring (AMR) structure and constriction. Their findings are both corroborative of and complementary to previous related studies.

      First, the authors provide a comparative analysis of the dynamic localization of these kinases at the bud neck, as well as a comparative analysis of defects in septin localization, splitting dynamics, AMR constriction rates, and cell morphology in kinase deficient cells. They find that septin localization and splitting kinetics, as well as AMR constriction rates, are significantly perturbed in elm1∆ and gin4∆ mutants but remain largely unaffected in hsl1∆ and kcc4∆. A similar trend is observed in terms of cell morphology and viability.

      Next, the authors focus on elm1∆ and gin4∆ cells, demonstrating that the residence time of the F-BAR protein Hof1 is significantly increased and defective in these mutants. Using yeast two-hybrid (Y2H) and in vitro binding assays, they show that the KA1 domain of Gin4 interacts with the F-BAR domain of Hof1, which may explain the cytokinesis-related functions of Elm1 and Gin4. Supporting this, they find that Gin4's role in septin localization, AMR constriction kinetics, and Hof1 bud neck localization is kinase-independent.

      The authors then conduct a series of artificial tethering experiments given their bud neck localization is mostly interdependent. They first demonstrate that artificially tethering Gin4 to the bud neck rescues the morphology defects of elm1∆ cells, with the strongest rescue observed when Gin4 was forced to interact with Hsl1-an effect that was also kinase-independent. Additionally, artificial tethering of Hsl1 to the bud neck restores the morphology of elm1∆ cells in a KA1 domain-dependent manner, suggesting that Hsl1 functions downstream of Elm1 to maintain normal cell morphology. Consistently, artificial tethering of Elm1 to the bud neck in gin4∆ cells rescues morphology defects, as well as defects in Myo1 localization and AMR constriction, but only in the presence of full-length Hsl1. The rescue fails in the absence of Hsl1 or when using a version of Hsl1 lacking the KA1 domain, which supports the role of Hsl1 downstream to Elm1 in cytokinesis.

      Strengths:

      Altogether, this study offers valuable insights into the mode of cytokinesis regulation mediated by the septin-related kinases, mainly Elm1, Gin4, and Hsl1, and would be an important contribution to the field of septins and cytokinesis after addressing current weaknesses.

      We thank the reviewer for the detailed summary and for highlighting the novel findings of our study.

      Weaknesses:

      (1) When assessing rescue of the elm1∆ phenotype, it needs to become clearer whether only morphology or also cytokinesis and septin organization are rescued.

      To clarify the extent of rescue observed in elm1Δ cells, we extended our analysis beyond morphological parameters by quantifying septin organization and AMR constriction dynamics. These analyses now show that artificial tethering partially restores septin organization and AMR constriction kinetics in elm1Δ cells in addition to improving cell morphology. These results are now described in detail in Fig. 5 and lines 373-406.

      (2) The quantification of the microscopy data does not always match up with the example images, and it's not always clear how the authors quantitatively analyzed their data.

      We revised the manuscript to clearly outline the quantification methods used for microscopy data analysis and specified the statistical tests, number of cells analyzed, and number of experimental replicates in the figure legends. We also clarified the criteria used for phenotype scoring and quantification in the Materials and Methods section. In addition, We replaced representative images where necessary to accurately reflect the quantified data throughout the revised manuscript.

      (3) The forced tethering data are key to the paper, but the lack of a summarizing table makes it difficult to grasp the full picture.

      We agree with the reviewer and have now included a new summary table (Table 1) that compiles the results of all artificial tethering experiments presented in this study, including the percentage of rescue in cellular morphology observed upon forced tethering of these kinases to the bud neck, thereby providing a clearer overview of these experiments.

      (4) Novel results and those confirming earlier results could be better distinguished.

      We have improved the overall clarity of the manuscript to distinguish novel findings from the results that corroborate previous studies, and have cited the appropriate literature throughout the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The study by Bhojappa et al. brings new and interesting elements about the stability of the septin ring and the crosstalk between septin and actomyosin ring assemblies. The study focuses on the four kinases associated with the septin ring, Elm1p, Gin4p, Hsl1p, and Kcc4p. Elm1 and Gin4 show strong knock-out phenotypes, whereas Hsl1p and Kcc4p show weak knock-out phenotypes. The Elm1p/Kccp1p and Gin4p/Hsl1p pairs show similar timing at the bud neck. While these kinases share redundant functions, Gin4 appears to have a unique interaction with the BAR domain protein Hof1, revealing a novel direct interaction between the septin and actomyosin rings. Interestingly, the kinase activity of Gin4 is not required for its role in septin organisation and AMR constriction. The last part of the manuscript shows an original protein tethering protocol used to show that Hsl1 and its membrane binding ability are required for phenotype rescue of gin4null cells.

      Strengths:

      The combination of genetics, cell imaging, and biochemical characterization of proteinprotein interactions is attractive.

      We thank the reviewer for recognizing the significance of our findings and for the helpful suggestions.

      Weaknesses:

      (1) Imaging and data analysis is the main weakness of this manuscript. The authors must avoid manual counting and selection when easy analysis software can be used to limit bias. Instead of presenting unclear statistics of "percentage phenotypes", they need to define clear metrics to offer meaningful phenotype analysis.

      We agree that improving the quantitative rigour of the image analysis is essential for this study. Accordingly, we implemented a semi-automated image analysis workflow in the revised manuscript that defines reproducible metrics, such as aspect ratio, and reduces reliance on subjective phenotypic scoring. The inclusion of these parametric measurements enables clearer and more objective comparison of the rescued phenotypes.

      (2) This manuscript examines a very complex mechanism with four kinases of overlapping function using new data and existing literature. A clearer picture/model at the end of the manuscript that synthesizes the current knowledge would be beneficial:

      We incorporated a new representative model (Fig. 9) that integrates current knowledge in the field with our findings. This model highlights crosstalk among Elm1, Gin4, and Hsl1 as a key mechanism coordinating septin architectural transitions with AMR constriction during cytokinesis and is discussed in lines 520-536 of the revised manuscript.

      We sincerely thank all the reviewers for their insightful comments. We incorporated new results in Fig. S3A, S3B, S4A-S4F, 5A-5F, 6A-6D, S6A-S6C, 8E, 8F, S10B and 9, along with Table 1 summarizing the artificial tethering experiments in the revised manuscript. We believe that these revisions have improved the rigor of our analyses and enhanced the overall clarity of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 70: " abnormal cytokinesis defects": this language is redundant, either "abnormal" or "defects" would suffice:

      We thank the reviewer for identifying the redundant wording. We have rephrased the final paragraph of the Introduction section.

      Please refer to line numbers 86-97.

      (2) Currently, the final paragraph of the Introduction is an extensive, detailed summary of the results. This is unnecessary, as the Abstract and the Results sections summarize the results. Better would be a short statement of the questions addressed in the manuscript:

      We thank the reviewer for this suggestion. We have revised the final paragraph of the Introduction to remove the detailed summary of results and instead outline the key questions addressed in this study. The revised paragraph now emphasizes the knowledge gap regarding how septin-associated kinases regulate septin organization and coordinate cytokinesis. Please refer to line numbers 86-97.

      (3) In the Results, the wording of the heading associated with the first section is confusing. "and their defects during the cell cycle": it is not expected that the kinases themselves will have defects; defects may be observed in cells upon mutation of the kinases, but that is not clear from this wording:

      We thank the reviewer for this comment. The heading has been revised from “and their defects during the cell cycle” to “and defects associated with their deletions” for improved clarity.

      Please refer to line numbers 99-100.

      (4) Lines 107-109 (" they may act as molecular signals for the transition and a trigger for crosstalk of cytokinesis") are redundant with earlier lines 104-105 (" suggesting them to be a possible trigger for septin remodelling"):

      We have rephrased the text to improve clarity and avoid redundancy.

      Please refer to line number 123-125.

      (5) Lines 121-122: "Septin-associated kinases are believed to play an essential regulatory role": is the role essential, or is it regulatory? Since these kinases are not individually essential for cytokinesis, "essential" doesn't seem appropriate here:

      This has been corrected in the revised manuscript.

      (6) Figure 1 panels C and F: the font size is exceedingly small for the labels and should be greatly increased. This is true for multiple panels in Figures 2-5 and the Supplemental Figures as well:

      We thank the reviewer for highlighting this issue. We have significantly increased the font sizes of the text and the x- and y-axis labels across all figures in the manuscript.

      (7) The timing of mitotic spindle breakdown was used as a timepoint for comparison but it is not made clear in the manuscript how this was determined. Presumably, the mRuby2-Tub1 marker was visualized and the timepoint when the mitotic spindle separated into two discrete entities was called the "breakpoint" timepoint, but it would be important to better describe (and ideally show an example) of how that timepoint was determined. The only images I find with labelled tubulin are Tub1-GFP and these do not show spindle breakdown:

      We thank the reviewer for raising this point. We used GFP-Tub1 (pAFS125-GFPTUB1) and mRuby2-Tub1 (pHIS3p:mRuby2-Tub1+3′UTR::URA3) plasmids to visualize spindle dynamics across the cell cycle, and in both cases defined the spindle breakpoint as time zero. As suggested, we have now included time-lapse images of Cdc3-mCherry and GFP-Tub1 in wild-type cells in Fig. S2D to illustrate the spindle breakpoint event used for temporal alignment. Corresponding changes have also been made in the Materials and Methods section to explicitly describe this analysis.

      Please refer to line numbers 760-761.

      (8) Line 165: "F-BAR protein Hof1, which senses and induces membrane curvature" and lines 170-172 "the F-BAR protein Hof1, which is known to be associated with septin hourglass and transit to AMR during split ring trigger". It is awkward to introduce the same protein twice, in two different ways, within a few lines of each:

      We thank the reviewer for noting the redundant description. We have rephrased the text to introduce the F-BAR protein Hof1 in a more concise and streamlined manner while retaining the relevant functional information.

      Please refer to line numbers 206-208.

      (9) Throughout, it would be helpful to introduce more line breaks and organize the Results into smaller paragraphs:

      We thank the reviewer for this suggestion. We have revised the layout of the Results section and introduced additional line breaks to improve readability.

      (10) This is somewhat of a personal preference, but in the interest of transparency (and with the understanding that the 0.05 value is entirely arbitrary), would authors be willing to show actual P values in the figure panels rather than "**" or "ns", for example? Some readers may wish to apply a different standard of significance than 0.05, and not showing the P values makes this impossible. Furthermore, some readers (like this one) may interpret differently a P value of 0.044 versus "*" or 0.051 vs "ns":

      We agree with the reviewer’s suggestion. To improve the transparency of the quantitative analyses, we have modified the graphs across all main and supplementary figures to display the “actual p-values” and specified the corresponding statistical tests in the figure legends, with significance indicated by asterisks. An example is provided in the attached image showing the residence time of Inn1-mNG in wild-type and gin4Δ cells complemented with kinase-active and kinase-dead Gin4 constructs (Fig. 3E in the revised manuscript). For this analysis, significance was assessed using the nonparametric Kruskal-Wallis statistical test.

      (11) It is written that the authors "performed a Yeast Two Hybrid screen to find novel interacting partners at the bud neck" but I do not find evidence anywhere of a "screen" being performed, i.e., an unbiased search of many proteins to find a few interactors. Instead, it appears that the authors performed a two-hybrid assay to visualize interactions between a small, specific set of proteins. "Screen" should be replaced with "assay", as is currently the case in the Methods section. It is probably also valuable to point out in the text that using a yeast two-hybrid assay to assess interactions between yeast proteins has the caveat that any interactions observed could be indirect, as they may be "bridged" by endogenous yeast proteins.

      We thank the reviewer for raising this point. Our Yeast Two-Hybrid experiments were performed using a small subset of septin-associated proteins rather than as an unbiased screen. Accordingly, we have replaced the term “screen” with “assay” throughout the revised manuscript and in the Materials and Methods section.

      Please refer to line 241.

      We also agree with the limitations inherent to this assay and have now explicitly stated it in the revised manuscript, please see lines 248-255.

      (12) Panel 2H: I do not understand what the middle (as opposed to the top and bottom) blot segment represents. It is labelled "anti-HIS", like the one above it, but it is not associated with any molecular weight/ladder marker and I do not know what other species in the binding reaction in that lane would be recognized by the anti-HIS antibody. Perhaps the top segment is an "input" sample, and below it (in the middle segment) is what was bound to the beads. The figure legend is uninformative in this regard. Also, panel I in this figure is unnecessary to show, assuming that the bands shown in H are what I think they are. The blot makes the point without the need for quantification.

      As suggested by the reviewer, we have added molecular weight markers for each blot panel showing the input and bead-bound fractions. The figure legend has also been updated accordingly to clearly describe the different blot segments and experimental conditions.

      Please refer to figure legend 2H.

      In addition, as suggested by the reviewer, we have removed the quantification graph corresponding to the in-vitro binding assay from the revised manuscript.

      (13) The results in Figure 2H demonstrate that the Gin4-KA1 fragment is not non-specifically "sticky", because it does not bind GST alone, but there is no demonstration that binding by the Hof1 fragment is specific because there is no equivalent negative control for binding:

      We thank the reviewer for this suggestion. We repeated the in-vitro binding assay using 6His-bdSUMO as a negative control alongside 6His-bdSUMO-Gin4<sup>KA1</sup> to demonstrate binding specificity of the Hof1 fragment. The Hof1 N-terminal F-BAR fragment did not pull down the control 6His-bdSUMO fragment but specifically pulled down 6His-bdSUMO-Gin4<sup>KA1</sup> under identical experimental conditions (Fig. 2H), confirming the specificity of the interaction between Hof1 F-BAR domain and Gin4KA1.

      The corresponding text and results have been updated in the revised manuscript.

      Please refer to Fig. 2H and lines 259-263.

      (14) Line 257-258: "Localisation via Hsl1 is necessary to rescue the morphological defects exhibited by Δelm1 cells partially": what does "partially" refer to here? To the rescue, or the defects?

      The term “partially” refers to the extent of rescue. Specifically, elongated cell morphology was rescued in 63.75% of the elm1Δ cell population, rather than in all cells, upon artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP.

      (15) Lines 292-293: "can restore the morphological defects": this wording is unclear. "Restore" means "return to a former condition", which in this case would be normal cellular morphology, not defective cellular morphology. Similarly, see line 304: "While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering": presumably the proper localization was restored, not the mislocalization:

      In Lines 292-293, by phrase “can restore the morphological defects” was intended to indicate rescue of the elongated/clumped morphology associated with gin4Δ cells upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP. We have now rephrased this sentence as: “can rescue the elongated/clumped phenotype exhibited by gin4Δ cells”.

      Please refer to line numbers 481-482.

      Similarly, in Line 304, the statement “While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering” referred to rescue of the Myo1 mislocalization phenotype observed in gin4Δ cells. We have rephrased this sentence as: “However, Myo1-3xmCherry localization was restored to the bud neck upon artificial tethering of Elm1-GFP in gin4Δ cells”.

      Please refer to lines 506-508 in the revised manuscript.

      (16) Lines 295-296: "We find that Elm1-GFP tethering via Shs1-GBP, Bud4-GBP, and Hsl1-GBP" this should be "or", not "and":

      Thank you for pointing this out. We have now corrected the text accordingly.

      Please refer to line number 485.

      (17) Lines 311 and 312 refer to "Inn1-3xmCherry lifetime" but previously Inn1 residence time was measured. Since fluorescence lifetime is a distinct kind of measurement/ assay, it seems important to clarify here what kind of experimental data are being referred to:

      The term “Inn1-3xmCherry lifetime” was intended to describe the residence time of Inn1 at the cell division site, defined as the interval between the initial appearance of the Inn1 fluorescence signal and its complete disappearance during cytokinesis. For clarity and consistency, we have replaced the term “lifetime” with “residence time” in the revised manuscript.

      Please refer to the line numbers 511 and 512.

      (18) Discussion: "Septins are considered as the fourth cytoskeletal elements due to their extensive structural and functional diversity." This sentence is confusing, as it seems to imply that what defines a protein as being "cytoskeletal" is structural and functional diversity rather than anything to do with forming filaments, etc. The rest of this first paragraph of the Discussion also sounds like a summary of the background, is quite redundant with the Introduction, and should be shortened:

      We thank the reviewer for this suggestion. We have rewritten the first paragraph of the Discussion to improve clarity, reduce redundancy with the Introduction, and better emphasize the main findings of the study.

      Please refer to the line numbers 538-552 in the Discussion section.

      (19) Lines 365-366: "Gin4 and Hof1 are synthetic lethal": this should be revised to "gin4∆ and hof1∆ are synthetic lethal". This is a good place to point out that standard yeast nomenclature inserts the ∆ symbol after the gene name, not before (as is done in E. coli genetics, for example):

      We thank the reviewer for this suggestion. We have replaced “Gin4 and Hof1 are synthetic lethal” with “gin4Δ and hof1Δ are synthetic lethal” in the revised manuscript.

      Please refer to line number 576.

      In addition, we have now consistently placed the Δ symbol after the gene throughout the manuscript in accordance with standard yeast nomenclature.

      (20) Line 399: "We also performed an extensive GFP-GBP screens": again, here "screen" implies that a large collection of genes/proteins were assayed, perhaps in an unbiased way, which does not accurately portray what was actually done, which was an extensive tethering study using GFP-GBP:

      We thank the reviewer for this suggestion. We have replaced the term “GFP-GBP tethering screen” with “GFP-GBP tethering assay” throughout the revised manuscript. In addition, we have included a brief description of the specificity and functionality of GBP nanobody and its application in the GFP-GBP tethering strategy, extensively used in this study.

      Please refer to line numbers 326-336.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) Analysis of morphological defects of elm1∆ does not directly reflect the defects in cytokinesis and septin organization. For example, the deletion of SWE1 rescues the morphological defects of elm1∆ cells but not the cytokinesis or septin mislocalization (Bouquin et al., 2000). Considering this, the authors should address whether artificial tethering of Hsl1 to Gin4 in elm1∆ cells rescues the septin and cytokinesis defects or just the morphology. Is the role of Hsl1 in cytokinesis dependent on its role in the morphogenesis checkpoint? Can the authors comment on how much the defects observed by Hsl1 tethering to the bud neck may be a result of bypassing the morphogenesis checkpoint?:

      We thank the reviewer for this important point. We performed time-lapse imaging of Cdc3-mCherry in strains where Gin4-GFP partially rescued the elongated phenotype of elm1Δ cells (63.75%) when tethered to the bud neck via Hsl1-GBP. Under these conditions, 64.29% of cells showed rescue of Cdc3-mCherry mislocalization, and Gin4 localization itself was restored to the bud neck in 57.85% of cells. We also examined Myo1-ymScarletI dynamics, while 76.92% of untethered elm1Δ cells displayed Myo1 mislocalization to the bud cortex, this was reduced to 11.36% upon Gin4-GFP tethering via Hsl1-GBP. Together, these results indicate that the morphological rescue observed in elm1Δ cells is accompanied by restoration of normal septin organization and AMR dynamics.

      Previous work (Bouquin et al., 2000) [4] showed that Swe1 deletion rescues cell elongation in elm1Δ cells without restoring septin organization. Consistent with this, we found that 60.66% of elm1Δ swe1Δ cells exhibited a round morphology, but tethering of Gin4-GFP to the bud neck via Hsl1-GBP in elm1Δ swe1Δ background did not further enhance morphological rescue. These results suggest that the rescue of cellular morphology observed in our tethering experiments may, atleast in part, depend on Hsl1-mediated regulation of the morphogenesis checkpoint.

      Importantly, despite the lack of additional morphological rescue, a clear restoration of septin localization was observed when Gin4-GFP was tethered to the bud neck via Hsl1-GBP in elm1Δ swe1Δ cells. Overall, these results suggest that while Hsl1-dependent morphogenesis checkpoint regulation may contribute to cell shape rescue, the restoration of septin organization is independent of Hsl1’s function in morphogenesis checkpoint and instead reflects a direct requirement for Gin4 and Hsl1 at the bud neck.

      Please refer to Figures 5, 6, and S6 of the revised manuscript for these additional data.

      (2) As suggested by the authors, the interaction of the Gin4-KA1 domain with the FBAR domain of Hof1 may explain the cytokinesis-related functions of Gin4. As an orthogonal approach, how does KA1 domain deletion of Gin4 affect cytokinesis and Hof1 bud neck localization?

      We thank the reviewer for this suggestion. We first examined the bud neck localization of Gin4-ka1Δ-GFP in comparison with full-length Gin4-GFP. We observed that the localization kinetics of Gin4-ka1Δ-GFP were significantly altered relative to the full-length protein, with reduced recruitment and earlier removal from the bud neck. We also analyzed the localization kinetics of Hof1-mNG in both gin4-ka1Δ and gin4Δ cells. Our results show that Hof1-mNG displays increased residence time and altered accumulation kinetics during cytokinesis in both genetic backgrounds. Thus, loss of the KA1 domain phenocopies loss of the full-length Gin4 and is consistent with disruption of the physical interaction between Gin4 and Hof1.

      Please refer to Figure S4 for these results in the revised manuscript.

      (3) The authors state that "Elm1 and Kcc4 were present at lower abundance at the bud neck (Fig S1A-D) compared to the higher abundance of Gin4 and Hsl1, as observed in their fluorescence intensities (Figures S1B-C) ". This is not evident in the figures. The authors should show a quantification of how they judged abundance at the bud neck:

      We thank the reviewer for this question. Quantification of septin kinase fluorescence intensity at the bud neck was performed using the established protocol for measuring protein accumulation kinetics described by Okada et. al. 2020 [5]. Time-lapse imaging for kinetic analysis of septin-associated kinases during bud emergence shown in Fig. S1A-S1D was carried out using a point-scanning confocal microscope with a 100×oilimmersion objective. Different laser intensities were required because the fluorescence signals of Kcc4 and Elm1 were comparatively weak and not readily detectable above cellular background under the imaging conditions used for Gin4 and Hsl1. The images shown in Fig. S1A-S1D are therefore displayed using differential contrast settings to facilitate visualization. We have now explicitly clarified this in the figure legend.

      A more direct comparison of septin kinase abundance at the bud neck is now provided in Fig. S1E-F, where localization kinetics during the HDR transition/septin remodelling stage were captured using a laser-scanning spinning-disk microscope under similar imaging conditions. We have additionally included raw fluorescence intensity profiles during the HDR transition to better illustrate the relative abundance of these kinases at the bud neck during cytokinesis.

      Please refer to Figures S1E and S1F in the revised manuscript for the updated images and quantitative analyses.

      (4) How did the authors determine G1 and M-phase in the experiments shown in Figures S1A-D? Can the authors mark these phases on the timelapse images? Also, How do the authors explain the different behaviour of Cdc3 in graphs S1A-D among different strains?

      We thank the reviewer for this comment. Cell cycle stages were initially inferred based on bud size, where kinase accumulation at the bud neck correspond to bud emergence (small bud, G1), and kinase disappearance coincided with septin splitting (large bud, M phase). However, we agree that accurate assignment of cell cycle stages would require specific cell cycle markers. To avoid confusion, we have removed the cellcycle-specific stage assignments from the Results section and describe the kinetics relative to t=0 (bud emergence).

      The differential dynamics observed in the Cdc3-mCherry kinetic profiles likely reflect heterogeneity within the the cellular population. To address this, we combined the normalized fluorescence intensity profiles of Cdc3-mCherry from the strains expressing GFP-tagged septin kinases during bud emergence and have included this data as reference (Author response image 1).

      Author response image 1.

      Plot showing spatiotemporal kinetics of Cdc3-mCherry in strains expressing either Elm1-GFP, or Gin4-GFP, or Hsl1-GFP, or Kcc4-GFP.

      (5) Forced tethering based experiments are one of the key sets of experiments for this work, but it is difficult to have a comprehensive understanding of all the data considering how large the data set is and how dispersed it is in the supplemental and main figures (Figures 3-4-5 and Figures S4-S5-S6). It would be helpful to provide a table summarizing the tested forced tethering’s and the phenotypic outcome in the tested yeast strains (Wt/mutant):

      We thank the reviewer for recognising the extensive dataset generated from the artificial tethering experiments and for suggesting the inclusion of a summary table. We have now added Table 1, which summarizes the proteins used in the GFP-GBP artificial tethering experiments, their genetic backgrounds, the total number of cells quantified across three independent replicates, and the phenotypic outcomes associated with bud neck tethering under each condition.

      Please refer to Table 1 and lines 342, 361, 451, 453, 456 and 492 in the revised manuscript.

      (6) In the introduction section, it would help the reader to provide more information on already known molecular roles of septin-associated kinases in septin organization and AMR. Later in the results section (i.e. Figure S1, S2, and S6), the authors extensively explain and show data that independently corroborate some earlier findings, which makes it difficult for the reader to distinguish novel findings from the repeated findings. I suggest shortening the text for the corroborative results, which will help to put more emphasis on their novel findings:

      We thank the reviewer for this suggestion. We have now included the canonical roles of these four septin-associated kinases in the Introduction section of the revised manuscript.

      Please refer to line numbers 79-85.

      We have also revised sections describing corroborative findings and explicitly cited previous studies wherever relevant in the Results section to better distinguish previously established observations from the novel findings presented in this work.

      Other minor comments:

      (1) In Figure 1D-1F, also show the data for hls1∆ and kcc4∆ - which are shown in S3AC in the current version:

      We have now included the Inn1-mNG residence time in hsl1Δ and kcc4Δ cells, alongside elm1Δ and gin4Δ in Fig. 1E.

      Please refer to Fig. 1E in the revised manuscript.

      (2) The authors should be more careful in interpreting their negative Y2H data in Figure 2F.

      We have now explicitly discussed the caveats and inherent limitations of the Yeast Two-Hybrid assay in the manuscript.

      Please refer to line numbers 248-255.

      (3) Please provide quantification for Figure 5A, Figure S6I:

      We have added quantitative analyses showing rescue of Cdc3-mCherry and Myo13xmCherry mislocalization in gin4Δ and gin4Δ hsl1Δ strains upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP.

      Please refer to Figures 8E, 8F, and S10B in the revised manuscript.

      (4) In Figure S6: label is missing "∆" in front of hsl1:

      We thank the reviewer for pointing out this error. We have corrected the labels accordingly.

      (5) As common consensus on yeast gene nomenclature, I suggest the use of "gene∆" instead of "∆gene":

      We thank the reviewer for this suggestion. We have now consistently placed the Δ symbol after deleted gene names throughout the manuscript in accordance with standard yeast nomenclature.

      (6) Lines (535-536): min(distribution) and max(distribution) in the formula is confusing. Clarify it or if possible use "minimum value", "maximum value" instead:

      We thank the reviewer for this suggestion. We have replaced the term “distribution” with “value” in the formula for protein accumulation kinetics analysis.

      Please refer to the updated formula in the Materials and Methods section (Lines 752753).

      (7) In line 86, "Dynamics of Septin-associated kinases and their defects during the cell cycle": Change the title as it is not clear what is meant by "their defects" given the discussed results under this title:

      We thank the reviewer for this suggestion. We have revised the section heading from “and their defects during the cell cycle” to “and defects associated with their deletions”.

      Please refer to line numbers 99-100.

      (8) On the Hof1-mNG image (Fig2C), show the line used for the line scan profile. Additionally, a similar line-scan profile could be useful in Figure S3G:

      As suggested by Reviewer 3, we removed the line-scan analysis from Figure 2 in the revised manuscript because Hof1 ring organization showed substantial heterogeneity across cells, making line-scan analysis difficult to interpret reliably.

      Reviewer #3 (Recommendations for the authors):

      Major points:

      (1) The % phenotype units are terrible. With no explanation, we do not really know whether they represent the percentage of cells that have a particular phenotype, or whether they correspond to a metric that measures some deviation between normal and extreme phenotypes. I would strongly recommend using precise quantitative metrics systematically (i.e. intensities, aspect ratios, division times, etc.) to properly quantify phenotypes:

      We thank the reviewer for suggesting the inclusion of precise quantitative metrics to assess phenotypic differences in the GFP-GBP tethering experiments. In the revised manuscript, we adopted quantification workflows that have been extensively validated and widely used in the literature, including those reported by Marquardt et al., 2024 (Fig. 7B and 7D) [2] from the Bi Lab. In response to the reviewer’s suggestion, we have now incorporated additional quantitative measurements, including cell area and aspect ratio (defined as the ratio of the cell’s major axis to the minor axis), for the experimental datasets presented in the manuscript.

      In the main figures, we now include aspect ratio quantification, while additional parameters are provided for the reviewer’s reference. We also quantified the fluorescence intensity of tethered proteins at the large bud neck and present these data together with the aspect ratio analysis in Figure 4 for elm1Δ cells in which Gin4GFP is artificially tethered to the bud neck via Hsl1-GBP. These quantitative analysis corroborates our qualitative observations and further strengthens our conclusions. Please refer to Figures 4D and 4E as representative examples.

      We have also changed the y-axis labels throughout the revised manuscript. For example, the y-axis in Fig. 7B is now labelled as “Cells exhibiting round morphology (%)”. Please refer to Fig. 4H, 4K, 7C, 8D, S5G, S6C, S7C, S9G and S9J for inclusion of aspect ratio quantification. Statistical analyses for the represented graphs were performed using Kruskal-Wallis nonparametric test, (N=3, n>150 cells/strain) (*: p<0.05, **: p<0.01, ****: p<0.0001, ns: p>0.05).

      Author response image 2.

      (2) Some quantitative analyses were performed manually where simple automated analysis should be performed to provide unbiased, accurate quantification:

      We fully agree with the reviewer that automated image analysis approaches, such as segmentation-based methods, are generally preferred for minimizing bias in morphological quantification. However, elm1Δ and gin4Δ cells exhibit severe phenotypes, including pronounced elongation and clumping, which makes reliable automated segmentation technically challenging for accurate quantification of parameters such as aspect ratio and cell area. For this reason, we used manual annotation for these analyses, as this approach enabled accurate delineation of individual cell and reliable measurements of morphological parameters such as cell area and size across the datasets despite being more time-consuming.

      Please refer lines 769-775 in the Materials and Methods section.

      (3) The "tethering" data also lack clear quantification. The authors should properly quantify the average intensity of Hsl1-GFP at the bud neck in each condition and correlate the results with cell aspect ratios or any other relevant parameters. For example, when comparing elm1null and elm1null Bud4-GBP with elm1null Kcc4-GBP, the visual impression is that as much Hsl1-GFP protein is recruited to the bud neck, whereas the phenotypes are dramatically different:

      We thank the reviewer for pointing this out. We have revised the image representation to facilitate clearer interpretation of the tethering experiments. In addition, we performed the key GFP-GBP tethering experiments using GBP-ymScarletI constructs, allowing direct visualisation of both the GFP-tagged protein and the GBP-tagged partner at the bud neck following tethering. We have included quantitative analyses of cellular morphology, raw fluorescence intensities of GFP-tagged proteins at the large bud neck, and the corresponding aspect ratio measurements for these updated datasets (Author response images 3, 4, 5). We have also included a summary table (Author response table 1) compiling these quantitative results for easier comparision.

      (4) I am very confused by Figure 3 which shows normal localization of Gin4-GFP in elm1null cells and seems to contradict other claims in the manuscript. This is very problematic for the interpretation of most of the "tethering" data:

      We understand the reviewer’s concern and have replaced the representative images of Gin4-GFP in elm1Δ cells in Figure 4B. Although Gin4-GFP is initially recruited to the presumptive bud neck during bud emergence in elm1Δ cells, it subsequently becomes mislocalized to the bud cortex during early cell cycle stages, resulting in reduced bud neck localization. As the cell cycle progresses, the Gin4-GFP signal at the bud neck decreases substantially in elm1Δ cells while remaining stable in wild-type cells until its departure prior to septin HDR remodelling (Fig. S5A-S5D).

      (5) Figure 6 is neither explained in the text nor in its legend. Could the authors explain the model and offer a comprehensive picture of the current knowledge?

      We have simplified the representative model to more clearly distinguish previously established knowledge from the findings presented in this study. Based on our results, We propose that Hsl1 functions both downstream of and in coordination with Elm1 and Gin4 to regulate septin stability and the timely execution of cytokinesis. Deletion of Elm1 disrupts the normal localization and crosstalk between Gin4 and Hsl1 at the bud neck, leading to septin mislocalization and misregulation of AMR dynamics, thereby revealing a previously uncharacterized role for Hsl1 in cytokinesis. The Results section has also been updated to reflect the revised model.

      Please refer to Figure 9 and lines 520-536.

      (6) Could the authors provide information about the double/triple mutant kinase phenotypes to clarify the overlap of functions among them?:

      Barral et al., 1999 [6] reported that individual deletions of Hsl1 and Gin4 result in mild cytokinetic defects, whereas deletion of Kcc4 does not produce any striking phenotype compared to wild-type cells. In contrast, the hsl1Δ gin4Δ kcc4Δ triple mutant remains viable but exhibits severe morphological abnormalities, including branched chains of elongated cells with defective cell separation. These mutants also display aberrant septin organization at the bud neck, characterized by irregular patch-like structures. Analysis of double mutants (hsl1Δ gin4Δ, gin4Δ kcc4Δ, and hsl1Δ kcc4Δ) revealed intermediate phenotypes between the corresponding single and triple mutants, with the hsl1Δ gin4Δ combination showing the strongest defects. Together, these findings suggest that the Nim1-related kinases function redundantly to regulate Swe1 activity and maintain septin architecture at the bud neck.

      Further supporting this model, Bouquin et al., 2000 [4] demonstrated that Elm1 operates independently of the Nim1-related kinases in controlling septin organization. The hsl1Δ gin4Δ kcc4Δ elm1Δ quadruple mutant exhibits severe septin localization defects and strong growth defects, in contrast to the elm1Δ single mutant, which primarily displays septin mislocalization from the bud neck to the bud cortex. These findings indicate that the combined activity of these kinases is essential for proper septin anchorage at the division plane and for assembly of the septin ring.

      Importantly, the progressively stronger phenotypes observed in double, triple and quadruple mutants also suggest that these kinases retain partially specialized functions at the bud neck. Consistent with this framework, our results support a model in which Nim1-related kinases function redundantly to regulate septin architecture and cytokinesis, likely through modulation of the AMR machinery. Because the localization of these kinases appears interdependent, as reported previously (Marquardt et al., 2020; Marquardt et al., 2024) [2,7] and corroborated by our findings, interpretation of mutant phenotypes remains complex and future studies will be required to delineate their individual contributions more precisely.

      Minor points:

      (1) Abbreviations are not defined in the manuscript.

      We have now expanded and defined all abbreviations throughout the manuscript.

      (2) Some of the writing in the figures is too small. Please make sure that a minimal size of letters/numbers is respected:

      We thank the reviewer for raising this issue. We have enlarged the figure labels and axis labels throughout the revised manuscript to improve readability.

      (3) Figure 2D. I am not sure that the line scans bring any useful information as the rings in mutant cells are quite heterogenous. This panel is also not cited in the text. Please make sure that every panel is cited at least once:

      We agree with the reviewer regarding the heterogeneity observed in the Hof1 ring organization in mutant cells and have therefore removed the line-scan analysis from the revised manuscript.

      (4) Knocked-out genes are written incorrectly. Please use the usual yeast nomenclature:

      We thank the reviewer for this suggestion. We have corrected the nomenclature for all deleted genes throughout the manuscript in accordance with standard yeast nomenclature.

      Author response image 3.

      Artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Gin4-GFP via Shs1-GBP-ymScarletI, Hsl1-GBP-ymScarletI, Bud4-GBPymScarletI and Bni5-GBP-ymScarletI in elm1Δ cells. DC*=Differential contrast. Scale bar5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple-comparison test (**: p<0.01, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=397, elm1Δ: n=408, elm1Δ-Shs1-GBPymScarletI: n=467, elm1Δ-Hsl1-GBP-ymScarletI: n=462, elm1Δ-Bud4-GBP-ymScarletI: n=401 and elm1Δ-Bni5-GBP-ymScarletI: n=317 cells). (C) Quantification of aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, n>170 cells/strain). (D) Graph depicting the raw fluorescence intensity of Gin4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=166, elm1Δ: n=177, elm1Δ-Shs1-GBP-YmScarletI: n=176, elm1Δ-Hsl1-GBPymScarletI: n=188, elm1Δ-Bud4-GBP-ymScarletI: n=185 and elm1Δ-Bni5-GBP-ymScarletI: n=151 cells).

      Author response image 4.

      Artificial tethering of Hsl1-GFP to the bud neck via septins or Nim1-related kinases rescues cellular morphology in elm1Δ cells. (A) Representative images showing the relocalization of Hsl1-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI and Gin4GBP-ymScarletI. Scale bar-5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=508, elm1Δ: n=418, elm1ΔShs1-GBP-ymScarletI: n=535 and elm1Δ-Gin4-GBP-ymScarletI: n=482 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (*: p<0.05, ****: p<0.0001), (N=3, n>165 cells/strain). (D) Quantification of raw fluorescence intensity of Hsl1-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=156, elm1Δ: n=166, elm1Δ-Shs1-GBP-ymScarletI: n=169 and elm1Δ-Gin4-GBP-ymScarletI: n=165 cells.

      Author response image 5.

      Targeted localization of Kcc4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Kcc4-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI, Hsl1-GBPymScarletI and Gin4-GBP-ymScarletI. Scale bar-5µm. (B) Quantitative analysis representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), oneway ANOVA Tukey’s multiple-comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=443, elm1Δ: n=489, elm1Δ-Shs1-GBP-ymScarletI: n=312, elm1Δ-Hsl1-GBP-ymScarletI: n=563 and elm1Δ-Gin4-GBP-ymScarletI: n=337 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (****: p<0.0001, ns: p>0.05), (N=3, n>165 cells/strain). (D) Quantification for the raw fluorescence intensity of Kcc4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (**: p<0.01, ***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=163, elm1Δ: n=168 elm1Δ-Shs1-GBP-ymScarletI: n=161, elm1Δ-Hsl1-GBP-ymScarletI: n=172 and elm1Δ-Gin4-GBP-ymScarletI: n=150 cells).

      Author response table 1.

      Summary table showing rescue of elongated morphology in elm1Δ cells upon forced recruitment of Nim1-related kinases via septins or its related kinases tagged with GBPymScarletI.

      Additional changes:

      The graph in Fig. S3D (revised preprint) has been updated to reflect a slight increase in the Chs2-mNG residence time in both the elm1Δ and gin4Δ strains, whereas our previous version indicated a delay only in the gin4Δ strain. Because the elm1Δ strain exhibited a more pronounced phenotype than the gin4Δ strain, we re-examined the analysis. The revised results show that the residence time of Chs2 during cytokinesis is modestly prolonged by approximately 2 minutes in both backgrounds. Accordingly, the graph and statistical analyses have been updated.

      References:

      (1) Mortensen, E.M., McDonald, H., Yates, J., and Kellogg, D.R. (2002). Cell Cycle-dependent Assembly of a Gin4-Septin Complex. Molecular Biology of the Cell 13, 2091-2105. 10.1091/mbc.01-10-0500.

      (2) Marquardt, J., Chen, X., and Bi, E. (2024). Reciprocal regulation by Elm1 and Gin4 controls septin hourglass assembly and remodeling. J Cell Biol 223. 10.1083/jcb.202308143.

      (3) Asano, S., Park, J.E., Yu, L.R., Zhou, M., Sakchaisri, K., Park, C.J., Kang, Y.H., Thorner, J., Veenstra, T.D., and Lee, K.S. (2006). Direct phosphorylation and activation of a Nim1-related kinase Gin4 by Elm1 in budding yeast. J Biol Chem 281, 2709027098. 10.1074/jbc.M601483200.

      (4) Bouquin, N., Barral, Y., Courbeyrette, R., Blondel, M., Snyder, M., and Mann, C. (2000). Regulation of cytokinesis by the Elm1 protein kinase in Saccharomyces cerevisiae. Journal of Cell Science 113, 1435-1445. 10.1242/jcs.113.8.1435.

      (5) Okada, H., MacTaggart, B., and Bi, E. (2021). Analysis of local protein accumulation kinetics by live-cell imaging in yeast systems. STAR Protoc 2, 100733. 10.1016/j.xpro.2021.100733.

      (6) Barral, Y., Parra, M., Bidlingmaier, S., and Snyder, M. (1999 Jan 15). Nim1-related kinases coordinate cell cycle progression with the organization of the peripheral cytoskeleton in yeast. Genes & Development 13. 10.1101/gad.13.2.176.

      (7) Marquardt, J., Yao, L.L., Okada, H., Svitkina, T., and Bi, E. (2020). The LKB1-like Kinase Elm1 Controls Septin Hourglass Assembly and Stability by Regulating Filament Pairing. Curr Biol 30, 2386-2394 e2384. 10.1016/j.cub.2020.04.035.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The extent to which P. falciparum liver stage parasites export proteins into the host cell is unclear. Most blood-stage exported proteins tested in liver stages were not exported. An exception is LISP2, which is exported in P. berghei but not P. falciparum liver stages. While the machinery for export is present in liver stages, efforts to demonstrate export have so far been mostly unsuccessful. Parasite proteins exported during the liver stage could be presented by MHC and thereby become the target of immune control, an incentive to study liver stage export and identify proteins exported during this stage. However, particularly for P. falciparum, it is very difficult to study liver stages.

      This work studies LSA3 in P. falciparum blood and liver stages. The authors show that this protein is exported into the host cell in blood stages, but in liver stages, no or only very little export was detected. A disruption of LSA3 reduced liver stage load in a humanized mouse model, indicating this protein contributes to efficient development of the parasites in the liver.

      The paper also studies the localization of LSA3 in blood stages and uses a known inhibitor to show that it is processed by plasmepsin 5, a protease important for protein trafficking. The work also shows that LSA3 is not needed for passage through the mosquito.

      Strengths:

      The main strength of this work is the use of the humanized mouse model to study liver stages of P. falciparum, which is technically challenging and requires specialized facilities. The biochemical analysis of LSA3 localization and processing by plasmepsin 5 is thorough and mostly overcame adverse issues such as a cross-reactive antibody and the negative influence of the GFP-tag on LSA3 trafficking. The mosquito stage analysis is also notable, as these kinds of studies are difficult with P. falciparum. However, there was no evidence for a function of LSA3 in mosquito stages.

      We thank the reviewer for their perspective on the strengths of the study.

      Weaknesses:

      The cross-reactivity of the antibody, together with the co-infection strategy, prevents reliable assessment of LSA3 localization in liver stages. Despite this, it seems LSA3 is not exported in liver stages, and the paper does not bring us closer to the original goal of finding an exported liver stage protein.

      While the localization analysis in blood stages is well done and thorough, the advance is somewhat limited. LSA3 may be in structures like J dots, but this hypothesis was not tested. Although parasites with a disrupted LSA3 were generated, the function of this protein was not explored. Given that a previous publication found some inhibitory effect of LSA3 antibodies on blood stage growth, a comparison of the growth of the LSA3 disruption clones with the parent would have been very welcome and easy to do. At this point, LSA3 is one more of many proteins exported in blood stages for which the function remains unclear.

      It might be possible to refine some of the conclusions. The impact on liver stage development is interesting, but which phase of the liver stage is affected, and the phenotype remains largely unknown. The co-infection (WT together with LSA3 mutant) has the advantage of a direct comparison of the mutant with the control in the same liver, but complicates phenotypic analysis if the LSA3 antibody is also cross-reactive in liver stages. This issue adds a question mark to the shown localization and precludes phenotypic comparisons. The authors write that they do not know if the cross-reactive protein is expressed at that stage. But this should be immediately evident from the mixed WT/mutant infection. If all cells are positive for LSA3, there is a cross-reaction. If about half of the cells are negative, there isn't. In the latter case, the localization shown in the paper is indeed LSA3, and morphological differences between WT and LSA3 disruption could be assessed without additional experiments.

      We thank the reviewer for their comments. While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Significance:

      The conclusion from the paper that "our study presents just the second PEXEL protein so far identified as important for normal P. falciparum liver-stage development and confirms the hypothesized potential of exported proteins as malaria vaccine candidates" is partially misleading. Neither LISP2 nor LSA3 seems to be exported in P. falciparum liver stages, and we can't confirm the potential of vaccines with proteins exported in this stage. LSA3 is still important and may still be the target of the immune response, but based on this work, probably not due to export in liver stages.

      We thank the reviewer for the comment. We would like to emphasize the possibility that proteins localized at the PVM may be considered exported ‘if’ part or all of the protein (eg, a domain) faces the host cell lumen from the hepatocyte. We have not shown this to be the case for LSA3 or LISP2 but that possibility remains open. Nonetheless, LISP2 is exported (by P. berghei liver stages) and LSA3 is exported (by P. falciparum blood stages); both are exported proteins.

      Reviewer #2 (Public review):

      Summary:

      Immunogenic Plasmodium falciparum proteins that could be targeted to prevent parasite development in the liver are of significant interest for novel anti-malarial vaccine development. In this study, McConville et al evaluate the trafficking and functional importance of LSA3, a protein expressed in the blood and liver stages and previously shown to provide protection in immunized chimpanzees. LSA3 contains a PEXEL motif, but the authors have previously shown that this protein does not appear to be exported beyond the PVM in the liver stage (McConville et al, PNAS 2024). However, LSA3 trafficking and functional importance have not been comprehensively evaluated across stages. In the present study, the authors find that blood stage LSA3 undergoes PEXEL processing, and a portion of the protein is exported into the erythrocyte, where it localizes to punctate structures distinct from Maurer's clefts. Using a knockout mutant, LSA3 is shown to be dispensable for blood and mosquito stages but important to liver-stage development. Collectively, these results validate LSA3 as a liver-stage target and place it among several other PEXEL proteins that display differential trafficking beyond the PVM in the erythrocyte but not the hepatocyte.

      Strengths:

      The authors present a thorough analysis of LSA3 trafficking in the blood stage. PEXEL processing by Plasmepsin 5 is clearly demonstrated through a combination of mini LSA3-GFP reporters and Plasmepsin 5 inhibitors. Importantly, an LSA3 knockout mutant is used to show that the LSA3-C anti-sera also react with additional, unidentified parasite proteins in the blood stage. Nonetheless, comparison between the WT and KO parasites clearly indicates that a portion of LSA3 is exported into the erythrocyte, which is further supported by protease-protection assays with fractionated iRBCs. This contrasts with the liver stage, where LSA3 does not appear to traffic beyond the PVM, similar to what has been observed for other PEXEL proteins in the rodent malaria model.

      This study provides the first direct analysis of LSA3 function by reverse genetics, showing this protein is important for liver stage development in chimeric human liver mice. Several PEXEL proteins in P. berghei have been shown to be exported into the host cell in the blood stage, but do not appear to cross the PVM in the liver stage. These observations reinforce that even without detectable export into the hepatocyte, PEXEL proteins play critical roles during liver stage development.

      We thank the reviewer for their feedback regarding the strengths of the paper. 

      Weaknesses:

      A previous study reported that anti-LSA3 antibodies inhibit blood-stage growth, suggesting a role for LSA3 during erythrocyte infection. While the authors carefully evaluate the LSA3 mutant in mosquito and liver stages, the impact on blood stage fitness is not tested. While the knockout shows LSA3 is not essential in the blood stage, its importance during erythrocyte infection remains unclear.

      The authors previously reported that anti-LSA3-C signal in the liver stage localizes within the parasite and at the parasite periphery but is not exported into the hepatocyte. In the present study, it is shown that anti-LSA3-C reacts with other parasite proteins beyond LSA3 in the blood stage, and this may also occur in the liver stage. However, since liver-stage IFAs were only performed on samples co-infected with both WT and ∆LSA3 parasites, non-specific anti-LSA3C reactivity at this stage could not be determined, and the localization of LSA3 in the liver stage remains somewhat unclear.

      We thank the reviewer for their comments. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Reviewer #3 (Public review):

      Summary:

      This manuscript provides a comprehensive characterization of the Plasmodium falciparum protein LSA3, combining biochemical, genetic, and in vivo approaches. The authors convincingly demonstrate that LSA3 is expressed during liver stage infection and that disruption of the gene leads to a modest but reproducible reduction in liver stage parasite load in humanized mice.

      Strengths:

      Their biochemical and cell biological analysis of blood stages provides strong evidence that LSA3 is exported to the infected erythrocyte, and the detailed analysis of its PEXEL motif processing is well executed.

      We thank the reviewer for their comments.

      Weaknesses:

      The study suggests LSA3 as one of only two known P. falciparum PEXEL proteins contributing to this stage, although there is no evidence for the export beyond the vacuolar membrane. Several key conclusions, particularly regarding antibody specificity, localization in liver stage parasites, and the interpretation of the phenotypic data, are not fully supported by the current experiments.

      We understand the reviewer’s points. We agree that there is no evidence provided that LSA3 is targeted beyond the PVM; whether any of the protein faces the hepatocyte cytosol is unknown (and challenging to conduct) but this possibility remains plausible. LISP2- and LSA3deficient liver stages are less fit than parental controls and thus we stand by the conclusion that they are the two so far identified P. falciparum PEXEL proteins that are important for liver-stage development.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 163 says: "Altogether, this demonstrates that LSA3 is important but not critical for blood stage growth of P. falciparum": this is based on the cited Morita et al., 2017. However, previously LSA3 was considered dispensable based on a knock out in 3D7 (Maier et al., 2008; PMID: 18614010). Given that the authors generated a mutant for this work, it would be straightforward to test growth and clarify the importance of LSA3 in blood stages. If important, the analysis of the location and transport of LSA3 in blood stages would immediately become more relevant.  Maybe the data for this is already in the paper: the number of stage V gams was similar between mutant and control (Figure 4A). If this was calculated from the total number of asexual starting parasitemia, it includes blood stage growth, and it can be assumed that there is no growth defect in the mutant in the blood stages. If the number of stage 5 gams was calculated from the number of committed schizonts/rings, nothing can be said about blood stage growth, and asexual blood stage growth should be tested in specific experiments.

      We thank the reviewer for raising the function of LSA3 in blood stages and agree it was an obvious omission, though for good reason - a separate, collaborative study was underway. While this eLife preprint was in revision, our accompanying manuscript on the blood stage was published, showing the characterization of our NF54 DLSA3 mutant during blood-stage growth (PMID:41135800). The findings are now summarized and the citation included in the revised version of this preprint.

      Manuscript line 105: "although, notably, functional characterization of lsa3 deletion mutants has not yet been reported to confirm an important function": at least in blood stages, it was reported to be dispensable, see above. The corresponding study (Maier et al., 2008, PMID: 18614010) could be cited in that context. 

      The citation of PMID18614010 and 39913589 have now been added and we thank the reviewer.

      (2) Some questions central to the conclusions of this paper remain because it was unclear whether the serum did indeed detect LSA3 in the liver or not. It would be easy to check if all cells from the WT/Mutant mix experiment show LSA3 signal (this would mean it cross-reacts) or if only about half are positive (the mutants would be negative if there is no cross-reaction). This would be important to mention for Figure 5 because, at present, it is not known that what is labeled by the LSA3-C antibody in these images is (only) LSA3. 

      We thank the reviewer for this point and completely understand. We did check this via microscopy of liver sections co-infected with LSA3 mutant and control liver-stage parasites as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5-day post-infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (3) It is also unclear which parasites were imaged in Figure 5. The text of the results states that NF54 liver stages were used, but later: "As we employed a co-infection strategy to assess the essentiality of LSA3 versus NF54 in mice, we could not perform IFAs on individually infected mice in this study to validate the specificity of LSA3-C at the liver-stage". The legend says NF54 sporozoites on day 5 post-infection were used. I suspect it was a WT/mutant mix, in which case the above applies, and in the absence of cross-reactivity, half of the cells should be LSA3-C negative. If this is not the case, the localization in the liver becomes dubious.

      We apologize for the confusion and have corrected this. In Figure 5, we utilized liver sections from NF54-infected humanized mice that were stored at -80 C from a previously published study (McConville et al, PNAS 2024). Ideally, we would validate the specificity of LSA3 antibodies at the liver-stage using liver sections containing only DLSA3 parasites however the number of mice available was limited and the samples available to us also contained the Control line for qRT-PCR analyses (the co-infection strategy). As mentioned above, we couldn’t distinguish between these two strains by IFA at the time point analysed and this precluded us unequivocally validating the LSA3-C specificity in the liver-stage; however it cannot be excluded that the signal observed at the PVM is indeed LSA3. We are currently focusing research efforts on obtaining more humanised mice to answer this.

      Minor:

      (1) Introduction: Before the part on the PEXEL motifs, there are almost no references; please add references for all statements.

      We have added references.

      (2) Figure 1B is unclear regarding which part of the gene was deleted. The system used would permit a complete gene deletion, but the homology flanks seem to be within LSA3. If parts of the gene are left, the 75 kDa on the western blots might be a degradation product arising from both the truncated and the full-length protein. Please clarify in the sketch exactly where the homology flanks are, with respect to the start and stop of the gene. 

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (3) Line 161: Replace was with were.

      Corrected.

      (4) Figure 2, 224: Why do the authors think LSA3 must be in the luminal leaflet of the PVM as opposed to the outer leaflet of the plasma membrane?

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte.

      (5) Line 245: GFP core "derived from digestion of the reporter in the food vacuole, which confirmed it was secreted from the parasite". I wonder if the amount of GFP "core" really can be used as evidence for secretion, and its amount can be compared between experiments. Did the author quantify this for the full-length protein to get a proportion per sample?

      Use of GFP core to measure defects in P. falciparum GFP reporter secretion has been described previously (for example PMID:23387285 and 35906227). The comparison the reviewer asked for is an interesting and important question: however the control would be to compare the ratio of GFP core to uncleaved in the control lanes as well, which is not possible to do since the full-length protein is digested by plasmepsin V in the native PEXEL versions of the experiments (mLSA3-GFP in the first blot, Vehicle in the second blot) leaving no full-length protein to compare to. It stands to reason that inhibition of N-terminal processing results in less protein removal from the membrane (ER or COPII vesicle or PM) resulting in less secretion out of the parasite for retrograde transport to the food vacuole with cytostomal vacuoles (analogous to plasmepsin II). In the food vacuole, the chimeras are in normal cases digested by proteases back to the GFP core that is resistant to cleavage and evident as GFP core on the immunoblots (PMID:10775264 and 14709539 and 19055692 and 20130643). 

      (6) Figure 3 has the word plasmid in two lanes. In Figure 3E, amend the labelling of the blots.

      We apologize for the formatting error in converting the figures to PDF during the original submission and thank the reviewer for the suggestion. This has now been corrected.

      (7) Lines 266/271/284: "live IFAs", live immunofluorescence assay. Does this mean an antibody was given to living   parasites?

      The correct term is live microscopy and this has been corrected.

      (8) Does Figure 6A fit with the data in Figure 6B? It seems 6B has a milder phenotype than 6A.

      We thank the reviewer for the question. Yes the data directly correspond to each other and are represented in two ways: Panel A shows the qRT-PCR raw data for liver load of each parasite strain per humanized mouse using a scientific scale on the y-axis. Panel B shows that magnitude of the DLSA3 defect as a percentage of the total liver load per mouse:

      % total parasite liver load  = ( strain 1 or strain 2 liver load ) x100

      sum of strain 1 + strain 2 liver loads

      The intent of showing both data is to convey the correct magnitude of the difference in two ways to assist the reader in understanding the true defect, both are accurate and both are statistically significant. In revision we detected mislabelling of humanized mouse 2 and 3 in the original graphs that has now been corrected and we sincerely thank the reviewer for helping us identify this error.

      (9) Line 482: Please add references for this debate. 

      These have been added.

      Reviewer #2 (Recommendations for the authors):

      Major Comments: 

      (1) In general, the authors have taken care not to overstate conclusions from their study. Nonetheless, while not technically inaccurate, the title might misleadingly suggest LSA3 is exported in the liver stage (this was my initial impression on reading it until I looked at the data). I suggest the authors revise the title to avoid confusion by clarifying that export was only observed in the blood stage.

      We sincerely appreciate the reviewer’s point. As this article was posted as a preprint that has now been cited several times, we have carefully weighed the comment and in the end decided to retain the current title for the above reason.

      (2) While the ability to generate the ∆LSA3 parasites clearly shows that the protein is not essential in the blood stage, the impact on parasite fitness is never tested but simply assumed (for instance, in lines 163-164: "...this demonstrates that LSA3 is important...for blood-stage growth..."). Do the ∆LSA3 parasites have a fitness defect in the blood stage consistent with the previous GIA data that would support this claim? Since the rabbit anti-LSA3-C antibodies produced by Morita et al did not have GIA activity against the blood stage, it is possible that the GIA observed with the human and mouse antibodies might have been due to reactivity with a different protein. If ∆LSA3 does cause a fitness defect, it would be interesting to know if the endogenous GFP-tagged line, which alters protein trafficking/membrane association, also produces this effect.

      We agree with the reviewer and would like to clarify that this omission was not intended to create confusion but was by design, due to a separate collaborative study that was underway to address such questions. While this eLife preprint was in revision, our accompanying manuscript on characterising NF54 DLSA3 at the blood stage was published (PMID:41135800). The findings are now summarized and the citation included in the revised version of this eLife preprint. In sum, LSA3 is not critical for erythrocyte invasion but its deletion perturbs the rate and efficiency of merozoite invasion, at the step(s) of resealing of the PVM/host cell, resulting in aberrant accole forms that protrude from the infected erythrocyte.

      (2) Figure 1D: While the images are compelling and I don't doubt the claim that LSA3 is exported in the blood stage (also supported by the fractionation/Pk experiments), the authors should provide quantification of the difference in exported signal between the WT and ∆LSA3 parasites in these IFAs to rigorously support this conclusion. Also, please include details about how many independent experiments are represented by the microscopy data throughout the manuscript (Figures 1, 2, 3, and 5).

      We understand the reviewer’s request and wish to indicate that the export signal was absent in all cells infected with DLSA3 that was imaged. The microscopy performed was from n=2-3 experiments except for Figure 5 which was from n=1 humanized mouse per time point in which multiple EEFs from the liver were imaged. This has been indicated in the figure legends. 

      (3) Careful inspection of the z-series images in Figure 5A shows that most of the LSA3-C signal seen outside the PVM (beyond the boundary delineated by EXP1) is closely associated with DAPI puncta, suggesting these are merozoites. Together with the prominent gap in the EXP1 signal, this suggests the schizont has already ruptured. Thus, anti-LSA3-C signal beyond the PV seems best explained as coming from merozoites or other material released by PV rupture, not from export across the PVM, and this should be added to the text in place of comments about localization to PV extensions or potential export (lines 358-359, 422-423).

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have added the comment as requested.

      Minor Comments:

      (1) The authors may want to denote the disordered repeat region in the LSA3 schematic in Figure 1A that is mentioned in the text.

      We have added the residue boundaries of the predicted domain from AlphaFold into both the schematic and the text and included a link to the LSA3 pages in PlasmoDB and

      AlphaFold in the Methods section.

      (2) The authors use rabbit anti-LSA3-C antibodies previously generated by Morita et al. These polyclonal antibodies were raised against a recombinant C-terminal region of LSA3 (residues 750-1433), but the schematic in Figure 1A indicates the antibodies recognize a smaller region between residues 1154-1433. Please adjust the figure accordingly, or if this is not the same antiLSA3-C antibody reported by Morita, please provide details about its production.

      The figure is corrected.

      (3) The authors use Alphafold to identify a region of LSA3 with similarity to the substrate binding domain of DnaK, but the data is not shown. Please include the Alphafold prediction in supplementary figures and provide information about how the predicted structural homology was determined.

      We have added a link to the AlphaFold page for PF3D7_0220000 in the methods.

      (4) The schematic in Figure 1B indicates that the DHFR cassette was inserted at an internal site within the lsa3 gene. If this is the case, it seems possible that an N-terminal portion of the protein is still expressed, but I was unable to find details about the boundaries of the homology flanks to determine the precise insertion site. Please clarify the knockout strategy and indicate the specific insertion site.

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (5) Line 162: I think this should read "antibodies that react with LSA3 were...".

      Corrected.

      (6) Figure 1D: The merge with the transmitted light channel is missing for the third panel in the ∆LSA3 IFAs. Also, please define the scale bar length in the legend.

      Corrected.

      (7) Lines 744-746: The IFA fixation panel order description (top, bottom) in the Figure 2A legend is reversed from what is shown in the actual figure. Also, please define the scale bar length. 

      Corrected.

      (8) Lines 184-186: Since the fractionation/PK protection assays suggest most of LSA3 is in the PV, it would be interesting to know if the strong peripheral/PV signal observed in the PFA-fixed IFAs in Figure 2A is also present in the ∆LSA3 parasites, or is this non-specific? 

      Thank you for the suggestion. We agree this would be an interesting result to know but do not have the capacity at the present time.

      (9) Lines 219-225: It is unclear to me why these results are interpreted to suggest that the majority of LSA3 is peripherally associated with the luminal leaflet of the PVM. Wouldn't an integral membrane configuration in the PVM (with the C-terminus facing the host cytosol) or PPM (with the C-terminus facing the parasite cytosol) also account for the data? Adding a carbonate extraction would help clarify this point.

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM-associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte. If the question is whether LSA3 is an integral PVM protein, we agree that use of carbonate in the future would answer that question.

      (10) Figures 3D and E: There are some problems with some of the text wrapping in these panels.

      We apologise, this was a formatting issue as the manuscript was converted to PDF.

      We have corrected this error.

      (11) Line 422-423: In fact, the Z-sections shown in Figure 5 appear to indicate that the LSA3-C signal is predominantly located within the parasite, not at the PVM.

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have corrected the final conclusion to be more accommodating of this.

      (12) Lines 468-470: Since cross reactivity of anti-LSA3-C is substantial in the blood stage but was not defined in the liver stage by analysis of unmixed infections, how do the authors know that they were not observing ∆LSA3 parasites in their IFAs? I think what they mean here is that parasites lacking anti-LSA3-C reactivity were not observed, which is an important distinction.

      The reviewer is correct and this has been corrected.

      (13) Lines 478-479: The authors should also mention that the P. berghei PEXEL proteins evaluated in Fougere et al are exported in the blood stage, similar to LSA3. Moreover, other studies have shown something similar for additional endogenous PEXEL proteins or reporters in P. berghei (PMIDs 22329949, 26347246, 34956312).

      We have added the additional text regarding export into the infected erythrocyte and the reference to IBIS1.

      (14) Line 491: The data here don't support that LSA3 is "required" for liver stage development, only that it is important to it. Since the authors have not defined the cross-reactivity of anti-LSA3C in unmixed infections, it is not clear that ∆LSA3 parasites are arrested early in the liver stage, only that they show a reduced number of genome copies relative to the parental control. 

      We have amended the sentence to “required for normal liver stage development”.

      (15) Line 530: I think NGF54 should be NF54.

      Corrected.

      Reviewer #3 (Recommendations for the authors):

      (1) Antibody specificity in liver stage IFA experiments:

      The specificity of the anti-LSA3 antiserum (LSA3-C) used in liver stage IFA is not fully convincing. While the KO parasites were used effectively to validate specificity in blood stages, the same is not true for liver stages. 

      (a) It is essential to repeat IFA with ΔLSA3 parasites in liver stage infections to determine whether the observed PVM staining is truly specific.

      We appreciate the reviewer’s point, however at a cost of over $5000 per humanized mouse, we do not have the capacity to conduct this experiment at the present time. We highlight that, as the blood stage IFAs confirmed the specificity of LSA3-C for LSA3, the possibility remains open that LSA3 is specifically recognized at the PVM.

      (b) If the antibody is the same polyclonal serum used in Morita et al. (2017), why did the authors not employ a monoclonal antibody, which they presumably have access to and which would provide greater specificity? 

      We have included new data confirming that LSA3 is exported using LSA3-T, in addition to LSA3-C.

      (c) Given that rabbit antisera often show non-specific staining at the PVM in liver stage parasites, co-localization with PVM markers is not sufficient. Inclusion of the ΔLSA3 parasites in liver stage IFA is critical. It will also show whether there is any cross-reaction of the antiserum in liver stage parasites, as seen by IFA for blood stage parasites. 

      We thank the reviewer for their feedback.

      (d) To validate the serum further, the authors should infect HC-04 cells in vitro with GFP-LSA3 parasites and stain with LSA3-C to confirm overlap between the tagged protein and the antibody signal.

      We thank the reviewer for their feedback.

      (e) For higher-resolution co-localization, expansion microscopy - now commonly used even in malaria research - would substantially improve the analysis. 

      We thank the reviewer for their feedback.

      (2) The localization of LSA3 in this study differs notably from Morita et al. 2017, who reported localization to dense granules in merozoites and staining in ring-stage parasites at the PVM. 

      (a) The authors confirm DG localization, but they do not examine ring-stage parasites. They should include the IFA of ring stages to clarify whether they can replicate the previous findings.

      We thank the reviewer for their feedback.

      (b) Additionally, the differences in Western blot banding patterns between the two studies should be addressed. Do the authors have an explanation for these discrepancies? 

      We thank the reviewer for their feedback.

      (3) The authors report a ~40% reduction in liver parasite load using qPCR, which is statistically significant. However, this phenotype is modest and should not be interpreted as showing that LSA3 is essential.

      (a) Please avoid terms like "required" or "essential" and instead describe the protein as "contributing to normal development" or "influencing fitness."

      We have used the term “required for normal liver stage development”.

      (b) Since the authors generated liver sections, they should take advantage of these to quantify the number and size of liver stage parasites, which would help determine whether the phenotype reflects fewer infected cells or reduced parasite growth.

      We did check this via microscopy of liver sections, but all mice were co-infected with LSA3 mutant and control liver-stage parasites, as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5 day post infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (c) It would also be valuable to include IFA from singly infected ΔLSA3 livers (rather than co-infected), and possibly at earlier timepoints, to identify the developmental window affected.

      We agree it would be valuable.

      (4) The manuscript suggests that LSA3 may be exported beyond the PVM into the hepatocyte, based on a small number of peripheral puncta.

      (a) This claim is not convincingly supported by the data. The punctate signals shown in Figure 5 are weak and may rather reflect PVM extensions or TVN. In fact, one punctum even overlaps with the DAPI signal (figure 5, middle panel), which raises further doubt about the localization.

      We appreciate the reviewer’s careful eye and caution and have added the comment regarding DAPI.

      (b) Given the lack of KO controls in these liver stage IFAs, the authors should not describe LSA3 as "exported beyond the PVM". The language should be revised to reflect that the protein localizes predominantly to the PVM, and any extra-PVM signal remains unconfirmed and could be non-specific. 

      (c) This is especially important given the well-known tendency of rabbit antisera to produce background PVM staining in liver stage parasites. 

      Corrected.

      (e) In an earlier report (McConville et al, 2024, PNAS), they clearly state that LSA3 is NOT exported beyond the PVM. Actually, the staining in the previous report looks quite different from the images provided for Figure 5. The authors might wish to comment on this. 

      We thank the reviewer for their feedback.

      Minor comments:

      In some sections, the manuscript uses "exported" to refer to trafficking to the PVM. This terminology should be used more carefully and consistently, since "export" often implies translocation into the host cytosol

      We understand that export involves a protein localizing within the host cell and so protrusion through the PVM may also be considered exported, however, we have not confirmed this for LSA3 in liver stages.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We greatly appreciate the reviewers for their efforts in reviewing our manuscript. We highlight that the key contributions of our paper are to provide a framework for calibration and validation of high-fidelity cardiac electromechanical models based on a diverse compilation of clinical datasets, and that we provide one example of such an evaluation of our own baseline electromechanical model. The comments raised by the reviewers were chiefly focused on the second goal, which is our specific model and the outcomes of the evaluation process, rather than on the evaluation framework itself. As such, we have made improvements to our model implementation and to provide additional confidence in our specific modelling framework through this review process. Specifically, we have strengthened the verification component of this evaluation, provided additional quantitative measures, and included a more in-depth discussion of the remaining limitations in our modelling framework. We hope that the updated version of the manuscript and our efforts to improve it are well-received by our reviewers and editors, as well as by the modelling and simulation community at large.

      eLife Assessment

      This is a potentially important study that explores the relevant range of parameter values for calibration and validation of cardiac electromechanics in ventricular models. Although much of the work presented is solid, the evidence provided to support the authors' key scientific claims is incomplete, especially as it relates to the emphasis on standardized validation and verification approaches. Notably, the level of model personalization presented in this work falls short of the threshold for what could reasonably be called a "digital twin", even by the relatively relaxed standards that have emerged in computational physiology and related fields in recent years.

      We appreciate the eLife assessment for identifying the potential importance of our study. Regarding the threshold for 'digital twin', we note that a cardiac digital twin is envisioned as a patient-specific computational model of the heart, personalised from multi-modal clinical data and continuously updated to support diagnosis, prognosis, and treatment planning, which is a goal that to our knowledge no published electromechanical study has simultaneously fulfilled. It is for this reason that the community refers to the 'digital twin vision' rather than its realisation, and we adopt this framing consistently throughout the manuscript.

      The primary contribution of this manuscript is the framework: a systematic application of ASME V&V40 standards to a fully coupled electromechanical model, spanning electrical, mechanical, and haemodynamic biomarkers within a single study. In our revision, we have clarified that the model evaluation presented here is an example application of that framework, which was designed not to certify a model as complete, but to provide a transparent audit of current capability that identifies where confidence is established and where further development is needed. We have updated the title and language throughout the manuscript to reflect this framing consistently.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study by Wang et al. investigates cardiac electromechanical modeling and simulation techniques, focusing on the calibration and validation of ventricular models according to ASME V&V40 standards. The researchers aim to calibrate model parameters to align with key biomarkers such as QRS duration and left ventricular ejection fraction, and validate the model against independent measurements such as displacement and strain metrics. The authors also examine the impact of parameter variations on deformation, ejection fraction, strains, and other biomarkers. The overarching aim of the study is to give "credibility to the underlying computational electromechanics framework" and to "pave the way towards credible cardiac electromechanical Digital Twins."

      Strengths:

      (1) The study presents a solid validation strategy for cardiac models based on independent data.

      (2) It integrates electrophysiological, mechanical, and hemodynamic biomarkers for sensitivity analysis and calibration.

      Weaknesses and Limitations:

      (1) Model Assumptions: The study employs simplified modeling assumptions that are not state-of-the-art, e.g.,

      (a) Isotropic scaling of the mesh to generate an unloaded reference geometry.

      (b) Simple afterload and preload models that fail to produce physiological results.

      (c) Simplified epicardial boundary conditions.

      While our model was able to broadly achieve physiological behaviour based on the calibration and validation datasets, it also contains several simplifications that can be expanded with more sophisticated techniques to allow explorations in specific areas. We have added a dedicated Limitations subsection to the Discussion section of the manuscript to address these and to provide references to relevant studies.

      (2) Numerical Framework:

      (a) The mesh resolution and/or the numerical framework used for the mechanical part appears to suffer from known numerical artifacts (locking effects), leading to overly stiff or inaccurate behavior in finite element analysis. This results in an artificially stiff response to deformation, which is compensated by setting active contraction to ten times the value reported in the literature. The authors attribute this to limitations in using ex vivo tissue measurements to represent in vivo function, although similar issues were not observed in previous works.

      We thank the reviewer for raising this point and have investigated it carefully. We have added a verification section as well as discussions to the manuscript to more comprehensively discuss this point. In short, through various tests against benchmark (Land) and comparing stress-strain curves in cube simulations with the same mesh resolution, we could not identify evidence of volumetric locking effects. The elevation in contractile force was also necessary in a simplified ellipsoid version of the model in a previous publication [ref 7, Levrero-Florencio, et al. 2020]. We note that these benchmarks were performed in the incompressible transversely isotropic regime; whether analogous locking effects exist in the dynamic orthotropic active contraction framework used in the full biventricular simulations remains an open question, which we have identified as a priority for future benchmarking, for example against the Arostica et al. 2025 benchmark.

      We agree that the explanation of this as ex vivo vs in vivo difference in contractile force is too simple, and other contributing factors are better understood through comparison with similar studies in the field. Strocchi et al. (2023) used a four-chamber model with explicit atrial mechanics, and in her history matching varied Tref within +- 33-55% of a reference value of 120-150 kPa, targeting a peak active tension of 160 +- 15 kPa, which was a considerably more modest adjustment than applied here, likely reflecting differences in model geometry, pericardial constraint, and circulatory model between the two studies. Gerach et al. (2021, Mathematics) applied manual parameter adjustments informed by in vivo active tension measurements of 120 – 150 kPa and achieved ejection fractions of approximately 63%; however, they reported that systolic pressures in both ventricles were too high for a healthy heart, and similarly reported elevated peak ejection rates compared to MRI measurements, a difficulty we also encountered. Notably, Gerach et al., report that atrial contraction contributes approximately 11-13% of end-diastolic volume, which in a biventricular-only model would directly reduce the achievable LVEF and necessitate compensating adjustments to active tension. Zingaro et al. (2024, Journal of Computational Physics), using an alternative active tension model (RDQ20), similarly found it necessary to increase contractility parameter (a_XB) to achieve sufficient ejection, and explicitly report that no single parameter configuration simultaneously achieved physiological peak ejection rate and LVEF, a fundamental tension we also encountered. Together, these comparisons suggest that the elevated Tref in our model most likely reflects a combination of the absence of atrial filling, simplifications in pericardial constraint, and the lack of poroelastic behaviour, rather than volumetric locking alone. Additional investigations are needed as explained in the manuscript, and these factors are identified as open priorities for future development within our framework.

      We added a comparison with these three studies to the discussion section of the manuscript and have updated the limitations section to reflect this more nuanced account of the factors contributing to the elevated active tension scaling. Further work will be required to address these points.

      (b) Further, the authors employ the monodomain model for the simulation of the electrical excitation and relaxation on a relatively coarse grid with an approximate edge length of 1mm. This resolution is known to be insufficient for reliable results in organ-scale electrophysiology modeling.

      Our ECG simulations are robust against coarse mesh resolutions since we use an Eikonal solution to prescribe the activation times on the endocardial surface and we tune the diffusivity parameters in the model such that the correct conduction velocities are reached, as performed in Camps et al. (2024) (ref 33) using the tuneCV tool in monoAlg3D (https://github.com/rsachetto/MonoAlg3D_C/tree/master/scripts/tuneCV), which is similar to the tool in openCARP (described here: https://opencarp.org/documentation/examples/02_ep_tissue/03a_study_prep_tunecv).

      While this mesh resolution may not be sufficient for simulations of more complex behaviour, such as re-entry and fibrillation patterns, it is sufficient for simulations of ECGs in this study. We have noted this in the methods section.

      (3) Geometrical model and digital twin: The geometrical model, taken from a public cohort and calibrated to an ECG of another individual along with population-averaged values from a databank (UK Biobank), and unrelated measurements from surgical procedures, can hardly be considered a digital twin. Further, validation of the model was then performed against data from yet another cohort.

      We thank the reviewer for this point and welcome the opportunity to clarify our dataset choices. The use of multiple data sources was a deliberate methodological decision. While an ideal dataset for electromechanical model evaluation would combine full biventricular geometry, 12-lead ECG, invasive pressure measurements, and myocardial strain data from a single individual, no such dataset currently exists in the public domain, and acquiring it routinely would be impractical in clinical settings. Multi-source integration therefore reflects the realistic deployment scenario for future clinical translation of these tools.

      The specific choice of geometry was principled: the mesh associated with the ECG dataset that was available to us was truncated at the base due to the clinical acquisition protocol, which would have prevented physiologically realistic basal boundary conditions. The female Rodero geometry we chose provides full ventricular coverage and was selected on that basis.

      Demonstrating that a coherent, systematically evaluated framework can be constructed from compiled multi-modal data is itself a contribution because it makes the tools accessible to the wider community without requiring a single ideally acquired dataset.

      (4) Calibration procedure: There are apparent flaws in the calibration procedure, or it is not described in sufficient detail. The authors dedicate significant effort to motivating parameter ranges, but in the end they use mostly other parameters for the calibration process, aiming to maximize left ventricular ejection fraction. It is not clear whether the chosen parameters result in, e.g., physiological calcium traces or calibrated parameters that are within physiological ranges.

      Thank you for raising this point, which we have now clarified in the manuscript. The parameters that were chosen for the calibration process were based on the results of the sensitivity analyses.

      In addition, we have supplemented results Figure 1 with a subfigure F showing that the calcium transient and action potential durations fall within physiological ranges after calibration.

      (5) Goodness of fits, e.g., a direct comparison of the measured and the simulated ECG, are not provided to assess calibration quality.

      The calibrated model achieves QRS duration of 89 ms and QT interval of 360 ms, both of which fall within the healthy reference ranges compiled in Table 2, providing a biomarker-level assessment of calibration quality (Figure 1A). A full quantitative goodness of fit analysis of the simulated ECG morphology was performed following the methodology of Camps et al. [52], in which the same beat-averaged ECG was processed; we direct the reader to that work for full details rather than reproducing the analysis here.

      (6) Due to these limitations and weaknesses, the authors fall short of achieving some of their goals, particularly establishing credibility for the underlying computational framework and in reproducing healthy pressure-volume loops, and in achieving physiological simulations while using physiological or reported ranges for the calibrated parameters.

      For example, a key physiological requirement is that the right and left ventricular stroke volumes are approximately equal in a heart beating at a limit cycle, as the blood pumped by the right ventricle into the pulmonary circulation must match the amount pumped by the left ventricle into the systemic circulation. This balance is not achieved in this study.

      We thank the reviewer for identifying the stroke volume imbalance. We acknowledge the physiological requirement that, in a steady-state limit cycle, the right ventricular stroke volume must approximately equal the left ventricular stroke volume. However, since our model does not explicitly prescribe volumes, to achieve this, we would need to either explicitly tune active tension for the left and right ventricles separately, such as done in https://www.frontiersin.org/journals/physiology/articles/10.3389/fphys.2021.716597/full or develop a more sophisticated circulatory model and employ a multistep procedure that sequentially tunes circulatory dynamics, passive mechanics, and active contraction, such as done in https://www.biorxiv.org/content/10.64898/2025.12.11.693778v1.full. Both of which are beyond the scope of this paper.

      We note that, despite the absence of explicit RV calibration, the RV volumetric measures and pressures remain within physiological ranges, suggesting that the coupled biventricular mechanics are broadly plausible. As such, we have noted this limitation in our discussion section, and sign-posted to other studies where the stroke volume match is achieved.

      (7) The conclusive claim that "the study paves the way towards credible electromechanical cardiac Digital Twins" is not supported. The model exhibits non-physiological behavior, requires unsupported parameter alterations (such as a 10-fold active stress scaling), and does not represent a digital twin, as model data are drawn from various unrelated, non-patient-specific sources.

      We thank the reviewer for this comment, which gives us the opportunity to clarify our use of the term 'digital twin'. A cardiac digital twin is envisioned as a patient-specific computational model of the heart, personalised from multi-modal clinical data and continuously updated to support diagnosis, prognosis, and treatment planning. This is a transformative goal for precision cardiology that the field is actively working towards, with credible, systematically validated electromechanical models as its essential foundation. To our knowledge, no published study in cardiac electromechanical modelling has simultaneously fulfilled all three requirements, and it is for this reason that the community often refers to the 'digital twin vision' rather than its realisation.

      The primary contribution of this manuscript is the framework: a systematic application of ASME V&V40 standards to a fully coupled electromechanical model, spanning electrical, mechanical, and haemodynamic biomarkers in a single study. The model evaluation presented here is an example application of that framework. Importantly, the framework is not designed to certify a model as complete, but to provide a transparent audit of current capability by identifying where confidence is established and where further development is needed. In this sense, the limitations surfaced through this evaluation are themselves a contribution: they define open problems and priorities for the field.

      We have added a definition of the digital twin concept and the roadmap towards its realisation to the introduction and have updated the language throughout the manuscript to consistently reflect the distinction between the framework contribution and the model evaluation. We maintain that this transparent approach represents a meaningful step towards the digital twin vision.

      The specific limitations of the current model implementation are addressed in detail in the relevant sections of this response and in the updated manuscript, where we have substantially strengthened the verification and discussion components.

      Conclusion:

      Overall, this reviewer considers that the study requires a major revision, including improvements in numerical methods, modeling choices, and checks for physiological behavior. Nevertheless, the provided tables with averaged values from the UK Biobank and the presented validation strategy could be valuable to the research community.

      Reviewer #2 (Public review):

      The authors present an interesting study on calibrating and validating a biventricular cardiac electromechanical model. This is an important contribution, but some questions remain about the quantitative validation and verification aspects of the study.

      Major comments:

      (1) The title and paper stress the importance of validation on several occasions. However, the actual validation performed is limited to the section in lines 427-439. Furthermore, it is entirely qualitative, making assessing the model's quality difficult. Most of the paper is focused on sensitivity analysis, which is also interesting but unrelated to validation. Can you include a quantitative comparison with deformation biomarkers? E.g., spatially quantify strain differences between simulation and in vivo data, or overlay the current configuration of the geometry with MRI in various views, and calculate a displacement error norm.

      We thank the reviewer for this comment.

      We have strengthened the quantitative aspect of the validation by reporting the peak simulated strain values for each component and comparing them against the physiological ranges compiled in Table 2. Specifically, the simulated peak strains were: E_ff ≈ -0.20, E_cc ≈ -0.15, E_rr ≈ +0.15, and E_ll ≈ -0.23. These show broad agreement with the in vivo reference ranges from Moulin et al. (2021), noting that the reference ranges are derived from a cohort of 30 subjects and therefore represent a relatively narrow population sample. Shortening strains (fibre and circumferential) are in good agreement, while radial strain is underestimated. We have noted this as a limitation. We have also indicated that a further validation would include a fully quantitative spatial comparison, such as a displacement error norm or voxel-wise strain difference map. This would require access to the raw image data and patient-specific geometry registration, which is beyond the scope of the current study.

      (2) You mention the ASME V&V40 standards throughout your paper. Yet, you only address the "second V" validation, ignoring the "first V" verification. How did you ensure that your computational models are implemented correctly?

      Thank you for raising this point. We have now included a section on model verification to the manuscript at where we perform benchmarking simulation using the Land (2015) passive inflation benchmark. We also provide a mesh subdivision analysis of the final calibrated model. Additional verifications and previous sensitivity analyses using the same numerical scheme with idealised ellipsoid geometries are also referenced in the verification section, to provide additionally confidence.

      (3) All parameters discussed in this publication are physical parameters. What is the sensitivity of your model outputs concerning computational parameters?

      Numerical analyses for the Alya solver used in this study has previously been published in works including Levrero et al (2021), which performed sensitivity analyses in a truncated ellipsoid geometry, and Santiago et al (2018), which demonstrated mesh convergence in a cantilever. We have updated the manuscript to point the reader to these studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major concerns:

      (1) Active stress scaling:

      The initial value for T_ref appears to be 120kPa * 10, which would be ten times the literature value fitted to human contraction data. Additionally, Table 2 lists a range of [1200-2400], which is 10 to 20 times the literature value.

      This discrepancy suggests that other model parameters, model assumptions, or the numerical scheme may be inadequate. In contrast, similar calibrations using comparable models (ToRORd-Land) in other works, such as Strocchi et al. [29], yielded T_ref values close to the literature value.

      We thank the reviewer for this comment. As discussed in our response to the public review comment 3a, the elevated T_ref scaling warrants explanation.

      We note that Strocchi et al. use a four-chamber geometry include atrial mechanics and a different pericardial constraint, any of which could contribute to differences in the required T_ref scaling. The elevated scaling in our model likely reflects a combination of factors including the absence of poro-elastic behaviour, simplifications in pericardial constraint, and the lack of atrial mechanics, rather than volumetric locking alone. We have added text to the discussion acknowledging this more explicitly and have flagged planned additional benchmarking of the dynamic orthotropic scheme as future work.

      (2) Non-physiological results, see Figure 1:

      In a healthy heart, RV stroke volume should approx. match LV stroke volume. This is clearly not the case in Figure 1B, where the RV EF is also notably low at 35%.

      Consequently, the study fails to reproduce healthy pressure-volume loops, undermining its claim to create a credible cardiac electromechanical digital twin. Hence, also the "Question of interest" posed in line 206 must be answered with a clear "No".

      Matching stroke volumes should be a primary calibration goal.

      We thank the reviewer for this comment. We agree that stroke volume balance is an important physiological criterion, and we have added it explicitly to the framework criteria in the updated manuscript, noting that our current model evaluation does not satisfy it. This is precisely the kind of transparent appraisal the V&V40 framework is designed to produce: a systematic accounting of which criteria are met and which require further development. A framework that only gets applied to models that pass all criteria would be selection-biased and less informative to the community.

      However, we respectfully disagree that the question of interest must be answered with a clear 'No'. We draw the reviewer's attention to the quantities of interest defined in the paper, which are predominantly left ventricular biomarkers, reflecting the intended scope of the calibration framework. The framework successfully reproduces these defined quantities of interest, and the LV pressure-volume loops, strain, volumes and ejection fraction are all within physiological ranges and well-matched to reference data. These quantities of interest were selected based on their clinical implications in cardiac diseases, as detailed in Table 3.

      We agree that stroke volume balance is an important physiological requirement for a fully calibrated biventricular model, and we have strengthened the future work and limitations section accordingly.

      (3) Inadequate numerical framework:

      (a) Monodomain model: The geometries from Rodero et al. [25] have an average edge length of 1mm. It is known that such a coarse resolution leads to inaccurate EP results. It is not mentioned if the authors refined that geometry to an appropriate resolution or used an Eikonal model to mitigate this issue.

      As explained earlier, our ECG simulations are robust against coarse mesh resolutions since we use an Eikonal solution to prescribe the activation times on the endocardial surface, as performed in Camps et al. (2024), and we tune the diffusivity parameters in the model such that the correct conduction velocities are reached. We have added a figure in the appendix of this manuscript to show that by increasing the mesh resolution by one subdivision, we get virtually identical ECG simulations. While this mesh resolution may not be sufficient for simulations of more complex behaviour, such as re-entry and fibrillation patterns, it is sufficient for this study. We have noted this in the methods section.

      (b) Material law:

      - recent publications show that an unsplit deformation gradient for the anisotropic contribution is beneficial to reduce locking effects, see, e.g., Gueltekin et al. Computational Mechanics 63, no. 3 (2019): 443-53. https://doi.org/10.1007/s00466-018-1602-9.

      - K_ct is a penalty parameter to enforce some degree of incompressibility. Results are highly dependent on the grid size and the finite element formulation due to locking effects.

      As the authors write: "In our simulations, we saw that the LVEF was strongly sensitive to changes in the incompressibility of the tissue (Kct), such that an increase in compressibility of the myocardial tissue helped to increase LVEF." Which exactly points to the issue of locking effects.

      So an option would be to use a finer grid or a more adequate numerical scheme with quadratic finite elements, as eg. in [5] Fedele et al., or [6] Gerach et al,. or stabilized elements as in Karabelas et al. CMAME 394 (2022) https://doi.org/10.1016/j.cma.2022.114887.

      Overall, this does not point to "limitations in using ex vivo tissue measurements to represent in vivo function" but to limitations in the numerical setup. In fact, with an adequate numerical scheme, the simulations should be largely insensitive to the choice of this penalty parameter K_ct. See, e.g., Karabelas et al. above, where the authors varied K_ct from 650kPa to infinity (representing an incompressible material), and there is no visible influence on the PV loops.

      We investigated this point using the Alya solver, and we found that the mesh resolution did not alter the LVEF, and our benchmark simulations against Land (2015) did not show the existence of the volumetric locking issue that the reviewer refers to. It is possible, however, that such an effect exists in the elastodynamic orthotropic framework but not in the incompressible and transversely isotropic framework that the Land (2015) benchmarks were set up in. Future analyses could focus on performing additional benchmarking against more recent elastodynamic benchmarks, such as presented in Arostica (2025). We have updated the limitations text in our manuscript to reflect this and to cite relevant literature on this issue.

      (4) Boundary conditions:

      "This was a simplified version of the method [28], which uses an exponential decay formulation at the 'edge' of the pericardial constraint rather than a step function": I don't really see this in the cited work [28] which gives a spatially varying Robin-type boundary condition at the whole epicardium (i.e. regional scaling of normal springs stiffness based on image-derived motion from CT images) and not only at the edge.

      This is motivated by the fact that the pericardial tissue is in contact with various organs of different material properties. Not using spatially varying pericardial parameters is a limitation that might lead to non-physiological deformations, see also Pfaller et al. Biomechanics and Modeling in Mechanobiology 18 (2019): 503-29. https://doi.org/10.1007/s10237-018-1098-4.

      We thank the reviewer for this point and we have corrected the manuscript accordingly. To clarify: our implementation applies a uniform Robin spring constraint along the majority of the epicardial surface with zero constraint at the base, which is conceptually similar to Strocchi et al. [28]. The key difference is that Strocchi et al. use a smooth gradient transition from uniform constraint to zero constraint near the base, whereas our implementation uses an abrupt step transition. We acknowledge that a smooth spatially varying transition would more accurately represent the frictionless pericardial contact and have noted this as a limitation in the manuscript with reference to Pfaller et al. [41].

      Also check:

      - line 117: Gamma_valve_epi is introduced but not used. Was there any boundary condition defined on this valve plane?

      - the third equation, maybe (0,T] missing.

      - line 120: epicardium instead of endocardium.

      These errors have been corrected in the updated manuscript. No boundary conditions were applied on the epicardial surface of the valve plugs, the reference to gamma_valve_epi has been removed.

      (5) Reference geometry:

      The choice to scale the mesh to a lower volume for the unloading procedure seems questionable. This approach does not ensure that the reloaded mesh aligns with the mesh derived from image data. As a result, the geometry used for the simulations is no longer truly patient-specific.

      This mismatch is a significant limitation, as there are established methods available to achieve a proper unloaded configuration, as, e.g., in

      Marx et al. Journal of Computational Physics 463 (2022): 111266. https://doi.org/10.1016/j.jcp.2022.111266, and

      Regazzoni et al. Journal of Computational Physics 457 (2022): 111083. https://doi.org/10.1016/j.jcp.2022.111083.

      As our study aimed at creating a framework for calibration and validation in data-scarce scenarios such as it is often the case in the clinical context, using a compilation of multi-modal data from difference sources, rather than a specific method of personalisation, we did not feel it appropriate to invest significant energy to identify a patient-specific resting geometry, but rather felt that it was important for the resting geometry to fall within population values in terms of diastasis volume. We have clarified this issue in the manuscript and softened claims to Digital Twins in this study. The limitation has been addressed in the updated manuscript, and future work could further address this point.

      (6) Calibration procedure:

      There are apparent flaws in the calibration procedure, or it is not described in sufficient detail.

      We thank the reviewer for raising this point and we have substantially revised the calibration description in the manuscript to clarify the rationale behind each step.

      (a) Step 1: "Sample..." Why? kws and Cal50 are not the most significant parameters in the sensitivity analysis. Kct is a penalty parameter dependent on the numerical framework as described above; "ejection pressure threshold" was never mentioned, is it "P ejection LV" in Table 2? Aiming just for the highest LVEF might neglect non-physiological responses to parameter changes.

      While kws and Cal50 are not the single most significant parameters for LVEF in isolation, they were grouped in Step 1 because they affect both LVEF and peak systolic pressure simultaneously through cross-bridge cycling rate and residual active tension, making it necessary to sample them jointly rather than sequentially. Kct was included because myocardial compressibility affects wall thickening and therefore stroke volume. The ejection pressure threshold is P_ejection_LV in Table 1 and has now been described explicitly in the methods section. Regarding the concern about non-physiological responses: the action potential duration and active tension were monitored throughout calibration and verified to remain within physiological ranges, as now noted in the manuscript.

      (b) Step 2: As systolic pressure is directly dependent on arterial resistance for a 2-element Windkessel model, a uniform sampling approach might not be the best choice here.

      We acknowledge that uniform sampling may not be the most efficient approach for Step 2. However, since arterial resistance influences not only peak systolic pressure but also stroke volume and therefore LVEF, a more targeted approach focusing solely on pressure matching could compromise the LVEF achieved in previous steps. Uniform sampling allowed us to select the value that best balanced both quantities simultaneously.

      (c) Step 3: The authors mention in line 527: "A four-fold increase in GCaL caused an eight-fold increase in cellular active tension peak". An increase in active tension peak results in higher LVEF. So this step is likely to yield the upper boundary of the GCaL interval.

      The reviewer is correct that Step 3 tends to yield a high GCaL value. This was intentional — GCaL was used as a last resort to achieve physiological LVEF after Steps 1 and 2, since the model consistently undershot the target. The upper boundary of the sampled GCaL interval corresponds to a two-fold increase, which remains within the physiological variability bounds applied in previous studies. The resulting action potential duration was verified to remain within physiological ranges.

      (d) Step 4: Why again k_ws? It is not the most significant parameter in the SA.

      kws was resampled in Step 4 not to increase LVEF further, but to specifically target peak ejection rate and dP/dtmax, which were not adequately matched after Step 3. kws is the dominant parameter affecting these ejection dynamics biomarkers in the sensitivity analysis. Resampling at this stage allowed fine-tuning of ejection dynamics while maintaining the LVEF achieved in previous steps.

      (e) Step 5: As far as I can tell, the "diastolic volume change parameter" was mentioned the first time here.

      The diastolic volume change parameter C_pLAV has now been described in the methods section in the Phase 5 passive filling description, where it appears as the inverse of the penalty term controlling the rate of return to diastasis volume in the left ventricle.

      The whole calibration procedure seems to aim for the highest LVEF, and final values of the calibration parameters are not given.

      We note that the calibration procedure does not aim solely for the highest LVEF. As described above, the sequential strategy targets multiple quantities of interest in order of clinical importance: LVEF, peak systolic pressure, peak ejection rate, and peak filling rate, with each step designed to improve a specific subset of biomarkers without compromising those already matched. The final calibrated parameter values are reported in Figure 1F of the revised manuscript.

      (7) Novel features in this paper are actually scarce. A way more advanced calibration strategy with a whole heart model, emulators, and also the ToRORd-Land model was already presented in the study by Strocchi et al. [29]. The calibration to ECGs was presented by some of the same authors in Camps et al. [15], and the analysis of cellular effects was already published in several studies by the same group and in other publications, e.g., by the groups of Severi et al.

      The systematic compilation of credibility criteria spanning ECG morphology, pressure-volume characteristics, strain and displacement represents a novel contribution in itself, providing the field with a reusable evaluation framework. Furthermore, the present study is designed to yield mechanistic insight into how parameters at different scales influence both electrical and mechanical outputs simultaneously. This goal was not tackled in previous publications, which covered individual components, including ECG calibration in Camps et al. [15] and global sensitivity analysis with whole-heart models in Strocchi et al. [29] with no ECG consideration.

      Thus, the work by Camps et al. on ECG calibration was purely electrophysiological and did not investigate the influence of mechanical or haemodynamic parameters on ECG morphology in a fully coupled electromechanical framework. While the effect of mechanical parameters on ECG has been explored by others (e.g. Favino, 2016), this has not previously been examined alongside the relative importance of cellular, mechanical and haemodynamic parameters on pressure-volume characteristics within a single coupled framework. While Strocchi et al. present an emulation strategy, they did not address ECG biomarkers. This distinction is now stated explicitly in the introduction, where we position the present study relative to Camps et al. and Strocchi et al.

      (8) How could the calcium sensitivity Cal50 have such a drastic effect on diastolic function, i.e., filling and end-diastolic volume? As far as I understand from the description, the simulation starts with Phase 0 (loading), Phase 1 (atrial filling), and then in Phase 2, electrical activation ensues and active contraction develops, see also the section starting in line 165. Based on this description, I would expect the end-diastolic volumes to be identical across all Cal50 values. Or are the PV loops shown actually limit cycles established over simulations with multiple beats? This point wasn't explicitly clarified in the manuscript.

      Calcium sensitivity (Cal50) affects not only systolic active tension development but also diastolic residual active tension, i.e. the degree to which the muscle remains partially activated at end diastole. Higher Cal50 values increase this residual tone, effectively stiffening the myocardium during diastolic filling and reducing end-diastolic volume. This mechanism is well established as a contributor to diastolic dysfunction in heart failure [88]. We have clarified this in the manuscript and also clarified that the PV loops shown are single-beat simulations, not limit cycles, with the end-diastolic volume determined by the prescribed filling pressure alongside the passive and residual active stiffness of the myocardium.

      Minor concerns:

      (9) Line 29: The values provided: LVEF of 51%, EDV of 110 mL, and ESV of 50 mL are inconsistent. If these values are all related to the LV, the calculated LVEF should be approximately 54.55%, not 51%.

      The values quoted in the original abstract were rounded approximation, this has been corrected to report EDV=105 mL and ESV=51 mL, which are consistent with the simulated LVEF of 51%.

      (10) "Electromechanical cardiac Digital Twins have had broad applicability ..."

      Many of the cited works here are not true "Digital Twins" but rather static, non-patient-specific models of cardiac electromechanics. In some cases, the geometry may be derived from patient data, but this alone does not qualify the model as a digital twin.

      This sentence in the introduction has been rephrased as ‘Electromechanical cardiac models have had broad applicability...’. Furthermore, as stated earlier, we have removed explicit claims of Digital Twin from the paper while retaining the fact that this study provides a significant step towards rigorous credibility assessment of the high-fidelity electromechanical models that make Digital Twin construction possible.

      (11) While in the abstract and the conclusion, the authors mention "uncertainty quantification", it is mostly a sensitivity analysis that was performed in the paper.

      We have updated the text to say ‘sensitivity analysis’ where appropriate in the abstract, results, and conclusion, and replaced ‘uncertainty ranges’ with ‘variability ranges’ throughout. However, since the sensitivity analyses were performed over biologically informed ranges derived from population variability in the literature, the results are informative about how uncertainty in model inputs propagates to uncertainty in simulated biomarkers. We have therefore retained the framing of sensitivity analysis as a first step towards uncertainty quantification in the abstract and conclusion, and have added a clarifying sentence to the methods to this effect.

      (12) Line 98: As far as I can tell, the conduction velocity assigned to the endocardial surface - intended to mimic the Purkinje fiber network - is never specified. In the section beginning at line 294, only the transmural conduction velocities are reported.

      The endocardial conduction velocity has been specified in the methods section: Purkinje-myocardial junctions were modelled using a fast endocardial activation layer with isotropic conduction velocity of 300 cm/s.

      (13) Line 198, Table1:

      (a) "21/02/2025 11:09:00 AM" on two occasions is maybe not intended

      This has been removed.

      (b) For easing up comparisons, units should be consistent between the initial value and the literature ranges, e.g., PV control parameters, heart rate.

      Units have been made consistent between the initial values and literature ranges throughout Table 1.

      (14) Line 232: "... have already been used to calibrate and validation ...".

      This has been corrected.

      (15) Line 279, Table 2: This table of variability ranges is not entirely clear and could be improved:

      Table 2 has been combined with Table 1 such that the variability ranges sit next to the literature values, for ease of comparison.

      (a) "21/02/2025 11:09:00 AM" is maybe not intended.

      This has been removed.

      (b) use of units should be improved; sometimes it's given in the first column, sometimes in the second column (arterial resistance, compliance), then for k_epi it should be either kPa or kPa/cm.

      Units have been made consistent and the units for k_epi has been added in Table 1.

      (c) units should also be consistent throughout the paper, e.g. in Figure 1 E arterial resistance is Barye.ms/mL while in Table 2 it is mmHg.ms/mL.

      Barye has been removed and replaced by corresponding kPa values throughout the manuscript. This was in the original manuscript since the Alya simulation software were in units of cm, s, g, Barye.

      (d) it is also not clear how variability ranges were chosen; e.g., for arterial compliance,e literature ranges are 0.2-2.73 while the chosen range is [0.1,0.2].

      The previous ranges were chosen to achieve better LVEF. We have now updated the variability ranges to be purely based on literature values and updated the sensitivity analysis results. The ranges are now presented in Table 1 alongside the literature values for ease of comparison.

      (e) For Kct, the initial value in Table 1 is 5000kPa, the literature values are between 10 and 3333, and then the variability range is [10,500]? I guess there is a typo in one of these values.

      This has been corrected in the new Table 1.

      (f) Table 1 and 2 are in parts redundant.

      Table 1 and 2 have been combined into a single new Table 1.

      (15) Line 290: It should be uvc_l for the longitudinal coordinate.

      This has been corrected.

      (16) Line 388, Table 3, regarding values for pressure volume from reference [49]:

      (a) the number of participants is 800, including males and females; not only females, see also Table 12 https://jcmr-online.biomedcentral.com/articles/10.1186/s12968-017-0327-9/tables/12

      (b) why using female values here while having mixed sex for most of the others? Because the model is female?

      The reviewer is correct that reference [49] reports values from a mixed-sex cohort of approximately 800 participants. We used the female-specific values from Table 12 of that reference because the biventricular mesh used in this study was derived from a female subject, making sex-matched reference values the most appropriate comparison. This has been clarified in the manuscript.

      (17) Figure 5: What is Jup; why did you choose 0.93 x Jup as reference? Also in Figure 4, why did you use 0.93 x GCal as a reference?

      J_up refers to the SERCA<sup2+</sup> reuptake current, which has been relabelled as SERCA throughout the manuscript for consistency. The reference value of 0.93× was used because the sensitivity analysis sampled parameters uniformly between 50% and 200% of baseline using a fixed number of samples, and no sample fell exactly at 1.0×. The closest sampled value was 0.93×, which was therefore used as the reference. This has been clarified in the figure caption.

      (18) Line 377: The link to the GitHub repository does not work.

      This link has now been made publicly available.

      (19) Line 397, Table 3: for the sake of completeness, all abbreviations should be included: e.g., SVL, ESP, EDV, ESV are not included.

      This has been written out in full in the new Table 2.

      (20) Tick marks in many figures are not readable, e.g., Figure 5 and all the Figures in the appendix.

      Tick mark sizes and line widths have been increased across Figures 4, 5, and all appendix figures. The figures have been replotted and updated in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Minor Comments:

      (1) The provided GitHub link https://github.com/jennyhelyanwe/Alya_input_setup/ does not work, potentially because the repository is private. It would be nice to see the repository during the review.

      This link has now been made publicly available.

      (2) Table 3: Can you include the simulation outputs obtained for validation (with an error indication)? This would summarize the validation that's currently spread out over the results section.

      A new Table 3 has been added to the manuscript under the validation section, summarising the simulated values for all deformation and strain biomarkers alongside their reference ranges. The calibration and validate datasets are now reported separately in Tables 2 and 3, respectively.

      (3) Figure 1: Add axis labels to all plots.

      Axis labels have been added to all subplots in Figure 1 in the revised manuscript. Simulated pseudo-ECG amplitudes are normalised and therefore dimensionless.

      (4) Figure 2: Simulated and in vivo strains with exactly the same axes (size, range, ticks) and add grid lines to enable a comparison. Add the mean values of each in the other plot.

      The revised Figure 2 now includes the median in vivo strain values from Moulin et al. overlaid as a red dashed reference line on the simulation panels, enabling direct visual comparison. The simulated mean could not be overlaid on the in vivo panels as the original Moulin et al. figure data are not publicly available for replotting. Exact axis matching was not applied as this would cause some simulated curves to fall outside the visible range, obscuring the model behaviour.

      (5) Figure 3: The thickness (relative importance of the connections) is impossible to see in this plot. Instead of having gray background connections, remove them entirely below a certain threshold. Make the differences in thickness more pronounced or introduce a continuous color scale for the magnitude of the positive or negative correlation. Alternatively, you could rank the parameters from least to most important in each subfigure A-D and/or provide some numeric values.

      Figure 3 has been updated. All non-significant connections (|r| < 0.6 or p > 0.05) have been removed entirely, and gray lines have been removed in each subfigure, making the significant relationships clearer. A continuous blue-to-red colour scale has been applied to indicate the direction of correlation (blue: negative, red: positive), with line thickness proportional to the magnitude of the r-value.

      (6) Figure 3 and Table 2: Why were material parameters b, bf, bs, and bfs omitted from this study (but included a, af, as, and afs)?

      The b parameters (b, bf, bs, bfs) appear in the exponent of the Holzapfel-Ogden constitutive law and are strongly coupled to the a parameters (a, af, as, afs), which carry units of kPa. In practice, the b parameters can only be reliably identified from ex vivo multiaxial stretch experiments, whereas the a parameters can be estimated from clinical imaging data. Since our study focuses on calibration and validation in a clinical data setting, we included only the a parameters in the sensitivity analysis, consistent with previous personalisation studies.

      (7) Figure 4: What do the dotted lines represent?

      The dotted lines in Figure 4E highlight the increased longitudinal shortening with increasing GCaL, showing the basal plane moving towards the apex while the apical position remains unchanged due to the pericardial constraint. This has been clarified in the figure caption.

      (8) Figures 4, 5, A2-45: Can you use a continuous color scale (e.g., from blue to red) for low to high parameter uncertainty?

      A continuous blue-to-red colour scale has been applied to Figures 4, 5, and all appendix figures A2–A6, where blue indicates the lowest parameter value and red indicates the highest. A colour bar has been added to each figure for reference.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Wang et al., recorded concurrent EEG-fMRI in 107 participants during nocturnal NREM sleep to investigate brain activity and connectivity related to slow oscillations (SO), sleep spindles, and in particular their co-occurrence. The authors found SO-spindle coupling to be correlated with increased thalamic and hippocampal activity, and with increased functional connectivity from the hippocampus to the thalamus and from the thalamus to the neocortex, especially the medial prefrontal cortex (mPFC). They concluded the brain-wide activation pattern to resemble episodic memory processing, but to be dissociated from task-related processing and suggest that the thalamus plays a crucial role in coordinating the hippocampal-cortical dialogue during sleep.

      The paper offers an impressively large and highly valuable dataset that provides the opportunity for gaining important new insights into the network substrate involved in SOs, spindles, and their coupling.

      Thank you for this encouraging assessment. We appreciate your recognition of the value of the dataset and of the questions it allows us to address. Below, we respond to each of your points directly and revise the manuscript accordingly.

      Comments on revisions:

      Re 1: The revised introduction now cites a couple of papers but discusses them only very superficially, lumping together several studies with very different key results. This is still not very informative for the reader and does not sufficiently acknowledge previously published work. Here are two examples to illustrate this:

      (a) "These studies have generally reported that slow oscillations are associated with widespread cortical and subcortical BOLD changes, whereas spindles elicit activation in the thalamus, as well as in several cortical and paralimbic regions." Several studies even showed e.g., a clear activation of the hippocampus and parahippocampal gyrus associated with spindles, not just the thalamus

      Thank you for this comment. We agree that our previous sentence was too broad and did not sufficiently reflect the range of findings in the sleep literature. We have therefore rewritten the Introduction to state explicitly that spindle-related BOLD changes have been reported not only in the thalamus, but also in cortical and paralimbic regions, including the hippocampus and parahippocampal gyrus (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).

      Introduction, Page 3-4, Lines 58-62

      “Consistent with this view, prior human EEG-fMRI studies have reported spindle-related activation not only in the thalamus, but also in the hippocampus and adjacent parahippocampal gyrus (Bergmann et al., 2012; Schabus et al., 2007). Spindle-related activity has also been linked to striatal engagement, suggesting a broader network that may support memory-related processing during sleep (Fogel et al., 2017).”

      Introduction, Page 4, Lines 71-78

      “Previous EEG-fMRI studies on sleep have examined both global sleep characteristics (Hale et al., 2016; Moehlman et al., 2019) and the neural correlates of specific waves, including slow oscillations and spindles. These studies have generally shown that slow oscillations are associated with widespread cortical and subcortical BOLD changes (Czisch et al., 2009; Ilhan-Bayrakcı et al., 2022; Picchioni et al., 2011), whereas spindles have been linked not only to thalamic activation but also to cortical and paralimbic regions, including the hippocampus and parahippocampal gyrus (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).”

      Introduction, Page 5, Lines 103-106

      “This coupling was associated with increased activation in both the thalamus and hippocampus, with functional connectivity patterns suggesting thalamic coordination of hippocampal-cortical communication, in line with prior EEG-fMRI studies of spindle-related activity (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).”

      (b) "Although these findings provide valuable insights into the BOLD correlates of sleep rhythms, they often do not employ sophisticated temporal modeling (Huang et al., 2024) [, ...]." - previous studies have used e.g., spindle event-related regressors with individual spindle amplitudes as parametric modulators, first and second order derivatives of the HRF function, as well as PPI connectivity analyses, which I would consider rather sophisticated temporal modelling.

      We agree that several previous studies have already employed sophisticated modelling approaches, including parametric modulation, HRF derivatives, and PPI analyses (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Picchioni et al., 2011). Our intention was not to suggest that such methods are absent from the literature.

      Rather, we aimed to highlight that most prior work has focused on modelling individual SO or spindle events, whereas explicit modelling of their temporal interaction (e.g., SO-spindle coupling) has been less commonly addressed. We have revised the sentence to clarify this point more precisely.

      Introduction, Page 4 Lines 78-82

      “Although these findings provide important insight into the BOLD correlates of sleep rhythms, most previous studies have focused on individual oscillatory events rather than explicitly modelling their temporal interaction (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Picchioni et al., 2011). Only a few recent studies have begun to examine coupling between rhythms directly, for example Huang et al. (2024).”

      Re 4+9: The short overall recordings in some subjects on the one hand and the large number of spindles and SOs detected in N1 sleep stages are still highly concerning, in fact even more so, now that the actual numbers have been provided in the Supplementary Tables. Either the sleep staging or the detection of SO and spindle events must be incorrect. I understand that for specific EEG analysis and fMRI modelling purposes sometimes slightly different thresholds are used as compared to clinical sleep staging, but several parameters here are alarmingly off.

      (a) Given that proper NREM sleep (N2+N3) is the relevant stage for the analyses conducted in this paper, some of the N2+N3 durations are very short (eg 7-8 min) while those subjects' results have the same impact on the group level analyses as those with >100 min of N2+N3. Either subjects with very little relevant data (not overall recording time but N2+N3 time) should be excluded or weighting subject data for the group analyses according to the amount od contributed data should be done.

      Thank you for the suggestion. It is true that participants with very little N2/3 sleep could contribute noisier subject-level estimates to the group analysis. We therefore checked this directly. Only three participants contributed less than 10 min of N2/3 sleep, and excluding them did not change the main results. For example, hippocampal activation during SO-spindle coupling remained significant after exclusion, t<sub>(103)</sub> =2.50, p = 0.0071, compared with t<sub>(106)</sub> = 2.50, p = 0.0070 in the full sample. We have added this control analysis to the Results so that the robustness of the group findings is explicit in the manuscript.

      Results, Page 11-12, Lines 238-250

      To ensure the results were not driven by individual differences or parameter selection, we conducted a series of control analyses. First, we excluded participants with less than 10 minutes of N2/3 sleep. Only three participants met this criterion, and their exclusion did not change the main results. For example, hippocampal activation during SO-spindle coupling remained significant (t<sub>(103)</sub> = 2.50, p = 0.0071), comparable to the full sample (t<sub>(106)</sub> = 2.50, p = 0.0070). Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st-80th percentile; Fig. S6). Third, to test whether the results depended on the use of a single lateralised frontal electrode, we repeated the EEG-informed fMRI GLM using events detected from Fz. Hippocampal activation during SO-spindle coupling again remained significant (t<sub>(106)</sub> = 2.47, p = 0.0076), closely matching the original F3-based result (t<sub>(106)</sub> = 2.50, p = 0.0070).

      (b) The authors argue that the SO and spindle detection algorithms are valid since widely used and that they were developed for N2+N3 stages, which is why they will also detect events in other stages: "While, because the detection methods for SO and spindle are based on percentiles, this method will always detect a certain number of events when used for other stages (N1 and REM) sleep data, but the differences between these events and those detected in stage N23 remain unclear." I do agree that with very liberal thresholds, also SO and spindle vents may be detected in other stages, but it shouldn't be that many. If the percentiles of amplitude thresholds were defined based on properly scored N2+N3 stages only, very few events should be detected (erroneously!) in N1, as the occurrence of K-complexes (isolated SOs) and spindles per definition makes it N2, and during REM sleep only very few spindles and SOs are allowed to occur, without scoring it NREM instead. For the first subject (just as example, but with similar numbers for the rest of the sample), reveals as many as 60 SOs and 31 spindles within 8 min of N1 sleep (Table S2) as well as 13 SOs and 7 spindles within 2 min of REM sleep (Table S4). These numbers are completely unrealistic and question the correctness of the sleep staging as well as the physiological relevance of the EEG graphoelements identified as SO and spindles. It also completely undermines the interpretability of the respective event regressors for the fMRI analyses.

      (c) Likely, given the large numbers of coupled SO-spindle events and the apparently very low amplitude criteria for event identification, also the number of SO-spindle couplings is likely severely overestimated.

      We thank the reviewer for raising this important point. We agree with you that the original stage-wise percentile thresholding could inflate the apparent number of SOs and spindles outside N2/3 sleep. In the original analysis, the thresholds were estimated separately within each sleep stage. As you point out, this procedure can force the detector to label a relatively large number of events in N1 and REM, even when those waveforms should not be interpreted as canonical N2/3 SOs or spindles. We have therefore revised the detection procedure. Following your concern and Reviewer 2’s suggestion, the SO and spindle thresholds are now defined only from N2/3 sleep within each participant, where SOs and spindles are most abundant and physiologically expected to occur. These fixed N2/3-derived thresholds were then applied unchanged to N1 and REM for descriptive reporting. This avoids the artificial normalisation of event detection across sleep stages that can arise when each stage has its own percentile threshold. And we have revised all relevant sections of the manuscript, including “[Results, Page 6-7 Lines 134-148]; [Fig. 1e]; [Results, Page 9 Lines 175-191]; [Fig. 2b]; [Methods, Page 25-27, Lines 567-604]; [Fig. S2-S4]; [Table S2, S4].”

      With this revised procedure, detections outside N2/3 are clearly lower than those in N2/3. The mean densities are 2.95 SOs/min, 2.71 spindles/min, and 0.75 coupling events/min in N1, and 2.07 SOs/min, 1.81 spindles/min, and 0.43 coupling events/min in REM. We agree with you that the remaining detections in N1 and REM should not be treated as physiological equivalents of canonical N2/3 SOs, spindles, or SO-spindle complexes. We therefore report them only as descriptive detector outputs obtained under a fixed N2/3-derived threshold (see Table S2, S4 in the revised manuscript). We do not use them to support any physiological claim about SO-spindle coupling in N1 or REM.

      This point is also important for the fMRI analyses. You are right that inflated N1 or REM detections would undermine the interpretability of event regressors if those detections entered the EEG-informed fMRI models. They did not. All EEG-informed fMRI GLM and PPI analyses were restricted to N2/3 sleep, where SOs, spindles, and their coupling are physiologically expected and where the detection thresholds were defined. Thus, the central fMRI event regressors were based only on N2/3 events, not on detections from N1 or REM.

      We also agree with you that the absolute number of detected SO-spindle couplings depends on the chosen detection threshold. For this reason, we tested whether the main EEG-fMRI result depended on the specific detector setting. Hippocampal activation during SO-spindle coupling remained significant when the SO detection threshold was varied between the 71st and 80th percentiles, as shown in Fig. S6. We therefore do not argue that the detector provides a uniquely correct absolute count of SOs, spindles, or coupling events in every sleep stage. Our conclusion is more specific. The main N2/3 EEG-fMRI finding is robust across a reasonable range of SO detection thresholds, detections in N1 and REM are reported only descriptively, and the physiological interpretation of SO-spindle coupling is restricted to N2/3 sleep.

      Results, Page 6-7 Lines 134-148

      “Each sleep stage is characterised by distinct spectral properties and rhythmic waveforms, serving as physiological markers (Fig. 1c). Because SO and spindle detection relies on amplitude-based percentile thresholds, we avoided estimating separate thresholds within each sleep stage. Instead, for each participant, the SO and spindle thresholds were defined from N2/3 sleep only, where these rhythms are most abundant and physiologically expected, and the same fixed thresholds were then applied to N1 and REM for descriptive comparison.”

      “Under this fixed N2/3-derived thresholding, detected SOs and spindles were larger and more frequent in N2/3 than in N1 or REM. SO and spindle amplitudes were significantly higher during N2/3 sleep (SO: 25.59 ± 1.49 μV; spindle: 7.39 ± 0.27 μV) than during N1 (SO: 20.15 ± 2.32 μV; spindle: 5.23 ± 0.27 μV) and REM sleep (SO: 19.84 ± 1.22 μV; spindle: 5.60 ± 0.22 μV; all p < 1e-4; Fig. 1e, Fig. S2). The corresponding event densities showed the same pattern, with 9.64 ± 0.25 SOs/min and 4.19 ± 0.10 spindles/min in N2/3, compared with 2.95 ± 0.16 SOs/min and 2.71 ± 0.14 spindles/min in N1, and 2.07 ± 0.17 SOs/min and 1.81 ± 0.14 spindles/min in REM (all p < 1e-4). We therefore report detections in N1 and REM only as descriptive outputs of the detector under a fixed N2/3-derived criterion, rather than as physiological equivalents of canonical N2/3 SOs or spindles.”

      Fig. 1 legend, Page 8, Line 166-172

      “e, Amplitudes (μV) of detected SOs (left) and spindles (right) across sleep stages. SO and spindle detection thresholds were defined from N2/3 sleep within each participant and then applied unchanged to N1 and REM for descriptive comparison. Detections in N1 and REM should therefore be interpreted as detector outputs under this fixed N2/3-derived criterion. The SO amplitudes were measured from the 0.16-1.25 Hz filtered EEG data, and spindle amplitudes were measured from the 12-16 Hz filtered EEG data. Each dot represents an individual participant. Error bars indicate SEM. *** p < 0.001.”

      Results, Page 9 Lines 175-191

      “SO-spindle coupling is considered important for sleep-dependent memory consolidation. In the current study, using the same N2/3-derived detection thresholds described above, we found that SO-spindle coupling occurred most frequently during N2/3 sleep (2.46 ± 0.06 events/min). Coupling density was significantly lower in N1 (0.75 ± 0.05 events/min, t<sub>(106)</sub> = 23.54, p < 1e-4) and REM sleep (0.43 ± 0.04 events/min, t<sub>(106)</sub> = 31.24, p < 1e-4; Fig. 2b, Table S2-S4), consistent with the expected predominance of SO-spindle coupling in NREM sleep (Ngo et al., 2013; Staresina et al., 2015). As with the individual SO and spindle detections, coupling events detected in N1 and REM were retained only for descriptive stage-wise reporting (see Table S2, S4). They were not used to support physiological claims about SO-spindle coupling in these stages, and they were not entered into the EEG-informed fMRI analyses. All subsequent fMRI GLM and PPI analyses were restricted to N2/3 sleep.”

      “After extracting all N2/3 EEG epochs in which SO-spindle coupling occurred, we analysed their spectral and phase characteristics. The spindles were most likely to occur slightly before the UP-state peak of SOs (Fig. 2a, e), aligning with results from both animal studies (Maingret et al., 2016) and human research (Staresina et al., 2015). In our data, this pattern was consistent across subjects (Fig. 2d, Rayleigh test: z = 9.51, p < 1e-4), with the peak of the spindle aligned at an SO phase of −41.61 ± 0.86° (the SO UP-state peak is 0°).”

      Fig. 2 legend, Page 10, Line 202-205

      “b, SO-spindle coupling density across sleep stages, using SO and spindle detections obtained with fixed N2/3-derived thresholds. Coupling events in N1 and REM are shown only for descriptive comparison. The EEG-informed fMRI analyses used N2/3 coupling events only.”

      Results, Page 11-12, Lines 242-247

      “Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st - 80th percentile; Fig. S6).”

      Methods, Page 25-26, Lines 567-575

      “Detection of SOs. Data were first bandpass-filtered between 0.16 and 1.25 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). After identifying all positive-to-negative zero crossings, potential SOs were defined based on the interval between consecutive zero crossings, ranging from 0.8 s to 3 s. For each potential SO, we calculated the amplitude range as the peak minus the trough. For each participant, the amplitude threshold was defined as the 75th percentile of candidate SO amplitude ranges observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. Only candidates exceeding this threshold were labelled as SOs, following previous work (Schreiner et al., 2021).”

      Methods, Page 26, Lines 576-583

      “Detection of sleep spindles. Detection of sleep spindles. Data were bandpass-filtered between 12 and 16 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). The root mean square (RMS) of the filtered signal was computed with a 200 ms sliding time window. For each participant, the spindle threshold was defined as the 75th percentile of RMS values observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. RMS segments exceeding this threshold for 0.5 s to 3 s were identified as spindles (Staresina et al., 2015).”

      Methods, Page 26, Lines 584-591

      “Detection of SO-spindle couplings. From the detected SOs and spindles, we identified the peak time of each spindle. Within each SO interval, we checked whether a spindle peak occurred; if so, that SO was labelled as an SO-spindle coupling event. For descriptive stage-wise summaries, coupling events were assigned to the sleep stage of the corresponding SO trough. For every SO-spindle coupling event, an epoch was created time-locked to the SO trough as the central reference, following Schreiner et al. (2021). We extracted data in a [−4 s to 4 s] window around this point, forming the epoch for each coupling event. For the EEG-informed fMRI analyses, only SO, spindle and SO-spindle coupling events detected during N2/3 sleep were used.”

      Methods, Page 26-27, Lines 592-604

      “The detection procedures described above were developed primarily for N2 and N3 sleep, where SOs, spindles and their coupling are physiologically expected and most reliably observed (Hahn et al., 2020; Helfrich et al., 2019; Helfrich et al., 2018; Ngo, Fell, & Staresina, 2020; Schreiner et al., 2022; Schreiner et al., 2021; Staresina et al., 2015; Staresina et al., 2023). Because percentile-based thresholds can otherwise force the detector to label events in every sleep stage, we did not estimate separate thresholds within N1 or REM. Instead, for each participant, all SO and spindle thresholds were defined from N2/3 sleep and then applied uniformly across the recording. Tables S1 and S3 report detailed statistical information on sleep rhythm and N2/3 events detection. The N1 and REM events detection reported in Tables S2 and S4, and illustrated in Fig. S2-S4, should therefore be interpreted as descriptive detector outputs under this fixed N2/3-derived criterion, rather than as evidence for canonical N2/3 SOs, spindles or physiological SO-spindle complexes in those stages. These detections were not used in the EEG-informed fMRI GLM or PPI analyses, which were restricted to N2/3 sleep.”

      Re 10: The rationale for using a lateralized frontal electrode (F3) for both SO (should have been at least bilateral or central) and spindle detection (should have been a centro-parietal electrode) is not convincing. Other EEG-fMRI spindle or SO papers have used a number of frontal (SO) or centro-parietal (spindles) electrodes averaged or even approaches including all EEG electrodes. Searching events with low thresholds at suboptimal recording sites does not dot this highly valuable dataset justice.

      We thank the reviewer for this important comment. We agree that this choice is more sensitive to frontal SOs than to the centro-parietal fast spindle component. Our choice of F3 was driven by the practical constraints of prolonged nocturnal EEG-fMRI recordings. In our MR-compatible EEG setup, FCz was used as the online reference. Central electrodes close to FCz can have reduced signal contrast relative to the reference, and electrodes near the vertex are also more vulnerable to prolonged pressure against the MRI head coil when participants sleep supine for several hours. In this setting, frontal electrodes provided more stable signal quality across the recording. Because the EEG events were used primarily as temporal markers for fMRI modelling, our priority was to obtain reliable event timing during N2/3 sleep rather than to estimate the full scalp topography of SOs and spindles.

      We also agree with your concern that this valuable dataset would ideally be analysed with multichannel detection strategies. To test whether the main result depended on the single lateralised F3 site, we repeated the main EEG-informed fMRI analysis using Fz, a midline frontal electrode. The result was unchanged. Hippocampal activation during SO-spindle coupling remained significant when events were detected from Fz, t<sub>(106)</sub> = 2.47, p = 0.0076, closely matching the original F3-based result, t<sub>(106)</sub> = 2.50, p = 0.0070. This control analysis does not remove the limitation that centro-parietal fast spindles may be underrepresented, and we do not claim that it does. It does show, however, that the main hippocampal fMRI finding is not driven by idiosyncratic detections from one lateralised frontal electrode. We have made this clearer in the revised manuscript.

      Finally, your concern about low thresholds is also important. As described in our response above, the revised analysis now defines SO and spindle thresholds from N2/3 sleep and applies these thresholds uniformly for descriptive comparisons across stages. We also tested the robustness of the hippocampal fMRI result across SO detection thresholds, and the effect remained significant across the 71st to 80th percentile range. We have therefore narrowed the interpretation in the revised manuscript. The main EEG-fMRI result reflects BOLD activity associated with frontal-channel-detected SO-spindle coupling during N2/3 sleep, rather than a full multichannel characterisation of all SO and spindle topographies.

      Results, Page 12, Lines 247-250

      “Third, to test whether the results depended on the use of a single lateralised frontal electrode, we repeated the EEG-informed fMRI GLM using events detected from Fz. Hippocampal activation during SO-spindle coupling again remained significant (t<sub>(106)</sub> = 2.47, p = 0.0076), closely matching the original F3-based result (t<sub>(106)</sub> = 2.50, p = 0.0070).”

      Discussion, Page 18, Lines 380-394

      “Third, sleep oscillation detection was based on a single frontal electrode. This choice improved signal stability and event timing in the prolonged simultaneous EEG-fMRI setting, but it did not exploit the full multichannel EEG information and cannot characterise the full spatial distribution of SOs and spindles. In particular, F3-based detection may be more sensitive to frontal SOs and frontal sigma activity than to the centro-parietal fast spindle component. We therefore interpret the EEG-informed fMRI results as reflecting BOLD activity associated with frontal-channel-detected SOs, spindles, and their coupling during N2/3 sleep. Future studies using multichannel or source-informed detection strategies, with separate treatment of slow and fast spindles, will be better suited to capture the spatial dynamics of these sleep oscillations. Fourth, the use of large anatomical ROIs may mask subregional contributions of specific thalamic nuclei or hippocampal subfields. Finally, without a memory task, we cannot establish a direct behavioral link between sleep-rhythm-locked activation and memory consolidation. Future studies combining ultra-high-field fMRI or iEEG with cognitive tasks, as well as multichannel or source-informed detection strategies that separately characterize slow and fast spindles, will be better suited to refine our understanding of subregional network dynamics and the functional significance of sleep oscillations.”

      Methods, Page 25, Lines 556-566

      “It is worth noting that the primary aim of EEG rhythm detection was to identify reliable event times for EEG-informed fMRI modelling. Detection was performed on the F3 electrode because this channel provided stable signal quality during prolonged nocturnal EEG-fMRI recordings. In our MR-compatible EEG setup, FCz was used as the online reference. Central electrodes close to this reference, and electrodes near the vertex that were in prolonged contact with the head coil during supine sleep, were more susceptible to reduced signal contrast, impedance drift, and pressure-related degradation of electrode-scalp contact. We therefore used F3 as a pragmatic choice to maximize reliable event timing in N2/3 sleep. This choice was not intended to characterise the full scalp topography of SOs or spindles, and it may underrepresent the centro-parietal fast spindle component. As a sensitivity analysis, we repeated the main EEG-informed fMRI GLM using Fz, a midline frontal electrode, with the same detection and modelling procedure.”

      Re 7: It is not clear to me why/how larger voxels would reduce susceptibility-related distortions and partial volume effects. Usually, the opposite is true. This should be elaborated.

      What we meant was that we chose a relatively large voxel size to preserve signal-to-noise ratio and whole-brain coverage within a feasible repetition time for a long overnight EEG-fMRI protocol. This choice is useful for maintaining BOLD sensitivity in sleep recordings, where head motion, physiological noise, and participant comfort are major practical constraints. We agree that it may not be accurate to describe it as reducing susceptibility-related distortion or partial volume effects.

      We have rewritten the Methods to state this trade-off directly. The voxel size of 3.5 × 3.5 × 4.2 mm<sup>3</sup> allowed whole-brain coverage with a TR of 2000 ms, which was important for modelling sleep-rhythm-related BOLD responses across the whole brain during prolonged nocturnal recordings. A smaller voxel size would have improved spatial specificity, but would also have required either a longer TR, reduced brain coverage, or lower SNR, none of which would have been ideal for the present EEG-fMRI sleep design. We now explicitly acknowledge the cost of this choice.

      Methods, Page 21 Lines 453-463

      “For the functional scans, whole-brain images were acquired using a T2*-weighted gradient echo-planar imaging (EPI) sequence sensitive to the BOLD contrast. The sequence parameters were as follows: 33 slices in interleaved ascending order, TR = 2000 ms, TE = 30 ms, voxel size = 3.5 × 3.5 × 4.2 mm<sup>3</sup>, FA = 90°, matrix = 64 × 64, gap = 0.7 mm. A relatively large voxel size was chosen to preserve signal-to-noise ratio while maintaining whole-brain coverage within a feasible repetition time. This compromise was important for the prolonged overnight EEG-fMRI sleep protocol, where head motion, physiological noise, participant comfort, and sustained acquisition stability are substantial practical constraints (Bodurka et al., 2007; Laufs et al., 2008). A smaller voxel size would have improved spatial specificity, but would have required either a longer repetition time, reduced brain coverage, or lower signal-to-noise ratio.”

      Reviewer #2 (Public review):

      In this study, Wang and colleagues aimed to explore brain-wide activation patterns associated with NREM sleep oscillations, including slow oscillations (SOs), spindles, and SO-spindle coupling events. Their findings reveal that SO-spindle events corresponded with increased activation in both the thalamus and hippocampus. Additionally, they observed that SO-spindle coupling was linked to heightened functional connectivity from the hippocampus to the thalamus, and from the thalamus to the medial prefrontal cortex-three key regions involved in memory consolidation and episodic memory processes.

      This study's findings are timely and highly relevant to the field. The authors' extensive data collection, involving 107 participants sleeping in an fMRI while undergoing simultaneous EEG recording, deserves special recognition. If shared, this unique dataset could lead to further valuable insights.

      Thank you for this encouraging assessment. We appreciate your recognition of the effort involved in collecting this simultaneous EEG-fMRI sleep dataset. Below, we respond directly to your remaining concern.

      Comments on revisions:

      The authors' efforts in revising the manuscript and addressing the reviewers' comments are certainly commendable. However, I remain concerned about potential issues in detecting sleep-related oscillations (SOs, spindles, and consequently coupled SO-spindle events), which may arise due to suboptimal parameter selection or inaccurate sleep staging, potentially impacting all subsequent analyses.

      A review of Supplementary Tables 1-4 reveals an unusually high number of detected SOs and spindles during sleep stage N1 and REM sleep. While the authors correctly note that a percentile-based detection approach will always identify a certain number of events across sleep stages, the particularly high counts in N1 and REM are concerning. To mitigate the limitations of this method, the authors could have performed event detection independently of sleep stages (i.e., across the entire dataset for each participant) and subsequently assigned the detected events to the corresponding sleep stages. If the event counts in N1 and REM remained disproportionately high, this would indicate a fundamental issue with the detection procedure.

      In the previous version, thresholds were estimated separately within each sleep stage. As you point out, this can force the detector to identify a relatively large number of SOs and spindles in N1 and REM, even when those waveforms should not be interpreted as canonical N2/3 events.

      We have therefore revised the detection procedure so that event detection is no longer based on separate stage-wise thresholds. Following the logic of your suggestion, we first defined a fixed threshold for each participant and then assigned the detected events to their corresponding sleep stages afterwards. We used N2/3 sleep to define the SO and spindle thresholds because this is the stage in which these events are physiologically expected and most reliably observed. These same N2/3-derived thresholds were then applied unchanged to N1 and REM. This avoids the circularity of forcing a percentile-defined number of detections within each sleep stage.

      With this revised procedure, detections outside N2/3 are clearly lower than those in N2/3. The mean densities are 2.95 SOs/min, 2.71 spindles/min, and 0.75 coupling events/min in N1, and 2.07 SOs/min, 1.81 spindles/min, and 0.43 coupling events/min in REM. We also agree with you that the remaining detections in N1 and REM should not be interpreted as physiological equivalents of canonical N2/3 SOs, spindles, or SO-spindle complexes. We now state this explicitly in the manuscript. They are reported only as descriptive detector outputs under the fixed N2/3-derived criterion. And we have revised all relevant sections of the manuscript, including “[Results, Page 6-7 Lines 134-148]; [Fig. 1e]; [Results, Page 9 Lines 175-191]; [Fig. 2b]; [Methods, Page 25-27, Lines 567-604]; [Fig. S2-S4]; [Table S2, S4].”

      We also would like to clarify our sleep staging procedure. The sleep staging was first performed using an established automated algorithm, YASA toolkit (Vallat & Walker, 2021), and then manually reviewed by two sleep experts. More importantly for the central results, all EEG-informed fMRI GLM and PPI analyses were restricted to N2/3 sleep. Thus, N1 and REM detections did not enter the event regressors used for the main fMRI analyses and do not affect the interpretation of the hippocampal or thalamic findings.

      Finally, we agree that the absolute number of SO-spindle coupling events depends on the detection threshold. We therefore tested whether the main fMRI result depended on the specific SO threshold. Hippocampal activation during SO-spindle coupling remained significant when the SO detection threshold was varied between the 71st and 80th percentiles, as shown in Fig. S6. We have made this clearer in the revised manuscript.

      Results, Page 6-7 Lines 134-148

      “Each sleep stage is characterised by distinct spectral properties and rhythmic waveforms, serving as physiological markers (Fig. 1c). Because SO and spindle detection relies on amplitude-based percentile thresholds, we avoided estimating separate thresholds within each sleep stage. Instead, for each participant, the SO and spindle thresholds were defined from N2/3 sleep only, where these rhythms are most abundant and physiologically expected, and the same fixed thresholds were then applied to N1 and REM for descriptive comparison.”

      “Under this fixed N2/3-derived thresholding, detected SOs and spindles were larger and more frequent in N2/3 than in N1 or REM. SO and spindle amplitudes were significantly higher during N2/3 sleep (SO: 25.59 ± 1.49 μV; spindle: 7.39 ± 0.27 μV) than during N1 (SO: 20.15 ± 2.32 μV; spindle: 5.23 ± 0.27 μV) and REM sleep (SO: 19.84 ± 1.22 μV; spindle: 5.60 ± 0.22 μV; all p < 1e-4; Fig. 1e, Fig. S2). The corresponding event densities showed the same pattern, with 9.64 ± 0.25 SOs/min and 4.19 ± 0.10 spindles/min in N2/3, compared with 2.95 ± 0.16 SOs/min and 2.71 ± 0.14 spindles/min in N1, and 2.07 ± 0.17 SOs/min and 1.81 ± 0.14 spindles/min in REM (all p < 1e-4). We therefore report detections in N1 and REM only as descriptive outputs of the detector under a fixed N2/3-derived criterion, rather than as physiological equivalents of canonical N2/3 SOs or spindles.”

      Fig. 1 legend, Page 8, Line 166-172

      “e, Amplitudes (μV) of detected SOs (left) and spindles (right) across sleep stages. SO and spindle detection thresholds were defined from N2/3 sleep within each participant and then applied unchanged to N1 and REM for descriptive comparison. Detections in N1 and REM should therefore be interpreted as detector outputs under this fixed N2/3-derived criterion. The SO amplitudes were measured from the 0.16-1.25 Hz filtered EEG data, and spindle amplitudes were measured from the 12-16 Hz filtered EEG data. Each dot represents an individual participant. Error bars indicate SEM. *** p < 0.001.”

      Results, Page 9 Lines 175-191

      “SO-spindle coupling is considered important for sleep-dependent memory consolidation. In the current study, using the same N2/3-derived detection thresholds described above, we found that SO-spindle coupling occurred most frequently during N2/3 sleep (2.46 ± 0.06 events/min). Coupling density was significantly lower in N1 (0.75 ± 0.05 events/min, t<sub>(106)</sub> = 23.54, p < 1e-4) and REM sleep (0.43 ± 0.04 events/min, t<sub>(106)</sub> = 31.24, p < 1e-4; Fig. 2b, Table S2-S4), consistent with the expected predominance of SO-spindle coupling in NREM sleep (Ngo et al., 2013; Staresina et al., 2015). As with the individual SO and spindle detections, coupling events detected in N1 and REM were retained only for descriptive stage-wise reporting (see Table S2, S4). They were not used to support physiological claims about SO-spindle coupling in these stages, and they were not entered into the EEG-informed fMRI analyses. All subsequent fMRI GLM and PPI analyses were restricted to N2/3 sleep.”

      “After extracting all N2/3 EEG epochs in which SO-spindle coupling occurred, we analysed their spectral and phase characteristics. The spindles were most likely to occur slightly before the UP-state peak of SOs (Fig. 2a, e), aligning with results from both animal studies (Maingret et al., 2016) and human research (Staresina et al., 2015). In our data, this pattern was consistent across subjects (Fig. 2d, Rayleigh test: z = 9.51, p < 1e-4), with the peak of the spindle aligned at an SO phase of −41.61 ± 0.86° (the SO UP-state peak is 0°).”

      Fig. 2 legend, Page 10, Line 202-205

      “b, SO-spindle coupling density across sleep stages, using SO and spindle detections obtained with fixed N2/3-derived thresholds. Coupling events in N1 and REM are shown only for descriptive comparison. The EEG-informed fMRI analyses used N2/3 coupling events only.”

      Results, Page 11-12, Lines 242-247

      “Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st - 80th percentile; Fig. S6).”

      Methods, Page 25-26, Lines 567-575

      “Detection of SOs. Data were first bandpass-filtered between 0.16 and 1.25 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). After identifying all positive-to-negative zero crossings, potential SOs were defined based on the interval between consecutive zero crossings, ranging from 0.8 s to 3 s. For each potential SO, we calculated the amplitude range as the peak minus the trough. For each participant, the amplitude threshold was defined as the 75th percentile of candidate SO amplitude ranges observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. Only candidates exceeding this threshold were labelled as SOs, following previous work (Schreiner et al., 2021).”

      Methods, Page 26, Lines 576-583

      “Detection of sleep spindles. Detection of sleep spindles. Data were bandpass-filtered between 12 and 16 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). The root mean square (RMS) of the filtered signal was computed with a 200 ms sliding time window. For each participant, the spindle threshold was defined as the 75th percentile of RMS values observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. RMS segments exceeding this threshold for 0.5 s to 3 s were identified as spindles (Staresina et al., 2015).”

      Methods, Page 26, Lines 584-591

      “Detection of SO-spindle couplings. From the detected SOs and spindles, we identified the peak time of each spindle. Within each SO interval, we checked whether a spindle peak occurred; if so, that SO was labelled as an SO-spindle coupling event. For descriptive stage-wise summaries, coupling events were assigned to the sleep stage of the corresponding SO trough. For every SO-spindle coupling event, an epoch was created time-locked to the SO trough as the central reference, following Schreiner et al. (2021). We extracted data in a [−4 s to 4 s] window around this point, forming the epoch for each coupling event. For the EEG-informed fMRI analyses, only SO, spindle and SO-spindle coupling events detected during N2/3 sleep were used.”

      Methods, Page 26-27, Lines 592-604

      “The detection procedures described above were developed primarily for N2 and N3 sleep, where SOs, spindles and their coupling are physiologically expected and most reliably observed (Hahn et al., 2020; Helfrich et al., 2019; Helfrich et al., 2018; Ngo, Fell, & Staresina, 2020; Schreiner et al., 2022; Schreiner et al., 2021; Staresina et al., 2015; Staresina et al., 2023). Because percentile-based thresholds can otherwise force the detector to label events in every sleep stage, we did not estimate separate thresholds within N1 or REM. Instead, for each participant, all SO and spindle thresholds were defined from N2/3 sleep and then applied uniformly across the recording. Tables S1 and S3 report detailed statistical information on sleep rhythm and N2/3 events detection. The N1 and REM events detection reported in Tables S2 and S4, and illustrated in Fig. S2-S4, should therefore be interpreted as descriptive detector outputs under this fixed N2/3-derived criterion, rather than as evidence for canonical N2/3 SOs, spindles or physiological SO-spindle complexes in those stages. These detections were not used in the EEG-informed fMRI GLM or PPI analyses, which were restricted to N2/3 sleep.”

      Reviewer #3 (Public review):

      Summary:

      Wang et al., examined the brain activity patterns during sleep, especially when locked to those canonical sleep rhythms such as SO, spindle, and their coupling. Analyzing data from a large sample, the authors found significant coupling between spindles and SOs, particularly during the up-state of the SO. Moreover, the authors examined the patterns of whole-brain activity locked to these sleep rhythms. The authors next investigated the functional connectivity analyses, and found enhanced connectivity between the hippocampus and the thalamus and the medial PFC. These results reinforced the theoretical model of sleep-dependent memory consolidation, such that SO-spindle coupling is conducive for systems-level memory reactivation and consolidation.

      Strengths:

      There are obvious strengths in this work, including the large sample size, state-of-the-art neuroimaging and neural oscillation analyses, and the richness of results. The results now inform hemodynamic neural activity that coincided with SO-spindle couplings.

      Weaknesses:

      My earlier comments were about the inability to make inferences on memory given the lack of memory tasks, and the weakness in using the open-ended cognitive state decoding.

      Comments on revisions:

      The current revision has addressed these major concerns. The authors expanded discussions regarding the theoretical implications of the work in a more nuanced manner.

      Thank you for taking the time to re-evaluate the manuscript. We are pleased that the revised Discussion now reads as more nuanced, especially in relation to the limits of the memory-related interpretation. Your earlier comments helped us sharpen both the claims and the framing, and we are grateful for that.

      References:

      Bergmann, T. O., Mölle, M., Diedrichs, J., Born, J., & Siebner, H. R. (2012). Sleep spindle-related reactivation of category-specific cortical regions after learning face-scene associations. Neuroimage, 59(3), 2733-2742.

      Bodurka, J., Ye, F., Petridou, N., Murphy, K., & Bandettini, P. A. (2007). Mapping the MRI voxel volume in which thermal noise matches physiological noise—implications for fMRI. Neuroimage, 34(2), 542-549.

      Caporro, M., Haneef, Z., Yeh, H. J., Lenartowicz, A., Buttinelli, C., Parvizi, J., & Stern, J. M. (2012). Functional MRI of sleep spindles and K-complexes. Clinical neurophysiology, 123(2), 303-309.

      Czisch, M., Wehrle, R., Stiegler, A., Peters, H., Andrade, K., Holsboer, F., & Sämann, P. G. (2009). Acoustic oddball during NREM sleep: a combined EEG/fMRI study. PloS one, 4(8), e6749.

      Fogel, S., Albouy, G., King, B. R., Lungu, O., Vien, C., Bore, A., Pinsard, B., Benali, H., Carrier, J., & Doyon, J. (2017). Reactivation or transformation? Motor memory consolidation associated with cerebral activation time-locked to sleep spindles. PloS one, 12(4), e0174755.

      Hahn, M. A., Heib, D., Schabus, M., Hoedlmoser, K., & Helfrich, R. F. (2020). Slow oscillation-spindle coupling predicts enhanced memory formation from childhood to adolescence. Elife, 9, e53730.

      Hale, J. R., White, T. P., Mayhew, S. D., Wilson, R. S., Rollings, D. T., Khalsa, S., Arvanitis, T. N., & Bagshaw, A. P. (2016). Altered thalamocortical and intra-thalamic functional connectivity during light sleep compared with wake. Neuroimage, 125, 657-667.

      Helfrich, R. F., Lendner, J. D., Mander, B. A., Guillen, H., Paff, M., Mnatsakanyan, L., Vadera, S., Walker, M. P., Lin, J. J., & Knight, R. T. (2019). Bidirectional prefrontal-hippocampal dynamics organize information transfer during sleep in humans. Nature Communications, 10(1), 3572.

      Helfrich, R. F., Mander, B. A., Jagust, W. J., Knight, R. T., & Walker, M. P. (2018). Old brains come uncoupled in sleep: slow wave-spindle synchrony, brain atrophy, and forgetting. Neuron, 97(1), 221-230. e224.

      Huang, Q., Xiao, Z., Yu, Q., Luo, Y., Xu, J., Qu, Y., Dolan, R., Behrens, T., & Liu, Y. (2024). Replay-triggered brain-wide activation in humans. Nature Communications, 15(1), 7185.

      Ilhan-Bayrakcı, M., Cabral-Calderin, Y., Bergmann, T. O., Tüscher, O., & Stroh, A. (2022). Individual slow wave events give rise to macroscopic fMRI signatures and drive the strength of the BOLD signal in human resting-state EEG-fMRI recordings. Cerebral Cortex, 32(21), 4782-4796.

      Laufs, H., Daunizeau, J., Carmichael, D. W., & Kleinschmidt, A. (2008). Recent advances in recording electrophysiological data simultaneously with magnetic resonance imaging. Neuroimage, 40(2), 515-528.

      Maingret, N., Girardeau, G., Todorova, R., Goutierre, M., & Zugaro, M. (2016). Hippocampo-cortical coupling mediates memory consolidation during sleep. Nature Neuroscience, 19(7), 959-964.

      Moehlman, T. M., de Zwart, J. A., Chappel-Farley, M. G., Liu, X., McClain, I. B., Chang, C., Mandelkow, H., Özbay, P. S., Johnson, N. L., & Bieber, R. E. (2019). All-night functional magnetic resonance imaging sleep studies. Journal of neuroscience methods, 316, 83-98.

      Ngo, H.-V., Fell, J., & Staresina, B. (2020). Sleep spindles mediate hippocampal-neocortical coupling during long-duration ripples. Elife, 9, e57011.

      Ngo, H. V., Martinetz, T., Born, J., & Molle, M. (2013). Auditory closed-loop stimulation of the sleep slow oscillation enhances memory. Neuron, 78(3), 545-553.

      Picchioni, D., Horovitz, S. G., Fukunaga, M., Carr, W. S., Meltzer, J. A., Balkin, T. J., Duyn, J. H., & Braun, A. R. (2011). Infraslow EEG oscillations organize large-scale cortical–subcortical interactions during sleep: a combined EEG/fMRI study. Brain research, 1374, 63-72.

      Schabus, M., Dang-Vu, T. T., Albouy, G., Balteau, E., Boly, M., Carrier, J., Darsaud, A., Degueldre, C., Desseilles, M., & Gais, S. (2007). Hemodynamic cerebral correlates of sleep spindles during human non-rapid eye movement sleep. Proceedings of the National Academy of Sciences, 104(32), 13164-13169.

      Schreiner, T., Kaufmann, E., Noachtar, S., Mehrkens, J.-H., & Staudigl, T. (2022). The human thalamus orchestrates neocortical oscillations during NREM sleep. Nature Communications, 13(1), 5231.

      Schreiner, T., Petzka, M., Staudigl, T., & Staresina, B. P. (2021). Endogenous memory reactivation during sleep in humans is clocked by slow oscillation-spindle complexes. Nature Communications, 12(1), 3112.

      Staresina, B. P., Bergmann, T. O., Bonnefond, M., van der Meij, R., Jensen, O., Deuker, L., Elger, C. E., Axmacher, N., & Fell, J. (2015). Hierarchical nesting of slow oscillations, spindles and ripples in the human hippocampus during sleep. Nature Neuroscience, 18(11), 1679-1686.

      Staresina, B. P., Niediek, J., Borger, V., Surges, R., & Mormann, F. (2023). How coupled slow oscillations, spindles and ripples coordinate neuronal processing and communication during human sleep. Nature Neuroscience, 1-9.

      Vallat, R., & Walker, M. P. (2021). An open-source, high-performance tool for automated sleep staging. Elife, 10.

    1. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This valuable study reports that the ALDH-abundant cells display stem cell properties and may play a key role in the endometrial epithelial development in the mouse. The data supporting the main conclusion are solid, although further improvements are needed to strengthen the conclusions. This work will be of great interest to reproductive biologists and biomedical researchers working on women's reproductive health.

      We thank the reviewers and editor for their critical reading and assessment of our manuscript. We carefully considered each of the points raised by the reviewers. In this document and in the edited manuscript and figures, we have carefully addressed each of the comments and requested modifications. In light of these changes, we expect that you will find that the manuscript has improved.

      We indicate our responses to the reviewers below in blue font and highlight the changes in the manuscript using the line numbers corresponding to the tracked version of the revised document.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Tang et al. characterizes the expression dynamics and functional roles of aldehyde dehydrogenase 1 activity in uterine physiology. Using a combination of in vivo lineage tracing and cell ablation coupled with organoid culture, the authors propose that Aldh1a1 lineage-marked cells contribute to uterine gland development and epithelial regeneration. The descriptive data will be of interest to reproductive biologists and clinicians and will build on established hypotheses in the field. The manuscript is well written and scientifically sound; however, several experimental limitations and interpretation caveats should be addressed.

      We thank the reviewer for their comments and expert assessment of our paper.

      (1) The methods surrounding the passage number and duration of culture following sorting prior to transcriptomic profiling should be clarified in the figure legends. Related to this, the representative images in Figures 1D and 1E do not appear consistent with the quantification presented in Figures 1F-H and should be reconciled.

      Thanks for this comment. We have now clarified this in the Figure 1 legend as follows,

      LINES 1026-1029: “Organoid formation assay performed immediately after luminal epithelial cell isolation and by plating equal numbers of viable ALDH<sup>LO</sup> (D) and ALDH<sup>HI</sup> (E) epithelial cells. ALDH<sup>LO</sup> and ALDH<sup>HI</sup> organoids were cultured for two weeks and passaged once prior to the organoid formation assays and transcriptomic analyses.”

      Regarding the second comment, we recognize that the images we showed may not have been the most representative of our quantification. As such, we replaced them with the organoid images so that they better reflect the quantification outlined in Figure 1F-H.

      (2) The conclusion that ALDH1A1+ cells are enriched in populations with stem cell characteristics relies primarily on transcriptomic analysis. Protein-level co-localization should be performed to strengthen this claim.

      We thank the reviewer for this comment. Unfortunately, the antibodies for many of these stem cell markers (such as LGR5, AXIN2, and SUSD2) are not well-suited for immunostaining. Others that have been proposed in human and are amenable to immunostaining are not suitable markers for mouse endometrial stem cells (such as CDH2). We hope that by showing that ALDH1A1 is expressed in patterns that are similar to the previously published stem cell markers LGR5 and AXIN2 (i.e., throughout the epithelium in the developing uterus and subsequently enriched in the tips of the endometrial glands of adult mice), along with transcriptomic studies, we can demonstrate its utility as a marker for mouse endometrial stem cells.

      (3) The overlap of 19 genes between the data set here and AXIN2 HI data is presented as evidence of shared stemness identity, but no statistical assessment of this overlap is provided. A hypergeometric test should be performed to determine whether this overlap is greater than expected by chance.

      Thank you for this suggestion. We have performed a hypergeometric test and determined that the reported shared genes between the two datasets are greater than is expected by chance. We have updated the results section to state the following:

      Lines 137-140: "We determined that the overlap between ALDH<sup>HI</sup> and Axin2<sup>+</sup> stemness marker genes was significantly greater than expected by chance for both upregulated (21/346 genes, 1.81-fold enrichment, p = 0.0067) and downregulated (19/674 genes, 1.67-fold enrichment, p = 0.021) gene sets (hypergeometric test, universe = 23,182 genes)."

      (4) The impact of tamoxifen injection on Aldh1a1 expression should be characterized in the neonatal uterus, as tamoxifen itself has known estrogenic activity that could confound interpretation of the lineage tracing results at early postnatal timepoints.

      Although we took measures to control for this possibility by using multiple time-points and models to trace the impact of Aldh1a1<sup>+</sup> cells in development and adulthood, we recognize the importance of this comment and acknowledge that this is a limitation in the design of our study. We have included the following text to the Discussion acknowledging this point:

      Lines 433-441: “Given the well-documented impacts of tamoxifen for lineage tracing studies, it is imperative to use doses of tamoxifen that will minimize estrogenic impacts and result in off-target effects (Rios et al., 2016). This often requires administration at doses that will achieve maximal recombination of the desired gene, while ensuring that the potential deleterious impacts of tamoxifen are minimized (Chen et al., 2023; Pimeisl et al., 2013). The cre/ERT2 tamoxifen inducible model is widely used to study uterine biology where it serves as a useful tool to interrogate the spatiotemporal impact of key genes, either through inactivation or for lineage tracing. Despite its widely documented utility across many tissue types and developmental timepoints, the use of tamoxifen and its impacts on the endometrium remain a limitation of our study, which we tried to address by implementing multiple timepoints, doses, and orthogonal assays in our experimental design.”

      (4b) Related to this, while low-dose tamoxifen is shown to label individual cells within 24 hours of injection, the translation dynamics of the label following Cre-mediated recombination can require up to 72 hours. The presence of only a few labeled clones at PND8 but multiple separate clones per cross-section at later timepoints warrants discussion and may reflect labeling kinetics rather than clonal expansion.

      The reviewer raises an important point. We agree that the 72hr-translation kinetics of the cre-mediated recombination is a legitimate consideration for interpreting our data and we have added the text below to the Discussion section acknowledging this point.

      We have addressed this by adding the following text to the discussion:

      Lines 417-422: We hypothesized that the singly labeled cells observed from one day tracing experiments expanded in a clonal fashion during the various timepoints we measured. We note that the translation kinetics of the labeled cells following cre-mediated recombination may contribute to the limited labeling observed at PND8/PND15 and there is a potential for delayed labeling of cells between 24 and 72 hours of tamoxifen administration. However, the continuous increase in labeled cells at the subsequent timepoints favors our interpretation of clonal expansion as the primary explanation.

      (5) It would strengthen the in vivo ablation data to validate the degree of cell death following diphtheria toxin treatment directly. It is possible that a general decrease in cell number rather than specific loss of a stem cell population is responsible for the observed reduction in gland number and FOXA2 expression (Tongtong et al 2017).

      We agree that this is an important control to incorporate into our experimental design. To rule out this possibility, we performed immunohistochemistry of cleaved caspase 3 in the uterine tissues of DTR<sup>flox/flox</sup> and DTR<sup>flox/flox</sup>;Aldh1a1<sup>cre/ERT2</sup> mice 4 days after administration of diphtheria toxin. The results indicate similar levels of cleaved caspase 3 detection in both genotypes, suggesting that the decrease in FOXA2+ cells is not due to non-specific cell death, but rather the result of ALDH1A1<sup>+</sup> cells. These data and the following text have been added to the manuscript:

      Lines 320-324: “We determined that the decreased in FOXA2<sup>+</sup> cells in the experimental mice was not the result of non-specific DT-mediated cell death, as similar levels of cleaved caspase 3-positive cells were detected in the DT-treated control ROSA26<sup>DTR/DTR</sup> and ROSA26<sup>DTR/DTR</sup>;Aldh1a1<sup>cre/ERT2/+</sup> mice 4 days post-diphtheria toxin administration (Figure S3G-H’).”

      (6) The lineage tracing data in the postpartum endometrium demonstrate that Aldh1a1-marked cells are present during regeneration, but it remains unclear whether these cells are preferentially activated or expanded in response to tissue injury. Coupling these studies with diphtheria toxin-mediated ablation during active regeneration would more directly test the proposed regenerative role of this population.

      This is a great point and one that we would be very interested in pursuing as follow-up studies in our future work. Regretfully, due to the long generation time and experimental procedures associated with these proposed studies, we are not able to include these experiments in the current manuscript. Thus, we have changed our wording and conclusions throughout the manuscript to be less definitive in terms of the role of Aldh1a1 in regeneration, since this will be the focus of future studies.

      The contribution of stromal Aldh1a1 lineage-positive cells is underexplored in the discussion, given the lineage tracing data showing stromal labeling across multiple timepoints and its potential relevance to mesenchymal-to-epithelial transition.

      Thank you for the suggestion. We have now expanded this section in the Discussion to include the following:

      Lines 496-504: We also found ALDH1A1<sup>+</sup> stromal cells were more prevalent when tracing began in adult mice. Other studies have shown that mesenchymal cells contribute to endometrial regeneration in the postpartum phase or after induced menses through a process of MET (Cousins et al., 2014; Kirkwood et al., 2022; Li et al., 2025). Similarly, lineage tracing studies have shown that MET is an active process and contributes to epithelial cell regeneration in the post-partum phase (Huang et al., 2012; Patterson et al., 2013). Although this is an area of active investigation in the field, with some contradicting reports, it is plausible to hypothesize that endometrial tissue has the capacity to undergo wound-healing and regeneration via several mechanisms (Ang et al., 2023; Ghosh et al., 2020). The process of MET in wound healing is widely documented in other organs, such as the kidney, liver and lung, where MET is associated with depletion of the resident epithelial cell pool (Bi et al., 2012; Niayesh-Mehr et al., 2024; Zeisberg et al., 2005).

      Finally, the word 'control' may overstate the functional evidence presented. 'Contribute' may be more accurate given the partial and context-dependent nature of the phenotypes observed.

      We agree with the reviewer’s point that control may overstate the evidence that we provide in the manuscript. To reflect this, we have edited the manuscript title and text to address this suggestion.

      Reviewer #2 (Public review):

      Tang et al. investigated the contribution of Aldh1a1+ cells, as putative stem/progenitor cells, to endometrial development, maintenance during the estrous cycle, and postpartum repair in mouse models. They employed in vitro organoid formation and in vivo lineage tracing models coupled with RNA-seq to test the stem-ness of Aldh1a1+ cells. They found that mouse endometrial cells with high ALDH activity (using the ALDEFLUOR assay) formed more and larger organoids and were enriched for stem/progenitor cell gene signatures. Similar results were shown using endometrial cells from a human patient sample. Epithelial ALDH1A1 expression was shown to be hormonally regulated, becoming more restricted to the glands, a putative epithelial stem cell niche, under estrogen stimulation. Using lineage-tracing initiated postnatally/prepubertally, Aldh1a1+ epithelial cells were shown to expand, contributing to both the luminal and glandular epithelium into adulthood, whereas adult initiation of labeling showed expansion of stromal Aldh1a1+ cells but not epithelial. Postnatal ablation of single-labeled Aldh1a1+ epithelial cells resulted in impaired gland development. Lastly, Aldh1a1-lineage traced cells (adult labeled) were present during postpartum endometrial repair as were epithelial/mesenchymal transitional cells.

      This study addresses an important area of research in the field of endometrial stem/progenitor cell biology. The authors are commended for their use of multiple complementary methods, including lineage tracing, DTR-mediated cell ablation, organoid assays, and RNA-seq in mouse and human models to assess the stem-like nature of Aldh1a1+ cells. The data support the stem/progenitor phenotype of Aldh1a1+ epithelial cells during endometrial development; however, there are noted discrepancies between organoid formation assays and lineage tracing experiments regarding the stemness of Aldh1a1+ epithelial cells in adults. Specifically, organoids were generated from adult cells and demonstrated in vitro stem cell activity; however, in vivo lineage-tracing of adult cells either during the estrous cycle or postpartum repair does not show expansion of Aldh1a1+ cells, suggesting they do not have stem/progenitor activity. Additionally, the stem-ness of epithelial vs stromal Aldh1a1+ cells is confounded in the study because epithelial cells were not purified for organoid experiments, epithelial cells were not exclusively lineage-traced as stromal cells were also labeled, and mesenchymal-epithelial transition was suggested to occur during postpartum repair. The following specific comments are presented to detail these concerns:

      We thank the reviewer for their critical reading of our manuscript and constructive comments.

      (1) The statement in the brief summary, "...critical for lifelong endometrial regeneration," is not supported by the data provided.

      We have edited the brief summary to exclude this statement, it now reads as follows:

      Lines 4-5: “We uncover ALDH1A1<sup>+</sup> cells as a group of hormone sensitive stem cells contributing to endometrial development and regeneration.”

      (2) AlDH1A1 is not restricted to the endometrial epithelium, and epithelial cells were not purified by flow cytometry for experiments in Figure 1. Figure 2 clearly shows the presence of mesenchymal cells, even using the described method for enriching for epithelial cells. Therefore, contaminating mesenchymal cells with high ALDH activity may confound the experimental results in Figure 1, either through promoting epithelial cell growth or through MET. The authors should provide clear evidence of epithelial purity in organoid experiments or that mesenchymal cells are not contained in the ALDHhi population. These comments also apply to the human organoid experiments in Figure 7.

      We thank the reviewer for raising this important point. Our group has been using the enzymatic method to routinely separate epithelial from stromal cell populations from the mouse uterus (see references dating back to 2015, PMID 26721398, 28324064, 34099644). In these experiments we typically obtain >98% purity in the epithelial and stromal cell compartments, respectively. We can directly observe this purity in the immunofluorescence images shown I Author response image 1 and Author response image 2, where mouse endometrial epithelial cells and stromal cells were enzymatically separated and immunostained with E-cadherin and vimentin antibodies to detect epithelial and mesenchymal cells in both cell preparations. The images show very few contaminating epithelial and stromal cells in either cell preparation. We have observed similar results when preparing epithelial and stromal cell preparation from the human endometrium, where the epithelial cell organoids display high purity with ~100% epithelial cell expression when we perform immunostaining.

      Author response image 1.

      Purity of mouse endometrial epithelial cells obtained via enzymatic and mechanical dissociation. A-B) Shows the epithelial (A) and stromal (B) cells plated on glass coverslips and immunostained with an epithelial cell marker (cytokeratin 8, red), a stromal cell marker (vimentin, green), and DAPI.

      Author response image 2.

      Human endometrial epithelial organoids were fixed and immunostained with cytokeratin 8 (green) and DAPI. The images are typical for our epithelial cell cultures and demonstrate that all epithelial cells are CK8-positive.

      (3) Lines 186-187: Susd2 was increased in EpSC clusters, yet this is a mesenchymal stem/progenitor marker in humans. The authors should discuss the implications of this.

      We thank the reviewer for highlighting this. We have now included the following in our Discussion to address this point:

      Lines 527-532: Clustering with this population of EpSCs were Susd2<sup>+</sup> cells, which are well-characterized mesenchymal progenitors that are enriched in the perivascular regions of the human endometrium (Darzi et al., 2016; Khanmohammadi et al., 2021). The presence of Susd2<sup>+</sup> cells, while unexpected in an epithelial stem cell niche, could indicate the presence of a transitional mesenchymal or perivascular cell that is differentiating into epithelium. Evidence for both mesenchymal and Nestin2<sup>+</sup> pericytes have been recently described in the mouse endometrial epithelium (Kirkwood et al., 2022; Li et al., 2025).

      (4) In Figure 5, RFP+ epithelial cells should be quantified as in previous figures to substantiate the statement in lines 279-280, "At PPD5, the proportion of RFP+ epithelial cells had expanded relative to PPD1 and PPD3 (Figure 5E-E')." Especially because in the low mag images (C-E), RFP+ epithelial cells appear to be most abundant at PPD1 and decrease at PPD3 and PPD5, suggesting that they may not be involved in endometrial regeneration/repair (contradicting the interpretation in line 285). Further, if there is in fact a decrease over postpartum repair, then regeneration should be removed from the title of the manuscript. RFP+ stromal cells should also be quantified.

      We appreciate this reviewer’s comment and agree that as stated, the conclusion is not fully supported by the data. To address this comment, we have edited the results so that they clearly indicate the results and remove any ambiguity:

      As requested, we quantified the number of RFP+ stromal and epithelial cells during the postpartum phase and noted that RFP+ cells were prominent in the stromal compartment of the endometrium. While RFP+ epithelial were also observed during these timepoints, they were less abundant than RFP+ stromal cells. Because the number of RFP+ cells did not significantly change over the postpartum phases in neither the stromal nor epithelial compartment, we have modified our conclusion to state that ALDH1A1+ cells are transiently detected in the regenerating endometrium.

      Results:

      Lines 287-294: “By analyzing the uterine tissues near the placental detachment site, we observed that RFP positive cells were prominent in the endometrial stromal cells that were adjacent to the luminal epithelium (Figure 5C-C’, green arrows). RFP<sup>+</sup> cells were also observed in the stromal cells near the placental detachment sites at PPD1 and PPD3 (Figure 5D’-E’, red & blue arrows) and in limited luminal epithelial cells (Figure 5D”,E”). Quantification of RFP+ cells throughout these postpartum phases indicated that stromal cells had more frequent ALDH1A1<sup>+</sup> stromal cells (360 ± 103, PPD1, n=3; 217 ± 107, PPD3, n=3; 254 ± 32, PPD5, n=4) than ALDH1A1<sup>+</sup> epithelial cells in the regenerating endometrium (65 ± 65, PPD1, n=3; 20 ± 10, PPD3, n=3; 114.25 ± 39, PPD5, n=4) (Figure S4).”

      Discussion:

      Lines 512-520: “We also noted that a majority of ALDH1A1<sup>+</sup> cells were localized to the active areas of endometrial regeneration near the placental detachment sites at PPD1 with a pronounced expression in the sub-epithelial stromal cells. As regeneration progressed, we continued to observe ALDH1A1<sup>+</sup> cells in the stromal compartment within the placental detachment sites at PPD3 and PPD5, with a progressive, but not statistically significant, increase in ALDH1A1<sup>+</sup> epithelial cells. Collectively, our data demonstrate that ALDH1A1<sup>+</sup> lineage cells participate in the restoration of endometrial architecture and functional compartments in the postpartum phase, even if their direct contribution is transient. Future detailed and mechanistic studies will be necessary to fully characterize their role in this process and their long-term consequence in postpartum regeneration.”

      (5) For Figure 7F, it should be clearly stated in the main text that the results are from one patient sample and the data presented are experimental replicates, so as not to be confused with biological replicates (the same for Supplementary Figure S4). Were B and G in Figure 7 also from one patient?

      Thanks for pointing this out. We have edited the figure legends in the main text and supplemental figures to indicate this.

      Lines 336-337: “…main figures show representative results from one patient sample performed in technical replicates, with additional patient samples included in the supplement…”

      (6) Lines 425-427: "Ovariectomized mice treated with 90-day E2 pellets, on the other hand, showed a complete restriction of ALDH1A1 to the glandular crypts." In Figure 2 S' ALDH1A1+ cells are visible in the LE (the staining is lighter than in the GE but looks real), contradicting this statement.

      This is an important distinction. We have now edited this part of the manuscript to state:

      Lines 458-461: “Ovariectomized mice treated with 90-day E2 pellets, on the other hand, showed enriched ALDH1A1 in the glandular crypts with weak luminal epithelial staining, while the ovariectomized controls had strong ALDH1A1 expression throughout the luminal and glandular epithelium.”

      (7) Lines 466-467: "In cycling mice, we found sporadic cells that expressed both stromal and epithelial markers in the ALDHA1+ cells." These data are not presented.

      We apologize for the confusion, this sentence has been removed from the discussion.

      (8) These data support the role of Aldh1a1+ cells in endometrial epithelial development, but conclusions about their role in repair/regeneration should be tempered as the data are much weaker here.

      We thank the reviewer for their overall assessment. To address this point, we have thoroughly edited the appropriate areas to temper the conclusions and ensure that they are strongly supported by our data. We have also edited the manuscript’s title to reflect this.

      Reviewer #3 (Public review):

      Summary:

      Tan et al demonstrated the importance of ALDH-high cells in the epithelial development in the mouse endometrium, and these cells displayed properties of stem cells.

      We thank the reviewer for their assessment of our manuscript.

      Strengths:

      The findings are solid, supported and validated through a combination of technical methods. I appreciated this combined use of mouse and human endometrial cells to strengthen the findings. Genomic results from a single-cell sequencing dataset were informative as they depicted the different stages of the estrus cycle during the regeneration process. Verification with immunostainings with various markers made it convincing for readers to visualize the cell's location, progression, and status at different timepoints. Utilizing human endometrial cells further demonstrated that the phenomenon observed in mice can be translated to humans.

      This work will greatly advance the understanding of endometrial regeneration for reproductive biologists.

      We thank the reviewer for their expert assessment and positive comments regarding our manuscript.

      Weaknesses:

      No major weaknesses were identified by this reviewer.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) As this study evaluated Aldh1a1+ cells in both the epithelium and stroma, it is recommended that the title and abstract be revised to reflect this.

      Thank you. Both the title and abstract title have been updated to reflect this comment.

      (2) Lines 46-47 in the abstract: "Aldh1a1+ cells expanded during postnatal development, estrus cycling, and following post-partum repair." It is recommended to clarify stromal vs epithelial Aldh1a1+ cell expansion, as only the stromal cells expanded during the estrous cycle. Also, RFP+ epithelial cells were not quantified during postpartum repair and visually appear to decrease (see comment below regarding Figure 5), so this statement is misleading.

      The abstract was edited following the suggestions so that it depicts our results. Similarly, we have addressed the comment regarding Figure 5 and the interpretation of the post-partum regeneration experiments (see comment above for the full explanation and edits to the manuscript).

      (3) Lines 65-69, 186-187: the authors should be clear when describing markers of putative epithelial vs mesenchymal stem/progenitor cells in the introduction (e.g., CDH2/SSEA1/SOX9 for epithelial and SUSD2 for mesenchymal).

      Thank you for this suggestion. The Introduction section now states the following:

      Lines 70-73: “These markers identify specific cell types in the endometrium, with CDH2<sup>+</sup>, SSEA1<sup>+</sup>, and SOX9<sup>+</sup> cells corresponding to endometrial epithelial cells, while SUSD2<sup>+</sup> cells corresponding to mesenchymal endometrial cells enriched in the perivascular regions (Cousins et al., 2021).”

      (4) Lines 78-80: "Studies tracing the fate, ablation, and proliferative capacity of Lgr5+ cells in the uterus identified an Lgr5+ niche that is enriched in the crypts of the glandular epithelium and promotes endometrial regeneration (Seishima et al., 2019)." This statement is incorrect regarding regeneration, as Lgr5 marks stem/progenitor cells in the developing uterus but not the adult, during which endometrial regeneration occurs. Please revise.

      Thank for you for this clarification. The statement has been revised and now reads as follows:

      Lines 82-84: “Studies tracing the fate, ablation, and proliferative capacity of Lgr5<sup>+</sup> cells in the uterus identified an Lgr5<sup>+</sup> niche that marked stem/progenitor cells in the developing uterus (Seishima et al., 2019).”

      (5) Recommend using free-form shapes to outline GE, LE, and EpSC in Fig 2C and stating in the text which clusters correspond to LE and GE (lines 167-169).

      Thank you, this has been edited in Figure 2C and in the text, which now reads as follows:

      Lines 176-177: “Clusters 5, 7, 21 and 16 were classified as glandular epithelial cells, and clusters 0, 2, 3, 10, 14, and 24 were classified as luminal epithelial cells (Figure 2C).”

      (6) "Estrus" refers to the specific stage of the "estrous" cycle. Estrus cycle is incorrect and should be estrous cycle.

      Thank you for pointing this out. It has been corrected throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Suggest increasing the font size for some of the labels in the figures.

      Thank you, we have increased font size in the figures to improve the quality.

      (2) Need to include more details of the human endometrial tissue used in this study: pathology, age, and stage of menstrual cycle.

      We thank the reviewer for the helpful suggestion. The details that are available to us have been included in Supplementary Table S5.

      (3) Line 65-67 - Different endometrial stem cell subsets - SUSD2+ reside in perivascular regions, while other markers are located in glandular epithelium, need revision.

      Thank you, this has been revised in the Introduction. The area now reads as follows:

      Lines 70-73: These markers identify specific cell types in the endometrium, with CDH2<sup>+</sup>, SSEA1<sup>+</sup>, and SOX9<sup>+</sup> cells corresponding to endometrial epithelial cells, while SUSD2<sup>+</sup> cells corresponding to mesenchymal endometrial cells enriched in the perivascular regions (Cousins et al., 2021).

      (4) Include a discussion about the interpretation of their current finding in relation to the dynamic regeneration observed in human endometrium due to menstrual bleeding/tissue breakdown, compared to the cycles of growth and regression that occur in mice.

      This is a great suggestion. We have added the following statement to our Discussion section:

      Lines 537-544: Additionally, our studies in human endometrium extend our characterization of ALDH1A1 as an adult endometrial stem cell marker and emphasize the importance of ALDH1A1+ in the regenerative potential of the endometrium. The conserved hormonal responses between human and mouse endometrium support the hypothesis that cycles of proliferation, differentiation, and regression, whether through resorption/autophagy in mice or menstrual breakdown in humans, are governed by concerted growth factor signaling and dedicated stem cell populations with the capacity to expand and differentiate across repeated cycles of repair. Our detailed studies in both mouse and human models indicate that ALDH1A1+ cells represent a dedicated cell type within the endometrium with the potential to drive repair during adulthood. Collectively, these findings advance our understanding of the mechanisms that control endometrial cycling and regeneration throughout the reproductive lifespan.

      Reference

      Ang, C.J., Skokan, T.D., and McKinley, K.L. (2023). Mechanisms of Regeneration and Fibrosis in the Endometrium. Annu Rev Cell Dev Biol 39, 197-221.

      Bi, W.R., Jin, C.X., Xu, G.T., and Yang, C.Q. (2012). Bone morphogenetic protein-7 regulates Snail signaling in carbon tetrachloride-induced fibrosis in the rat liver. Exp Ther Med 4, 1022-1026.

      Chen, M.Y., Zhao, F.L., Chu, W.L., Bai, M.R., and Zhang, D.M. (2023). A review of tamoxifen administration regimen optimization for Cre/loxp system in mouse bone study. Biomed Pharmacother 165, 115045.

      Cousins, F.L., Murray, A., Esnal, A., Gibson, D.A., Critchley, H.O., and Saunders, P.T. (2014). Evidence from a mouse model that epithelial cell migration and mesenchymal-epithelial transition contribute to rapid restoration of uterine tissue integrity during menstruation. PLoS One 9, e86378.

      Cousins, F.L., Pandoy, R., Jin, S., and Gargett, C.E. (2021). The Elusive Endometrial Epithelial Stem/Progenitor Cells. Front Cell Dev Biol 9, 640319.

      Darzi, S., Werkmeister, J.A., Deane, J.A., and Gargett, C.E. (2016). Identification and Characterization of Human Endometrial Mesenchymal Stem/Stromal Cells and Their Potential for Cellular Therapy. Stem Cells Transl Med 5, 1127-1132.

      Ghosh, A., Syed, S.M., Kumar, M., Carpenter, T.J., Teixeira, J.M., Houairia, N., Negi, S., and Tanwar, P.S. (2020). In Vivo Cell Fate Tracing Provides No Evidence for Mesenchymal to Epithelial Transition in Adult Fallopian Tube and Uterus. Cell Rep 31, 107631.

      Huang, C.C., Orvis, G.D., Wang, Y., and Behringer, R.R. (2012). Stromal-to-epithelial transition during postpartum endometrial regeneration. PLoS One 7, e44285.

      Khanmohammadi, M., Mukherjee, S., Darzi, S., Paul, K., Werkmeister, J.A., Cousins, F.L., and Gargett, C.E. (2021). Identification and characterisation of maternal perivascular SUSD2(+) placental mesenchymal stem/stromal cells. Cell Tissue Res 385, 803-815.

      Kirkwood, P.M., Gibson, D.A., Shaw, I., Dobie, R., Kelepouri, O., Henderson, N.C., and Saunders, P.T.K. (2022). Single-cell RNA sequencing and lineage tracing confirm mesenchyme to epithelial transformation (MET) contributes to repair of the endometrium at menstruation. Elife 11.

      Li, S.Y., Whiteside, S., Li, B., Sun, X., and DeFalco, T. (2025). Mesenchymal-to-epithelial transition of perivascular cells contributes to endometrial re-epithelialization. Nat Commun 16, 10174.

      Niayesh-Mehr, R., Kalantar, M., Bontempi, G., Montaldo, C., Ebrahimi, S., Allameh, A., Babaei, G., Seif, F., and Strippoli, R. (2024). The role of epithelial-mesenchymal transition in pulmonary fibrosis: lessons from idiopathic pulmonary fibrosis and COVID-19. Cell Commun Signal 22, 542.

      Patterson, A.L., Zhang, L., Arango, N.A., Teixeira, J., and Pru, J.K. (2013). Mesenchymal-to-epithelial transition contributes to endometrial regeneration following natural and artificial decidualization. Stem Cells Dev 22, 964-974.

      Pimeisl, I.M., Tanriver, Y., Daza, R.A., Vauti, F., Hevner, R.F., Arnold, H.H., and Arnold, S.J. (2013). Generation and characterization of a tamoxifen-inducible Eomes(CreER) mouse line. Genesis 51, 725-733.

      Rios, A.C., Fu, N.Y., Cursons, J., Lindeman, G.J., and Visvader, J.E. (2016). The complexities and caveats of lineage tracing in the mammary gland. Breast Cancer Res 18, 116.

      Seishima, R., Leung, C., Yada, S., Murad, K.B.A., Tan, L.T., Hajamohideen, A., Tan, S.H., Itoh, H., Murakami, K., Ishida, Y., et al. (2019). Neonatal Wnt-dependent Lgr5 positive stem cells are essential for uterine gland development. Nat Commun 10, 5378.

      Zeisberg, M., Shah, A.A., and Kalluri, R. (2005). Bone morphogenic protein-7 induces mesenchymal to epithelial transition in adult renal fibroblasts and facilitates regeneration of injured kidney. J Biol Chem 280, 8094-8100.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This paper investigates how heparan sulfate (HS) engagement functions in the cellular entry of SARS-CoV-2. A prevailing model that has been developed over the last five years by work from many laboratories using a variety of biochemical, structural, and microscopic approaches is that HS acts a co-receptor for SARS-CoV-2; its binding to SARS-CoV-2 both concentrates virus on the surface of target cells and allosterically alters the spike protein to promote an "up/open" RBD conformation that enables engagement of the proteinaceous receptor human ACE2 on the cell surface (PMID: 32970989, 35926454, 38055954, 39401361, 40548749). These two events enable plasma membrane fusion (after a cleavage event promoted by plasma membrane TMPSS2) or endocytosis and subsequent pH-dependent fusion (which requires a cathepsin L-mediated cleavage of the spike).

      The authors in this study used a series of microscopy techniques, labeled pseudoviruses and authentic SARS-CoV-2 strains, and cells lacking or expressing HS and/or hACE2 to re-examine the specific stage(s) HS and hACE2 function in the entry process. They suggest that HS mediates SARS-CoV-2 cell-surface attachment and endocytosis, and that hACE2 functions "downstream" of this to facilitate productive infection. Their results also suggest that SARS-CoV-2 binds clusters of HS molecules projecting 60-410 nm, which act as docking sites for viral attachment. Blocking HS binding with pixantrone, a drug under clinical evaluation for cancer (due to its anti-topoisomerase II activity), inhibited SARS-CoV-2 Omicron JN.1 variant from attaching to and infecting human airway cells. The authors conclude that their work establishes a revised entry paradigm in which HS clusters mediate SARS-CoV-2 attachment and endocytosis, with ACE2 acting at some stage downstream. They speculate this idea might apply broadly to other viruses known to engage HS and has translational implications for developing antiviral agents that target HS interactions.

      The strengths of the interesting and technically well-executed study include the use of multiple high-resolution microscopy modalities, the tracking of labelled viruses, the use of both pseudoviruses and authentic SARS-CoV-2, and the use of primary airway cells. Nonetheless, there are issues that need to be addressed to buttress the proposed model compared to earlier ones. These include: (a) the distinction between macropinocytosis and receptor-mediated endocytosis and what this might mean for productive SARS-CoV-2 infection; (b) the need to account for TMPRSS2 expression and plasma membrane fusion; (c) addition of genetic studies in which hACE2 is expressed in cells lacking HS; (d) an unclear picture of exactly where downstream hACE2 functions; and (e) and a need for comparative/additional study of earlier SARS-CoV-2 variants, which preferentially fuse at the plasma membrane.

      We thank the reviewer for the strong support of this manuscript. We addressed the reviewer’s concerns in the Recommendations to the authors. We did not distinguish whether the endocytic route is macropinocytosis or receptor-mediated endocytosis, because it is a separate study beyond the scope of the present work. We did not examine earlier SARS-CoV-2 variants because we considered it a study beyond the scope of the present work, but a good idea that we may work on in the future. For detail on how we address the remaining concerns, please see our response to the reviewer’s Recommendations for the authors.

      Reviewer #2 (Public review):

      In this manuscript by Han et al, the authors assess the binding of SARS-CoV-2 to heparan sulfate clusters via advanced light microscopy of viral particles. The authors claim that the SARS-CoV-2 spike (in the context of pseudovirus and in authentic virus) engages heparan sulfate clusters on the cell surface, which then promotes endocytosis and subsequent infection. The finding that HSPGs are important for SARS-CoV-2 entry in some cell types is well-described, but the authors attempt to make the claim here that HS represents an alternative "receptor" and that HS engagement is far more important than the field appreciates. The data itself appears to be of appropriate quality and would be of interest to the field, but the overly generalized conclusions lack adequate experimental support. This significantly diminishes enthusiasm for this manuscript as written. The manuscript is imprecise and far overstates the actual findings shown by the data. Additional controls would be of great benefit.

      Further, it is this reviewer's opinion that the findings do not represent a novel paradigm as claimed. HS has been well described for SARS-CoV-2 and other viruses to serve as attachment factors to promote initial virus attachment. While the manuscript provides new insight into the details of this process, the manuscript attempts to oversell this finding by applying new words rather than new molecular details. The authors would be better served by presenting a more balanced and nuanced view of their interesting data. In this reviewer's opinion, the salesmanship significantly detracts from the data and manuscript.

      We thank the reviewer for pointing out that our manuscript is of interest to the field. However, we do not think that we oversell our data. hACE2 has been widely considered the receptor (or the binding partner) that mediates SARS-CoV-2 cell-surface attachment, whereas HS is considered only an attachment factor that facilitates SARS-CoV-2 binding with hACE2 at the cell surface. In the present work, we found that HS, but not hACE2, is the cell-surface attachment receptor (or binding partner), whereas hACE2 is not essential for attachment, but acts downstream of virus endocytosis to facilitate viral genome expression. This finding suggests significant modification of the current model by replacing the attachment receptor (or binding partner) from hACE2 to HS, treating HS as a primary receptor rather than an attachment factor, and relocating the hACE2 action site from the cell surface to the endosome. For these reasons, we do not consider these statements overselling our data. However, as the reviewer suggested in his/her specific comments, we revised the manuscript to ensure that we did not overgeneralize our findings (see our responses to the reviewer’s Recommendations to the authors).

      Major Comments:

      The authors need to rigorously define a "receptor" vs an "attachment factor." They also should avoid ambiguous terms such as "receptor underlying ...attachment" and "attachment receptor" (or at least clearly define them). Much of their argument hinges on the specific definition of these terms. This reviewer would argue that a receptor is a host factor that is necessary and sufficient for active promotion of viral entry (genome release into the cytoplasm), while an attachment factor is a host factor that enhances initial viral attachment/endocytosis but is neither necessary nor sufficient. The evidence does NOT implicate HS as a receptor under this fairly textbook definition. This is proven in Figure 1 (and elsewhere) in which ACE2 is absolutely required for viral entry.

      The authors should genetically perturb HS biosynthesis in their key assays to demonstrate necessity. HS biosynthesis genes have been shown to be important for SARS-CoV-2 entry into some cells but not others (Huh7.5 cells PMID 33306959, but not in Vero cells PMID 33147444, Calu3 cells 35879413, A549 cells 33574281, and others 36597481. The authors need to discuss this important information and reconcile it with their data and model if they want to claim that HS is broadly important.

      Is targeting HS really a compelling anti-viral strategy? The data show a ~5-fold reduction, which likely won't excite a drug company. The strengths and limitations of HS targeting should be presented in a more balanced discussion. Animal data showing anti-viral activity of PIX is warranted. This would enhance this claim and also provide key evidence of a relevant role for HS in a more physiologic model.

      The authors provide little discussion of the fact that these studies rely exclusively on cell lines (which also happen to be TMPRSS2-deficient). The role of proteases in the role of HS should be tested in the cell lines and primary cells used, as protease expression is a key determinant of the site of fusion.

      The claim that "SARS-CoV2 JN.1 variant binds to heparan sulfate, not hACE2, in primary human airway cells" is extraordinary and thus requires extraordinary evidence.

      First, PIX reduces attachment by 5-fold, which is not the same as "nearly abolished." Also, anti-ACE2 "nearly abolished" entry in 7D, while PIX did not. If the authors want to make these claims, an alternative method to disrupt HS (other than PIX) is needed in primary airway cells. A genetic approach would be much more convincing. The authors should also demonstrate whether entry in their primary cell assays is TMPRSS2 vs Cathepsin L dependent (using E64d and camostat, for instance) as mentioned above.

      Each figure should clearly state how many independent experiments and replicates per experiment were performed. What does "3 experiments" mean? Are these three independent experiments or three wells on one day?

      In the well-accepted current model, hACE2 is considered the receptor mediating SARS-CoV-2 cell-surface attachment, entry into cells, and infection, whereas HS is an attachment factor that facilitates SARS-CoV-2 binding to hACE2 at the cell surface. The present work revises this view: HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, with ACE2 acting downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined.

      We made this point clearer throughout the newly revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docks at HS clusters.

      The cited CRISPR-screen literature supports context-dependent host-factor usage. However, the absence of HS biosynthesis genes from a given screen does not prove that HS is irrelevant in that cell type; it only indicates that HS biosynthesis was not detected as a genetic dependency under that assay’s conditions. Such negative results can reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency. In the revised manuscript, we included the following in the Discussion:

      “While some studies using genome-wide CRISPR screening to identify genes involved in SARS-CoV-2 reveal genes for HS biosynthesis, others do not (45-50). The negative result, which might reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency, needs to be verified with specific gene knockout.”

      The ~5-fold reduction is likely due to the inhibitor not completely abolishing HS-virus binding. We revised the Discussion to strengthen the suggestion that targeting the virus cell-surface attachment by interfering HS binding is a therapeutic strategy to prevent and treat COVID-19, as in the following:

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding. However, this strategy has not been the focus for developing methods to prevent and treat COVID-19, likely because HS is considered only a regulator that is not essential for SARS-COV-2 entry. Our finding that HS is the attachment receptor re-emphasizes the importance of perturbing virus-HS binding, the first step of the viral entry, to efficiently block SARS-CoV-2 infection. Further supporting this view, inhibition of HS binding with a clinically used HS-binding agent, pixantrone, inhibits authentic SARS-CoV-2 JN.1 subvariant binding with HS on the cell surface and infection in primary human airway cells (Figs. 6, 7). These results suggest a combinatorial anti-SARS-CoV-2 strategy: early HS blockade to prevent attachment combined with ACE2 targeting to inhibit post-attachment steps”

      We include a sentence in the Discussion that our suggestions are limited to the cells we examined as below.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      Three experiments refer to three independent experiments. We added “independent” accordingly throughout the manuscript.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors define a new paradigm for the attachment and endocytosis of SARS-CoV-2 in which cell surface heparan sulfate (HS) is the primary receptor, with ACE2 having a downstream role within endocytic vesicles. This has implications for the importance of targeting virion-HS interactions as a therapeutic strategy.

      Strengths:

      The authors show that viruses are internalized via dynamin-dependent endocytosis and that endocytic internalization is the major pathway for pseudotyped SARS-CoV-2 genome expression. They show that HS-mediated viral attachment is a critical step preceding viral endocytosis and also subsequent genome expression. Further, they show that hACE2 acts downstream of endocytosis to promote viral infection, and may be co-internalised with virions after HS attachment. Pseudotyped virus and authentic SARS-CoV-2 provide similar results. In addition, the authors demonstrate that remarkable clusters of multiple HS chains exist on the cell surface, visualised by a number of elegant microscopy methods, and that these represent the docking sites for virions. These visualisations are an important general contribution in themselves to understanding the nanoscale interactions of HS at the cell surface.

      The use of a complementary range of methods, virus constructs, and cell models is a strength, and the results clearly support the conclusions.

      Overall, the results convincingly demonstrate a different model to the currently accepted mechanism in which the ACE2 protein is regarded as the cell surface receptor for SARS-CoV-2. Here, the authors provide compelling evidence that cell surface clusters of HS are the primary docking site, with ACE2 interactions occurring later, after endocytosis (whilst still being essential for viral genome expression). This is an exciting and important landmark evidence which supports the view that HS-virion interactions should be viewed as a key site for anti-viral drug targeting, likely in strategies that also target the downstream ACE2-based mechanism of viral entry within endosomes.

      We thank the reviewer for the strong support of the present work.

      Weaknesses:

      This reviewer identified only minor points regarding citing and discussing other studies and typos, which can be corrected.

      We have addressed these points in the revised manuscript. For detail, please see our response to the reviewer’s Recommendations to the authors.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Pathway of internalization.

      The authors show clearly that labeled SARS-CoV-2 (pseudovirus or authentic virus) can become internalized in cells lacking hACE2, and this process depends on HS. However, they also show that this pathway is non-productive with regard to infection. Are the entry vesicles mediated by HS alone, HS + hACE2, and hACE2 alone the same? Or does the combination of co-receptor (HS + hACE2) drive SARS-CoV-2 into endocytic vesicles, whereas HS alone promotes macro- or micropinocytosis (lines 361-362). If HS alone directed SARS-CoV-2 into a non-productive entry vesicle, then hACE2 likely would be acting concurrently with HS and not downstream. A more detailed analysis of the different entry vesicles/pathways that occur with HS alone, HS + hACE2, and hACE2 alone is needed.

      During endocytosis, we did not detect a difference in the size distribution of virus-containing vesicles between BHK (HS alone) and BHK<sub>hACE2</sub> cells (HS+hACE2) (Fig. 2D). The similarity in the vesicle size suggests a similar endocytic path with HS alone or with HS + hACE2. In the revised manuscript, we added the following sentence.

      “Third, 3D-STED imaging showed that A490-labeled vesicle’s full-width-at-half-maximum (W<sub>H</sub>) was 363 ± 17 nm (n = 55) in BHK cells, similar to that (333 ± 13 nm, n = 70) in BHK<sub>hACE2</sub> cells (Fig. 2B-D), supporting a similar endocytic path regardless of hACE2 presence or not.”

      (2) TMPRSS2 and plasma membrane fusion.

      Although the authors allude to membrane fusion as an alternate mechanism of entry, their mechanistic experiments do not address the roles of HS and hACE2 in this process, possibly because their BHK and other cells do not co-express significant levels of TMPRSS2. While many Omicron variants preferentially enter cells via endocytosis (relative to antecedent strains in the pandemic) because of spike mutations that reduce cleavage by TMPRSS2 (PMID: 35104837, 36625591, 35145066), plasma membrane fusion can still occur. The authors should add experiments with co-expression of TMPRSS2/hACE2 [with or without HS] and earlier SARS-CoV-2 variants to establish the role of HS in plasma membrane fusion. Also, are there differences in entry pathways if viruses are prepared in cells expressing TMPRSS2?

      We thank the reviewer for this important comment and agree that our mechanistic experiments were not designed to address TMPRSS2-supported plasma membrane fusion. The reviewer’s suggestion for direct testing of HS function in TMPRSS2-supported plasma membrane fusion, including hACE2/TMPRSS2 co-expression and comparison with earlier SARS-CoV-2 variants, will require dedicated experiments and detection of the fusion pathway that we have not yet designed. It is beyond the scope of the present work. In the revised manuscript, we clarify that the observed ACE2-independent uptake and the predominance of endocytic entry refer to the tested cell systems and do not exclude TMPRSS2-dependent plasma membrane fusion in other cell types, as in the following.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Experiments with hACE2 in cells lacking HS.

      Apart from drug treatment (heparinases or pixantrone) studies shown, the current studies do not directly address whether expression of hACE2 on human cells can allow for endocytosis and productive infection in the complete and genetic absence of HS. The only experiments that use genetically deficient cells are the CHO [hamster] cell studies, and these cells lack hACE2 expression. The authors should knock out a key HS biosynthesis gene (e.g., B4GALT7) in more relevant human cells (e.g., A549-hACE2; ideally with sorted subpopulations having different levels of surface hACE2 expression) and assess endocytosis and infection. This is important given studies in the literature by others suggesting that KO of HS expression reduces but does not abrogate SARS-CoV-2 infection.

      We thank the reviewer for these comments. We showed that virus endocytosis is independent of hACE2 (Fig. 1). The reviewer’s question is whether hACE2 alone can allow for endocytosis of viruses. We have shown that in either BHK (without hACE2) or BHK<sub>hACE2</sub> cells (BHK cells expressed with hACE2), heparinase I/II/III mixture (HPRase) nearly abolished cell-surface immunolabelled HS (Fig. 3D), reduced cell-surface virus attachment by ~83-85% (Fig. 3E), reduced viral uptake by ~80% (Fig. 3F, 3G). These results suggest that hACE2 is not essential for viral attachment and endocytosis. We did not test whether hACE2 alone (without HS) plays a minor role for viral attachment and endocytosis, because to our knowledge, HS is present in nearly every cell. Under this physiological condition, it is HS, not hACE2, that plays an essential role in viral cell-surface attachment and endocytosis. In the revised manuscript, we added a sentence admitting that we did not test whether hACE2 alone is sufficient to support viral uptake and productive infection, as in the following.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis.

      (4) hACE2 function in entry.

      In many places, the authors suggest that hACE2-spike functional interaction occurs "downstream" of HS-dependent binding and endocytosis (e.g., lines 25, 33, 48, 210, 309, 312, 318, 333, 346). However, in their model, it is not clear where exactly this interaction occurs. Are the authors suggesting that this spike binds hACE2 on the cell surface, but this has nothing to do with endocytosis, or that the interaction with hACE2 is occurring at a post-entry step? Can they experimentally demonstrate the stage at which hACE2 is functioning? Is it the same or different in cells lacking HS? What about when TMPRSS2 is present?

      We showed that viral attachment and endocytosis are independent of hACE2, whereas entry as determined by viral gene expression, depends on hACE2. We also showed that most virions bind to HS, not hACE on the cell surface. Based on these results, we propose a model that hACE2 functions downstream of virion endocytosis. We cannot rule out the possibility that a small subset of viruses can also bind to hACE2 after their binding with HS at the cell surface.

      We have not been able to design an experiment to visualize hACE2 mediated virion fusion in endosomes, where hACE2 may facilitate virus fusion and delivery of viral genomes to the cytosol. Productive infection requires only a limited number of successful virion–hACE2 engagement events. While many internalized virions can be visualized, the specific virion or vesicle that ultimately gives rise to productive infection cannot be identified from the present imaging data. This makes it difficult to trace the precise stage or compartment in which the functionally relevant spike–hACE2 interaction occurs. In the revised manuscript, we added a paragraph discussing this limitation as below.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis. Although our data suggest that hACE2 functions downstream of endocytosis to facilitate viral fusion at the endosome for genome delivery to the cytosol, we do not know whether hACE2 binding with the virus occurs at the cell surface or endosomes. The binding may occur in both places, but not essential for virus attachment and endocytosis.”

      (5) Other comments.

      (a) Figure 1A. "Antibody" is misspelled.

      Corrected. Thank you.

      (b) The imaging experiments with pseudoviruses and authentic viruses lack any information on the multiplicity of infection or the number of virions added per cell. If this is particularly high and non-physiological (e.g., >100), is it possible that such conditions might enable viruses to enter [dominantly] through secondary [non-infectious] pathways?

      To address the reviewer’s concern, we used flow cytometry to measure cell-associated VSV-S signal as we diluted the virus by ~600-fold. We found that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHK<sub>hACE2</sub> cells over a ~600-fold dilution of the virus (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations. In the revised manuscript, we included the following sentence and Fig. S7 (Supplementary Information).

      “Flow cytometry also showed that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHKhACE2 cells over a ~600-fold dilution of the virus concentration (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations.”

      (c) Figure 1C and elsewhere. Most of the internalization studies rely on various imaging modalities to demonstrate the pseudovirus or virus on or in the cell. The experiments would be strengthened by inclusion of data from orthogonal binding/internalization assays that measuring virion-associated viral RNA on the surface [4oC binding assay] or inside the cell [after a 37oC temperature shift and exogenous proteinase K and RNAse A treatment]) - such assays can be performed at much lower MOI (e.g., <1, addressed comment #2 above) an also allow more objective quantitation and kinetic analyses of virus internalization (e.g., 0, 5, 15, 30 min at 37oC).

      We demonstrate virion attachment and uptake using multiple approaches, including confocal, STED, and EM analysis, showing virions with the expected morphology at the cell surface and in the cytosol. Furthermore, flow cytometric analysis provides population-level quantitation supporting the same overall conclusion. Thus, while we appreciate and agree that an RNA-based binding/internalization assay would provide additional information, we do not consider it essential to the main conclusion of this work.

      (d) Figure 2. (i) Is there any indication of which vesicles the bath dye is in? Is most of this fluid taken up by micropinocytosis? Are these the same vesicles where the virus that is destined for productive infection (HS/hACE2 engaging) transits? (ii) In all panels, can the authors clearly indicate/label which cells are being used (BHK or BHK-hACE2)? (iii) For the studies with dynasore or dominant-negative dynamin-2-K44A, the readout is at 24 h, a late timepoint, which also could affect virus egress and spread. Can the studies be repeated at much earlier time points (e.g., 15 min to 2 h) to demonstrate that viruses are internalized via dynamin-dependent endocytosis in these cells?

      (i) The bath dye A490 was used as a fluid-phase marker for endocytic uptake, rather than as a marker for a specific vesicle class or intracellular compartment. In principle, any vesicle that takes up extracellular fluid could become labelled by this approach. Since nearly all viruses are in the A490-containing vesicles, productive virus infection must come from some of these vesicles.

      (ii) In the revised Fig. 2 legends, we explicitly indicate which cells are used for each panel.

      (iii) To address the reviewer’s concern, we examined earlier time points for dynasore treatment and found that the virus uptake and genome expression were already markedly reduced at 1 h and 8 h after virus incubation. In the revised manuscript, we described these results as below and in Fig. S5.

      “Fourth, dynasore or dominant-negative dynamin 2-K44A overexpression, which inhibits fission of dynamin-dependent endocytosis [28-30], substantially reduced V-A647 internalized 1-24 h after viral incubation (Figs. 2F-G, S5).

      In addition to inhibiting V-A647 endocytosis, dynasore or dynamin 2-K44A inhibited V-EGFP expression 8-24 h after virus incubation by ~66-77% (Figs. 2F-G, S5), suggesting that endocytosis is the main route for viral genome expression.”

      (e) Line 225. "Envelop" should be "envelope".

      Corrected, thank you.

      (f) Line 235. The authors should clarify that they conclude that the "Omicron variant" of SARS-CoV-2 enters "BHK" cells indistinguishably from VSV-S.

      Thank you for pointing this out. We have rephrased the conclusion as “…omicron variant of SARS-CoV-2 enters BHK cells indistinguishably to VSV-S.”

      (g) Line 278. What happens to virus binding if the authors ectopically express hACE2 in CHO-K1 WT and CHO-pgsA-745 cells?

      We did not perform this experiment (see also our response to major comment 3 above).

      (h) Lines 280-281 and elsewhere (line 635). The authors state "pixantrone (PIX), a drug under clinical trial that binds HS to inhibit HS binding with proteins...." The authors should clarify that the drug is under clinical evaluation for cancer treatment because of its DNA intercalating activity (and not its HS binding activity) and cite any relevant ongoing trials. Also, in line 635, is reference #46 correct?

      As suggested, we modified this sentence as “pixantrone (PIX), a drug under clinical trial for cancer treatment due to its DNA intercalating activity, which can bind HS to inhibit HS binding with proteins”

      (i) Line 281-282. The authors should confirm in a Supplementary Figure that the anti-hACE2 antibody used blocks SARS-CoV-2-JN.1 binding to ACE2.

      In Figure 7D, we showed that PIX and anti-hACE2 antibody block SARS-CoV-2-JN.1 infection, suggesting that anti-hACE2 blocks SARS-CoV-2-JN.1 binding with hACE2.

      (j) Figure 7B. Can hACE2 co-localization be added to this panel?

      We did not perform this experiment. We addressed the role of ACE2 in these airway cells in subsequent panels of Fig. 7.

      (k) Figure 7C. The quantitative data show a 50% reduction in binding signal with pixantrone, whereas the microscopy images appear to show a much greater effect. Can more representative images be shown so that the data better corresponds?

      A ~50% effect is not as visually obvious as the current Fig. 7C. Therefore, we chose not to change the images. However, the statistics in Fig. 7C (right) clearly indicate an average effect of about 50%, as the reviewer pointed out.

      (l) In the Discussion, it is not necessary to use Figure callouts (as done in the Results). Please remove, with the exception of reference to the model.

      We prefer to call out Figures in the Discussion so that we can remind the readers where to find the data. The readers may choose to neglect these callouts. But some readers may read most the discussion part without going through the results carefully. In this case, the figure callouts may help these readers.

      (m) Please delete all references to "new" or "novel" models. It is unnecessary.

      As the reviewer suggested, we deleted “new” and “novel” throughout the revised manuscript.

      (n) Figure legends. Please make sure each panel indicates the # of independent experiments performed. This is included for some but not all panels. Also, a few panels use an unpaired t-test where an ANOVA with multiple comparisons is required (e.g., Figure 1G and S1).

      We agree that, for the three-group sub-comparisons shown within Fig. 1G and Fig. S1, the relevant analyses should account for multiple comparisons. In the revised manuscript, we therefore analyzed these predefined three-group subsets using ordinary one-way ANOVA followed by Dunnett’s multiple-comparisons test, with BHK or Vero used as the reference group as appropriate. The two-group comparisons were analyzed using unpaired two-tailed t-tests.

      Reviewer #2 (Recommendations for the authors):

      (1) It is well established that ACE2 is the receptor for SARS-CoV-2. The authors should not downplay this by saying it is "widely assumed", "typically thought", etc. The specific molecular details at various stages of entry (i.e, the role of HS) remain a bit unclear, but it is disingenuous to imply ACE2 is not the bona fide receptor by any conventional definition.

      The present work does not challenge the well-established view that ACE2 is the receptor for SARS-CoV-2 entry/infection, but suggests that HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, whereas ACE2 acts downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined. We made this point clearer throughout the revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docked at the HS clusters.

      As the reviewer suggested, we removed “assumed” and “typical” and clarify that our findings do not challenge this concept. For example, we modified the abstract

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely assumed as the receptor.”

      as

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely considered the receptor for cell-surface attachment and subsequent cell entry.”

      (2) When the authors state pseudovirus internalization is independent of ACE2, they should clarify that this is the case in cells not expressing TMPRSS2. Most physiologically relevant cell types express TMPRSS2, which will facilitate entry at the plasma membrane.

      As the reviewer suggested, we included the following sentence in the Discussion section: “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Line 130: "Endocytic internalization is the main viral infection pathway" and Line 180-181 is not precise and should be rephrased to include the cell types described in the figure. This may be true in BHK-ACE2 cells, but the evidence in this section does not show that this is universally or broadly true.

      We agree and have revised these sentences to limit the conclusions to the experimental context directly supported by our data. Specifically, our results support endocytic uptake as the major route leading to pseudovirus genome expression in the pseudovirus assays and cell types examined here, rather than as a universal entry mechanism for SARS-CoV-2 across cell types. We have therefore modified the subsection title and the relevant sentence in the Results to explicitly refer to the tested cells/assays.

      Across the revised manuscript, we have accordingly revised the text to distinguish initial virion docking/attachment from productive entry, to acknowledge ACE2 as the established receptor for productive infection, and to limit our mechanistic conclusions to the cellular systems directly tested here.

      (4) All bar plots should show individual dots (i.e., Figure 1G) to better reveal the variance of each dataset.

      While we respect the reviewer’s suggestion, this is not required in the journal style. We prefer plotting bar graphs without individual data points, which often makes it difficult to see the mean values.

      (5) Line 57: This is not accurate. HIV uses a receptor and a co-receptor, for instance.

      We thank the reviewer for noting this inaccuracy. We agree that viral entry frequently involves coordinated engagement of multiple host factors rather than a single receptor, for example, HIV requires both a primary receptor and a co-receptor. We have revised the statement in the Introduction (Line 57–58) to reflect that entry can involve receptors together with co-receptors and/or attachment factors, which collectively facilitate membrane fusion or endocytic uptake.

      In the Introduction (Line 57), we replaced the sentence with “Viral entry is often initiated by engagement of host receptors and associated co-factors that together facilitate subsequent viral membrane penetration.”

      (6) Line 60: "most" --> "many"

      As suggested, we have changed “most” to “many”.

      (7) Remove "clinically relevant" in reference JN.1, as JN.1 is not circulating currently. A more appropriate term could be "full-length" or "authentic", or "wild-type".

      As suggested, we changed it to “authentic”.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors omit to mention the work of Zhang et al, 2023 Nature Comms. "Host heparan sulfate promotes ACE2 super-cluster assembly and enhances SARS-CoV-2-associated syncytium formation". These authors also use PIXN and MTN compounds and define different mechanisms based on ACE2 clustering for virus entry. The authors should mention this work in the Discussion and try to reconcile the different findings.

      As suggested, we include the following discussion in the revised manuscript.

      “Consistent with this possibility, HS may promote spike-dependent ACE2 super-cluster assembly at the cell surface and enhance SARS-CoV-2–associated syncytium formation, suggesting that HS may organize ACE2 nanoscale architecture in a cell–cell fusion context [43].”

      (2) The authors should strengthen their case for the validity of HS-virion interactions as a therapeutic target by mentioning studies showing effectiveness of interference with HS-Covid interactions by heparin and other investigational drugs eg. first study to demonstrate heparin inhibition of SARS CoV2 attachment, Mycroft-West et al, Thromb Haemostatis, 2020; recent report of successful clinical trials of nebulized heparin, The Lancet, Sept 2025; and the superior efficacy of HS mimetic Pixatimod/PG545 compared to heparin (Guimond et al 2022 ACS Chemical Sciences).

      We thank the reviewer for this suggestion and add the following paragraph with citations the reviewer mentioned in the Discussion section.

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding.”

      (3) Figure 1a: incorrect label for antibody.

      Corrected, thank you.

      (4) Some misspellings noted in the manuscript, e.g., MINFLLUX, so please recheck the manuscript for typos.

      We have rechecked the manuscript and corrected the typos.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      There are two main criticisms:

      (1) It is not clear how much the factors uncovered here are true beyond B6 mice. B6 mice, compared to humans, are known to be very Th1-skewed, and Tbet is a strong inhibitor of Th17-specific T cells. Many people make IL-17-producing T cells in response to Mtb infection.

      We appreciate the point that not all findings in mice are directly translatable to humans. The B6 mouse is widely used as a model organism for tuberculosis due to its tractability and the wealth of genetic tools available for this strain. While it is true that many individuals do produce Th17 cells after infection with Mtb, humans are still very Th1-dominant, and not all infected individuals produce Th17 cells. We can speculate that the mechanisms outlined in this paper may contribute to the reasons that Th17 responses are not more robust in humans, a finding that may be useful in guiding vaccine design in the future.

      (2) Very few novel insights are mechanistically revealed about how Th17 induction is restricted by Mtb. Tbet induction is known to restrict Th17 development, and this is a T-cell intrinsic mechanism. In contrast, the IL-23 association revealed seems to be extrinsic to T cells and to act on T cells. How, if at all, are these factors related to each other in restricting Th17 induction? Also, the conclusion that it is not a result of attenuation is not completely convincing.

      While it is established that Th1 differentiation can inhibit Th17 differentiation, we believe that rigorously demonstrating this genetically in the context of Mtb infection remains important. Moreover, it cannot be assumed that IL-17 elicited by dampening the Th1 response can lead to enhanced control of infection. We view addressing this as a significant contribution. Furthermore, we also show that the ESX-1 and PDIM virulence factors are functionally linked by suppression of IL-17 responses. The effect is unlikely to be simply due to attenuation of the strains as an equally attenuated control strain does not elicit Th17 cells. We believe that these insights are both novel and important for understanding immune responses to Mtb.

      Other points:

      (1) The authors show that mice infected with a deficiency in ESX-1 have more IL-17-producing CD4 T cells in response to stimulation with an ESAT-6 peptide pool (Figure 3B). Because ESAT-6 is encoded by ESX-1, why do mice infected with this Mtb mutant have any ESAT-6-specific T cells? Is it an incomplete knockdown?

      The ESX-1 knock-out M. tuberculosis Erdman strain is a ΔEccC1 mutant. This strain can produce Esat-6 but cannot secrete Esat-6 out of the bacterial cell. Thus Esat-6 protein is present and able to be processed for MHC-II presentation. We also use Ag85b peptide pool stimulation and report similar effects as Esat-6 peptide pool stimulation.

      (2) The manuscript states, "Under the conditions where Th17s are highly induced, mice infected with either ΔESX-1 or PDIM lacking Mtb, the Il17a-/- mice had ~3-5 fold higher CFU than WT mice (Figures 3F-G). These results indicate that the induction of Th17s is not dependent on the attenuation of Mtb in general, but instead Mtb utilizes ESX-1 and PDIM to suppress the induction of a Th17 response that enhances protection against Mtb infection." I don't think the last sentence is necessarily true. I can imagine a scenario in which the induction of the Th17s is, in fact, due to the attenuation, and the Th17 induction still contributes to protection.

      We tested another attenuated M. tuberculosis strain with no known relationship with ESX-1 or PDIM, ΔMmpL4. This attenuated mutant fails to induce IL-17A–producing CD4 T cells to the same extent as observed in mice infected with ESX-1-deficient or PDIM-deficient strains, which is strong evidence that simple attenuation of virulence does not result in higher numbers of Th17 cells being elicited.

      (3) ESX-1, PDIM, and mmpl4 mutants all have similarly reduced CFUs in the lung, but what about the LN? The bacterial burden in the LN may be more important for regulating T-bet, IL-23, and Th17 differentiation, since the LN is where T cell priming occurs, than the CFU in the lung. Perhaps ESX-1 and PDIM mutants have reduced CFU in the LN, but mmpl4 does not. This difference in LN burdens may be the primary driver of Th17 priming, as high avidity interactions are thought to be an important driver of T-bet induction.

      We acknowledge that this is a formal possibility, however we maintain that the phenotype is specific to ESX and PDIM mutants, rather than MmpL4 mutants. Even if this phenotype arises from a tissue-specific attenuation of ESX/PDIM mutants, it remains a specific phenotype of these mutants, and not all attenuated mutants, albeit less directly. More importantly, the observation that these mutants induce higher levels of the Th17-polarizing cytokine IL-23 from infected cells ex vivo suggests that this is not an indirect phenomenon.

      (4) Do LN cDC1 and high levels of IL-12 p35 in mice infected with the mmpl4 mutant? Likewise, LN cDC2's express low levels of IL-12 p19 (akin to those infected with WT Mtb)? If these observations for ESX-1 and PDIM mutants are mechanistically linked to the increased numbers of Th17 cells, then you would expect mice infected with mmpl4 mutants to be more like those infected with WT Mtb than those infected with ESX-1 and PDIM mutants.

      Because ΔMmpL4 and complemented strains resulted in T cell profiles that were not different from the wild-type, we did not measure mediastinal lymph node dendritic cell expression of IL-12 p35 and IL-23 p19 in infections with these mutants.

      (5) ESX-1 and PDIM are very different virulence factors - a protein secretory pathway and cell wall lipid, respectively? Mechanistically, how would mutants in these pathways give very similar outcomes regarding Th17 cells unless it was simply as an aspect of their attenuation? Perhaps, mmpl4 mutants simply differ in some aspects of their attenuation, such as bacterial burdens in LNs, or their interaction with cDCs?

      We are not the first to link phenotypes of ESX-1 and PDIM. Both systems have both been shown to be important for M. tuberculosis permeabilization of the host cell phagosome after phagocytosis, and for suppression of type I IFN responses, among other responses. Thus, these seemingly different virulence factors clearly work together to support specific virulence traits during infection. The exact mechanism of how ESX-1 and PDIM interact is not completely understood and is an area for future investigation.

      Reviewer #2 (Public review):

      The following conclusions and interpretations should be revisited, rephrased, and re-evaluated:

      (1) The manuscript neglects to analyze T cell responses in the dLN, which is the critical site where these responses are initiated (only DC cytokine production is measured in the dLN). The differences in the lungs could reflect trafficking of T cells to the lungs, local lung T cell responses, or durability of the T cell responses in the lungs. The authors state in the last results section that "These results indicate that the ESX-1 and PDIM virulence factors impact naïve T cell differentiation at the draining mediastinal lymph node..." but T cell responses are never measured in the dLN.

      Due to the limited size of the mediastinal lymph node at 3 weeks post infection, we were unable to obtain enough cells for both myeloid cell analysis and T cell analysis, as we perform staining for these panels separately due to the decrease in viability of myeloid cells observed during T cell restimulation. In addition, because T cells in the lung are the population of cells most critical for mediating the outcome of infection, we believe analyzing the T cell response in the lymph nodes though interesting, is not crucial for this study. We have edited the manuscript to be clearer, as suggested by the reviewer.

      (2) Figure 2: The authors state that "Importantly, IFN-γ deficient mice did not exhibit elevated levels of IL-17A producing CD4 T cells demonstrating that IFN-γ production is not the mechanism by which Th1 T cells limit a Th17 response during Mtb infection", but the difference is significantly different and even more obvious in Panel B. In fact, if the Panel D y-axis was on a log scale, the Ifng-/- would likely look more like Tbet-/- than WT. Based on this data, it seems like IFNg is having an effect and should not be completely discounted. Does the deletion of Ifng affect the number of Tbet+ T cells?

      We agree that the IFN-γ<sup>-/-</sup> have only 5x more IL-17 producing CD4 T cells than WT mice while Tbet<sup>-/-</sup>mice exhibit a 25-fold increase compared to WT. We have added this information to the text, and now point out that IFN-γ production is not the sole mechanism by which Th1 T cells limit a Th17 response during Mtb infection.

      In addition, the deletion of Tbet results in an increased number of IFNg+IL-17+ double positive T cells (Figure 2B), in addition to a sizable IFNg single positive T cell population maintained in the Tbet-/- mice (10x the negative control of Ifng-/-). Is this why Tbet deletion is not as severe as Ifng deletion, because T cells are still making IFNg?

      It is possible that the residual IFN-γ produced by T-bet-deficient animals contributes to their relatively modest susceptibility to infection. However, our data show that deletion of IL-17 in this background renders T-bet–deficient mice nearly as susceptible as IFN-γ deficient mice, arguing that the remaining IFN-γ is not a major protective factor.

      Along these lines, the statement in the text that, "Tbet-/-Il17a-/- mice completely lacked both IFN-γ producing...." T cells is not supported by the data in Figure 2C. Tbet-/-Il17a-/- mice look to have more gamma-producing T cells than Tbet-/- mice (which is already 10x the negative control of Ifng-/- in panel 2B if one includes the gamma single positive and IFNg/IL-17 double positive).

      We have amended the language in the text to be more consistent with the data.

      (3) In the Results sections describing Figures 3, 4, and 5, the authors equate IL-17 production by T cells with TH17 responses and IFNg expression with TH1, but Tbet and RORgt expression in the T cells should be measured to make conclusions about TH1 and TH17. Or the authors can rephrase their findings to specifically state the observations as IFNg or IL-17 expressing CD4+ T cells.

      We believe that calling a CD4 T cell in the lung that is producing IFN-γ (and not IL-17) a Th1 cell is appropriate. Potentially confounding cells include those which also produce IL17, which we have ruled out, or T<sub>FH</sub> cells that may be common in lymph nodes but are not common in lungs at this time point and under these conditions.

      (4) Conceptually, do the authors think that ESX1/PDIM promotes TH1 responses and this blocks TH17 or are ESX1/PDIM blocking TH17 responses directly, allowing for increased TH1 responses? It would be helpful to clarify the model in this regard, describe how the data supports one model or the other, and then make sure the language is consistent throughout. Can these effects on T cell responses be tested and recapitulated in vitro using infected APC and T cell co-cultures?

      While it is possible that PDIM and ESAT-6 suppress Th17 through promotion of Th1 differentiation, we do not have data to support this model currently. However, we have added a comment making this point to the discussion.

      Reviewer #3 (Public review):

      Weaknesses:

      (1) The authors should acknowledge and reference key findings from the literature that have identified suppression of Th17 differentiation as an Mtb virulence mechanism, e.g., the role of the Hip1 protease and CD40 signaling (Madan-Lala JI 2014, Sia Plos Path 2017, Enriquez iScience 2022) and Khader JI 2005, showing the requirement of IL-23 for Th17 responses in vivo in a TB mouse model.

      We thank the reviewer for pointing these references out and have added them to the discussion section of the manuscript.

      (2) Addressing several questions related to the Tbet KO mouse experiments would strengthen the study. Do the Tbet KO mice have elevated IL-4/5/13 (which has been previously reported in non-TB studies) in addition to IL-17? The lack of Th17 cells in the IFNg KO compared to the Tbet KO may be due to a difference in timing, since only 3-week data are shown; earlier and later time points would provide better interpretation. The authors do not present any data on neutrophil infiltration in WT vs Tbet KO vs IFNg KO mice. Since IL-17 is known to be important for recruiting neutrophils to the lung, data on neutrophils are important for clarifying the mechanism for the CFU outcomes.

      We agree that it is surprising that, in the context of TB, Th17 responses are protective whereas excessive neutrophil recruitment is detrimental to the host. In IFN-γ–deficient mice, neutrophils are recruited and contribute to the increased susceptibility of this strain (PMID: 21967766). In separate work from our lab, we have shown that the phenotype of neutrophils recruited to the lungs during Mtb infection influences disease outcome (PMID: 40937719). It is possible that differences in the host environment and the timing of the response shape the effects of neutrophils on the host; these and the other questions raised by the reviewer will be the subject of future studies.

      (3) While IL-23 is important for sustaining IL-17 production, IL-6, TGF-b and/or IL-1β are necessary for Th17 polarization. What were the levels of these cytokines in DCs in the lung? (Figure 5). Additionally, Tbet-deficient DCs exhibit impaired activation of antigen-specific Th1 cells and have reduced IL-12 production. Given the data showing higher IL-17 levels in Tbet KO mice, the authors should provide information on the DC phenotype (IL-23, IL-6, etc.) in the Tbet KO experiments.

      While these are interesting points, investigating mechanisms of Tbet-dependent suppression of IL-17 is beyond the scope of this study.

      (4) The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is not clear. While data showing a role for ESX-1 and PDIMs in inhibiting Th17 responses is interesting, there is no insight into the potential mechanism of action. Figure 3 showing reduction in IFNg+ CD4 T cells after infection with eccC1 and fadD28 mutants suggests that this outcome is due to a lower bacterial load relative to WT Mtb at the 3-week time point. Since IFNg is known to suppress IL-17, the higher levels of Th17 cells could be due to the reduction in IFNg due to the attenuated growth of the mutants. Additionally, what was the level of Type I IFNs elicited by these mutants?

      We included the MmpL4 knockout Mtb Erdman strain as a control to ensure that attenuation of mutants is not the cause of the increase in IL-17. We also showed that eliminating type I IFN signaling by deleting its receptor has minimal impact on Th17 differentiation, even in the context of a host that produces excess type I IFN. Therefore we do not believe that type I IFN elicited by these mutants is explanatory for the phenotype.

      (5) Since macrophages have been implicated in the reduced cytokines seen in the ESX-1 mutant, IL-23 and other cytokine data on lung macrophages would complement the DC data.

      Because dendritic cells are primarily responsible for priming CD4 T cell responses, we believe that this result in macrophages would not substantially alter our conclusions. That said, it was demonstrated previously that macrophages infected with ESX-1 mutants produce less IL-12p40, a subunit of IL-23.

      (6) Figure 5. There are many fewer DCs overall in the eccC1 and fadD28 mutant groups, which could account for the increased % IL-23p19 in DCs (5D). What were the levels of IL-23 in DC1s?

      The amount of IL-23 p19+ in type I conventional dendritic cells (cDC1s) was near zero as shown in supplementary figure 6A. cDC1s are known to not express IL-23 p19 in mice.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) What do the authors mean by the alternative secretion part of "ESX1 Type VII alternative secretion system" that they refer to?

      Bacterial alternative secretion systems facilitate the export of proteins from the bacterial cell independent of the canonical Sec-dependent secretion system required for export of most secreted bacterial proteins across the inner membrane. However to avoid confusion, we have removed the word alternative.

      (2) Not sure naïve fits in this sentence at the end of the introduction: "Furthermore, we observe a strong Th17 response during infection with ΔESX-1 or PDIM lacking Mtb in naïve mice....".

      We have removed the word Naïve.

      (3) Figure legend for 1A-C says analysis performed at 21 dpi, but the figure shows the time course.

      We have corrected this error.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1 should show the non-stimulated flow plot.

      We have added the unstimulated samples.

      (2) The % IL-17 in the flow plots is not consistent across Figures 1, 2, and 3. Not sure why the scales for the Y-axis for IL-17 differ so much between Figures 1 and 2/3. IS there a technical issue with compensation?

      We did not experience any difficulties with compensations. These experiments were done over several years of work. For every experiment, new single-color controls were used and gating was done with the FMO gating strategy. Minor variation such as we see here is not surprising.

      (3) Discuss Yeh et al J Neuroimmunol 2014- show that IFNγ inhibits Th17 differentiation and function via Tbet-dependent and Tbet-independent mechanisms.

      We have added this reference to the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      The strength of this study is its use of a simple behavioral parameter, TOWA, and also a simple design of behavior, WAFO. The importance of the behavioral assay is reproducibility and comparability. In fact, the author demonstrated a summary of comparisons where different treatments result in scalable behavioral changes in WAFO and TOWA.

      We appreciate this assessment and fully agree that the simplicity of the assay and the demonstration of its scalability and reproducibility are a strength.

      Weaknesses:

      The weakness of the study is the lack of further experiments to support their assumption related to TOWA. The authors suggested that TOWA can be interpreted as a behavioral proxy for exogenously induced arousal. However, it could be interpreted as higher activity, although the authors argued that the circadian clock increasing locomotor activity around ZT0 and ZT12 does not affect TOWA, and therefore TOWA is not related to the locomotor activity per se. As the author cited, flies lose locomotor activity in the circular arena of 6.6 cm in diameter, whereas they continuously move during a 1-h recording in the authors' arena of 1 cm in diameter.

      I would agree that the arena of 1 cm in diameter, but not 6.6 cm in diameter, serves as an exogenous stimulus inducing arousal, and TOWA is manifested by arousal. However, TOWA would also be affected by other behavioral parameters, including the activity, motivation for exploration, or perception of the space. Therefore, it could be reasonable to re-examine some of the flies tested in this study in the circular arena of 6.6 cm in diameter. If arousal is biased by the components presented in Figure 6 and TOWA can assess mainly exogenously induced arousal, the treatment altering TOWA in the arena of 1 cm in diameter would not affect their behavior in the arena of 6.6 cm in diameter. My concern is that Figure 6 may demonstrate too simplistic a diagram to interpret the results. I would suggest adding the experiments using the arena of 6.6 cm diameter or softening the argument.

      We are grateful that you prompted us to investigate the relation between TOWA and arousal and different arena diameters in more depth. Based on your comments, we compared naïve and stressed behaviour between arena diameters of 1, 2.2 and 5.8 cm. The sizes were chosen as to optimally comply with the camera field of view in our setup. Naïve flies showed stable locomotor activity throughout the 60 min of recordings in the different arenas (new Figure 1 – figure supplement 1 A-B’’). Moreover, no significant difference in TOWA over the first 10 min was found between the different arena sizes (new Figure 1 – figure supplement 1 C). This suggests to us that arenas with a diameter smaller than 6.0 cm (and not only with a 1 cm diameter) induce some form of activity that resembles stimulated activity as defined by Meehan and Wilson (1987). A mechanical shock (shake) resulted in significantly increased WAFO in all arena sizes (new Figure 1 – figure supplement 1 D-D’’). We did not test the effects in larger arenas > 5.8 cm, as we found that flies do not longer show persistent and quantifiable wall following. We started to see that also in few flies in the 5.8 cm arena – these few flies were excluded from our analysis presented in new Figure 1 – figure supplement 1. Unlike the naïve response, the stress-induced response in WAFO and TOWA appears to be transient (new Figure 1 – figure supplement 2), which is in line with the definition of emotions as a transient state.

      Reviewer #2 (Public review):

      Summary:

      Strengths:

      The main strength of the paper is the rigorous use of several stressful or aversive treatments and their subsequent removal to show that WAFO is a robust proxy for stresslike emotional primitives across multiple stimuli. The pharmacological, molecular, and neuronal activity manipulations, although more limited in scope, lend further credence to the authors' central claim.

      We are glad about this assessment and share your opinion.

      Weaknesses:

      The conceptual advance of this research is unclear, as previous work (Mohammad et al., 2016, Curr Biol.) carried out similar treatments and manipulations and reached largely similar conclusions.

      Thank you very much for bringing this up. We rewrote respective parts of the introduction (second last paragraph) and discussion to more clearly outline the advances over the previous work by Mohammad et al. 2016. While our study builds upon Mohammad et al. 2016, the conceptual advance and novelty is that we constitute and treat TOWA as a second and independent dimension equal to WAFO in the OFT. Mohammad et al. had measured locomotor activity (reported as average speed in their paper (total distance walked/time of recording), but primarily to test the dependency of WAFO on locomotor activity. They found that WAFO metrics were poorly correlated with average walking speed, showing a significant degree of independence of both measures – a finding that our results confirm. However, unlike us, they did not consider average speed/TOWA further for their analysis, possibly because they focused on anxiety-like behaviour while our study looked broader on emotion-like behaviour in general. We further used a round (not square) arena to exclude “cornering” in order to reduce the complexity of the assay, which may explain differences of observed speed/TOWA between our studies.

      Moreover, while WAFO is a good proxy for 'stress', I am not convinced that TOWA necessarily represents an emotional state in all cases. Indeed, as the authors themselves acknowledge, changes in total walking may be associated with other factors, such as starvation-induced hyperactivity, physical exhaustion after sleep deprivation, increased sex drive after mating, alcohol sedation, etc.

      Your comment raises a question in comparative research on emotions which is very difficult if not impossible to conclusively answer. At first sight, the most conservative stance seems to be to completely disregard the idea of emotions and affective experiences in animals. This, however, would mean that we cannot use animal models to study the basics of emotions (= emotion primitives) and would ignore that by all likelihood emotions are a product of evolution and hence should exist at least in more basic forms in animals. Obviously, we have no means to ask flies or any other animal whether they connect “hunger” or “mating” to a feeling or an emotional state (which must not be conscious) but can only observe the behaviour. We here adopt the often-cited “Pankseppian” view (based on the book of Jaak Panksepp: “Affective Neuroscience”) and firmly believe that – in order to fully understand how the brain drives behaviour- we also need to take affective states into account that bias behaviour towards adaptive responses.

      In short, we are unfortunately unable to give a clear and definite answer to your comment whether TOWA represents an emotional state in all cases. Perhaps you are right. We believe, however, that a “Pankseppian” view is adequate, and we may ask what evidence exists that shows that starvation-induced hyperactivity or post-mating is not associated with an affective emotion-like state in the fly or any other animal.

      Another unclear point is the interpretation of some unexpected results, such as the finding that both serotonin transporter overexpression and its knockdown give the same phenotype.

      Thank you very much for this comment, which we also received by reviewer #1. As suggested by the other reviewer, a reasonable mechanism that may underly the similarity of effect after knockdown or overexpression of SerT may be a differential effect on the different serotonin receptor subtypes expressed in the brain. The possibility that the concentration-dependent effect of a biogenic amine follows a U-shape is further reasonable and has been demonstrated for dopamine. For example, the relationship between cognitive performance or working memory and dopamine levels in primates follows an inverted U-shape (see e.g. Cools and D’Esposito 2011 10.1016/j.biopsych.2011.03.028, Desimone 1995 10.1038/376549a0). Also in Drosophila, both reduced and increased dopamine levels lead to increased male-to-male courtship behaviour (Liu et al. 2008 10.1523/JNEUROSCI.5290-07.2008, Liu et al. 2009 10.1371/journal.pone.0004574). While we are unaware of similar examples for serotonin, we note that serotonin levels must be kept at optimum level during development – both higher and lower than optimum levels result in behavioural impairments in adults (see e.g., Shah et al. 2018 10.3389/fnbeh.2018.00114). We have now extended the discussion accordingly.

      Finally, there are some issues with the use of the OFT in rodent research (e.g., inconsistent effects of anxiolytic drugs; see Rosso et al., 2022, Neurosci Biobehav Rev., for a meta-analysis). These should be explained to place the Drosophila findings in their appropriate context.

      Thank you very much for bringing this systematic review to our attention which assessed the usefulness of various behavioural tests including the OFT to study the effect of anxiolytic drugs in rodents. Overall, the review casts “serious doubt on both construct and predictive validity” of behavioural tests for anxiolytics. While diazepam (the only drug used in our study) was the drug with the most consistent effects across the analysed behavioural assays, only 59% of the OFTs revealed significant effects. We were already aware of earlier findings in the same direction (Prut and Belzung 2003 10.1016/s0014-2999(03)01272-x), but as we only used one drug did not include a discussion in the manuscript. Unfortunately, the number of studies employing the OFT in flies is very small and does not yet allow for a similar comparison. We now changed the respective sentences in the discussion:

      “In rodents, diazepam mostly but not consistently leads to an anxiolytic response in the OFT behaviour which questions the usefulness of the OFT for testing anxiolytic drugs (see (Prut and Belzung 2003; Rosso et al. 2022)).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Overexpression of SerT suppressed increased WAFO after electric shocks. However, knockdown of SerT did similar. It is worth reporting this finding, but possible mechanisms could be mentioned in the manuscript. For example, given that flies carry five serotonin receptor genes, upregulation of overall serotonin level may affect the specific serotonin receptor, but downregulation of it may affect other receptors, leading to unexpected outcomes. Exploring any other possibilities could be better to add to guide future research.

      Thank you very much for this comment and the suggestion of a reasonable mechanism that may underly the similarity of effect after knockdown or overexpression of SerT. The possibility that the concentration-dependent effect of a biogenic amine follows a U-shape is further reasonable and has been demonstrated for dopamine. For example, the relationship between cognitive performance or working memory and dopamine levels in primates follows an inverted U-shape (see e.g. Cools and D’Esposito 2011 10.1016/j.biopsych.2011.03.028, Desimone 1995 10.1038/376549a0). Also in Drosophila, both reduced and increased dopamine levels lead to increased male-to-male courtship behaviour (Liu et al. 2008 10.1523/JNEUROSCI.5290-07.2008, Liu et al. 2009 10.1371/journal.pone.0004574). While we are unaware of similar examples for serotonin, we note that serotonin levels must be kept at optimum level during development – higher or lower levels result in behavioural impairments in adults (see e.g. Shah et al. 2018 10.3389/fnbeh.2018.00114). We have now extended the discussion accordingly.

      Reviewer #2 (Recommendations for the authors):

      (1) The advance over Mohammad et al., 2016, Curr Biol. must be clearly and emphatically articulated in the Introduction and/or Discussion. It is otherwise impossible to appreciate what is conceptually novel about this work.

      We have now rewritten parts of the two last paragraphs in the introduction to make the advances clearer. As outlined above, we conceptually advanced the analysis of OFT behaviour by integrating TOWA as a second and independent dimension in our analysis. During the revision process, we have spent great effort to better characterise the nature of the locomotor activity encountered in the OFT (Figure 1 – supplementary figures 1 and 2). Also, this is an advancement over previous studies, including Mohammad et al. 2016. We further included new results on the general effect of neuropeptides (silver mutants, impaired in neuropeptide processing).

      (2) The most important metric for a stress-like emotional primitive is WAFO. Therefore, figures should be revised in a way that highlights WAFO differences. Figure 4 is a good example of this. In contrast, in Figures 1-3, the WAFO box-and-whisker plots are very small and obscured under the raw WAFO-TOWA plots, which are difficult to see (especially given the light blue background of all plots) and redundant. I strongly recommend just showing the WAFO box-and-whisker plots for the sake of visibility, clarity, and brevity.

      Although we understand your reasoning, we would like to stick with the old figures as we consider TOWA as important to characterise the OFT response as WAFO (see our comments above regarding the conceptual advances of our study over Mohammad et al. 2016). It is true that Figures 1-3 are small, but at least in our print-out well legible. Further, it appears that in the current version of the eLife system the resolution is downsampled. In addition, we anticipate that in the version of record figures can be enlarged online as in other eLife articles.

      (3) Given the criticisms against the OFT in rodent research, one of which is inconsistent effects of anxiolytic drugs (Rosso et al., 2022, Neurosci Biobehav Rev.), it may be useful to expand pharmacological treatments beyond diazepam.

      As our focus is not on the testing of anxiolytic drugs and since it was already very difficult to be granted access to diazepam (we are not at a medical institution), we refrained from testing further drugs. Moreover, as rightfully mentioned by you, the OFT may not be the best test for the efficacy of anxiolytic drugs. On the other hand, diazepam was the most consistent anxiolytic in the OFT in rodents (see Rosso et al. 2022).

      (4) The authors should show results of the effects of at least some stressors/punishments on WAFO/TOWA of female flies to understand if observed effects are sex-specific.

      Thank you very much for bringing this topic to our attention. To test whether the effects are sex-specific, we now performed several new experiments. First, we compared the naïve OFT response of mated and unmated males and females (see new Figure 6). This revealed that without prior stress treatment, the WAFO response is independent of sex and mating status. In contrast, the naïve TOWA response turned up to be sex- and mating state-specific (see new Figure 6). To test whether the mated females show a different stress-induced OFT response to males, we applied mechanical stress (shake) that we had also used to assess the effects of arena diameter (Figure 1 – supplementary Fig. 1 D-D’’). After a first round of shaking, females showed increased TOWA, but WAFO was unaffected. A second round of shaking, however, led to a significant increase in WAFO and TOWA. This suggests that the OFT response is qualitatively similar between the sexes and mating status, yet the threshold for elicited responses differs between males and females. We now added a respective paragraph to the main text in the results section plus a new figure (Fig. 6).

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an interesting study that addresses whether mitochondrial DNA (mitoDNA) variants impact telomere length (TL), which may be relevant to potential maternal inheritance of TL in offspring. The study addresses this question using a cybrid model approach in which mitochondria from donor platelets from 7 individuals that vary in TL and differ in mitoDNA variants are introduced into 143B cells that lack mitochondria. MitoDNA variants that exhibited reduced complex I activity showed telomere shortening in cybrids and increased telomere dysfunction. Interestingly, these phenotypes could be reduced with NAC antioxidant and NAD+ supplementation, suggesting that ROS and oxidative DNA damage at telomeres contributed to the telomere shortening. They further showed that cybrids with lower levels of ROS correlated with longer TL in the lymphocytes of the mitochondrial donors.

      Strengths:

      This study provides compelling evidence that mtDNA variants influence TL through a mechanism involving mitochondrial-derived ROS, potentially causing telomeric oxidative damage. The data are robust, and the manuscript is well written. However, the study could be strengthened by addressing the following questions and minor weaknesses below.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      (1) Introduction. Line 81, the relationship between TL and the risk of lymphoid and myeloid leukemia is not straightforward. POT1 variants associated with long TL increase the risk for lymphoid and myeloproliferative neoplasms (see PMID: 41564438 for example).

      We appreciate the reviewer raising the important link between pathogenic POT1 variants and lymphoid malignancies driven by elongated telomeres. We would like to clarify, however, that our introduction focused not on the pathological telomere attrition characteristic of telomere biology disorders, but rather on the natural, non-pathological variations observed in individuals with baseline telomere lengths on the shorter end of the spectrum.

      Nevertheless, we will include this observation regarding patients with long telomeres in the introduction to underscore the complex relationship between telomere length and tumorigenesis.

      The two following publications by the Armanios lab will be added:

      DeBoy EA et al. Familial clonal hematopoiesis in a long telomere syndrome. 2023. N Engl J Med 388, 2422-2433.

      Davidson-Swinton HR et al. Lymphoid malignancy and clonality in the POT1-mediated long telomere syndrome. 2026. Blood 147, 2226-2237.

      (2) Figure 1. Since sex also influences TL, it would be good to know the sex of the selected individuals or explain why this is not necessary.

      Because this study investigated the potential maternal inheritance of TL via the mitochondrial genome, our cohort consisted predominantly of female donors (6/7). Donor #5, the husband of Donor #7, was the only male included. We will update Figure 1D to include the sex of each donor and clarify this rationale in the text.

      Notably, our analysis revealed minimal influence of sex on TL, which cannot account for the observed differences between the extreme groups.

      (3) Please include a description of the 143B cells that were used for cybrid formation in the Results section when introducing the cybrids.

      We will do so.

      (4) Lines 155-156. The authors note that cybrids from donors 1 and 2 show "pronounced" telomere damage. This result indicates an increase in 53BP1-positive telomeres, which could be indicative of telomere dysfunction or damage. Quantification of the increased chromosome end fusions for cybrids 1 and 2 would strengthen the result.

      We thank the reviewer for this suggestion. Following their advice, we used our metaphase spread FISH analyses to quantify chromosome end fusions in the parental 143B Rho0, Cybrid 1 and Cybrid 6 cells. However, because Cybrid 2 metaphase spreads were of insufficient quality for adequate chromosome analysis, we will restrict our telomere fusion comments exclusively to Cybrid 1 and include the quantifications in our revised manuscript.

      Author response image 1.

      The number of fusions and total chromosomes analyzed are indicated on each bar.

      Do the increased fusions correlate with an increase in telomere signal-free ends? These should be apparent in the telomere FISH images of metaphase chromosomes.

      While this is a strong argument, the telomeres within these cybrid models are critically short. Consequently, the FISH signal intensity falls below the threshold required for reliable quantification of telomere-free ends.

      (5) Lines 168-169. What is the evidence that the "in vitro metabolic shift" causes acute oxidative stress?

      The reviewer is correct that we have not formally demonstrated this in our experimental system. Instead, our assumption was based on the fact that the initial phase of Rho0 cell repopulation involves a temporary ROS burst, partly driven by incompletely assembled ETC supercomplexes that are known to elevate ROS levels (Maranzana et al, 2013). This is supported by our experimental observation that the NAC antioxidant, combined with NR, potently inhibits telomere shortening in cybrids with low CI activity. We propose to add this explanation in the revised manuscript.

      (6) Why did the elevated ROS in cybrid #3 (Figure 4C) not translate to shorter telomeres in the cybrid (Figure 2A)? Perhaps there is a difference between factors that determine TL in the cybrid vs the donor's lymphocytes?

      We cannot fully explain the discrepancy, but we indeed suspect in vivo oxidative stress differs significantly from cell cultures (21% O<sub>2</sub>). Additionally, early cybrid replenishment involves an unknown telomere elongation step that may offset mitochondrial ROS-induced shortening.

      To further investigate this question, we measured telomere length across varying population doublings (PDs) and observed the following shortening after 66-67 PDs:

      - Cybrid 3: about 1.6 kb reduction

      - Cybrid 6: about 1.2 kb reduction

      - Cybrid 7: about 0.8 kb reduction

      These results suggest that the rate of telomere shortening in culture may be higher in Cybrid 3 cells, possibly due to increased ROS levels.

      We propose to include (as Supplementary figure) and discuss these data in the revised manuscript.

      Author response image 2.

      In Figure 4B, it appears that the statistical comparisons for mitochondrial superoxide are all relative to Cyb3. If so, why are the comparisons not with the parental 143B rho0 cell line? Please clarify.

      We excluded Rho0 cells from our superoxide measurements because these cells were grown in a different culture medium. Instead, we focused on comparing mitochondrial ROS across cybrids to accurately correlate these values with the telomere length of the corresponding donors’ lymphocytes. Consequently, Rho0 cell measurements would not have contributed to this correlation analysis.

      We propose to discuss this in the revised manuscript.

      (7) Given the heterogeneity in TL and mtDNA variants in the human population, the conclusions could be further strengthened by increasing the number of donors and cybrids analyzed. However, there are admittedly practical factors. Overall, these findings are compelling and provide a solid foundation for expanding this analysis in the future. This is more of a comment than a weakness.

      We thank the reviewer for this constructive feedback and entirely agree that including more donors would have added depth to our findings. While we acknowledge this limitation, pursuing this further is currently impossible without acquiring new ethical approvals and establishing fresh collaborations with clinicians.

      Reviewer #2 (Public review):

      Summary:

      The authors aim to determine whether mitochondrial genotype influences telomere length. By generating cybrids harboring different mitochondrial backgrounds, the authors seek to establish a mechanistic link between mitochondrial status and telomere biology.

      Strengths:

      A major strength of the study is the use of cybrid technology, which provides a great approach to investigate the role of mitochondrial DNA independently of the nuclear genome. The authors also employ multiple complementary assays to assess telomere-related phenotypes associated with mitochondrial dysfunction. Together, these experiments generate an interesting dataset that will be of value to researchers interested in the intersection between mitochondrial biology, genome stability, aging, and development. These results also build on previous work supporting roles for ROS/mitochondria in driving telomere shortening.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      The data support the conclusion that mitochondrial background is associated with differences in telomere length and telomere-related phenotypes. However, some of the mechanistic interpretations would benefit from additional evidence. In particular, the manuscript discusses mitochondrial influences on telomere shortening, yet telomere length in some experiments is assessed at a single time point. Consequently, the current data do not directly address the rate of telomere attrition. Differences observed between cybrid lines could potentially arise from events occurring during cybrid formation, clonal selection, or subsequent cell expansion. Longitudinal analyses across multiple passages, ideally beginning immediately after cybrid generation and controlling for population doublings, would help establish whether mitochondrial function directly affects telomere shortening dynamics. Some experimental results would also benefit from additional quantification, clarification, and some biological replicates are missing.

      We limited our TL measurements to the earliest viable time point after cybrid formation to avoid the confounding effects of cellular adaptation in culture. For instance, the early telomere shortening observed in Cybrid 1 and Cybrid 2 was later alleviated—likely due to the upregulation of the NAD+ salvage pathway genes NAMPT and NAPRT1.

      To accurately capture the effects specific to cybrid formation, we isolated 4 independent clones per donor, all of which showed highly consistent TL values, as shown in Figure 2A and S3D.

      While we acknowledge the reviewer's point about multi-passage longitudinal analyses, we did, in fact, measure telomere length across varying population doublings (PDs) in Cybrid 3 (high mito ROS levels), 6 and 7 (low mito ROS levels) and observed the following shortening after 66-67 PDs:

      - Cybrid 3: about 1.6 kb reduction

      - Cybrid 6: about 1.2 kb reduction

      - Cybrid 7: about 0.8 kb reduction

      These results suggest that the rate of telomere shortening in culture may be higher in Cybrid 3 cells, possibly due to increased ROS levels.

      We propose to include a Supplementary figure and discuss these data in the revised manuscript.

      We further propose to carefully edit the manuscript so as to clarify the text and, whenever possible, add quantifications. Among others, as suggested by Reviewer #1, we will add the quantification of chromosome end fusions in the parental 143B Rho0, Cybrid 1 and Cybrid 6 cells in our revised manuscript (See Author response image 1).

      Overall, this study provides interesting evidence linking mitochondrial background to telomere biology. The cybrid models represent a useful resource for the field, and the work raises important questions regarding mitochondria-telomere communication.

      Reviewer #3 (Public review):

      Strengths:

      Mahieu and colleagues address an interesting and underexplored question: whether non-pathogenic variation in the mitochondrial genome contributes to the inter-individual variability of human telomere length (TL). Using a Belgian Flow-FISH reference cohort (n=491) to identify donors at TL extremes, they generate transmitochondrial cybrids from platelets of seven donors of distinct mtDNA subhaplogroups and characterize the resulting cells with a broad and well-executed toolkit (TRF, TeSLA, ddTRAP, EPR-based mitoROS, Seahorse with permeabilized-cell ETC dissection, LC-MS metabolomics, telomeric PAR-FISH). The most compelling finding is that cybrids derived from donors with low complex I (CI) activity undergo rapid telomere shortening during the glycolysis-to-OXPHOS transition of cybrid formation, and that this is largely prevented by co-treatment with NAC and the NAD⁺ precursor nicotinamide riboside, supporting a model in which CI sustains the NAD⁺ pool required for PARP1-mediated repair of oxidative damage at telomeres. The authors further report an inverse correlation between donor lymphocyte TL and mitoROS in the corresponding cybrids, and provide preliminary evidence that the K1a-defining ATP6 A177T variant (m.G9055>A) may be enriched in long-telomere individuals.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      (1) Statistical support and donor sampling for the central in vivo correlation (Figure 4C).

      The inverse correlation between donor lymphocyte TL and cybrid mitoROS (R<sup>2</sup>=0.794, p=0.007) is the principal in vivo claim of the paper, but it is built on seven donors deliberately selected from the extremes of the Flow-FISH distribution. Sampling at the tails of the outcome variable can substantially inflate apparent correlation strength and significance. I would encourage the authors to (i) explicitly state this sampling structure where the correlation is introduced, (ii) report a leave-one-out sensitivity analysis to confirm the relationship is not driven by one or two donors (Cyb3 and Cyb6 appear to anchor the line), and (iii) where feasible, extend the analysis to additional donors with intermediate TL to test whether the relationship holds across the full distribution. Even a modest expansion (e.g., 4 to 5 additional donors at P25 to P75) would substantially strengthen this central claim.

      We thank the reviewer for this insightful comment. As suggested, we will explicitly describe the sampling structure upon introducing the correlation analysis. Furthermore, we have conducted the requested leave-one-out analysis:

      - removing Cyb3: R<sup>2</sup>=0.730; p=0.0302

      - removing Cyb6: R<sup>2</sup>=0.782; p=0.0194

      - removing Cyb1: R<sup>2</sup>=0.740; p=0.0278

      While we agree that additional donors would enhance the study, further experiments are however currently impossible without new ethical clearances and additional clinical partnerships.

      (2) Reconciling the cybrid CI / TL relationship (Fig 3B) with the absence of a CI / TL relationship in donor lymphocytes (Figure 4A).

      Figure 3B shows a strong correlation between CI activity and TL in cybrids (R<sup>2</sup>=0.87), while Figure 4A shows no correlation between donor CI activity (measured in the same cybrids) and donor lymphocyte TL. The authors acknowledge this, but the manuscript subsequently builds toward a CI-centric model of in vivo TL regulation, which seems to outrun the data. The most internally consistent interpretation is that the cybrid CI phenotype reports a sensitized in vitro response to the acute oxidative stress of the metabolic shift, rather than a steady-state determinant of leukocyte TL. I would suggest reframing the abstract, significance statement, and Discussion to make this distinction clearer. The in vitro CI / NAD⁺ / PARP1 axis is a strong finding on its own, while the in vivo role of CI activity (as opposed to ROS more broadly) is not yet established here. Donor #1's profile (very long lymphocyte TL, low CI activity, severe shortening in cybrids, no telomere inheritance in offspring) is informative in this regard and could be discussed more directly as a case that helps delineate where the cybrid model does and does not recapitulate in vivo biology.

      We acknowledge that our study does not establish the in vivo role of CI activity in TL regulation. Our abstract specifically highlights an in vitro phenomenon: “Under the specific conditions of cybrid formation, which involve a metabolic shift from glycolysis to oxidative phosphorylation, mtDNA variants associated with reduced CI activity induced rapid telomere shortening, …”.

      We are nevertheless happy to revise the text to clearly separate our in vitro results from in vivo biology as requested.

      (3) The K1a / ATP6 A177T inheritance claim.

      The proposal that K1a (and specifically ATP6 A177T) contributes to maternal inheritance of long telomeres is intriguing but currently rests on three pedigrees (one of which, donor #1, does not support the hypothesis) and a chi-square test that does not reach significance (p=0.153, Figure 4F). The supporting evidence is also limited by the fact that platelet-mediated mitochondrial transfer delivers donor mitochondrial proteins, lipids, and residual mtRNA in addition to mtDNA, making it difficult to attribute the cybrid phenotype of donor #6 specifically to the ATP6 A177T variant. I would recommend either: (a) extending the genotyping screen to additional unrelated donors and, if feasible, confirming the effect of ATP6 A177T through an isogenic approach (e.g., mtDNA base editing in a clean background), or (b) softening the relevant statements to "suggestive trend warranting larger studies," and presenting the K1a observation as hypothesis-generating rather than supportive. The Ashkenazi-centenarian connection raised in the Discussion is an excellent direction for follow-up and could be framed accordingly.

      We agree that the evidence for the AT6 A177T inheritance claim remains inconclusive. To clarify, we do not argue that this mitochondrial variant is solely responsible for longer telomeres; indeed, the mtDNA genome of donor #1 suggests otherwise. Furthermore, the phenotypic impact of such mtDNA variants likely depends on nuclear variants in other telomere-related genes (e.g., hTERT or hTR), meaning AT6 A177T may not consistently result in elongated telomeres. Unfortunately, our ethical protocol precludes screening additional unrelated donors. We will revise the text to soften our statements accordingly.

      New references:

      DeBoy EA et al. Familial clonal hematopoiesis in a long telomere syndrome. 2023. N Engl J Med 388, 2422-2433.

      Davidson-Swinton HR et al. Lymphoid malignancy and clonality in the POT1-mediated long telomere syndrome. 2026. Blood 147, 2226-2237.

      Maranzana E et al. Mitochondrial respiratory supercomplex association limits production of reactive oxygen species from complex I. 2013. Antioxid Redox Signal 19, 1469-1480.

    1. Author response:

      We thank the reviewers for their assessment of our work and their comments. We are grateful for their evaluation of our findings as fundamental and convincingly supported, and for their appreciation of the relative scope of this manuscript and of future work. The most direct requests for new experimental data are from reviewer #2, who asks for direct assessment of the effects of CHOP deletion on expression of GADD34 and on protein synthesis. We agree that these are important experiments to conduct for the revision.

      The reviewers requested more clarity on the experimental logic of the paper and on the place of our findings in the broader context of ER stress signaling, which we will be happy to provide in a revised manuscript. These revisions will include a more explicit consideration of how the regulation of metabolic genes by CHOP contributes to its effects in the liver independently of its role in regulating eIF2a dephosphorylation.

      In particular, there were concerns about the logic of the time points chosen that we feel are important to also address here. For analysis of ChopHKO animals, all experiments were carried out 8 hours after ER stress challenge. This is because, as we show in Fig. 1B and also in our previous paper on CHOP (1), this is the time point at which CHOP expression is at its maximum. Thereafter, hepatocytes become heterogeneous with respect to whether they do or do not express CHOP. This is an interesting finding because it suggests that CHOP is part of a cellular switch, and potentially even an effector of that switch—a point currently raised in the Discussion but worth further highlighting in a revision. At the practical level, it means that discerning the contribution of CHOP to ER stress signaling and adaptation at subsequent time points will require sophisticated single cell analyses that can discriminate cells that express CHOP from cells that do not, which are an important future direction.

      In contrast, for Atf6aHKO animals, all experiments were carried out 48 hours after ER stress challenge. As we have previously shown (2), at short time points after a stress challenge, such as 8 hours, there is very little difference in ER stress signaling between wild-type animals and those lacking ATF6a. The reason for this lack of distinction is that the major targets of ATF6a are ER chaperones and the like. Because adaptation to ER stress in the early phases of the response depends more on non-transcriptional mechanisms such as inhibition of protein synthesis and IRE1-dependent mRNA decay (RIDD), the failure to fully upregulate ATF6a targets is initially of little consequence. It is only at later time points when wild-type animals restore ER homeostasis and largely silence ER stress signaling. In contrast, at these same later points, animals lacking ATF6a show evidence of persistent ER stress, most notably in the form of persistent Xbp1 mRNA splicing and profound suppression of metabolic genes. Although the 8 hour time point for experiments in ChopHKO animals differs from the 48 hour time point for Atf6aHKO animals, the two lines of experimentation are united by the persistence of ongoing ER stress and of ISR signaling despite diminished eIF2a phosphorylation at the points when the presence of CHOP or the absence of ATF6a are of the greatest impact. A revised manuscript will present this logic more clearly.

      References

      (1)  Liu K, et al., EMBO Reports 25, 228 (2024)

      (2)  Rutkowski DT, et al., Dev. Cell 15, 829 (2008)

    1. Author response:

      We would like to express our deepest gratitude to the Editors and Reviewers for their highly rigorous and constructive evaluation of our manuscript. We are greatly encouraged by the recognition of our study’s ambition, the unique value of the in vivo intrathecal contrast MRI dataset, and the conceptual novelty of linking macroscopic glymphatic physiology with neural activity and regional proteopathy.

      We fully agree with the thoughtful limitations and methodological concerns raised in the eLife Assessment and the Public Reviews. In our upcoming revised manuscript, we are implementing a comprehensive set of revisions to address these points. Specifically, our planned revisions focus on the following key areas:

      - Tempering Causal Interpretations: We agree that our cross-sectional design precludes definitive causal inferences. We are systematically revising the manuscript to soften causal language (e.g., replacing "drives" with "is spatially associated with"). We will explicitly frame our findings as macroscopic spatial associations and discuss the potential influence of joint physiological confounders.

      - Tightening Terminology and Imaging Physics: We are refining our terminology to more accurately reflect our MRI measurements. We will replace assertive terms like "direct glymphatic flow" with precise descriptors such as "imaging proxies for tracer enhancement and retention." Furthermore, we are expanding the Limitations section to explicitly acknowledge the confounding effects of Partial Volume Averaging (PVE), systemic tracer redistribution, and renal clearance kinetics.

      - Conducting Supplementary Imaging & Robustness Analyses: To address concerns regarding cohort heterogeneity and the sample size of the rs-fMRI subgroup (n=15), we are performing a series of rigorous supplementary analyses. This includes conducting sensitivity analyses (e.g., excluding the motor neuron disease subgroup) and applying leave-one-out cross-validation to rigorously assess the subject-level stability and robustness of the spatial coupling between neural activity and tracer clearance.

      - Clarifying the Conceptual Model and "Mismatch" Index: To improve readability, we are moving the anatomical definitions of the cortical gradients directly into the Results section. Additionally, we are introducing schematic diagram to intuitively explain the mathematical formulation and biological interpretation of the "activity-clearance mismatch" index.

      - Re-framing External Dataset Analyses: We are carefully re-framing the interpretations of the Allen Human Brain Atlas (AHBA) transcriptomic data and the external PiB-PET amyloid dataset, emphasizing that these reflect spatial correspondences of intrinsic regional vulnerability across groups, rather than individual-level direct interactions.

      We believe these revisions will significantly enhance the scientific rigor, clarity, and precision of our study.

  2. Jul 2026
    1. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public Review):

      This work provides a new dataset of 71,688 images of different ape species across a variety of environmental and behavioral conditions, along with pose annotations per image. The authors demonstrate the value of their dataset by training pose estimation networks (HRNet-W48) on both their own dataset and other primate datasets (OpenMonkeyPose for monkeys, COCO for humans), ultimately showing that the model trained on their dataset had the best performance (performance measured by PCK and AUC). In addition to their ablation studies where they train pose estimation models with either specific species removed or a certain percentage of the images removed, they provide solid evidence that their large, specialized dataset is uniquely positioned to aid in the task of pose estimation for ape species.

      The diversity and size of the dataset make it particularly useful, as it covers a wide range of ape species and poses, making it particularly suitable for training off-the-shelf pose estimation networks or for contributing to the training of a large foundational pose estimation model. In conjunction with new tools focused on extracting behavioral dynamics from pose, this dataset can be especially useful in understanding the basis of ape behaviors using pose.

      We thank the reviewer for the kind comments.

      Since the dataset provided is the first large, public dataset of its kind exclusively for ape species, more details should be provided on how the data were annotated, as well as summaries of the dataset statistics. In addition, the authors should provide the full list of hyperparameters for each model that was used for evaluation (e.g., mmpose config files, textual descriptions of augmentation/optimization parameters).

      We have added more details on the annotation process and have included the list of instructions sent to the annotators. We have also included mmpose configs with the code provided. The following files include the relevant details:

      File including the list of instructions sent to the annotators:

      OpenMonkeyWild Photograph Rubric.pdf

      Mmpose configs:

      i) TopDownOAPDataset.py

      ii) animal_oap_dataset.py

      iii) init.py

      iv) hrnet_w48_oap_256x192_full.py

      Anaconda environment files:

      i) OpenApePose.yml

      ii) requirements.txt

      Overall this work is a terrific contribution to the field and is likely to have a significant impact on both computer vision and animal behavior.

      Strengths:

      Open source dataset with excellent annotations on the format, as well as example code provided for working with it.

      Properties of the dataset are mostly well described.

      Comparison to pose estimation models trained on humans vs monkeys, finding that models trained on human data generalized better to apes than the ones trained on monkeys, in accordance with phylogenetic similarity. This provides evidence for an important consideration in the field: how well can we expect pose estimation models to generalize to new species when using data from closely or distantly related ones?

      Sample efficiency experiments reflect an important property of pose estimation systems, which indicates how much data would be necessary to generate similar datasets in other species, as well as how much data may be required for fine-tuning these types of models (also characterized via ablation experiments where some species are left out).

      The sample efficiency experiments also reveal important insights about scaling properties of different model architectures, finding that HRNet saturates in performance improvements as a function of dataset size sooner than other architectures like CPMs (even though HRNets still perform better overall).

      We thank the reviewer for the kind comments.

      Weaknesses:

      More details on training hyperparameters used (preferably full config if trained via mmpose).

      We have now included mmpose configs and anaconda environment files that allow researchers to use the dataset with specific versions of mmpose and other packages we trained our models with. The list of files is provided above.

      Should include dataset datasheet, as described in Gebru et al 2021 (arXiv:1803.09010).

      We have included a datasheet for our dataset in the appendix lines 621-764.

      Should include crowdsourced annotation datasheet, as described in Diaz et al 2022 (arXiv:2206.08931). Alternatively, the specific instructions that were provided to Hive/annotators would be highly relevant to convey what annotation protocols were employed here.

      We have included the list of instructions sent to the Hive annotators in the supplementary materials. File: OpenMonkeyWild Photograph Rubric.pdf

      Should include model cards, as described in Mitchell et al (arXiv:1810.03993).

      We have included a model card for the included model in the results section line 359. See Author response image 1:

      Author response image 1.

      It would be useful to include more information on the source of the data as they are collected from many different sites and from many different individuals, some of which may introduce structural biases such as lighting conditions due to geography and time of year.

      We agree that the source could introduce structural biases. This is why we included images from so many different sources and captured images at different times from the same source—in hopes that a large variety of background and lighting conditions are represented. However, doing so limits our ability to document each source background and lighting condition separately.

      Is there a reason not to use OKS? This incorporates several factors such as landmark visibility, scale, and landmark type-specific annotation variability as in Ronchi & Perona 2017 (arXiv:1707.05388). The latter (variability) could use the human pose values (for landmarks types that are shared), the least variable keypoint class in humans (eyes) as a conservative estimate of accuracy, or leverage a unique aspect of this work (crowdsourced annotations) which affords the ability to estimate these values empirically.

      The focus of this work is on overall keypoint localization accuracy and hence we wanted a metric that is easy to interpret and implement, in this case we made use of PCK (Percentage of Correct Keypoints). PCK is a simple and widely used metric that measures the percentage of correctly localized keypoints within a certain distance threshold from their corresponding groundtruth keypoints.

      A reporting of the scales present in the dataset would be useful (e.g., histogram of unnormalized bounding boxes) and would align well with existing pose dataset papers such as MS-COCO (arXiv:1405.0312) which reports the distribution of instance sizes and instance density per image.

      We have now included a histogram of unnormalized bounding boxes in the manuscript, see Author response image 2:

      Author response image 2.

      Reviewer #2 (Public Review):

      The authors present the OpenApePose database constituting a collection of over 70000 ape images which will be important for many applications within primatology and the behavioural sciences. The authors have also rigorously tested the utility of this database in comparison to available Pose image databases for monkeys and humans to clearly demonstrate its solid potential.

      We thank the reviewer for the kind comments.

      However, the variation in the database with regards to individuals, background, source/setting is not clearly articulated and would be beneficial information for those wishing to make use of this resource in the future. At present, there is also a lack of clarity as to how this image database can be extrapolated to aid video data analyses which would be highly beneficial as well.

      I have two major concerns with regard to the manuscript as it currently stands which I think if addressed would aid the clarity and utility of this database for readers.

      (1) Human annotators are mentioned as doing the 16 landmarks manually for all images but there is no assessment of inter-observer reliability or the such. I think something to this end is currently missing, along with how many annotators there were. This will be essential for others to know who may want to use this database in the future.

      We thank the reviewer for pointing this out. Inter-observer reliability is important for ensuring the quality of the annotations. We first used Amazon MTurk to crowd source annotations and found that the inter-observer reliability and the annotation quality was poor. This was the reason for choosing a commercial service such as Hive AI. As the crowd sourcing and quality control are managed by Hive through their internal procedures, we do not have access to data that can allow us to assess inter-observer reliability. However, the annotation quality was assessed by first author ND through manual inspections of the annotations visualized on all of the images the database. Additionally, our ablation experiments with high out of sample performances further vaildate the quality of the annotations.

      Relevant to this comment, in your description of the database, a table or such could be included, providing the number of images from each source/setting per species and/or number of individuals. Something to give a brief overview of the variation beyond species. (subspecies would also be of benefit for example).

      Our goal was to obtain as many images as possible from the most commonly studied ape species. In order to ensure a large enough database, we focused only on the species and combined images from as many sources as possible to reach our goal of ~10,000 images per species. With the wide range of people involved in obtaining the images, we could not ensure that all the photographers had the necessary expertise to differentiate individuals and subspecies of the subjects they were photographing. We could only ensure that the right species was being photographed. Hence, we cannot include more detailed information.

      (2) You mention around line 195 that you used a specific function for splitting up the dataset into training, validation, and test but there is no information given as to whether this was simply random or if an attempt to balance across species, individuals, background/source was made. I would actually think that a balanced approach would be more appropriate/useful here so whether or not this was done, and the reasoning behind that must be justified.

      This is especially relevant given that in one test you report balancing across species (for the sample size subsampling procedure).

      We created the training set to reflect the species composition of the whole dataset, but used test sets balanced by species. This was done to give a sense of the performance of a model that could be trained with the entire dataset, that does not have the species fully balanced. We believe that researchers interested in training models using this dataset for behavior tracking applications would use the entire dataset to fully leverage the variation in the dataset. However, for those interested in training models with balanced species, we provide an annotation file with all the images included, which would allow researchers to create their own training and test sets that meet their specific needs. We have added this justification in the manuscript to guide the other users with different needs. Lines 530-534: “We did not balance our training set for the species as we wanted to utilize the full variation in the dataset and assess models trained with the proportion of species as reflected in the dataset. We provide annotations including the entire dataset to allow others to make create their own training/validation/test sets that suit their needs.”

      And another perhaps major concern that I think should also be addressed somewhere is the fact that this is an image database tested on images while the abstract and manuscript mention the importance of pose estimation for video datasets, yet the current manuscript does not provide any clear test of video datasets nor engage with the practicalities associated with using this image-based database for applications to video datasets. Somewhere this needs to be added to clarify its practical utility.

      We thank the reviewer for this important suggestion. Since we can separate a video into its constituent frames, one can indeed use the provided model or other models trained using this dataset for inference on the frames, thus allowing video tracking applications. We now include a short video clip of a chimpanzee with inferences from the provided model visualized in the supplementary materials.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Please provide a more thorough description of the annotation procedure (i.e., the instructions given to crowd workers)! See public review for reference on dataset annotation reporting cards.

      We have included the list of instructions for Hive annotators in the supplementary materials.

      An estimate of the crowd worker accuracy and variability would be super valuable!

      While we agree that this is useful, we do not have access to Hive internal data on crowd worker IDs that could allow us to estimate these metrics. Furthermore, we assessed each image manually to ensure good annotation quality.

      In the methods section it is reported that images were discarded because they were either too blurry, small, or highly occluded. Further quantification could be provided. How many images were discarded per species?

      It’s not really clear to us why this is interesting or important. We used a large number of photographers and annotators, some of whom gave a high ratio of great images; some of whom gave a poor ratio. But it’s not clear what those ratios tell us.

      Placing the numerical values at the end of the bars would make the graphs more readable in Figures 4 and 5.

      We thank the reviewer for this suggestion. While we agree that this can help, we do not have space to include the number in a font size that would be readable. Smaller font sizes that are likely to fit may not be readable for all readers. We have included the numerical values in the main text in the results section for those interested and hope that the figures provide a qualitative sense of the results to the readers.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This rigorous and creative study uses an elegant combination of metabolomics, transcriptomics, and budding yeast molecular genetics to discover that (i) activating AMPK to maintain mitochondrial respiration fueled by cytosolic Acetyl CoA and (ii) increasing fatty acid synthesis independent of respiration drive independent pathways that increase the fitness of replicatively-aged budding yeast cells, albeit without increasing their lifespan. This work will be of interest to scientists in the field of aging and metabolism. Some clarifications in the text would address the following concerns, which would increase the impact of the study:

      (1) What does activation of AMPK (via PGDP-Sak1 expression) do to the replicative lifespan? How many bud scars, in general, do the subpopulations that are older - yet have less Tom70 (increased mitochondrial fitness) - have, after the 48 hrs timepoint that they are examining? How many divisions occurred in this 48hr time period - i.e. is it long enough to have all cells reach the end of their replicative lifespan? This information is important to rule out that a subset of the mutant cells just divided faster and hence had more divisions within 48 hrs (growing faster and living longer are different things). Having identical growth curves doesn't indicate per se that they all divide at the same rate, as there may be a subpopulation that divides faster and a subpopulation that doesn't grow so well.

      Increasing AMPK activity increases replicative lifespan [PMID: 25869125], but given our finding that AMPK activation splits the population, such replicative lifespan assays are hard to interpret. Bud scar counts have a similar issue. Hence we restricted the lifespan and bud scar analyses to wt and A2A which are more homogenous (Figures S2 B and E). A2A cells at 48 h have ~25% more bud scars than wt cells. Yes, by 48 h most of the cells have lost viability (Figure 2E). The reviewer is correct that you can't properly compare the lifespan curves if the cells divide at different rates, hence our follow-up test of wt at 48 h vs A2A at 40 h viability after we had confirmed that these time points captured cells at equivalent replicative ages (Figure 2D, E). This shows that viability of A2A is slightly lower than wt at matched age, indicating a slightly shorter lifespan. 

      (2) A2A cells do not have an extended replicative lifespan (RLS) but show an increase in the "low senescence" population (Figure 2). If the cells are not becoming senescent, why don't they have longer RLS? Not having a longer lifespan seems inconsistent with the statement that "bud scar counting confirmed that A2A cells reach a higher age than wild type", which comes back to how many times the cells can divide in the 48hr timepoint studied and their rate of cell division? Also, the lifespan curve shown is plotted against time, not cell division number, which does not take into account different division times of cells within the population (described above). It would be much more useful to show standard lifespan curves showing cell division numbers per lifespan per cell.

      Our observation that cells can reach the end of life without senescing is consistent with other studies that have studied the life course of individual cells by microscopy [PMID: 31291577, 32675375]. These studies always highlight some proportion of the cells that reach the end of life with no or minimal senescence, though this fraction varies with the experimental system. The question of why cells lose viability without senescing is a complete unknown in the field, and reflects a wider lack of consensus as to why yeast lose viability with replicative age.

      In liquid culture we can only assess viability over time, not cell division number, which we agree is not optimal and we are wary about making strong statements on lifespan for exactly the reasons the reviewer notes. Unfortunately, it is clear from the comparison of liquid and solid media lifespans performed by the Gottschling lab [PMID: 19652178] that culture system has a huge effect on lifespan, with cells in classical plate-based microdissection assays living far longer than the same strains do in liquid. This means that lifespans determined by microdissection-based assays are of questionable relevance to ageing studies performed in liquid culture. Senescence cannot be assayed on plates, while microfluidic systems lack the throughput necessary and preclude key techniques like RNA-seq, so liquid culture assays were the only option for this work. We agree that this leaves an unsatisfactory approximation for lifespan measurements, but we consider it critical that everything is measured in the same system. We therefore restricted our conclusion on lifespan to simply say that lifespan of A2A cells is not extended which our data in Figures 2D, E, S2B does support (see also answer to Q1), and therefore with the majority of A2A cells showing low senescence marks and high fitness at 48 h we can conclude that lifespan and fitness loss must be separable.

      We have added a note of these limitations of lifespan measurements in the materials and methods section of the manuscript.

      (3) Increased "fitness" of the old cells is implied from the increased size of the colonies that the old cells can make. However, this is a measure of the fitness of the daughters per se, not the old mother cells. Are the old mothers just passing on healthier mitochondria and more lipids to the daughters, such that they can divide more times? If the aged cells have an "increased fitness", why don't they divide more times themselves (i.e. live longer?).

      Yes, colony growth speed is defined by daughter cell replication, but as long as the daughters and subsequent generations divide at the same rate irrespective of whether they come from a young or old mothers then the size of the colony after 24 hours varies based on the time it took the initial mother to produce a daughter. This is what the assay really measures. We note that aged wildtype mothers often do not divide at all in the first 24 hours after being put on an agar plate (hence the tiny reported colony size), even though they do eventually produce a daughter which then forms a colony, whereas A2A cells tend to produce the first daughter rapidly whether young or old. It is known that daughters of aged wildtype mothers also divide slower, as to some extent do grand-daughters (PMID: 2644196), which will also contribute to differences in colony size, and this may well result from a lipid and/or mitochondrial contribution, but the primary driver of colony size in 24 hours is the time the mother took to initially divide. We have added this detail to the materials and methods section of the manuscript.

      As noted above, the mechanistic basis of lifespan is unknown, but although senescence can shorten lifespan, our work and that of others shows that lifespan is still limited in the absence of senescence.

      (4) The statement is made that "these experiments define two classes of aging cells with distinct metabolic needs, coherent with the model of two aging trajectories previously proposed (referencing Nan Hao's work)". However, the big difference here is that in Nan Hao's work, their two aging trajectories influenced the length of lifespan, but that does not appear to be the case here. That distinction should be made clear. Perhaps the authors could also speculate as to why the A2A yeast stops dividing after presumably the same number of cell divisions, even though they have an activated AMPK and activated fatty acid synthesis pathway.

      Yes, this is a good point and we have added this distinction to the Discussion:

      “Here we have characterised two classes of ageing cells seemingly differentiated by high and low availability of cytosolic Acetyl-CoA, consistent with a previous demonstration that ageing follows two trajectories in yeast though it should be noted that in this previous report, the two trajectories also differed in replicative lifespan (6).”

      We would love to speculate on why the A2A cells don't have an extended lifespan, but at this point we don't have a strong hypothesis. We have come up with many theories for this, but none that we haven’t managed to disprove experimentally. One thing worth considering is that many cells which lose replicative viability in liquid culture and probably in plate assays remain intact – for example, DNA and RNA integrity is not compromised over 24- 48 h – so those cells are probably not dead per se. But we also detect apoptosis-sized DNA fragments, which must come from dead cells, so there is clearly not a single mechanism defining the end of replicative lifespan.

      (5) I am a bit confused by the use of the word "senescence" by this lab here and in their previous growth on galactose studies. If yeast don't senesce, which is usually defined as an irreversible arrest of the cell cycle where cells stop dividing, shouldn't the yeast that do not senesce still be dividing and hence have a longer lifespan? Should a different term be used rather than senescence? Such as "fitness late in life". The authors giving their definition of senescence may help reduce this apparent contradiction.

      We completely agree, this is confusing and noted this distinction in the Introduction. Use of the term senescence to mean a loss of fitness late in life in yeast stems from the classical definition of senescence as applied to whole organisms. However, the term senescence as applied to cells has a more specific meaning in terms of the cell cycle as the reviewer notes. As an individual S. cerevisiae is both a cell and an organism, the terminology clashes. However, the marker we largely employ (Tom70-GFP) which in our hands is a very good proxy for fitness was originally defined as marking the senescence entry point (SEP), so overall we feel we can't avoid the term.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate how cytosolic acetyl-CoA metabolism influences replicative aging in budding yeast. They propose that acetyl-CoA regulates aging through three major pathways: (1) mitochondrial transport to support mitochondrial function, (2) fatty acid synthesis, and (3) global protein acetylation. The data show that AMPK activation promotes mitochondrial import of acetyl-CoA and partially mitigates mitochondrial decline in a subset of aging cells.

      Furthermore, the engineered A2A strain, which enhances mitochondrial acetyl-CoA utilization while relieving inhibition of fatty acid synthesis, increases the proportion of cells exhibiting a "low senescence" phenotype.

      Overall, this is a thoughtful and potentially impactful study that advances our understanding of metab to olic control of aging. Addressing the points below, particularly by refining interpretations and, where feasible, incorporating additional analyses, will further strengthen the manuscript and its conclusions.

      Strengths:

      The study has several notable strengths. It addresses an important question by shifting the focus from lifespan to preservation of late-life fitness, which is highly relevant to aging biology. The work integrates metabolic, genetic, and functional analyses to link cytosolic acetyl-CoA flux with distinct aging outcomes, and the engineering of the A2A strain provides a clear and elegant demonstration of how coordinated pathway modulation can improve cellular fitness.

      Weaknesses:

      (1) While the manuscript focuses on mitochondrial transport and fatty acid synthesis, cytosolic acetyl-CoA is also a key regulator of histone acetylation and chromatin silencing. It would strengthen the study to consider whether acetyl-CoA depletion contributes to improved fitness through enhanced rDNA silencing. Given the well-established role of rDNA instability in yeast aging, additional experiments examining rDNA silencing and stability would be valuable. For example, monitoring rDNA copy number changes (not necessarily ERCs) under AMPK activation, oleic acid supplementation, and in the A2A strain, similar to approaches used in the authors' prior work, would help clarify whether chromatin regulation contributes to the observed phenotypes.

      We have added data addressing these points to the manuscript and Supplemental Figures 2, 3 and 4, though the outcomes are complex. Histone acetylation changes chromatin accessibility and could therefore alter global gene expression; in accord with this, RNA-seq shows that P<sub>GPD</sub>-SAK1 reduces known age-linked gene expression dysregulation. However, A2A does not further reduce the effect, meaning either that another driver exists in addition to cytosolic acetyl-CoA, or that age-linked gene expression dysregulation is unrelated to cytosolic acetyl-CoA. Oleic acid has little effect on age-linked gene expression dysregulation despite rescuing fitness. With regard to rDNA silencing, transcription of the rDNA intergenic spacer non-coding RNAs promotes ERC formation; we have added data showing that ERC accumulation is not reduced in A2A but slightly higher coherent with the higher replicative age of A2A at 48 h, which suggests silencing is not better in A2A. By RNA-seq, these intergenic spacer transcripts are massively upregulated with age, but this will be a consequence of the increased genomic copy number on ERCs; the upregulation is less in A2A than other conditions, but this arises because the log phase spacer transcript levels are higher and so does not reflect better rDNA silencing. We have previously assayed for heritable changes in rDNA copy number arising during ageing and found (to our surprise) absolutely nothing, so we don't expect any changes under these conditions. The upregulation of transcripts from Sir2-repressed telomeric and MAT loci with age is decreased in P<sub>GPD</sub>-SAK1 and A2A, but the effect size is not different from any other low-expressed genes so we do not think there is a particular effect at loci subject to chromatin silencing (see our previous study Zylstra et al PMID 37643194 for evidence that Sir2-mediated gene silencing is not affected by age). We have added our conclusions from these experiments to the Discussion.

      (2) The current data do not fully distinguish whether AMPK activation and oleic acid supplementation act on distinct subpopulations of aging cells. An alternative explanation is that oleic acid supplementation enhances mitochondrial function and acts additively with AMPK activation, thereby increasing the fraction of cells in the "low senescence" state. Since this distinction is not central to the main conclusions, I suggest softening the language around subpopulation specificity. Emphasizing instead that the A2A strain coordinately modulates multiple branches of acetyl-CoA metabolism to improve late-life fitness would maintain the strength of the central message without over interpretation.

      We respectfully disagree with the reviewer on this point. We show that P<sub>GPD</sub>-SAK1 rescues senescence in ~half the population by a Cat2/Mls1 dependent mechanism (Figure 1F). We then show that in A2A, which rescues most cells, deletion of CAT2/MLS1 restores senescence in ~half the cells (Figure 3F/G). This cannot be explained by an additive mechanism as this would either result in all cells being partially rescued in the P<sub>GPD</sub>-SAK1 and in the A2A cat2Δ mls1Δ mutants, which is definitely not the case either by Tom70-GFP or fitness. Instead the population splits into high/low senescence and fit/unfit cells in the different assays.

      On the specific point of whether lipid synthesis additively increases mitochondrial function, we have added oxygen consumption rate data showing that A2A cells respire more than P<sub>GPD</sub>-SAK1 at 48h but only by a relatively small amount (Figure S3D), so there is indeed an additive improvement in mitochondrial function, but too little to explain the difference in population fitness in our opinion.

      We realise that the reviewer is asking more specifically about oleic acid, but again in the flow data, Figure 4C, what changes with oleic acid or P<sub>GPD</sub>-SAK1 is the proportion of cells in the low Tom70 / high WGA sector. Under an additive effect model, oleic acid or P<sub>GPD</sub>-SAK1 individually would partially reduce Tom70 and partially increase WGA, but the population in the low Tom70 / high WGA sector has the same average Tom70/WGA values in oleic acid, P<sub>GPD</sub>-SAK1 or P<sub>GPD</sub>-SAK1+oleic acid. It is the proportion of cells in this population that changes. Furthermore, under an additive model, wildtype cells aged with oleic acid would not have highest fitness than P<sub>GPD</sub>-SAK1 or A2A (Figure 4D) as these individual cells would lack the mitochondrial upregulation from P<sub>GPD</sub>-SAK1.

      (3) The manuscript proposes that lipid starvation and excess acetyl-CoA are major drivers of senescence in distinct subpopulations of wild-type aging cells. This conclusion is not yet fully supported by the presented data. Direct measurements of age-dependent divergence in acetyl-CoA and fatty acid levels at the single-cell level would be needed to substantiate this model. Based on the current evidence, a more conservative interpretation would be that aging cells exhibit differential sensitivity to perturbations in acetyl-CoA and lipid metabolism. Accordingly, I recommend revising the statement in the Abstract ("We further implicate lipid starvation and excess acetyl coenzyme A availability as major drivers of senescence...") and the corresponding discussion text to better align with the data.

      We agree and have adjusted the abstract to make it clearer that the lipid starvation / excess acetyl-coA interpretation is a model.

      “Our findings support a model in which lipid starvation and excess acetyl-coenzyme A availability are major drivers of senescence in replicatively aged wild-type yeast.”

      Reviewer #3 (Public review):

      Summary:

      These findings suggest that PGPD-SAK1 yeast show a subpopulation with lowered TOM70-GFP expression in high bud scar staining aged cells. Deletion of CAT2 or MLS1 reduces this effect. A PGPD-SAK1 acc1S1157A double mutant (called "A2A" here) shows an even larger effect of lowered tom70 expression in high bud scar staining aged cells. Utilization of various additional mutants involved in acetyl-CoA transport, carnitine shuttle, respiration, etc., leads the authors to conclude that these shifts in TOM70-GFP in aged cells are linked to the AMPK-fatty acid metabolic regulatory system.

      Strengths:

      These extensive and clearly described experiments reveal interesting changes in TOM70-GFP intensity in subsets of aged yeast in several mutants eventually identified as linked to the AMPK-fatty acid metabolic regulatory system.

      Weaknesses:

      (1) 3 biological replicates for mRNASeq is low.

      Thank you for pointing this out. We performed another replicate after posting the initial preprint to confirm the finding but didn’t update the figure in the eLife-reviewed version. We have added this to the scatter plots and analysis in Figure 1, there are minor changes but the set of genes we followed up are still highly significant. For ageing experiments, we sequence to n=3 as a first pass which is sufficient to detect widespread age-linked gene expression effects, and add more replicates if required to solidify findings for specific sets of genes. Hence, the additional RNAseq experiments we have added to the manuscript to Address Reviewer 2’s comments on widespread gene expression effects are also n=3-4.

      (2) While "Traditional conceptions of ageing implicate a progressive accumulation of damage leading to systemic degradation in performance until death, with evolutionary pressures acting to maximise early life fitness and fecundity at the expense of ageing health." is tangential perhaps to the data and conclusions of the study, both claims of this sentence are at best controversial, and the manuscript is no weaker for their omission.

      We would prefer not to remove this sentence, which we see as important to a major message of the manuscript: that ageing does not have to involve a loss of fitness before death. Outside the ageing biology field, ageing is often described as the progressive wearing out of components leading to decline and death (‘like an old car’ is a common analogy); in the ageing field this is certainly controversial, but outside the field it remains the normal understanding. This is what we mean by traditional conceptions, and it is important to consider the contradiction between this widely held viewpoint and our findings (and of course those of many others in the ageing field).

      The second part of the sentence about evolutionary pressures alludes to antagonistic pleiotropy, which we have now made explicit. Antagonistic pleiotropies as a driving mechanism for ageing, while not universally accepted, are as far as we can tell the most widely accepted type of theory in the ageing field. Our interpretation that yeast are bet-hedging as a population growth strategy and this drives ageing in the long term is a classic antagonistic pleiotropy and we need to raise this concept in the introduction.

      (3) The statement that "Here, we determine the basis of senescence and fitness loss in replicatively ageing yeast" is a bit strong as a summary of the present careful work presented here. If the authors had created yeast mutants that retained fitness indefinitely, this would be a more appropriate strength of claim to summarize the work.

      We agree and have moderated this sentence:

      “Here, we show that senescence and fitness loss in replicatively ageing yeast can be almost completely avoided without extension of lifespan by rewiring the conserved AMPK-fatty acid metabolic regulatory system.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The labelling of Figure 3G horizontal axis needs to be realigned with the data.

      Fixed – thank you.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 3G, the x-axis labels appear misaligned and should be corrected for clarity.

      Fixed – thank you.

      (2) Figures S3B and S3C appear to be mislabeled and should be revised.

      Fixed – thank you.

      (3) On page 6 (3rd paragraph), the statement that the beneficial impact arises from acetyl-CoA removal "rather than a benefit of respiration" may be overstated. The data support a role for acetyl-CoA removal but do not fully exclude a contribution from respiration. A more balanced phrasing would improve accuracy.

      We have revised this sentence and also added data:

      “Working in sip2Δ to avoid an increase in AMPK activity due to reduced Acetyl-CoA availability, we observed that ald6Δ increased the low senescence population through decreasing Tom70-GFP (S3C), and therefore the beneficial impact of PGPD-SAK1 on this pathway arises primarily through Acetyl-CoA removal. It is possible that respiration is adding to this benefit, and we detect a significant increase in Oxygen Consumption Rate in aged PGPD-SAK1 cells, but the further increase in A2A is smaller and we consider that this cannot fully explain the effect of acc1S1157A.”

      Reviewer #3 (Recommendations for the authors):

      This manuscript is clearly written, and the data are clearly presented. While 3 biological replicates is inadvisably low for mRNASeq, the subsequent experiments motivated by the genes identified there nevertheless stand on their own as presented.

      Thank you.

    1. Author response:

      The following is the authors’ response to the original reviews.

      This is a summary of the changes that have been made to the Reviewed Preprint:

      (1) The data from RettBASE which was analysed in the manuscript has been added in the form of four supplementary tables. Supplementary Table 1 contains the download of all MeCP2 mutations contained in RettBASE. Supplementary Tables 2-4 contain subsets of this data which were used in Figure 2B and Figure 3 SF2. Supplementary Table 4 also has the HGVS nomenclature for both e1 and e2 isoforms and the ClinVar Variation ID for each allele. Wording has been changed to clarify that analysis in the manuscript used this data from RettBASE and not the information that was deposited in ClinVar.

      (2) Similarly, Supplementary Tables 5-8 contain the gnomAD data that was used in the preparation of Figure 2, Figure 2 SF1 and Figure 3 SF1. The “high confidence” alleles in Supplementary Table 8 have been annotated with their HGVS names.

      (3) The criteria for selecting “high confidence” RettBASE and gnomAD alleles have been more explicitly stated in both the Results and Materials and Methods sections.

      (4) An additional “high confidence” RTT mutation (c.1152_1195) has been added to Figures 2B and 3B.

      (5) Figure 3 Supplementary Figure 2 has been added to show reading frame data for all frameshifting deletions in the C-terminal deletion-prone region (CT-DPR), showing that +2 frameshifts predominate in this larger data set, not just in the “high confidence” set. This has necessitated changing the previous Fig. 3 SF2 to Fig. 3 SF3.

      (6) A summary of the genetic alterations described in the manuscript, and their outcomes, has been added as Figure 7.

      (7) A simple flow chart which assists in the classification of human CTDs as “likely benign” or “likely pathogenic” has been added as Figure 8. This will aid future assessment of novel mutations in this region.

      (8) Additions have been made to the Materials and Methods section to comply with reporting guidelines.

      (9) Minor changes have been made to the text to correct typographical errors and to clarify meaning.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors scrutinized differences in C-terminal region variant profiles between Rett syndrome patients and healthy individuals and pinpointed that subtle genetic alternation can cause benign or pathogenic output, which harbors important implications in Rett syndrome diagnosis and proposes a therapeutic strategy. This work will be beneficial to clinicians and basic scientists who work on Rett syndrome, and carries the potential to be applied to other Mendelian rare diseases.

      Strengths:

      Well-designed genetic and molecular experiments translate genetic differences into functional and clinical changes. This is a unique study resolving subtle changes in sequences that give rise to dramatic phenotypic consequences.

      Weaknesses:

      There are many base-editing and protein-expression changes throughout the manuscript, and they cause confusion. It would be helpful to readers if authors could provide a simple summary diagram at the end of the paper.

      We have added a summary diagram, as suggested (Figure 7). We have also provided a flowchart which shows how to classify human CTDs as “likely benign” or “likely pathogenic” based on their location.

      Reviewer #2 (Public review):

      Summary:

      This study by Guy and Bird and colleagues is a natural follow-up to their 2018 Human Molecular Genetics paper, further clarifying the molecular basis of C-terminal deletions (CTDs) in MECP2 and how they contribute to Rett syndrome. The authors combine human genetic data with well-designed experiments in embryonic stem cells, differentiated neurons, and knock-in mice to explain why some CTD mutations are disease-causing while others are harmless. They show that pathogenic mutations create a specific amino acid motif at the C-terminus, where +2 frameshifts produce a PPX ending that greatly reduces MeCP2 protein levels (likely due to translational stalling) whereas +1 frameshifts generating SPRTX endings are well tolerated.

      Strengths:

      This is a comprehensive and rigorous study that convincingly pinpoints the molecular mechanism behind CTD pathogenicity, with strong agreement between the cell-based and animal data. The authors also provide a proof of principle that modifying the PPX termination codon can restore MeCP2-CTD protein levels and rescue symptoms in mice. In addition, they demonstrate that adenine base editing can correct this defect in cultured cells and increase MeCP2-CTD protein levels. Overall, this is a well-executed study that provides important mechanistic and translational insight into a clinically important class of MECP2 mutations.

      Weaknesses:

      The adenine base editing to change the termination codon is shown to be feasible in generated cell lines, but has yet to be shown in vivo in animal models.

      This work is the obvious next step and is in progress. However, with the rise in pre- and neonatal genetic testing we felt it was important to disseminate our findings as soon as possible. The family pedigree in Figure 3C is a clear illustration of this need

      Reviewer #3 (Public review):

      Summary:

      Guy et al. explored the variation in the pathogenicity of carboxy-terminal frameshift deletions in the X-linked MECP2 gene. Loss-of-function variants in MECP2 are associated with Rett syndrome, a severe neurodevelopmental disorder. Although 100's of pathogenic MECP2 variants have been found in people with Rett syndrome, 8 recurrent point mutations are found in ~65% of disease cases, and frameshift insertions/deletions (indels) variants resulting in production of carboxy-terminal truncated (CTT) MeCP2 protein account for ~10% of cases. Many of these occur in a "deletion prone region" (DPR) between c.1110-1210, with common recurrent deletions c.1157-1197del (CTD1) and c.1164_1207del (CTD2). While two major protein functional domains have been defined in MeCP2, the methyl-binding domain (MBD) and the NCoR interacting domain (NID), the functional role of the carboxy-terminal domain (CTD, beyond the NID, predicted to have a disordered protein structure) has not been identified, and previous work by this group and others demonstrated that a Mecp2 "minigene" lacking the CTD retains MeCP2 function suggesting that the CTD is dispensable. This raises an important question: If the CTD is dispensable, what is the pathogenic basis of the various CTT frameshift variants? Prior work from this group demonstrated that genetically engineered mice expressing the CTD1 variant had decreased expression of Mecp2 RNA and MeCP2 protein and decreased survival, but those expressing the CTD2 variant had normal Mecp2 RNA and protein and survival. However, they noted that differences between the mouse and human coding sequences resulted in different terminal sequences between the two common CTD, with CTD1 ending in -PPX in both mouse and human, but CTD2 ending in -PPC in human but -SPX in mouse, and in the previous paper they demonstrated in humanized mouse ES cells (edited to have the same -PPX termination) containing the CTD2 deletion resulted in decreased Mecp2 RNA and protein levels. This previous work provides the underlying hypotheses that they sought to explore, which is that the pathological basis of disease causing CTD relates to the formation of truncated proteins that end with a specific amino acid sequence (-PPX), which leads to decreased mRNA and protein levels, whereas tolerated, non-pathogenic CTD do not lead to production of truncated proteins ending in this sequence and retain normal mRNA/protein expression.

      In this manuscript, they evaluate missense variants, in-frame deletions, and frame shift deletions within the DPR from the aggregated Genome Aggregated Database (gnomAD) and find that the "apparently" normal individuals within gnomAD had numerous tolerated missense variants and in-frame deletions within this region, as well as frameshift deletions (in hemizygous males) in the defined region. All of the gnomAD deletions within this region resulted in terminal amino acid sequences -SPRTX (due to +1 frameshift), whereas nearly all deletion variants in this region from people with Rett syndrome (from the Clinvar copy of the former RettBase database) had a terminal -PPX sequence, due to a +2 frameshift. They hypothesized that terminal proline codons causing ribosomal stalling and "nonsense mediated decay like" degradation of mRNA (with subsequent decreased protein expression) was the basis of the specific pathogenicity of the +2 frameshift variants, and that utilizing adenine base editors (ABE) to convert the termination codon to a tryptophan could correct this issue. They demonstrate this by engineering the change into mouse embryonic stem cell lines and mouse lines containing the CTD1 deletion and show that this change normalized Mecp2 mRNA and protein levels and mouse phenotypes. Finally, they performed an initial proof-of-concept in an inducible HEK cell line and showed the ability of targeted ABE to edit the correct adenine and cause production of the expected larger truncated Mecp2 protein from CTD1 constructs.

      The findings of this manuscript provide a level of support for their hypothesis about the pathogenicity versus non-pathogenicity of some MECP2 CTT intragenic deletions and provide preliminary evidence for a novel therapeutic approach for Rett syndrome; however, limitations in their analysis do not fully support the broader conclusions presented.

      Strengths:

      (1) Utilization of publicly available databases containing aggregated genetic sequencing data from adult cohorts (gnomAD) and people with Rett syndrome (Clinvar copy of RettBase) to compare differences in the composition of the resulting terminal amino acid sequences resulting from deletions presumed to be pathogenic (n+2) versus presumed to be tolerated (n+1).

      (2) Evaluation of a unique human pedigree containing an n+1 deletion in this region that was reported as pathogenic, with demonstration of inheritance of this from the unaffected father and presence within other unaffected family members.

      (3) Development of a novel engineered mouse model of a previously assumed n+1 pathogenic variant to demonstrate lack of detrimental effect, supporting that this is likely a benign variant and not causative of Rett syndrome.

      (4) Creation and evaluation of novel cell lines and mouse models to test the hypothesis that the pathogenicity of the n+2 deletion variants could be altered by a single base change in the frameshifted stop codon.

      (5) Initial proof-of-concept experiments demonstrating the potential of ABE to correct the pathogenicity of these n+2 deletion variants.

      Weaknesses:

      (1) While the use of the large aggregated gnomAD genetic data benefits from the overall size of the data, the presence of genetic variants within this collection does not inherently mean that they are "neutral" or benign. While gnomAD does not include children, it does include aggregated data from a variety of projects targeting neuropsychiatric (and other conditions), so there is information in gnomAD from people with various medical/neuropsychiatric conditions. The authors do make some acknowledgement of this and argue that the presence of intragenic deletion variants in their region of interest in hemizygous males indicates that it is highly likely that these are tolerated, non-pathogenic variants. Broadly, it is likely true that gnomAD MECP2 variants found in hemizygous males are unlikely to cause Rett syndrome in heterozygous females, it does not necessarily mean that these variants have no potential to cause other, milder, neuropsychiatric disorders. As a clear example, within gnomAD, there is a hemizygous male with the rs28934908 C>T variant that results in p.A140V (p.A152V in e1 transcript numbering convention). This pathogenic variant has been found in a number of pedigrees with an X-linked intellectual disability pattern, in which males have a clear neurodevelopmental disorder and heterozygous females have mild intellectual disability (see PMIDs 12325019, 24328834 as representative examples of a large number of publications describing this). Thus, while their claim that hemizygous deletion variants in gnomAD are unlikely to cause Rett syndrome, that cannot make the definitive statement that they are not pathogenic and completely benign, especially when only found in a very small number of individuals in gnomAD.

      We have included the possibility that mutations found in gnomAD may give rise to less severe neurological conditions in the discussion.

      (2) The authors focus exclusively on deletions within the "DPR", they define as between c.1110-1210 and say that these deletions account for 10% of Rett syndrome cases. However, the published studies that are the basis for this 10% estimate include all genetic variants (frameshift deletions, insertions, complex insertion/deletions, nonsense variants) resulting in truncations beyond the NID. For example, Bebbington 2010 (PMID: 19914908), which includes frameshift indels as early as c.905 and beyond c.1210. Further specific examples from RettBase are described below, but the important point is that their evaluation of only frameshift variants within c.1110-1210 is not truly representative of the totality of genetic variants that collectively are considered CTT and account for 10% of Rett cases.

      The vast majority of C-terminal truncating mutations do occur within the “CT-DPR”, likely due to its C-rich nucleotide sequence and the presence of microhomologies within the region. Looking at frameshifting deletions in RettBASE that start after the NID, a large proportion of these end within the CT-DPR and result in a -PPX ending. We decided to restrict our analysis to the c.1110-1210 region to avoid including the rarer examples that may have a different reason for their pathogenicity. We do not assert that all C-terminal truncations are pathogenic due to this mechanism, but current evidence suggests that most are.

      (3) The authors say that they evaluated the putative pathogenic variants contained within RettBase (which is no longer available, but the data were transferred to Clinvar) for all cases with Classic Rett syndrome and de novo deletion variants within their defined DPR domain. Looking at the data from the Clinvar copy of RettBase, there are a number (n=143) of c-terminal truncating variants (either frameshift or nonsense) present beyond the NID, but the authors only discuss 14 deletion frameshift variants in this manuscript. A number of these variants have molecular features that do not fall into the pathogenic classification proposed by the authors and are not addressed in the manuscript and do not support the generalization of the conclusions presented in this manuscript, especially the conclusion that the determination of pathogenicity of all c-terminal truncating variants can be determined according to their proposed n+2 rule, or that all of the 10% of people with Rett syndrome and c-terminal truncating variants could be treated by using a base editor to correct the -PPX termination codon.

      It is important to state here that we did not use the data in ClinVar for our analysis, but the original information that was held in RettBASE. We have clarified this in the manuscript and have now included supplementary tables containing the data we downloaded from both RettBASE and gnomAD. Table 1 contains all the RettBASE entries with MECP2 mutations, while Tables 2-4 contain subsets of this data pertaining to CTDs. We have extended our “high confidence” set of RTT alleles to contain one more that was previously overlooked (c.1152_1195) to bring the total number of alleles to 15. Taken together these alleles account for 158 individual entries in RettBASE. We have now included an analysis of all frameshifting mutations in the CT-DPR (Supplementary Table 3, Fig. 3 SF2) which covers 69 different mutations and 260 individual entries. Of these, +2 frameshifts make up the large majority, in contrast to the gnomAD data shown in Supplementary Table 8 and Fig. 3 SF1.

      (4) The HEK-based system utilized is convenient for doing the initial experiments testing ABE; however, it represents an artificial system expressing cDNA without splicing. Canonical NMD is dependent on splicing, and while non-canonical "NMD-like" processes are less well understood, a concern is whether the artificial system used can adequately predict efficacy in a native setting that includes introns and splicing.

      We disagree with this opinion. We show that the loss of protein and mRNA seen with knock in mouse and human alleles is recapitulated when using a cDNA-based transgene in the HEK system, demonstrating that the mechanism of loss does not involve factors bound at splice junctions etc. We also demonstrate the effect of the A to G change at the stop codon is the same whether we do this by base editing our cDNA transgene in T-REx cells or by making the CTD1 X>W knock-in mouse. Both result in increased levels of a slightly extended but still truncated protein.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The phrase in the title, "an alternative therapeutic approach" is only insinuated in the manuscript, making it rather inappropriate to be in the title.

      In this study we use adenine base editing to modify RTT-causing CTD mutations in T-REx cells which is clearly a precursor to developing a therapy, utilising the new findings in this study. We therefore feel that the use of “an alternative therapeutic approach” can be justified.

      Reviewer #2 (Recommendations for the authors):

      I have a few minor comments for the authors to consider:

      (1) Please double-check Figure 2 Supplementary Figure 1, as the allele count for E394K does not appear to be in the thousands; rather, E397K seems to be the variant shown in the graph.

      Yes, this was an error and has been corrected to E397K in the text. Thank you for spotting it.

      (2) On page 12, the phrase “common DNA sequence features shared by all CTDs that give rise to RTT” might be better described as “amino acid sequence features.”

      This has been altered in the text as suggested.

      (3) On page 3, the sentence "analysis of patient mutations and experimental data from mouse models support a role in transcriptional repression" cites Gabel et al. 2015 and Kinde et al. 2016, which focus on null alleles but not patient mutations. It would be appropriate to also cite Johnson et al. 2017, which analyzed MeCP2 T158M and R106W patient mutations.

      This has now been cited.

      (4) In the same sentence, Bajikar et al. 2025 are described as studying the "acute loss of MeCP2," but gene expression was analyzed after a week or longer period of time, not minutes to hours as in degron-mediated degradation systems.

      The word “acute” has been removed from the text.

      (5) On page 7, the statement that "E394K is common... who were later found to have additional pathogenic MECP2 mutations" should include a supporting reference.

      We have now included the references Moncla et al (2002) and Wan et al (1999) to address this.

      (6) Similarly, on page 10, the sentence "these findings question the validity of two cases where individuals presented with classical Rett..." is missing a reference to the case report mentioned.

      The references Bienvenu et al (2000) and Philippe et al (2006) have been added.

      (7) While the manuscript is well written and full of detail, if space is an issue, the authors might consider tightening sections that reiterate findings from their 2018 HMG paper.

      A section discussing the CTD2 allele from the 2018 HMG paper has been removed from the results section.

      Reviewer #3 (Recommendations for the authors):

      (1) Overall, the manuscript is rather dense and potentially challenging to follow easily, especially for a non-expert reader.

      Minor edits have been made to the text which will hopefully make it easier to follow. We have also added two new figures (Figures 7 and 8) to summarize the different alleles and edits which appear in the paper, and to show how to determine whether a CTD in the region is likely to be benign or pathogenic.

      (2) The introduction of data presented in Figure 1 within the manuscript introduction seems inappropriate and should be moved to the results section.

      We would say that this is unconventional rather than inappropriate, and is referred to in the introduction, so we would prefer to leave it as it is.

      (3) Providing specific, common nomenclature for genetic repository variants (rs numbers, gnomAD IDs, etc) somewhere would be beneficial. This is an issue because of the complexity of numbering (either coding or protein) for MECP2 due to the different transcript-based numbering systems.

      This nomenclature is now included in the supplementary tables of data from RettBASE and gnomAD.

      (4) As described in the public comments, there are a number of MECP2 genetic variants listed in the Clinvar copy of RettBase, resulting in c-terminal truncations that are not mentioned or discussed within the manuscript. Without the level of detail present in the original form of RettBase (number of events, de novo, etc) in the currently available Clinvar iteration, it is unclear why a number of variants, even within the limited DPR region, were not mentioned. A supplementary file including the more complete information from the RettBase version, with a complete listing of all c-terminal truncating variants, and an explicit rationale for the exclusion of variants would be helpful.

      We have now included supplementary tables with our download of all MECP2 mutations which were held in RettBASE. We have further added tables with the subset of mutations that we have analysed and have more explicitly stated our criteria for defining the “high confidence” sets of mutations. We did not download the data relating to “evidence of pathogenicity” (ie de novo?, absent from parents etc) from RettBASE, but annotated our list of CTDs with this information while RettBASE was still available. This was used in Supplementary Table 3.

      (5) A discussion of the limitations, notably that the fact that the focus exclusively on deletion variants within a restricted region (c.1110-1210) does not truly represent all genetic variants that cumulatively account for 10% of Rett cases, is needed. Furthermore, as pointed out, not all frameshift variants, even those that are n+2, result in the -PPX termination that is presented as the pathogenic basis of c-terminal truncations and amenable to correction by ABE. This should be noted in the discussion, as well as consideration of the late nonsense variants that cause c-terminal truncations (some of which would be very similar to the deletion variants discussed but without the proposed primary pathogenic driver, -PPX).

      We do not claim to explain the pathogenic mechanism of all C-terminal frameshift mutations found in cases of Rett syndrome. There will certainly be some that do not fit our explanation. However, we believe we have shown evidence that a large proportion of CTDs in RTT will be amenable to the therapy we propose.

      (6) Regarding point 3 in the public review, specifically:

      (a) n=7 nonsense variants (S360X, K363X, E397X, R453X, E455X) that do not carry the destabilizing -PPX sequence.

      (b) n=136 frameshift indel variants beyond the NID.

      (i) n=11 that have indels that extend past the native stop codon, n=4 of which start within the DPR domain (c.1110-1210) but would have a different terminal sequence than their proposed pathogenic -PPX sequence.

      (ii) n=125 frameshift indels with terminal breakpoint before the native stop codon

      (c) n=89 that have start or stop points within c.1110-1210

      (d) n=72 not mentioned within the manuscript.

      (e) n=27 are n+1, with 26/27 having what the authors term as the "tolerated" -SPRTX ending, but 1/27 having a frameshift beyond this region (c.1133_1361)

      (f) n=45 are n+2, with 32/45 ending in -PPX (supporting authors conclusion), but 10/45 will use the frameshift stop codon preceding the -PPX and have a different terminal sequence, and 3/45 result in a frameshift termination beyond the -PPX sequence.

      (g) n=36 have breakpoints either before c.1110 or after c.1210

      (h) n=16 start before c.1110, with 11/16 ending before c.1110. 4 of these 11 are n+2, but would use the earlier frameshift stop codon and not have -PPX terminal sequence. For 5/16, the indel extends past c.1210, with the n+2 leading to frameshift termination codons beyond the -PPX sequence.

      (i) n=20 indels start beyond c.1210, with 8/20 being n+1 and 12/20 being n+2, with neither leading to the -SPRTX or --PPX termination sequences characterized in this manuscript.

      As mentioned previously, we have used data taken from RettBASE, not from ClinVar. Both RettBASE and ClinVar will contain MECP2 mutations found in cases of RTT which are not the causative mutation. Databases of this kind contain sequencing errors and mutations that have been mistakenly assigned as causative. It is therefore imperative that the publicly available information is screened to only include mutations that meet stringent criteria. This is why we chose to start by looking at high confidence sets of mutations, with our conclusions supported by analysis of all such mutations in our region of study.

      As mentioned in response to point 5 above, we do not claim to explain every mutation in the region, but believe this study reveals an important disease mechanism for a large proportion of CTDs, leading to a potential therapy. It also contains significant information for predicting the likely prognosis of individuals with CTDs, who may remain healthy but are currently informed that their mutation is likely pathogenic. At present it seems that this is often based solely on the presence of a frameshifting mutation with similarities to bona fide RTT CTDs, without strong evidence of pathogenicity.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Noirot-Gros et. al. presents a herculean effort to map the protein-protein interactome of the c-di-GMP signaling network in Pseudomonas fluorescens (Pf). C-di-GMP, the key driver of biofilm formation in bacteria, is controlled by a highly complex network of synthesis, degradation and effector proteins. Pf is no exception as it encodes dozens of such proteins. The authors use a Yeast Two-Hybrid approach genome-wide screen with 10 diguanylate cyclase (DGC) enzymes as bait to assess protein-protein interactions in this network. The results identify over one hundred such interactions with several different hubs, including c-di-GMP signaling, other signaling systems, membrane proteins, etc. The authors then explore the original bait proteins as well as identify interactors on biofilm formation-related phenotypes and swarming using a high-throughput CRISPRi expression knockdown approach. The amount of data generated is quite impressive. Much of the manuscript uses statistical-based network analysis to group different proteins based on their interactions or impact on phenotypes, which is a high-level analysis that can catalyze further study into this system. The authors chose three specific proteins to assess their impact on cell morphology, DNA repair, and protein localization. Overall, in my view, this is perhaps the best analysis of a c-di-GMP protein-protein interactome, and it provides a multitude of hypotheses to be tested. However, therein lies the weakness of the manuscript in that very few of these hypotheses are actually tested. But such is not the goal of this network analysis type of approach. Overall, I think the work will be highly impactful to those in the c-di-GMP field, and it provides a template for others attempting such analyses of protein-protein interactions.

      Strengths:

      The manuscript is impressive in the sheer scale of the protein-protein interactions identified, network analysis, and phenotypic analysis of specific proteins in the network. It is an impressive amount of work that could be very useful to the field. It is also statistically rigorous in its analysis of significant interactions or network nodes.

      Weaknesses:

      The weakness of the manuscript is that, with three exceptions, very few of the hypotheses are actually tested. For example, BifA is shown to be a network hub protein that interacts with many other diguanylate cyclases, and this is hypothesized to be through GGDEF heterodimerization. I appreciate that experimentally testing such a hypothesis is probably another entire manuscript, but some early forays into such ideas could be undertaken using AlphaFold structural modeling of protein-protein interactions compared with GGDEFs that don't form heterodimers. Also, an inherent weakness is that such detailed analyses of a c-di-GMP signaling network, in which each diguanylate cyclase and phosphodiesterase may respond to a unique cue, is that the network identified and the conclusions made are highly specific to the experimental conditions in which the work was done. Therefore, it is unclear how broadly these conclusions (i.e. BifA is the central regulator of c-di-GMP signaling) apply to other conditions. But it is impossible to get around such a limitation, and this work can lead to testing the robustness of the identified network in other environments.

      We would like to thank the reviewer sincerely for their positive comments on our manuscript and for their constructive feedback. We recognize the limitations arising from the lack of extensive knowledge regarding the environmental cues that trigger the regulation of all CDG activities in P. fluorescens. We hypothesize that DipA acts as a central local hub that positively or negatively regulates the activity of its interacting CDG partners throughout the cell life cycle, lifestyle transitions and environmental signals. Testing this hypothesis would indeed require extensive biochemical and omics approaches. However, strengthening the significance of DipA complexes in silico using AlphaFold is a very appealing proposition and we are currently considering including this analysis in the revised version of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Noirot-Gros and coworkers investigated the network of c-di-GMP associated protein complexes in Pseudomonas fluorescens. They did so by using a genome-wide yeast two-hybrid screen, and that was further probed by phenotypic screening that focused on biofilm and motility phenotypes. From this network map, they discovered that the phosphodiesterase DipA interacts with the GGDEF domains of many c-di-GMP-binding proteins.

      Strengths:

      (1) Broadness of screen led to identification of new interactions: The genome-wide yeast two-hybrid screening approach permitted broad investigation of c-di-GMP-associated protein-protein interactions. These interactions included some previously validated interactions as well as newly discovered interactions.

      (2) Complementary experimental validation: The proposed network was experimentally validated, including by using a CRISPRi-based approach in which the expression of genes encoding proteins identified in the network was systematically suppressed, and then the impact on the biofilm and motility phenotypes was assessed.

      Weaknesses:

      The findings would have been strengthened by further biochemical analysis, but this is likely beyond the scope of the paper.

      We would like to express our gratitude to the reviewer for their positive evaluation assessment, and for taking into account the limitations of the study's scope.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Noirot-Gross et al take an open-ended approach to elucidate the c-diGMP-associated protein complexes in Pseudomonas fluorescens. Starting with 10 cyclic d-GMP putative proteins, they use a combination of genome-wide two-hybrid system followed by CRISPRi-mediated exploration of phenotypes to describe the cyclic di-GMP-associated regulation of biofilm formation, and how it relates to other functions. Overall, this work presents an excellent example of how genome annotations can be further confirmed with the use of integrated functional genomic approaches. Some areas of improvement can be applied to this manuscript to enhance readability and provide a clearer distinction between confirmatory results and new findings, which are provided below:

      Strengths:

      (1) The authors have explored their findings extensively and provide a comprehensive view of the topic.

      (2) The combination of genome-wide explorations of protein-protein interactions with the more focused phenotypic exploration of the interactions found provides a solid framework for the work presented.

      Weaknesses:

      (1) Overall goal of the work:

      While articles that describe open-ended approaches can be comprehensive and descriptive in nature, the authors should have a main overall goal, which can guide the reader through the main and most compelling findings at the end. As written, the overall goal is not clear. The network perspective is interesting, and the focus on biofilm formation appears in the title. Why P. fluorescens? How is cyclic di-GMP-mediated regulation of biofilm formation in P. fluorescens different from P. aeruginosa? Why would it be studied? (Positive or negative regulation of biofilm formation?)

      We would like to express our appreciation to the reviewer for their thorough evaluation of our manuscript and for the constructive feedback they provided. The overall goal of this study will be further refined, and outlined in the introduction in the revised version of the manuscript.

      (2) Abstract:

      The abstract is very well written and guides the reader to the DipA as a hub protein in the network. From further reading, the article could clarify whether this finding is confirmatory or novel (does DipA play a similar role in P. aeruginosa?) It would be appropriate to mention the role of DipA in other Pseudomonas species from the beginning, and not only in the discussion session.

      (3) Introduction:

      The introduction is nicely written. An area of improvement could be giving more attention to protein interactions as relevant to c-di-GMP. The authors could consider an independent paragraph starting with line 84-85 "Protein-protein interactions involving DGCs, PDEs, and target effectors are crucial in establishing localized signalling through the generation of local pools of c-di-GMP", expanding on this particular aspect with an example of localized signal, after explaining that localization could help decipher specific function within the network of DGCs and PDEs. Then go into connecting biofilms with c-di-GMP and protein-protein interactions, using the example of GcbC and LapD.

      We propose highlighting the example to the local signalling cascade formed by the tripartite system YdaM, YciR and MlrA. This will be addressed in the revised version of the manuscript.

      (4) The rationale of choosing 10 PDEs could be clarified. The nice diagrams shown in the supplementary table could be used as part of Figure 1, so the reader understands why these proteins were used, and what is known about them (for example, add them as Figure 1a).

      We propose to include a specific section in the supplementary file to explain the whole rationale behind choosing these CDGs. These proteins were selected based on their involvement in different steps of biofilm formation in Pseudomonas, as well as their role in the ability of P. fluorescens strains to colonize plant roots.

      (5) Figures 1b and 2 convey the same information as in Figure 1a. They could be removed without affecting the understanding of the article.

      Figure 2 will be transferred in Supplementary as part of the Figure S1

      (6) CRISPRi and Figure 3. Figure 3 shows the methodology of CRICPR phenotypic screening. A diagram showing the CRISPRi system in P. fluorescens could help the non-expert reader. While the choice of 23 proteins related to the emerging hub DipA is clear, the choice of the other 33 genes could be better explained. Are these proteins already related to biofilm formation? Where are they part of the network detected? How about the other 14 SBW25 genes? The authors could clarify the rationale of the choices. Figure 4 could be combined with Figure 3 or moved to the supplementary material.

      A better description of the rationale behind the choice of tested interacting protein partners will be provided. We also agree to combine Figure 4 with Figure 3.

      (7) Figures 5, 6 and 7 represent solid network analysis of the findings. Still, they could be improved in clarity on the main findings. The authors conclude at the end of section 3.2.3 that there are networks that exert a "positive role" and a "negative role". The authors could show that in the figures, explaining what those roles are: more biofilm structural coding genes? positive or negative regulation of biofilm formation?)

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors claim that bacteria are guided by diffusiophoresis. They perform experiments of bacterial motility in microfluidic channels with salt gradients. The data show that P. putida bacteria swim towards higher sodium chloride concentrations, but there is no evidence that this is due to diffusiophoresis.

      Weaknesses:

      It is well known that bacteria perform chemotaxis in salt gradients (see e.g., PNAS 86, pp. 8358-8362, 1989). The underlying mechanism based on chemoreceptors is widely accepted, but the authors do not mention this possibility. I recommend a control experiment where the chemotaxis genes are knocked out. Even if this mechanism can be ruled out, the current data show no evidence for a mechanism based on diffusiophoresis.

      We thank the reviewer for raising this important comment. We agree that receptor-mediated salt taxis is a well-established mechanism in some bacteria, including the classic study by Qi and Adler (PNAS, 1989). We note, however, that the original manuscript did discuss this possibility and cited Qi and Adler in line 117: “We also note that we did not observe any significant difference in the tumble rates between the control and NaCl gradient cases (Figure 3h; Figure S2, SM), suggesting that NaCl gradients do not interfere with chemoreceptors (Qi and Adler, 1989).” In Figure S2, the run-time distributions show no significant difference between the no-gradient condition and the NaCl-gradient condition, with fitted tumble rates of λ = 0.41 s<sup>−1</sup> and λ = 0.38 s<sup>−1</sup>, respectively. We also do not observe a directional bias in run duration, namely longer runs up the salt gradient and shorter runs down the gradient, which would be expected for canonical chemoreceptor-mediated taxis.

      The basis for assigning the observed migration to diffusiophoresis is that the NaCl gradient produces a directional drift of the bacterial body without a measurable change in the run-and-tumble statistics. This behavior is consistent with our previous work [1], where non-motile bacteria were shown to undergo diffusiophoretic migration toward higher salt concentration. Because that migration occurred in non-motile cells and across different bacterial types and morphologies, it supports the interpretation that native bacterial surface charge can drive a non-specific diffusiophoretic response in salt gradients.

      That said, we agree with the reviewer that a genetic control would provide a stronger test against receptor-mediated chemotaxis. We will therefore perform additional experiments using a ∆cheA strain. Because CheA is required for canonical chemotactic signal transduction, observing the same directional migration in the ∆cheA mutant would directly test whether the NaCl gradient response persists in the absence of receptor-mediated chemotaxis. We will include these new data and revise the manuscript to more explicitly distinguish diffusiophoretic drift from chemoreceptormediated salt taxis.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate how salt gradients influence the transport of Pseudomonas putida in confined microfluidic environments. They report that salt gradients enhance directional migration, increase run persistence, and promote transport toward contaminant-rich regions. To explain these observations, the authors propose a physical steering mechanism in which differential diffusiophoretic mobilities of the cell body and flagellar bundle generate an aligning torque that reorients cells along the salt gradient.

      Strengths:

      The study addresses an interesting question at the interface of microbiology, complex fluids, and active matter. Their experiments suggest that salt gradients influence bacterial transport behavior and lead to more persistent, directional motion. Once confirmed, the proposed mechanism would broaden our understanding of how environmental gradients can shape microbial migration through physical interactions in addition to more traditional sensing-based pathways.

      Weaknesses:

      The main limitation of the current study is that the proposed steering mechanism is not directly demonstrated. The evidence for the diffusiophoretic torque is largely inferred from trajectory statistics and theoretical modeling. While the observed transport behavior is convincing, the causal link between the observed migration patterns and the proposed reorientation mechanism remains less well established. In particular, the manuscript focuses primarily on cell trajectories and transport properties, whereas the proposed mechanism fundamentally involves changes in cell orientation. Additional evidence connecting orientation dynamics to the proposed torque mechanism would strengthen the conclusions.

      We thank the reviewer for the constructive comments. We agree that the proposed steering mechanism should be supported by evidence that directly connects the salt gradient response to bacterial orientation dynamics, not only to trajectory-level transport statistics.

      We would like to clarify that the original manuscript already includes an orientation-based analysis in Figure 4f,g in the main text. The corresponding methodology and results are described in lines 173–183 and in the Supporting Information. Specifically, we quantified cell steering by measuring the change in body angle, ∆θ, along individual run trajectories as a function of arc length, s, using the orientation correlation ⟨cos(∆θ)⟩<sub>s</sub>. In the absence of salt gradients, the orientation correlation decays slowly with arc length, indicating persistent swimming along the initial

      Author response image 1.

      Instantaneous angular velocity as a function of heading angle relative to the salt gradient orientation. (a) Experimental and (b) simulated mean angular velocity of cells as a function of heading angle θ (measured relative to the gradient direction; θ = 0° points toward the gel/high-salt side) in the absence (blue) and presence (red) of a NaCl gradient. Positive and negative values indicate counterclockwise and clockwise rotation, respectively, with arrows showing rotation direction. Under the gradient, both experiments and simulations show a signed, angle-dependent rotation rate that is largest near θ = ±90° and approaches zero near θ = 0° and 180°, consistent with a restoring torque that steers cell heading toward the gradient direction. Simulations reproduce this behavior with comparable magnitude to the experimental measurements, and in the absence of a gradient, angular velocity remains relatively small in the no-gradient case with no consistent directional bias across heading angles.

      run direction. Under a salt gradient, the correlation decays more rapidly, indicating stronger directional reorientation during runs. Because the tumble statistics do not change significantly between the no-gradient and salt-gradient conditions, this enhanced orientational decorrelation is not attributed to increased tumbling or rotational noise. Instead, it is consistent with continuous deterministic steering during runs, as expected from a diffusiophoretic torque acting on the cell body–flagellar bundle system.

      To further address the reviewer’s concern, we performed an additional orientation-dynamics analysis using the same dataset shown in Figure 4. Following the approach used by Stehnach et al. [2], we calculated the instantaneous angular velocity during individual runs as a function of the cell heading angle relative to the salt gradient direction. The cell orientation was obtained from the run trajectories, and tumble events were excluded because they produce large transient angular velocity spikes that are not representative of continuous steering during runs.

      The new analysis is shown in Author response image 1. Under the no-gradient condition, the angular velocity remains small and nearly independent of heading angle. In contrast, under the salt gradient condition, the angular velocity becomes strongly heading-dependent. The angular velocity is largest when cells swim nearly perpendicular to the salt gradient, where a steering torque is expected to be maximal. Moreover, the sign of the angular velocity indicates rotation toward alignment with the gradient direction. This behavior is consistent with the proposed diffusiophoretic torque mechanism and provides a direct link between the observed transport behavior and salt-gradient-induced reorientation dynamics.

      We plan to add this angular velocity analysis to the revised manuscript and revise the relevant text to make the connection between trajectory statistics, orientation dynamics, and the proposed torque mechanism clearer.

      A related concern is whether alternative physical mechanisms associated with the imposed salt gradients have been fully excluded. For example, weak flow-mediated effects or other hydrodynamic influences could potentially contribute to the observed transport behavior. The manuscript would benefit from a more thorough discussion of such possibilities and a clearer justification for why the proposed diffusiophoretic mechanism should be regarded as the dominant explanation.

      We thank the reviewer for raising this important point. We agree that alternative physical mechanisms associated with the imposed salt gradient should be considered explicitly. In the revised manuscript, we will add a more detailed discussion explaining why flow-mediated or other hydrodynamic mechanisms are unlikely to account for the observed steering behavior.

      First, the characteristic diffusiophoretic drift velocity in our experiments is approximately u<sub>d</sub> ≈ 1 µm/s, corresponding to only about 1–5% of the typical swimming speed of P. putida. Thus, the proposed mechanism does not require externally driven advection of the cells. Instead, the salt gradient produces a weak but persistent differential diffusiophoretic slip on the cell body and flagellar bundle, which can generate a reorienting torque during active swimming.

      We also considered whether diffusio-osmotic flow along the channel walls could generate sufficient shear to induce rheotaxis. In a dead-end channel, the diffusio-osmotic velocity profile can be estimated as [3]

      which gives a wall shear rate (at z = h) of

      This value is below the shear rate threshold reported by Marcos et al. [2], where rheotactic drift becomes negligible for S < 0.1 s<sup>−1</sup>. Therefore, the shear generated by diffusio-osmotic wall flow in our experiments is too weak to explain the observed directional reorientation. This effect would be even smaller for P. putida, whose thin flagellar bundle is expected to experience weaker shear-induced alignment than organisms with larger flagellar structures.

      We further considered viscotaxis as a possible mechanism. However, the viscosity difference between 1 mM and 100 mM NaCl solutions is marginal. This contrast is far smaller than the approximately 4–5-fold viscosity difference reported to induce strong viscophobic turning in bacteria [4]. Thus, salt-gradient-induced viscosity variations are insufficient to account for the measured steering response.

      Taken together, these estimates indicate that hydrodynamic shear, rheotaxis, and viscotaxis are too weak under our experimental conditions to explain the observed migration and orientation dynamics. In contrast, the proposed nonuniform diffusiophoretic mechanism naturally accounts for the key observations: directional migration up the salt gradient, enhanced curvature during runs, heading-dependent angular velocity, and the absence of significant changes in tumble statistics. We will include this analysis in the revised manuscript to clarify why diffusiophoretic steering is the dominant mechanism under the present conditions.

      The manuscript would also benefit from a clearer positioning within the broader literature on physically induced microbial transport and swimmer reorientation. Previous studies have demonstrated directed migration arising from rheotaxis (Marcos et al., 2012, PNAS) and viscosity-gradient-induced steering (Stehnach et al., 2021, Nature Physics). While the mechanism proposed here appears distinct, a more explicit discussion of how the present work relates to these earlier studies would help readers better understand the specific conceptual advance being made.

      We thank the reviewer for this helpful suggestion. We agree that the manuscript should more clearly position the proposed mechanism within the broader literature on physically induced microbial transport and swimmer reorientation.

      In the revised manuscript, we have added a discussion comparing our results with prior studies on rheotaxis and viscosity-gradient-induced steering. Specifically, we now discuss the work of Marcos et al. [4], which showed that shear flow can generate a torque on the helical flagellum of Bacillus subtilis, producing rheotactic alignment independent of chemical sensing. We also discuss the work of Stehnach et al. [2], which showed that viscosity gradients can steer Chlamydomonas reinhardtii through asymmetric viscous drag on its two flagella, producing viscophobic turning down the viscosity gradient.

      The mechanism proposed in the present work is distinct from both of these cases. Unlike rheotaxis, it does not require externally imposed shear flow. Unlike viscophobic turning, it does not rely on a substantial viscosity contrast. Instead, we propose that a salt concentration gradient generates differential diffusiophoretic motion of the cell body and flagellar bundle, producing a torque that continuously reorients swimming cells along the gradient. This mechanism therefore identifies salt gradients as a distinct physical cue capable of steering bacteria through surface-mediated transport rather than through flow, viscosity contrast, or canonical chemoreceptor signaling.

      We have added the following text to the revised manuscript: ”Our findings also relate to other physical mechanisms of microbial reorientation. Bacterial rheotaxis, arising from a torque generated by shear flow acting on the helical flagellum, steers Bacillus subtilis independent of any chemical gradient [4]. Viscosity gradients similarly drive a viscophobic turning in the alga Chlamydomonas reinhardtii, where uneven viscous drag on its two flagella produces a torque that reorients cells down the gradient [2], a behavior confirmed by measuring angular velocity as a function of heading angle and revealing a sinusoidal form, ω(θ) = −ω <sub>visc</sub> sin(θ), which resembles what we report here for diffusiophoresis in (Figure S4). This identifies salt gradients, independent of flow or viscosity, as a distinct physical route by which swimming cells can be steered and guided, broadening the set of known nonchemoreceptor mechanisms for directed microbial transport.” References

      (1) V. S. Doan, P. Saingam, T. Yan, S. Shin, A trace amount of surfactants enables diffusiophoretic swimming of bacteria, ACS Nano 14 (10) (2020) 14219–14227.

      (2) M. R. Stehnach, N. Waisbord, D. M. Walkama, J. S. Guasto, Viscophobic turning dictates microalgae transport in viscosity gradients, Nature Physics 17 (8) (2021) 926–930.

      (3) S. Shin, E. Um, B. Sabass, J. T. Ault, M. Rahimi, P. B. Warren, H. A. Stone, Size-dependent control of colloid transport via solute gradients in dead-end channels, Proc. Natl. Acad. Sci. 113 (2) (2016) 257–261.

      (4) Marcos, H. C. Fu, T. R. Powers, R. Stocker, Bacterial rheotaxis, Proceedings of the National Academy of Sciences 109 (13) (2012) 4780–4785.

    1. Author response:

      General Statements:

      We appreciate the reviewers for the critical review of the manuscript and the valuable comments. We have carefully considered the reviewer’s comments and have revised our manuscript accordingly.

      Point-by-point description of the revisions:

      Reviewer #1 (Evidence, reproducibility and clarity):

      Major comments

      (1) This study leaves out lipid metabolism as a major energy metabolism pathway relevant to AD. The authors themselves cite the significance of acylcarnitines and CPT1A in AD (pg. 3, lines 32-33, pg. 4, lines 1-2). Lipid metabolism and homeostasis is known to be disrupted in AD1. Fatty acid oxidation is a known energy source in the prefrontal cortex2 and will also generate acetyl coA, which this study reveals is a significant decreased metabolite in AD. Furthermore, sphingomyelin emerges as one of the major decreased DEMs as well. Thus, lipid metabolism should be highlighted in Figure 3 and discussed throughout the manuscript; otherwise its omission should be clearly stated and justified.

      We appreciate the reviewer’s insightful comment regarding a critical role of lipid metabolism in AD. We recognize that lipid metabolism is a metabolic pathway deeply involved in AD pathology (Baloni et al., 2022, 2020; Varma et al., 2021). Accordingly, we have revised the Limitations section to more strongly emphasize its role as a vital energy source (pg. 13, lines 15-17). Regarding the visualization of lipid metabolism, we extracted lipid-related pathway from the trans-omic network but found that the regulatory relationships among DEPs and DEMs were excessively complex and interconnected. Thus, interpreting this regulatory network seemed to be more challenging compared to the other energy production pathways presented in our manuscript. Therefore, we have concluded that the pathway analysis in our trans-omic network may not be suitable for deeply elucidating the lipid dysregulation in AD. We have added a statement acknowledging this as a limitation of our current methodology in the revised manuscript (pg. 13, lines 13-22).

      (2) The covariates used for differential analysis should be discussed and justified. Notably, age is used as a covariate for transcriptomic analysis but not proteomic and metabolomic analysis, with no justification. Additionally, given the known importance of lipid metabolism in AD and the putative role of APOE in lipid homeostasis3, APOE genetic status should be considered as a covariate, or its omission should be justified.

      We appreciate the reviewer’s comment regarding the included covariates in differential analyses of our study. The reason we did not include other variables, such as age at death and RIN, is that these data were not available for each sample. Thus, we referred to the original research articles from which proteomic or metabolomic datasets used in our study were derived. Regarding the metabolomic dataset, in the original article (Batra et al., 2023), only two metabolites, 1-methyl-5-imidazoleacetate and N6-carboxymethyllysine, were significantly associated with age. In addition, no metabolites were significantly associated with sex, BMI, and years of education. Regarding the proteomic dataset, in the original article (Johnson et al., 2020), age at death, PMI, and sex were included as covariates in the analyses, though these variables were not found to strongly influence the data (Extended Data Fig.2 in (Johnson et al., 2020)).

      (3) The authors make a conclusion statement that suggests intervention: "Collectively, our data suggests that preserving or improving the ability to produce ATP and early intervention in the process of nitrogen metabolism are candidates for the prevention and treatment of dementia" (pg. 12, lines 12-14). This claim is not well-supported by the evidence provided in the study. There are a few limitations: (a) This was an observational, not interventional study; (b) The study did not establish whether the metabolic disruptions are causes or effects in AD; and (c) ATP or other bioenergetic indicators were not directly measured. Therefore, any statements about potential interventions should be removed or qualified as highly speculative.

      We agree with the reviewer that the statement regarding potential interventions was not sufficiently supported by our analyses. Accordingly, we have removed the sentence regarding prevention and treatment from the revised manuscript (e.g., we have deleted final paragraph of the previous manuscript).

      (4) In conjunction with the last point, the main conclusion of the study is that energy production is down in AD. The data presented in Figure 3 are consistent with this conclusion, but it is far from definitive due to limitations stated above in comments 3a and 3b. The authors should offer additional support for this conclusion: experimental follow-up, flux modeling, analysis of alternative datasets with ATP measurement, causal inference.

      We sincerely thank the reviewer for this valuable and constructive suggestion. Regarding flux modeling, we agree that metabolic flux analysis could provide important mechanistic insight. Indeed, previous studies have applied flux modeling in the context of lipid metabolism in Alzheimer’s disease (Baloni et al., 2022). We also attempted to perform flux modeling focusing on energy metabolism. However, we found it difficult to obtain biologically meaningful and robust results and therefore decided not to include these analyses in the current manuscript.

      With respect to ATP measurements, we fully agree that direct evidence of altered ATP levels would further strengthen our conclusion. However, to the best of our knowledge, there are currently no publicly available large-scale datasets that directly measure ATP levels in human postmortem brain tissues. This limitation makes it challenging to incorporate validation in the present study.

      Regarding experimental follow-up, we agree that functional validation is essential to confirm the mechanistic implications of our findings. We are actively considering follow-up experimental studies. However, we consider the present work to be a multi-omic integrative analysis aimed at identifying key molecular alterations and generating biologically important hypotheses. We have revised the Limitation section to more clearly position this manuscript as an observational systems-level analysis (pg. 13, lines 20-22).

      (5) The validation analysis did not sufficiently show the generalizability of this study's results. The authors demonstrated a correlation of 0.53 to the MSBB transcriptomics data and 0.60 to the AMP-AD DiverseCohorts proteomics data. Beyond these correlation coefficients, no meaningful comparison between the datasets is offered. How concordant are the differentially expressed features (or pathways) between the datasets? How robust would the trans-omic network be if incorporating the alternate datasets? Is the main conclusion (energy metabolism is down in AD) supported by the validation datasets? We think this analysis should be expanded and described in the main text.

      Although the results for external metabolomics datasets are reported in Fig S2C, correlation coefficients with the external data are not reported. The authors state, "Note that each study used different definitions for AD and CT groups, had variations in measurement methods and brain regions analyzed." We appreciate these limitations. However, the external data should be re-analyzed using the same definitions of AD and CT, if possible. The limitations and results (which DEMs are shared between datasets) should be discussed in the main text.

      We thank the reviewer for this important comment regarding the generalizability of our findings. In the revised manuscript, we have expanded the validation analyses and summarized the results in Figure S2. First, at the transcriptomic level, Figure S2B and S2C show the overlap between up- and downregulated genes in AD identified in our ROSMAP-derived analyses and those reported in a previously published large-scale meta-analysis of 2,114 postmortem samples across seven brain regions (Wan et al., 2020). A substantial proportion of DEGs were shared, supporting cross-cohort and cross-region robustness to some extent. At the proteomic level, Figure S2E shows a comparison between the ROSMAP and the AMP-AD DiverseCohorts datasets. We highlighted the subset of enzymes involved in the energy metabolism analysis shown in Fig. 3 and calculated a separate correlation coefficient for this subset (Pearson coefficient = 0.86, p-value = 1.5e-7), further supporting our main conclusion. In addition, to assess the concordance between the two datasets in a threshold-independent manner, we additionally performed Rank-Rank Hypergeometric Overlap (RRHO) analysis (Figure S2E). RRHO analysis (Cahill et al., 2018; Plaisier et al., 2010) enables the comparison of ranked protein lists without relying on arbitrary differential expression cutoffs and has been used for cross-dataset comparison in several previous studies (Fröhlich et al., 2024; Maitra et al., 2023). The RRHO heatmaps demonstrated significant enrichment in the concordant quadrants, confirming systematic agreement between datasets beyond simple correlation coefficients. For metabolomics, Figure S2G shows RRHO analyses comparing the ROSMAP metabolomic data with other datasets measured by the same UPLC-MS/MS platform (Batra et al., 2024; Novotny et al., 2023), demonstrating significant concordance in ranked metabolite changes in AD.

      (6) The glycolysis analysis and discussion needs more development. Glycolysis and gluconeogenesis share many of the same enzymes, but they are not the same pathway and should not be discussed as such. To make a claim about the overall influence of enzyme and metabolite levels on glycolysis, the authors should focus on the energetically committing steps of glycolysis (hexokinase, phosphofructokinase, pyruvate kinase) in Figure 3A, and include the full/current version of the figure in the supplement. Gluconeogenesis-specific enzymes (pyruvate carboxylase, PEPCK) are not mentioned at all - are they among the DEPs/DEGs?

      We appreciate the reviewer’s comment regarding the distinction between glycolysis and gluconeogenesis pathway. Among the gluconeogenesis-specific enzyme proteins, G6PC1, FBP1, PC, and PCK2 were measured in our dataset, but none of them were identified as DEPs. In addition, gluconeogenesis is a process that occurs primarily in the liver and kidney rather than the brain. Given this biological context and the lack of significant changes in relevant enzymes, we have revised the terminology throughout the manuscript, replacing “glycolysis/gluconeogenesis pathway” with “glycolysis pathway” in the revised version.

      (7) Given that there wasn't good concordance between the DEGs and DEPs, did including the mRNA and transcription factor layers in the network really add anything useful? It seems like the main conclusions of the manuscript were driven by the protein and metabolite layers only. How many of the DE metabolic enzymes were coregulated at the transcript and protein level? It would be useful to include the 5-layer trans-omic network in the supplement to display these results. Given your network, at what level does it appear that energy metabolism is regulated?

      It is true that our primary conclusion regarding the regulation of energy metabolism is driven by the changes in protein and metabolite abundance. However, we consider the low concordance between mRNA and protein expression itself to be an important feature of AD pathology, as also reported in previous studies (Johnson et al., 2022; Tasaki et al., 2022). Although we did not perform a further analysis of this discordance, we believe that including the TF and mRNA layers into the metabolic trans-omic network strengthens a system-wide view of metabolic dysregulation in AD.

      Regarding the mRNA changes corresponding to the DEP enzymes, please refer to Figure S7A.

      (8) Comment further on the results from Figure 2D. What can be learned from identifying metabolites with the greatest degree centrality? What pathways other than energy metabolism are highlighted by the trans-omic network?

      We assume that some energetic indicators, including AMP and acetyl-CoA, and nitrogen metabolism-related metabolites, Glu, 2-oxoglutarate, and urea, can be potential key regulators of dysregulated metabolism in AD.

      (9) (Suggestion) We suggest the authors leverage their trans-omic network in additional ways beyond giving a snapshot of a few energy metabolism pathways. The analysis of top DEMs could go further. What pathways are impacted beyond energy metabolism? Among the metabolic reactions allosterically regulated by top DEMs, what metabolic pathways are enriched?

      We identified the enriched metabolic pathways that were allosterically regulated by DEMs in AD using Fisher’s exact test. Alanine, aspartate, and glutamate metabolism pathways were significantly enriched in 2-oxoglutarate, glutarate, alanine, and glutamate-regulating metabolic reactions. Arginine and proline metabolism pathway was enriched in N-methyl-L-arginine and putrescine-regulating metabolic reactions. Arginine biosynthesis pathway was enriched in arginine-regulating metabolic reactions. Glycerophospholipid metabolism pathway was enriched in CDP-ethanolamine-regulating metabolic reactions. Glycine, serine, and threonine metabolism pathway was enriched in serine-regulating metabolic reactions. Purine metabolism pathway was enriched in AMP-regulating metabolic reactions. Pyrimidine metabolism pathway was enriched in deoxyuridine and thymidine-regulating metabolic reactions. Sphingolipid metabolism pathway was enriched in sphingosine-regulating metabolic reactions. However, this analysis did not yield sufficiently valuable insights into the regulatory relationships among biomolecules in AD. Thus, we did not include these results in the revised manuscript.

      (10) (Suggestion) Figure 3 shows that most differential signal in AD points to lower energy production due to the combination of differentially expressed metabolites and enzymes, but we are not given much context about the strength of these among all the differential signals. We would suggest including volcano plots where the features of interest, i.e. DE enzymes and metabolites, are colored differently (or a similar figure).

      We thank the reviewer for this constructive suggestion. To provide better context regarding the importance of the differential signals, we have added volcano plots for mRNAs, proteins, and metabolites in Figure S4A, B, and C.

      (11) (Suggestion) The PPI network could be better leveraged to understand metabolic changes in AD. If nodes are grouped into subnetworks (e.g. by Louvain / Leiden clustering) and tested for pathway enrichment, could you find functional subnetworks of coordinately up- and down- regulated metabolic enzymes? This could yield some pathways of interest beyond the energy metabolism pathways already highlighted.

      We appreciate the reviewer’s suggestion to utilize the PPI network for subnetwork analysis. However, it is important to note that the proteomic dataset analyzed in this study is derived from the original work of (Johnson et al., 2020). In that paper, the authors already performed a Weighted Gene Co-expression Network Analysis (WGCNA) across several datasets to identify co-expressed modules and functional pathways.

      Given this, we assumed that applying additional clustering methods to the same dataset would be unlikely to yield significant biological insights beyond the established findings.

      Minor comments

      (1). "All genes" and "all metabolites" should not be the background for the proteomic and metabolic pathway enrichment analysis by Metascape and MetaboAnalyst. The background should be limited to the proteins and metabolites that were measured.

      We fully agree with the reviewer that using “all gene” or “all metabolites” as a background is not suitable for enrichment analyses. As suggested, we have revised the enrichment analyses using the measured proteins and metabolites as a background in both Metascape and MetaboAnalyst (Fig. S4D).

      (2) Highlight the metabolic enzymes in Fig S2B. Calculate a separate correlation coefficient for the enzymes extracted in the energy metabolism analysis from Fig 3.

      We appreciate the reviewer’s suggestion to refine the correlation analysis. As requested, we have revised Fig. S2D to explicitly highlight the subset of enzymes involved in the energy metabolism analysis shown in Fig. 3. We calculated a separate correlation coefficient for the subset (Pearson coefficient = 0.86, p-value = 1.5e-7).

      (3) Use a multiple hypothesis adjusted p-value or q-value in Figure S3.

      We agree with the reviewer regarding the necessity of correcting for multiple comparisons. Accordingly, we have revised Fig. S4D using q-values.

      (4) Describe the methods used to calculate the logFC values from the validation dataset.

      We have revised the Methods to include a detailed description of the procedure used to calculate the log2FC values for the validation datasets (pg. 21, lines 13-15).

      (5) It is difficult to read Figure 3. We would recommend really emphasizing to the reader to refer to Fig S7B as a "key" to this figure. The description of the red/blue arrows and nodes in the methods section (pg. 24, lines 21-36, pg 25, lines 1-4) were also helpful, but very lengthy. We recommend putting an abridged version of this description into the Fig S7 figure legend.

      We appreciate the feedback regarding the readability of Fig. 3. As recommended, we have revised the manuscript to explicitly direct readers to Fig. S8B as an essential “key” for interpreting the network visualization (pg. 8, lines 28). Furthermore, we have added an abridged description of the network elements to the legend of Fig. S8B.

      (6) The S7 figure legend should refer to panels A and B, not E and F.

      We apologize for this oversight. We have corrected the legend of Fig. S8.

      (7) (Suggestion) Are any of the differentially expressed metabolites allosteric regulators of the DE transcription factors? This could be interesting to discuss.

      We appreciate the reviewer’s insightful suggestion about the potential allosteric regulation of the DETFs by DEMs. We conducted an extensive literature search to identify any reports related to this perspective. However, to the best of our knowledge, no such direct interactions have been reported to date.

      Reviewer #1 (Significance):

      The study's strength lies in leveraging three omics modalities across large patient cohorts (n ~ 150-240) to identify coherent signals between transcriptomics, proteomics, and metabolomics in postmortem DLPFC tissue. It was encouraging to see that the main result, showing downregulation for TCA, oxidative phosphorylation, and ketone body metabolism, emerged from consistent signals across both proteomics and metabolomics. This result was consistent with previous findings in other models cited by the author4,5 and other studies 6,7 demonstrating deficiency in energy-producing pathways in AD.

      Another strength of the study is the application of thoughtful methodology to connect differentially expressed proteins and metabolites via an intermediate data layer of metabolic reactions. The authors leverage the KEGG and BRENDA databases and apply sound logic to estimate the effects of enzyme level and metabolite level on pathway activity, with metabolites serving as substrate, product, or allosteric regulator for reactions. This trans-omic network methodology was developed in previous studies cited by the author8,9.

      However, as written, this study is limited in its contribution of new knowledge to the AD research field. The main conclusion (energy production is down in AD, due to regulatory disruption of energy metabolism) is not strongly supported (see comments 1, 3, and 4 for elaboration). The evidence could be improved by orthogonal approaches: further experimentation, further integration of external datasets, causal modeling, or flux modeling. Alternatively, even in the absence of new experimental and computational approaches, the story could be made more complete by further leveraging the trans-omic network to provide insights into (a) the regulation of energy metabolism; and (b) the impacts of key disrupted metabolites (see comments 7-9).

      The study is also limited in its demonstrating the power of these methodologies to provide integrative insights. As mentioned above, the integration of enzyme levels and metabolite levels is clearly useful (Figure 3). In contrast, the utility of the mRNA and transcription factor layers was not evident. The study did not appear to improve or expand upon trans-omic network methodology described in the previous works. Finally, the various analyses (analyzing the trans-omic network for nodes with the highest degree centrality, the PPI analysis, and viewing the energy metabolism pathways in the network) provided disparate results that were only tenuously connected in the discussion section.

      Reviewer #2 (Evidence, reproducibility and clarity):

      Summary

      This manuscript integrates public transcriptomic, proteomic, and metabolomic datasets from ROSMAP DLPFC samples to construct a multi-layer metabolic trans-omic network in Alzheimer's disease. By linking transcription factors, enzyme mRNAs, proteins, metabolic reactions, and metabolites, the authors report coordinated downregulation of the TCA cycle, oxidative phosphorylation, and ketone body metabolism, along with mixed regulatory signals in glycolysis/gluconeogenesis. They interpret these patterns as indicative of broad energetic dysfunction and alterations in amino-acid/nitrogen metabolism in AD. While the framework is conceptually appealing, much of the analysis remains descriptive, and several biological interpretations extend beyond what the data can robustly support. The reliance on bulk tissue without accounting for cell-type composition, limited covariate adjustment, and the absence of validation or sensitivity analyses reduce confidence in the mechanistic conclusions. Overall, the study provides a preliminary systems-level overview, but additional rigor is needed before the proposed trans-omic regulatory insights can be considered convincing.

      Major Comments

      (1) Interpretation requires more cautious phrasing, and validation is essential. The manuscript frequently asserts that specific pathways are "inhibited" or that energetic deficits are "compensated," but these conclusions extend beyond what the descriptive, bulk-level data can support. Because no metabolic flux, causality, or direct functional measurements are included, the results should be framed as putative regulatory shifts, not confirmed impairments. Critically, key claims about pathway inhibition would require flux modeling, perturbation analyses, or experimental validation to be convincing. Without such validation, the mechanistic interpretations remain speculative.

      We thank the reviewer for this crucial comment. We fully agree that, given the descriptive and bulk-level nature of our analysis, mechanistic interpretations must be made with caution. In the absence of direct metabolic flux measurements or experimental validation, our findings should be interpreted as putative regulatory shifts rather than confirmed functional impairments. Accordingly, we have revised the manuscript to temper mechanistic claims. We have replaced definitive statements with more speculative phrasing (e.g., “Our analysis revealed a putative coordinated downregulation …” instead of “Our analysis revealed a coordinated downregulation …” in Abstract section; “we demonstrate the systems-level view of the potential dysregulated energy production …” instead of “we demonstrate the systems-level view of the dysregulated energy production …” in pg. 10, lines 25-26).

      (2) Although the authors acknowledge this in the limitations, bulk-level differences may primarily reflect altered proportions of neurons, astrocytes, microglia, and oligodendrocytes rather than true within-cell-type regulation. Incorporating a cell-type deconvolution or performing a sensitivity analysis would substantially improve interpretability. This issue also impacts the trans-omic network: if the molecules included originate from different cell types, the inferred regulatory relationships may not reflect true intracellular processes.

      We appreciate the reviewer’s point that bulk-level differences can reflect altered proportions of different brain cell types, subsequently affecting the inferred trans-omic network analysis. To assess the changes in cell type proportions of the samples that we used in our study, we additionally used public single-cell transcriptomic datasets, which were obtained from DLPFC tissue of 465 subjects in the ROSMAP cohort (Green et al., 2024). For each omic data that we used in our analyses, we matched the same subjects and calculated the following cell type proportions, astrocytes, excitatory neurons, inhibitory neurons, microglias, oligodendrocytes, and OPCs. Then, we statistically compared the cell type proportions between control subjects and patients with AD (Fig. S3). In the transcriptomic data, we confirmed that the proportion of inhibitory neurons in the AD group was smaller than in the CT group, and that the proportion of oligodendrocytes in the AD group was larger than in the CT group. In the proteomic data, we did not observe any statistically significant changes in the cell type proportion between the two group. In the metabolomic data, we found that the proportion of inhibitory neurons in the AD group was smaller than in the CT group (pg. 6, lines 8-11).

      (3) Differential analysis covariates. For the differential expression analyses, only gender and PMI were included as covariates. Additional variables, such as age at death, RIN, neuropathological measures, and comorbidities, can strongly influence molecular profiles and should be considered to ensure that the observed differences reflect AD-related biology rather than confounding pathological or technical factors.

      We appreciate the reviewer’s comment regarding the included covariates in differential analyses of our study. The reason we did not include other variables, including age at death and RIN, is that these data for each sample were not available. Thus, we referred to original research articles from which proteomic or metabolomic datasets used in our study were derived. Regarding the metabolomic dataset, in the original article (Batra et al., 2023), only two metabolites, 1-methyl-5-imidazoleacetate and N6-carboxymethyllysine, were significantly associated with age. In addition, no metabolites were significantly associated with sex, BMI, or education. Regarding the proteomic dataset, in the original article, age at death, PMI, and sex were included as covariates in the analyses, though these variables were not found to strongly influence the data (Extended Data Fig.2 in (Johnson et al., 2020)).

      (4) Network stability and sample non-overlap. Proteomic, transcriptomic, and metabolomic data come from partially overlapping individuals. The authors should test whether the reconstructed network is robust to: different significance thresholds, restricting analyses to overlapping samples and alternative definitions of AD vs control.

      We appreciate the reviewer’s comment for the trans-omic network stability. In our study, the number of individuals for whom all omic modalities were measured was relatively small (n=25 in CT and n=35 in AD). This limited overlap reduces statistical power and can affect the downstream network construction. We have acknowledged this limitation in the revised manuscript and clarified that the reconstructed networks should be interpreted with caution regarding reproducibility and generalizability (pg. 13, lines 13-23).

      Minor Comments

      (1) Some TF enrichment and regulatory inferences lack explicit mention of multiple-testing correction.

      We apologize for the lack of clarity in our original description. We have corrected for multiple-testing for the TF inference. Thus, we have revised the Methods section to explicitly describe the correction method used and the threshold applied (pg. 23, lines 23-24).

      (2) The limitations section is strong but should explicitly discuss the influence of postmortem interval on metabolite levels.

      We appreciate the reviewer’s comment about the effect of postmortem interval on changes in metabolite levels. Accordingly, we have added the description of this perspective in our revised manuscript (pg. 13, lines 1-5).

      Reviewer #2 (Significance):

      The study extends a trans-omic integration framework, originally applied to metabolic disease, into the context of Alzheimer's pathology. Although the biological findings largely confirm known alterations in mitochondrial and energy metabolism, the network-based approach offers a structured way to view cross-layer regulatory changes. Its main advance is conceptual rather than biological, providing a unified framework rather than uncovering fundamentally new mechanisms. This work will primarily interest researchers in neurodegeneration and systems biology, as well as computational groups developing multi-omics integration methods.

      Reviewer #3 (Evidence, reproducibility and clarity):

      This study leverages existing transcriptomic, metabalomic and proteomic datasets from prefrontal cortex (PFC) to assess metabolic dysregulation in Alzheimer's disease (AD). They found a downregulation of multiple metabolic pathways, including TCA cycle, oxidative phosphorylation, and ketone metabolism, that may explain bioenergetic alterations in AD.

      The study used matching ROSMAP omics datasets from the DLPFC that have allowed more robust data integration. However, the datasets are all generated using bulk tissue, which makes data interpretation difficult. For example, the AD changes they observed may be due to shifts in cell type proportion with disease (e.g. cell death, neuron inflammation). Did the authors account for any potential shifts in cell type proportion in their analysis?

      If the assumption is that the changes in AD are cell intrinsic, which cell types are likely to be impacted? Can the authors integrate any existing single-cell analysis to infer which cell types may be driving the signals they detect, and whether this accounts for some of the antagonistic regulatory effects that were detected?

      We thank the reviewer for their insightful comments. We agree that the use of bulk tissue datasets cannot account for cell-type heterogeneity. As noted in our Limitations section (pg. 12, lines 24-27), we recognize that previous studies have found that the Braak stage is correlated positively with microglia and astrocyte proportions and negatively with oligodendrocyte proportion (Hannon et al., 2024; Shireby et al., 2022). Regarding the integration of single-cell analysis, we have referenced recent snRNA-seq findings (Mathys et al., 2024) in our Limitations section (pg. 12, lines 28-32) to deconvolve our bulk signatures.

      Furthermore, in our revised manuscript, we additionally used public single-cell transcriptomic datasets, which were obtained from DLPFC tissue of 465 subjects in the ROSMAP cohort (Green et al., 2024). For each omic data that we used in our analyses, we matched the same subjects and calculated the following cell type proportions, astrocytes, excitatory neurons, inhibitory neurons, microglia, oligodendrocytes, and OPCs. Then, we statistically compared the cell type proportions between control subjects and patients with AD (Fig. S3). In the transcriptomic data, we confirmed that the proportion of inhibitory neurons in the AD group was smaller than in the CT group, and that the proportion of oligodendrocytes in the AD group was larger than in the CT group. In the proteomic data, we did not observe any statistically significant changes in the cell type proportion between the two groups. In the metabolomic data, we found that the proportion of inhibitory neurons in the AD group was smaller than in the CT group (pg. 6, lines 8-11).

      Reviewer #3 (Significance):

      The manuscript provides multimodal insight into metabolic dysregulation in AD in the PFC. Given that metabolic dysfunction is likely to play a major in disease pathogenesis, this is a study of importance. However, the findings lack granularity at the cell type level, which limits the impact of the study.

      Reference

      (1) Baloni, P., Arnold, M., Buitrago, L., Nho, K., Moreno, H., Huynh, K., Brauner, B., Louie, G., Kueider-Paisley, A., Suhre, K., Saykin, A. J., Ekroos, K., Meikle, P. J., Hood, L., Price, N. D., Alzheimer’s Disease Metabolomics Consortium, Doraiswamy, P. M., Funk, C. C., Hernández, A. I., … Kaddurah-Daouk, R. (2022). Multi-Omic analyses characterize the ceramide/sphingomyelin pathway as a therapeutic target in Alzheimer’s disease. Communications Biology, 5(1), 1074.

      (2) Baloni, P., Funk, C. C., Yan, J., Yurkovich, J. T., Kueider-Paisley, A., Nho, K., Heinken, A., Jia, W., Mahmoudiandehkordi, S., Louie, G., Saykin, A. J., Arnold, M., Kastenmüller, G., Griffiths, W. J., Thiele, I., Alzheimer’s Disease Metabolomics Consortium, Kaddurah-Daouk, R., & Price, N. D. (2020). Metabolic Network Analysis Reveals Altered Bile Acid Synthesis and Metabolism in Alzheimer’s Disease. Cell Reports. Medicine, 1(8), 100138.

      (3) Batra, R., Arnold, M., Wörheide, M. A., Allen, M., Wang, X., Blach, C., Levey, A. I., Seyfried, N. T., Ertekin-Taner, N., Bennett, D. A., Kastenmüller, G., Kaddurah-Daouk, R. F., Krumsiek, J., & Alzheimer’s Disease Metabolomics Consortium (ADMC). (2023). The landscape of metabolic brain alterations in Alzheimer’s disease. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 19(3), 980–998.

      (4) Batra, R., Krumsiek, J., Wang, X., Allen, M., Blach, C., Kastenmüller, G., Arnold, M., Ertekin-Taner, N., Kaddurah-Daouk, R., & Alzheimer’s Disease Metabolomics Consortium (ADMC). (2024). Comparative brain metabolomics reveals shared and distinct metabolic alterations in Alzheimer’s disease and progressive supranuclear palsy. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 20(12), 8294–8307.

      (5) Cahill, K. M., Huo, Z., Tseng, G. C., Logan, R. W., & Seney, M. L. (2018). Improved identification of concordant and discordant gene expression signatures using an updated rank-rank hypergeometric overlap approach. Scientific Reports, 8(1), 9588.

      (6) Fröhlich, A. S., Gerstner, N., Gagliardi, M., Ködel, M., Yusupov, N., Matosin, N., Czamara, D., Sauer, S., Roeh, S., Murek, V., Chatzinakos, C., Daskalakis, N. P., Knauer-Arloth, J., Ziller, M. J., & Binder, E. B. (2024). Single-nucleus transcriptomic profiling of human orbitofrontal cortex reveals convergent effects of aging and psychiatric disease. Nature Neuroscience, 27(10), 2021–2032.

      (7) Green, G. S., Fujita, M., Yang, H.-S., Taga, M., Cain, A., McCabe, C., Comandante-Lou, N., White, C. C., Schmidtner, A. K., Zeng, L., Sigalov, A., Wang, Y., Regev, A., Klein, H.-U., Menon, V., Bennett, D. A., Habib, N., & De Jager, P. L. (2024). Cellular communities reveal trajectories of brain ageing and Alzheimer’s disease. Nature, 633(8030), 634–645.

      (8) Hannon, E., Dempster, E. L., Davies, J. P., Chioza, B., Blake, G. E. T., Burrage, J., Policicchio, S., Franklin, A., Walker, E. M., Bamford, R. A., Schalkwyk, L. C., & Mill, J. (2024). Quantifying the proportion of different cell types in the human cortex using DNA methylation profiles. BMC Biology, 22(1), 17.

      (9) Johnson, E. C. B., Carter, E. K., Dammer, E. B., Duong, D. M., Gerasimov, E. S., Liu, Y., Liu, J., Betarbet, R., Ping, L., Yin, L., Serrano, G. E., Beach, T. G., Peng, J., De Jager, P. L., Haroutunian, V., Zhang, B., Gaiteri, C., Bennett, D. A., Gearing, M., … Seyfried, N. T. (2022). Large-scale deep multi-layer analysis of Alzheimer’s disease brain reveals strong proteomic disease-related changes not observed at the RNA level. Nature Neuroscience, 25(2), 213–225.

      (10) Johnson, E. C. B., Dammer, E. B., Duong, D. M., Ping, L., Zhou, M., Yin, L., Higginbotham, L. A., Guajardo, A., White, B., Troncoso, J. C., Thambisetty, M., Montine, T. J., Lee, E. B., Trojanowski, J. Q., Beach, T. G., Reiman, E. M., Haroutunian, V., Wang, M., Schadt, E., … Seyfried, N. T. (2020). Large-scale proteomic analysis of Alzheimer’s disease brain and cerebrospinal fluid reveals early changes in energy metabolism associated with microglia and astrocyte activation. Nature Medicine, 26(5), 769–780.

      (11) Maitra, M., Mitsuhashi, H., Rahimian, R., Chawla, A., Yang, J., Fiori, L. M., Davoli, M. A., Perlman, K., Aouabed, Z., Mash, D. C., Suderman, M., Mechawar, N., Turecki, G., & Nagy, C. (2023). Cell type specific transcriptomic differences in depression show similar patterns between males and females but implicate distinct cell types and genes. Nature Communications, 14(1), 2912.

      (12) Mathys, H., Boix, C. A., Akay, L. A., Xia, Z., Davila-Velderrain, J., Ng, A. P., Jiang, X., Abdelhady, G., Galani, K., Mantero, J., Band, N., James, B. T., Babu, S., Galiana-Melendez, F., Louderback, K., Prokopenko, D., Tanzi, R. E., Bennett, D. A., Tsai, L.-H., & Kellis, M. (2024). Single-cell multiregion dissection of Alzheimer’s disease. Nature, 632(8026), 858–868.

      (13) Novotny, B. C., Fernandez, M. V., Wang, C., Budde, J. P., Bergmann, K., Eteleeb, A. M., Bradley, J., Webster, C., Ebl, C., Norton, J., Gentsch, J., Dube, U., Wang, F., Morris, J. C., Bateman, R. J., Perrin, R. J., McDade, E., Xiong, C., Chhatwal, J., … Harari, O. (2023). Metabolomic and lipidomic signatures in autosomal dominant and late-onset Alzheimer’s disease brains. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 19(5), 1785–1799.

      (14) Plaisier, S. B., Taschereau, R., Wong, J. A., & Graeber, T. G. (2010). Rank-rank hypergeometric overlap: identification of statistically significant overlap between gene-expression signatures. Nucleic Acids Research, 38(17), e169.

      (15) Shireby, G., Dempster, E. L., Policicchio, S., Smith, R. G., Pishva, E., Chioza, B., Davies, J. P., Burrage, J., Lunnon, K., Seiler Vellame, D., Love, S., Thomas, A., Brookes, K., Morgan, K., Francis, P., Hannon, E., & Mill, J. (2022). DNA methylation signatures of Alzheimer’s disease neuropathology in the cortex are primarily driven by variation in non-neuronal cell-types. Nature Communications, 13(1), 5620.

      (16) Tasaki, S., Xu, J., Avey, D. R., Johnson, L., Petyuk, V. A., Dawe, R. J., Bennett, D. A., Wang, Y., & Gaiteri, C. (2022). Inferring protein expression changes from mRNA in Alzheimer’s dementia using deep neural networks. Nature Communications, 13(1), 655.

      (17) Varma, V. R., Wang, Y., An, Y., Varma, S., Bilgel, M., Doshi, J., Legido-Quigley, C., Delgado, J. C., Oommen, A. M., Roberts, J. A., Wong, D. F., Davatzikos, C., Resnick, S. M., Troncoso, J. C., Pletnikova, O., O’Brien, R., Hak, E., Baak, B. N., Pfeiffer, R., … Thambisetty, M. (2021). Bile acid synthesis, modulation, and dementia: A metabolomic, transcriptomic, and pharmacoepidemiologic study. PLoS Medicine, 18(5), e1003615.

      (18) Wan, Y.-W., Al-Ouran, R., Mangleburg, C. G., Perumal, T. M., Lee, T. V., Allison, K., Swarup, V., Funk, C. C., Gaiteri, C., Allen, M., Wang, M., Neuner, S. M., Kaczorowski, C. C., Philip, V. M., Howell, G. R., Martini-Stoica, H., Zheng, H., Mei, H., Zhong, X., … Logsdon, B. A. (2020). Meta-Analysis of the Alzheimer’s Disease Human Brain Transcriptome and Functional Dissection in Mouse Models. Cell Reports, 32(2), 107908.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary

      Large language models (LLMs) have been developed rapidly in recent years and are already contributing to progress across scientific fields. The manuscript tries to address a specific question: whether LLMs can accurately infer signaling networks from gene lists. However, the evaluation is inadequate due to four major weaknesses described below. Despite these limitations, the authors conclude that current general-purpose LLMs lack adequate accuracy, which is already widely recognized. Its key contribution should instead be to provide concrete recommendations for the development of specialized LLMs for this task, which is completely absent. Developing such specific LLMs would be highly valuable, as they could substantially reduce the time required by researchers to analyze signaling networks.

      Strengths

      The manuscript raises a good question: whether current LLMs can accurately generate signaling networks from gene lists.

      Weaknesses:

      (1) The authors evaluate LLM performance using only three signaling networks: "hypertrophy", "fibroblast", and "mechanosignaling". Given the large number of well-established signaling pathways available, this is not a comprehensive assessment. Moreover, the analysis need not be restricted to signaling networks. Other network types, including metabolic and transcriptional regulatory networks, are already accessible in well-known databases such as KEGG, Reactome, BioCyc, WikiPathways, and Pathway Commons. Including these additional networks would substantially strengthen the evaluation.

      We agree with the reviewer that our evaluation of LLM performance is not comprehensive of all signaling networks, and that the benchmarking was previously limited to signaling networks. The purpose of this study is to benchmark LLMs against peer-reviewed computational models that make testable predictions and are highly validated experimentally, of which these three signaling networks are strong examples. KEGG, Reactome, Biocyc, WikiPathways, and Pathway Commons are databases that contain collections of individual reactions or pathways but are not computational models themselves and have not been experimentally validated in that sense.

      While this study focuses primarily on signaling networks, we agree that it would be useful to evaluate how well LLMs perform in generating networks of another type, for which predictive and validated computational models are available. Therefore, in new Figure 3 we now test the ability of LLMs to reconstruct the E. Coli core metabolic network, as well as test its ability to predict growth on metabolic substrates using flux balance analysis. We find that Claude Opus 4.6, GPT 5.2 Pro, and Gemini 3 Pro Preview perform well at reconstructing reactions from E. Coli core metabolism, but these reactions are not sufficient to predict growth on a variety of substrates. Given this expansion of scope, we replaced “signaling networks” in the title with “biochemical networks”.

      (2) In LLM evaluation, the authors use the gene lists that exactly match those in their "ground truth" networks, thereby fixing the set of nodes and evaluating only the predicted edges. However, in practical research, the relevant genes or nodes are not fully known. A more realistic assessment would therefore include gene lists with both genes present in the ground-truth network and additional genes absent from it, to evaluate the ability of the LLM to exclude irrelevant genes.

      We agree with the reviewer that evaluating the capacity of these LLMs to exclude additional genes is interesting. But because biological networks are always incompletely known, there is no “ground truth” of genes absent from a given network. Therefore, for the most rigorous benchmarking against a “ground truth”, we examine the positive predictive ability of LLMs. However, in response to this comment and point 3 below, we further examine additional measures of performance that include “false positives”.

      (3) The authors report only the recall/sensitivity of the LLM, without assessing specificity. In practical applications, if an LLM generates a large number of incorrect interactions that greatly exceed the correct ones, researchers may be misled or may lose confidence in the LLM output. Therefore, a comprehensive evaluation must include both sensitivity and specificity. Furthermore, it would be informative to check whether some of the "false positives" might in fact represent biologically plausible interactions that are absent from the manually curated "ground truth". Manually generated "ground truth" can overlook genuine interactions, and the ability of LLMs to recover such missing edges could be particularly valuable. This may even represent one of the most important potential contributions of LLMs.

      We agree with the reviewer that additional metrics could inform the evaluation of LLM performance. Therefore, as recommended, we calculated sensitivity, specificity, precision, negative predictive value (NPV), accuracy, and F1 score for each of the network models (hypertrophy, fibroblast, and mechanosignaling). These new results are summarized in confusion matrices shown in a new Supplementary Figures 3, 4, and 5. We performed this additional benchmarking across the 10 replicates for each LLM.

      One limitation of this approach is the substantial class imbalance within these confusion matrices. Because we interpret actual negatives as connections that are not found between any nodes of the ground truth models, there will be >10x more true negatives than any of the other classes. This makes specificity, accuracy, and the negative predictive value less informative.

      To illustrate this point, consider the ground-truth hypertrophy network which contains 191 connections between 106 nodes. The total number of possible connections between any two nodes is 106<sup>2</sup> = 11,236. Given that there are 191 actual positives, that leaves 11,045, actual negatives (as illustrated in the null predictor of Supplementary Figure 3A). As the number of node-to-node connections predicted by LLMs is on the order of a few hundred, the number of true negatives is always in the thousands, often outweighing the true positives in the specificity or accuracy calculations.

      The effects of this class imbalance are illustrated with the results of the “null predictors” in Supplementary Figure 4A, which are hypothetical models that fail to predict any connection between nodes (have predicted positive values of 0). These null predictors have sensitivities of 0 and specificities of 1 and high accuracies and NPVs because of the high true negative rates.

      The precision and F1 scores calculated using these confusion matrices are robust to these class imbalances. Indeed, there is substantial heterogeneity in the number of “false positives” connections generated by the LLMs as illustrated by the precision and F1 scores. We are hesitant to unequivocally label these novel, predicted connections as false positives because, as the reviewer points out, these connections could represent true molecular interactions that were undiscovered at the time of the ground truth models’ conception but have since been demonstrated experimentally and published.

      (4) It is widely known that applying differential equation models to highly complex biological networks, such as the three networks in the manuscript, is meaningless, because these systems involve a large number of parameters whose values can drastically alter the results. As Richard Feynman once said: "with four parameters I can fit an elephant, and with five I can make him wiggle his trunk." Thus, the evaluation of LLMs on "logic-based differential equation models" does not make much sense.

      Differential equation models have been the primary mathematical framework for studying complex systems for decades. We refer readers unfamiliar with differential equations to the Nobel prize-winning work of Hodgkin/Huxley (Physiology or Medicine1963, action potential of neurons), Prigogine (Chemistry 1977, non-equilibrium thermodynamics and pattern formation), John Nash (Economics 1994, game theory and Nash Equilibrium), Merton/Scholes (Economics 1997, dynamics of financial derivatives), and Manabe/Hasselmann (Physics 2021, dynamic modeling of atmosphere and oceans).

      We are confused by the quote of a joke by Richard Feynman about fitting equations to data in the shape of an elephant. While this famous joke is amusing, it is both misattributed (it was a recollection by Enrico Fermi in 1953 of a joke once made by John von Neumann) and deliberately hyperbolic (see https://en.wikipedia.org/wiki/Von_Neumann%27s_elephant). Regardless, the relevance of this joke to our study is unclear, because we are not fitting equations to data. As described in the text, in previous studies we validated the predictions of these three logic-based network models with experimental data that was not used to develop the models.

      Reviewer #1 (Recommendations for the authors):

      (1) All figures are in very poor resolution.

      Thank you for identifying this. We have fixed this issue, which was due to SVG embedding. We now embed as higher resolution PNG and provide full resolution files separately.

      (2) The manuscript does not include data availability or code availability.

      As described in the Methods, all code and data is now available via GitHub (https://github.com/saucermanlab/LLM-network-generation).

      Reviewer #2 (Public review):

      (1) Information on the accuracy of directionality of interaction would help understand if there is a bias towards either a positive or negative association.

      To assess if LLMs are biased in predicting either stimulatory or inhibitory connections, we examined the proportion of stimulatory and inhibitory connections for each ground truth model along with the prediction sets from the different LLMs (Author response table 1). These findings suggest that any directional bias is minimal and not conserved across the different ground truth models. These tables were not included in the revised manuscript.

      Author response table 1.

      Proportion of stimulatory and inhibitory connections present in each ground truth model and in the sets of connections predicted by each LLM (Claude, GPT, and Gemini).

      The primary subset of reaction types that the LLM’s tend to miss are often downstream, cell-type specific, and/or gene regulatory connection (e.g. Figure 1B). This observation is consistent with the fact that the ground truth models were constructed using experimental evidence from specific publications involving defined experimental models and cell types whereas the LLM’s presumably draw from the entire corpus of published literature.

      (2) Do all LLMs capture similar information, or are some LLMs better at capturing certain information than others? Further to this, it would be interesting to look into whether amalgamating information across all three LLMs results in a more accurate network.

      The reviewer asks interesting questions that can be qualitatively answered in Figure 1B and in the network visualizations (Supplementary Figures 1 and 2). These diagrams illustrate predicted connections that are shared between the nodes. To include a more quantitative, comprehensive assessment of this overlap, we have included Venn diagrams showing the extent to which LLMs capture shared information (Supplementary Figures 3-5). One such Venn diagram for the hypertrophy model is included in Supplementary Figure 3C. Indeed, it seems that there is substantial overlap in the “false positive” connections predicted by Claude and GPT. It could be interesting to evaluate if these connections are reflective of newly discovered molecular interactions, as referenced in our response to reviewer one point 3.

      While amalgamating information across all three LLMs might generate a more accurate network, these Venn diagrams illustrate that there remain connections within the ground truth models that are not represented in any of the prediction sets from the LLMs. Indeed, the union of all predicted connections for the hypertrophy network made by any of the 10 replicates from the different LLMs would still lack 33 ground truth connections (Supplemental Figure 3C).

      (3) Would it be possible to retrieve a confidence value of the interactions from the LLMs and conduct Precision, recall rate, AUPR and calibration analyses? These metrics would also help with the benchmarking process.

      We agree with the reviewer that including additional evaluative metrics is instructive. Therefore, for each reference network, we have included precision, specificity, recall, accuracy, and F1 score calculations (Supplementary Figures 3-5). Additionally, we conducted calibration analyses to illustrate the LLM’s reliability. While this latter analysis is interesting, we note that the calibration curves were generated using the frequency of predictions among the 10 replicates as a proxy for confidence. This frequency is highly sensitive to the temperature parameter of the LLMs. Different temperature settings could substantially influence the trajectories of these confidence curves and the distributions of the associated histograms. See Supplemental Figure 3B for calibration analyses of the hypertrophy model.

      We considered performing analyses resembling AUPR as suggested by the reviewer, but we did see a defensible way to vary a “threshold” for the classification of a connection as positive or negative. In this study, a connection predicted by the LLM either does or does not exist within the ground truth mode.

      Typos:

      (1) Generated networks to predict THE CLASSIC "fetal gene program" gene expression.

      Thank you for catching this error. We have corrected it.

      (2) A manually curated network has A functional accuracy of

      Thank you for catching this error. We have corrected it.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states, but also enabled the real time monitoring of protein conformational changes precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results and conclusions supported by the experimental data. Authors have determined the conformational dynamics of TmrAB across different ATP concentrations including physiological ones and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies. Authors have also mentioned limitations in the study.

      Comments on revised version.

      Authors have worked on most of the revisions stated in previous feedback and included in the newer version, which has been significantly improved. Other comments have been described to be out of scope from this study.

      Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATPbound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I had three major concerns with the original version, all of which have been addressed by the authors in this revised version.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I mentioned that the final section of the Results part seems like an afterthought, especially since the heading suggests a broader scope.

      Reply: We appreciate this comment. We have revised the final section of the Results to improve its structure and ensure that the scope indicated by the heading is fully reflected in the content. This section now more clearly integrates kinetic and thermodynamic aspects of the transport cycle.

      The changes made to the section do not align with the wording of the reply. Please consider modifying it further.

      We appreciate the positive feedback and this final comment. We have revised the final section of the Results to better reflect the scope indicated by the heading. In addition to clarifying the kinetic analysis, we now explicitly relate our kinetic observations to previously determined thermodynamic measurements, showing that the rapid interconversion of ATP-bound conformations observed during steady-state turnover is consistent with a thermodynamic landscape characterized by a near-zero free-energy difference and entropy–enthalpy compensation. This revision more clearly integrates the kinetic and thermodynamic aspects of the transporter cycle.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript aims to test the idea that visual recognition (of faces) is hierarchically organized in the human ventral occipital-temporal cortex (VOTC). The paper proposes that if VOTC has a hierarchical organization, this should be seen in two independent features of the VOTC signal. First, hierarchy assumes that signals along the hierarchy increase in representational complexity. Second, hierarchy assumes a progressive increase in the onset time of the earliest neural response at each level of the hierarchy. To test these predictions, the authors extract high-frequency broadband signals from iEEG electrodes in a very large sample of patients (N=140). They find that face selectivity in these signals is distributed across the VOTC with increasing posterior-anterior face selectivity, hence providing evidence for the first prediction. However, they also find broadband activity to occur concurrently, therefore challenging the view of a serial hierarchy.

      Strengths:

      (1) The hypothesis (that VOTC is hierarchically organized) and predictions (that hierarchy predicts increases in representational complexity and increases in onset time) were clearly described.

      (2) The number of subjects sampled (140) is extremely large for iEEG studies that typically involve <10 subjects. Also, 444 face selective recording contacts provide a very nice sampling of the areas of interest.

      We would like to thank this reviewer for their positive comments and evaluation of our manuscript.

      Weaknesses:

      (1) A control analysis where areas have known differences in response onset should be performed to increase confidence that the proposed analyses would reveal expected results when a difference in response onset was present across areas. From Figure 3, it can be seen that many electrodes are placed in earlier visual areas (V1-V3) that have previously been shown to have earlier broadband responses to visual images compared to VOTC (e.g. Martin et al., 2019, JNeurosci https://doi.org/10.1523/JNEUROSCI.1889-18.2018). The same analyses as in Figures 4 and 5 should be used comparing VOTC to early visual areas to confirm that the analyses would detect that V1-V3 have earlier onsets compared to VOTC.

      First, we would like to mention that the analyses performed in our paper are commonly accepted analyses to extract time-domain information.

      Yet, the reviewer is right that, considering our claim, evidence that our approach does indeed allow revealing ‘true’ differences in onset latencies would provide further support for our claims. The solution proposed by the reviewer is interesting but the number of face-selective recording contacts/sites in posterior occipital cortex (colored disks in figure 3) is too small for any meaningful comparison. Moreover, while absolute responses to visual images should indeed emerge earlier in early visual cortex than association cortex, it may not be the case for category-selective responses to visual images (which is what our claim is about).

      To address the reviewer’s concern, we used face-selective responses from regions known to have different response onset latencies: occipital and posterior temporal lobe vs. medial temporal lobe structures (e.g., Mormann et al., 2008, https://pmc.ncbi.nlm.nih.gov/articles/PMC2676868/). Waveforms and onset latencies for these regions (OCC, PTL, MTL) are shown in Author response image 1. Despite the small number of contacts showing significant face-selective activity in the MTL (N=20) and the lower SNR in this region, the onset latency differences between OCC/PTL and MTL are significant using all 4 methods of latency estimation (see methods in the main manuscript), with medium to large effect sizes. This was despite noisy latency estimates for the MTL (in particular, the ‘delta slope’ method could not be used to get meaningful Cohen’s d when comparing MTL to OCC). Latency estimates are also slightly higher for the ‘% of peak’ method than in the manuscript, as we estimated the latency at 25% of the peak (instead of 20% in the manuscript), again to allow meaningful estimations for the MTL.

      Author response image 1.

      In addition to this, we performed a simulation analysis where we statistically compared the measured PTL signals to ATL signals that have been artificially, incrementally, shifted forward in time. Author response image 2 shows (top row) the measured onset latencies differences between the 2 regions (PTL minus ATL) estimated using 4 different approaches as a function of the temporal shift applied to ATL, as well as the associated p-values (bottom row). As shown in Author response image 2, the ‘original’ unshifted data yields no significant difference between the 2 regions. The difference however becomes significant when ATL signals is shifted forward in time 10 to 30 ms, depending on the method used.

      Author response image 2.

      These 2 observations provide evidence that our approach does indeed allow revealing ‘true’ differences in onset latencies, further supporting our claims.

      Last, as also suggested by reviewer 2, we conducted a thorough equivalence testing using ROPE and Bayesian factor to support the lack of differences between regions. Equivalence bounds and region of practical equivalence (ROPE) were defined to account for physiological variability corresponding to a small effect (Cohen’s d = 0.199, i.e. standard in equivalence testing) and axonal conduction delays between regions (i.e. ATL is further away from early visual cortex than OCC). These analyses, now reported in the result section of the revised manuscript (Table 1), largely confirm the hypothesis of concurrent onset latencies across VOTC.

      (2) It is unclear why correlating mean timeseries helps understand how much variance is shared between regions (Figure 4). Any variance between images is lost when averaging time series across all images, and this metric thus overestimates the variance shared between areas. Moreover, the finding that correlating time domain signals across VOTC areas does not differ from correlating signals within an area could be driven by this averaging. For example, if the same analysis was done on electrodes in left and right V1 when half of the images had contrast in the left hemifield and the other half had contrast in the right hemifield, the average signals may correlate extremely well, while this correlation falls apart on a trial-by-trial basis. These analyses therefore need to be evaluated on a trial-by-trial basis.

      This is an interesting comment. We agree that variance between images is lost when averaging time-series across all images. However, to use the reviewer’s analogy, in order to support the claim that left and right V1 would show the same onset times and time-course (i.e., no hemispheric lateralization) for lateralized presentations (= the same kind of claim that we make in our paper), it’s the average response across images that should be compared, not a correlation run on a trial-by-trial basis (which would indeed falls apart because of a lack of response in the ipsilateral V1).

      Moreover, we would like to emphasize that the goal of this analysis in our paper is not to make claims about the variance shared between regions. In fact, this is not a key analysis in our paper, the outcome of which is not strictly necessary for the main argument made. Finally, if we were to perform a (time-consuming) image-by-image analysis in our study, correlations would be weak due to low signal-to-noise ratio (each face image appears only 1.6 times per stimulation sequence on average) and the fact that each face image appears after a different non-face image across presentations.

      (3) Previous studies on visual processing in VOTC have shown that evoked potentials are more predictive of the onset of visual stimuli than broadband activity (e.g. Miller et al., 2016, PLOS CB, https://doi.org/10.1371/journal.pcbi.1004660). Testing the prediction from a hierarchical representation that signals along the VOTC increase in onset time should therefore include an evaluation of evoked potential onsets in addition to broadband signals.

      We have used HFB responses in our study as these signals tend to be easier to characterize in the time domain than evoked potentials, and they are more local given their reduced SNR compared to evoked potentials (Jacques et al., 2022; https://pmc.ncbi.nlm.nih.gov/articles/PMC9457683/). Moreover we have previously shown highly correlated time courses across HFB and low frequency evoked potentials in the same paradigm (Jacques et al., 2022, eLife).

      Yet, to address this reviewer’s concern, we replicated the main analyses on low-frequency event-related potential signal, identifying contacts exhibiting significant face-selective responses in the same manner as in Jacques et al (2022). Namely, we start from bipolar-referenced sequences of recording corresponding to the full visual stimulation sequences (~70 s). For each recorded intracerebral contact, we average sequences in the time-domain, crop the average to contain an integer number of face frequency (1.2 Hz) cycles, run an FFT on the cropped sequences and identify the significant contacts with a Z-score procedure identical to that used for HFB signals. We then notch-filter out the visual response (6 Hz and harmonics) from the full length sequences, extract short epochs from the filtered sequences around the onset of each face image, average across epoch for each recording contact, subtract the mean signal measured in the baseline (-0.166 to 0 s relative to face onset) and take the absolute value (to be able to average across contacts despite differences of morphology and polarity). Significant contacts are then subjected to the same analyses as for the HFB signal.

      Results from these analyses are presented as supplementary material (Figure S9, Table S4) in the revised manuscript (referenced in lines of the main manuscript). While we were not able to obtain reliable latency estimates using the z-score method with the same parameters as for HFB signal, these analyses with ERP signal indicate similar onset latencies for ERPs as for HFB activity and largely replicate observations made with HFB. In particular, onset latencies were in a very similar range (~100 to 140 ms) with similar patterns across regions or along VOTC and between-region signal correlations. There were also a few significant face-selective responses over posterior ventro-medial occipital cortex, likely overlapping ‘early visual cortex’ (V1,V2v,V3v,hV4), probably due to limited low-level contributions in this paradigm (see Or et al., 2019, JOV; https://jov.arvojournals.org/article.aspx?articleid=2734585). Over these regions, onset latency was systematically earlier (up to 40ms) than in slightly more anterior regions, (i.e. anterior to -80 mm) where very little variability in onset latency was found up to the ATL region. We have acknowledged this in the revised manuscript.

      (4) Testing the second prediction, that the onset time of processing increases along the VOTC posterior to anterior path, is difficult using the iEEG broadband signal, because from a signal processing perspective, broadband signals are inherently temporally inaccurate, given that they are filtered. Any filtering in the signal introduces a certain level of temporal smoothing. The manuscript should clearly describe the level of temporal smoothing for the filter settings used.

      The reviewer is right that HFB signals are temporally smoothed, potentially yielding slightly underestimated onset latencies. However, our time-frequency analyses parameters ensured a minimal degree of smoothing. In fact, the original submission already contained a description of the expected temporal smoothing resulting from the wavelet transform. This is what we wrote in the original submission: “The number of cycles (i.e., central frequency) of the wavelet was adapted as a function of frequency from 2 cycles at the lowest frequency to 9 cycles at the highest frequency. The temporal smoothing resulting from the wavelet transform was minimal: wavelets had 20 ms of full width at half maximum (FWHM) across the frequency range (i.e. median of FWHM computed at each frequency bin), ensuring that onset timing information is accurate up to 10 ms (i.e half of the FWHM).”

      In the revised manuscript we further elaborate as follows:

      “The number of cycles (i.e., central frequency) of the wavelet was adapted as a function of frequency from 2 cycles at the lowest frequency to 9 cycles at the highest frequency. The temporal smoothing resulting from the wavelet transform was minimal: wavelets had a temporal spread of 20 ms (full width at half maximum - FWHM) across the frequency range (i.e. median of FWHM computed at each frequency bin). A simulation of HFB signals with a constant abrupt onset time and realistic signal-to-noise ratio indicates that the potential underestimation of onset latency due to the wavelet analysis is around 5-10 ms, which is on par with the value of the half width at half maximum (= FWHM/2 = 20/2 ms).”

      Author response image 3 displays simulated HFB signal (using identical wavelet parameters than in our manuscript) in an ideal scenario with a response starting at 150 ms in all trials (N=150 trials), reaching maximum 10 ms later. This provides a theoretical estimate of the slight underestimation of onset latency due to the wavelet transform. It shows onset latency estimates are at most 12 ms underestimation of true onset time.

      Author response image 3.

      That being said, given the physiological noise in the data, the fact that the response onset likely varies slightly from trial to trial, with a variable slope in activity increase, these wavelet parameters (within a certain margin) have likely little influence on the actual latency estimation.

      (5) The onsets of neural activity in VOTC are surprisingly early: around 80-100 ms. This is earlier than what has previously been reported. For example, the cited Quian Quiroga et al. (2023) found single neuron responses to have the earlier onset around 125 ms (their Figure 3). Similarly, the cited Jacques et al., 2016b and Kadipasaoglu et al., 2017 papers also observe broadband onsets in VOTC after 100 ms. Understanding the temporal smoothing in the broadband signal, as well as showing that typical evoked potentials have latencies compared to other work, would increase confidence that latencies are not underestimated due to factors in the analysis pipeline.

      In the revised manuscript, as suggested by reviewer 2, we have modified the data resampling strategy (using hierarchical bootstrap and permutation test that respects the nested structure of the data) to estimate onset latencies, confidence interval and permutation tests. Moreover, since the absolute onset latency estimates depend on the methods used, we now provide estimates using 4 different methods. The overall absolute onset latencies differ slightly across the 4 methods but all median onset latencies vary between 95 ms and 130 ms, which is similar to what was reported in some of the participants in Kadipasaoglu et al., 2017 (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0188834; note that latencies reported at the level of single sites or individual are usually higher due to lower signal-to-noise ratio). Moreover, these latencies for face-selective responses are actually very similar to those measured for non-selective/absolute responses to visual stimuli (e.g. Jacques et al., 2016: around 90-100 ms for faces in FG; Yoshor et al., 2007: ~100 ms in posterior Fusiform Gyrus; Regev et al. 2018: 90-120ms in posterior to middle FG). Other studies, measuring non-selective responses using more ‘conservative’ onset detection methods, find slightly later onset latencies (e.g. Cao et al. 2025: 139 ms in posterior FG; Martin et al., 2019: ~150 ms in ventral occipital -VO- regions).

      Cao R, Zhang J, Zheng J, Wang Y, Brunner P, Willie JT, Wang S. 2025. A neural computational framework for face processing in the human temporal lobe. Current Biology. DOI: https://doi.org/10.1016/j.cub.2025.02.063

      Martin AB, Yang X, Saalmann YB, Wang L, Shestyuk A, Lin JJ, Parvizi J, Knight RT, Kastner S. 2019. Temporal dynamics and response modulation across the human visual system in a spatial attention task: An ECoG study. Journal of Neuroscience 39:333–352. DOI: https://doi.org/10.1523/JNEUROSCI.1889-18.2018, PMID: 30459219

      Regev TI, Winawer J, Gerber EM, Knight RT, Deouell LY. 2018. Human posterior parietal cortex responds to visual stimuli as early as peristriate occipital cortex. European Journal of Neuroscience 48:3567–3582. DOI: https://doi.org/10.1111/ejn.14164, PMID: 30240547

      Yoshor D, Bosking WH, Ghose GM, Maunsell JHR. 2007. Receptive fields in human visual cortex mapped with surface electrodes. Cerebral Cortex 17:2293–2302. DOI: https://doi.org/10.1093/cercor/bhl138, PMID: 17172632

      As an important note, in the revised manuscript, we have removed data from 3 recording contacts from 1 participant that were located in the upper bank of the Calcarine Sulcus, which is actually outside of our VOTC region of interest. The 3 contacts being located in dorsal V1 or V2 were showing very early responses and were biasing our latency estimates for the OCC region.

      (6) Understanding the extent to which neural processing in the VOTC is hierarchical is essential for building models of vision that capture processing in the human brain, and the data provides novel insight into these processes.

      For additional context, a schematic figure of the hierarchical view and a more parallel system described in the paragraph on models of visual recognition (lines 553) would help the reader interpret and understand the implications of the paper.

      Our observations in the current study clearly indicates concurrent face-selective processing in the VOTC, which is incompatible with a serial hierarchical model. While we discuss how such concurrent activity could be implemented in the cortex (e.g., via direct input from ‘early visual cortex’ to different VOTC face-selective clusters), our data do not allow to provide more evidence in that respect to what already exists in the literature. Moreover, we are not providing data regarding connectivity (feedforward or re-entrant) either between face-selective regions or between these regions and ‘early visual cortex’.

      Author response image 4 shows a very simplified versions of standard hierarchical/serial versus concurrent/parallel models.

      Author response image 4.

      Reviewer #2 (Public review):

      Summary:

      This very ambitious project addresses one of the core questions in visual processing related to the underlying anatomical and functional architecture. Using a large sample of rare and high-quality EEG recordings in humans, the authors assess whether face-selectivity is organised along a posterior-anterior gradient, with selectivity and timing increasing from posterior to anterior regions. The evidence suggests that it is the case for selectivity, but the data are more mixed about the temporal organisation, which the authors use to conclude that the classic temporal hierarchy described in textbooks might be questioned, at least when it comes to face processing.

      Strengths:

      A huge amount of work went into collecting this highly valuable dataset of rare intracranial EEG recordings in humans. The data alone are valuable, assuming they are shared in an easily accessible and documented format. Currently, the OSF repository linked in the article is empty, so no assessment of the data can be made. The topic is important, and a key question in the field is addressed. The EEG methodology is strong, relying on a well-established and high SNR SSVEP method. The method is particularly well-suited to clinical populations, leading to interpretable data in a few minutes of recordings. The authors have attempted to quantify the data in many different ways and provided various estimates of selectivity and timing, with matching measures of uncertainty. Non-parametric confidence intervals and comparisons are provided. Collectively, the various analyses and rich illustrations provide superficially convincing evidence in favour of the conclusions.

      We thank the reviewer for their positive comments on our manuscript.

      Weaknesses:

      (1) The work was not pre-registered, and there is no sample size justification, whether for participants or trials/sequences. So a statistical reviewer should assess the sensitivity of the analyses to different approaches.

      Pre-registration of fundamental research in a clinical context is quite uncommon for intracranial data, owing, for instance to the time needed to accumulate data, or to the uncertainty of cortical sampling location in a given participant. Nevertheless, in the current study, sample size is much higher than in typical intracranial studies (usually 5-20 participants), in fact much higher than most typical Cognitive Neuroscience research. The same is true for the number of recording contacts (>10000 site here), and the number of trials considered for analysis. Each participant had a minimum of 164 face trials and an average of 262 trials (i.e. an average of 3.2 stimulation sequences of 82 trials), which is higher than most standard human electrophysiological studies.

      In the revised manuscript, to unsure that we have the maximum available power, and because our hypothesis is independent of hemisphere, we collapsed data across hemispheres for all analyses. We nevertheless provide analyses split by hemispheres as supplementary material.

      In addition, since onset latency estimations depend on the methods used, we now report onset latencies from 4 different methods (2 statistical and 2 non-statistical).

      (2) Frequentist NHST is used to claim lack of effects, which is inappropriate, see for instance:

      Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. https://doi.org/10.1007/s10654-016-0149-3

      Rouder, J. N., Morey, R. D., Verhagen, J., Province, J. M., & Wagenmakers, E.-J. (2016). Is There a Free Lunch in Inference? Topics in Cognitive Science, 8(3), 520-547. https://doi.org/10.1111/tops.12214

      Please see reply to the next comment (3).

      (3) In the frequentist realm, demonstrating similar effects between groups requires equivalence testing, with bounds (minimum effect sizes of interest) that should be pre-registered:

      Campbell, H., & Gustafson, P. (2024). The Bayes factor, HDI-ROPE, and frequentist equivalence tests can all be reverse engineered-Almost exactly-From one another: Reply to Linde et al. (2021). Psychological Methods, 29(3), 613-623. https://doi.org/10.1037/met0000507

      Riesthuis, P. (2024). Simulation-Based Power Analyses for the Smallest Effect Size of Interest: A Confidence-Interval Approach for Minimum-Effect and Equivalence Testing. Advances in Methods and Practices in Psychological Science, 7(2), 25152459241240722. https://doi.org/10.1177/25152459241240722

      We thank the reviewer for pointing this out. In the revised manuscript we conduct and report a thorough examination of equivalence using ROPE and Bayesian factor to support the lack of differences between regions. We did not use TOST procedures as these have low power and require huge samples be meaningful (Riesthuis, 2024). Instead we relied on Bayes factor and descriptive proportion in ROPE. Equivalence bounds and region of practical equivalence (ROPE) were defined to account for physiological variability corresponding to a small effect (Cohen’s d = 0.199, i.e. standard in equivalence testing) and axonal conduction delays between regions (i.e. ATL is further away from ‘early visual cortex’ than OCC). These analyses, now reported in the result section of the revised manuscript, along with effect sizes, confirm the hypothesis of concurrent onset latencies across VOTC.

      Riesthuis P. 2024. Simulation-Based Power Analyses for the Smallest Effect Size of Interest: A Confidence-Interval Approach for Minimum-Effect and Equivalence Testing. Advances in Methods and Practices in Psychological Science 7.

      Detailed methods are reported as well:

      “In addition, to statistically assert whether onset latencies measured across main VOTC regions (OCC, PTL, ATL) were consistent with a concurrent (parallel) face-selective activation, we used Bayesian equivalence testing, relying on two separate metrics: (1) the percentage of differences in region of practical equivalence (ROPE), and (2) the Bayes factor using Cauchy prior. Equivalence bounds for ROPE were defined by combining two components: (1) a component of physiological variability and (2) a component reflecting expected delays in response onset between regions attributed to neural conduction delay, given the differential distances separating early visual cortex (EVC) from posterior face-selective regions (e.g. IOG) vs. anterior regions (ATL) and assuming signal mainly travels between VOTC regions through major postero-anterior axis fiber bundles of the Inferior longitudinal fasciculus (ILF) or the inferior fronto-occipital fasciculus (IFOF). Physiological variability corresponded to expected measurement noise and between-subject variability. Physiological equivalence bound was obtained by multiplying a Cohen’s d of 0.2 (conventional threshold for a negligible effect) with the (pooled) between-subjects variability in onset latency (computed for each region using a jackknife procedure and a correction factor of [N-participants – 1] to the jackknife standard deviation). For conduction delay bounds, the expected latency difference between regions under parallel activation depends on: (1) the distance between each region (considering the estimated origin of the signal is the same for all regions - EVC), and (2) neural conduction velocity. For inter-region distance we determined, for each region, the 5 and 95 percentile of the Talairach y-coordinate distribution and defined the maximal distance bounds as the distance between the y-coordinate corresponding to 5% of region 1 (e.g. OCC) to the coordinate corresponding 95% of region 2 (e.g. PTL). This resulted in the following maximum distance values: OCC-PTL: [54]mm; PTL-ATL: [54]mm; OCC-ATL: [83] mm. For conduction velocity, we used a constant value of 3.5 m/s, based on median axonal conduction velocity for cortico-cortical connections (Lemarachal et al., 2022; Van Blooijs et al., 2023), which was more conservative than using a range of values (e.g. 1.7 to 5.3 m/s based on Lemarachal et al., 2022). Maximum expected conduction delay was computed as [maximum distance / conduction speed] (e.g. for OCC to PTL: 0.054 / 3.5 = 15 ms). Resulting equivalence bounds (ROPE) were asymmetrical given than one region (e.g. OCC) is always closer to the source (EVC) than the other region (e.g. PTL) and were defined as [-1*physiologial_bound +1*physiologial_bound+max_conduction_delay]. For instance, using the z-score method to measured onset latencies, ROPE was [-12 to 27] ms for OCC to PTL, meaning that under equivalence, OCC can be activated up to 27 ms earlier than PTL (maximum conduction delay + noise), while allowing for some instances where OCC activates later (up to 12ms, due to noise only).

      For each pair of region compared, we used a bootstrap procedure to (1) define the percentage of differences between regions that fall within the ROPE, (2) compute the bayes factor using a Cauchy distribution (scale = 0.5) to estimate the proportion of the prior distribution in ROPE, and the bootstrap distribution to estimate the proportion of the posterior distribution in ROPE. The bootstrap distribution was obtained using a hierarchical stratified bootstrap procedure that naturally respects the nested structure of the data, that accommodates for unequal numbers of participants, contacts, and trials across regions, as well as partially overlapping participants samples across regions.

      For each bootstrap iteration, with first sample participants with replacement within each stratum (i.e. paired vs unpaired participants samples). For paired participants, sampling was performed jointly across regions to preserve the dependency structure, whereas unpaired participants were sampled independently within each region. Within each sampled participant, contacts were then resampled with replacement, and within each contact, trials were resampled with replacement. For paired participants, trial resampling was performed using identical trials across sampled contacts with a participant to preserve trial-level covariance. Resampled trials were averaged at the contact level, contact-level signals were averaged within participant, and participant-level signals were averaged to obtain a region-level response. Onset latency was then estimated from this averaged signal for each region using one of the 4 methods defined above (‘HFB response timing parameters). This procedure was repeated across 2000 bootstrap iterations to obtain a distribution of latency estimates for each region that respects the structure of data at iteration-level. Latency differences between regions were computed at each iteration, yielding a bootstrap distribution of differences which was used to compute percentage of differences in ROPE and posterior distribution for the Bayes factor.”

      (4) The lack of consideration for sample sizes, the lack of pre-registration, and the lack of a method to support the null (a cornerstone of this project to demonstrate equivalence onsets between areas), suggest that the work is exploratory. This is a strength: we need rich datasets to explore, test tools and generate new hypotheses. I strongly recommend embracing the exploration philosophy, and removing all inferential statistics: instead, provide even more detailed graphical representations (include onset distributions) and share the data immediately with all the pre-processing and analysis code.

      Data will be shared upon publication of the manuscript (see OSF repository in https://osf.io/2qzym). While we agree the dataset is large and could be explored in many ways, we do not consider the current study to be exploratory in nature. While our measurements could have turned out to clearly support hierarchical processing in human VOTC, our point in this manuscript is that the evidence derived from this large dataset unequivocally points instead toward concurrent activation of face-selective regions along the VOTC from IOG to antFG+ (i.e. a ~90 mm portion of cortex), with potential small variability accounted for by variability in axonal conduction velocity, signal-to-noise ratio or simple physiological variability. Other likely sources of variability such as type and density/size of fiber bundles across regions cannot easily be modeled with the current data set.

      (5) Even if the work was pre-registered, it would be very difficult to calculate p-values conditional on all the uncertainty around the number of participants, the number of contacts and the number of trials, as they are random variables, and sampling distributions of key inferences should be integrated over these unknown sources of variability. The difficulty of calculating/interpreting p-values that are conditional on so many pre-processing stages and sources of uncertainty is traditionally swept under the rug, but nevertheless well documented:

      Kruschke, J.K. (2013) Bayesian estimation supersedes the t test. J Exp Psychol Gen, 142, 573-603. https://pubmed.ncbi.nlm.nih.gov/22774788/

      Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p values. Psychonomic Bulletin & Review, 14(5), 779-804. https://doi.org/10.3758/BF03194105 https://link.springer.com/article/10.3758/BF03194105

      All analyses and preprocessing stages are identical between regions and the number of trials is large enough not to be a constraining factor. As indicated above and below, we now report detailed equivalence testing and effect sizes and recomputed all statistics, taking into account the structure of the data as suggested by the reviewer.

      (6) Currently, there is no convincing evidence in the article to clearly support the main claims.

      Bootstrap confidence intervals were used to provide measures of uncertainty. However, the bootstrapping did not take the structure of the data into account, collapsing across important dependencies in that nested structure: participants > hemispheres > contacts > conditions > trials.

      Ignoring data dependencies and the uncertainty from trials could lead to a distorted CI. Sampling contacts with replacement is inappropriate because it breaks the structure of the data, mixing degrees of freedom across different levels of analysis. The key rule of the bootstrap is to follow the data acquisition process, and therefore, sampling participants with replacement should come first. In a hierarchical bootstrap, the process can be repeated at nested levels, so that for each resampled participant, then contacts are resampled (if treated as a random variable), then trials/sequences are resampled, keeping paired measurements together (hemispheres, and typically contacts in a standard EEG experiment with fixed montage). The same hierarchical resampling should be applied to all measurements and inferences to capture all sources of variability. Selectivity and timing should be quantified at each contact after resampling of trials/sequences before integrating across hemispheres and participants using appropriate and justified summary measures.

      The authors already recognise part of the problem, as they provide within-participant analyses. This is a very good step, inasmuch as it addresses the issue of mixing-up degrees of freedom across levels, but unfortunately these analyses are plagued with small sample sizes, making claims about the lack of differences even more problematic--classic lack of evidence == evidence of absence fallacy. In addition, there seem to be discrepancies between the mean and CI in some cases: 15 [-20, 20]; 8 [-24, 24].

      In light of the reviewer’s comment, we recomputed all timing analyses using a stratified hierarchical approach to evaluate confidence intervals (using bootstrapping), statistical comparisons (using permutation tests) and equivalence testing.

      This is what we wrote in the revised methods:

      “The first two timing parameters of face-selective response, onset and offset latencies, were quantified per main VOTC region using a hierarchical bootstrapping approach to respect the nested structure of the data (region > participants > contacts > trials). For each bootstrap iteration and each region, we first sampled participants with replacement. Within each sampled participant, we then sampled contacts and then trials within sampled contacts, with replacement. Resampled trials, then contacts within participants, then participants within a region, were successively averaged to obtain a bootstrapped region-level response from which we derived onset latency (4 different methods) and offset latency. We obtained bootstrap distributions of onsets/offsets using 2000 bootstrap iterations per region, allowing to compute the median and 95% confidence interval for these 2 parameters.”

      Then, later about permutation tests:

      “Statistical significance of latency differences between main VOTC regions was assessed using a hierarchical permutation test. We use a stratification approach to partition participants into a paired set (i.e. participants that had recording contacts in the two regions compared) and unpaired set (participants with contacts in a single region). For paired participants, the region labels were randomly swapped within subject (i.e., exchanging the signals from the two regions), thereby preserving all participant-, contact-, and trial-level structure while breaking the association between region and latency estimates. For unpaired participants, participants were randomly reassigned between regions while preserving the original group sizes, to generate pseudo-groups under the null hypothesis of no regional difference. In each permutation, signals were averaged across trials, then contacts, then participants within each permuted group and latency was computed and stored from the resulting region-level signals. We performed 10000 permutations to obtain a distribution of regional differences of latencies under the null hypothesis and determine the p-value as the fraction of the null distribution larger or smaller than the observed (non-permuted) difference.”

      And then about equivalence testing:

      “For each pair of region compared, we used a bootstrap procedure to (1) define the percentage of differences between regions that fall within the ROPE, (2) compute the bayes factor using a Cauchy distribution (scale = 0.5) to estimate the proportion of the prior distribution in ROPE, and the bootstrap distribution to estimate the proportion of the posterior distribution in ROPE. The bootstrap distribution was obtained using a hierarchical stratified bootstrap procedure that naturally respects the nested structure of the data, that accommodates for unequal numbers of participants, contacts, and trials across regions, as well as partially overlapping participants samples across regions.

      For each bootstrap iteration, with first sample participants with replacement within each stratum (i.e. paired vs unpaired participants samples). For paired participants, sampling was performed jointly across regions to preserve the dependency structure, whereas unpaired participants were sampled independently within each region. Within each sampled participant, contacts were then resampled with replacement, and within each contact, trials were resampled with replacement. For paired participants, trial resampling was performed using identical trials across sampled contacts with a participant to preserve trial-level covariance. Resampled trials were averaged at the contact level, contact-level signals were averaged within participant, and participant-level signals were averaged to obtain a region-level response. Onset latency was then estimated from this averaged signal for each region using one of the 4 methods defined above (‘HFB response timing parameters). This procedure was repeated across 2000 bootstrap iterations to obtain a distribution of latency estimates for each region that respects the structure of data at iteration-level. Latency differences between regions were computed at each iteration, yielding a bootstrap distribution of differences which was used to compute percentage of differences in ROPE and posterior distribution for the Bayes factor.”

      (7) Three other issues related to onsets:

      (a) FDR correction typically doesn't allow localisation claims, similarly to cluster inferences: Winkler, A. M., Taylor, P. A., Nichols, T. E., & Rorden, C. (2024). False Discovery Rate and Localizing Power (No. arXiv:2401.03554). arXiv. https://doi.org/10.48550/arXiv.2401.03554

      Rousselet, G. A. (2025). Using cluster-based permutation tests to estimate MEG/EEG onsets: How bad is it? European Journal of Neuroscience, 61(1), e16618. https://doi.org/10.1111/ejn.16618

      In fairness, we do not understand or share the reviewers’ concern here. Hundreds of fMRI or EEG studies use FDR or cluster tests to make inference about spatial or temporal location. We use FDR correction in one of the onset latency estimation method and only consider one-sided differences. Other methods in the revised manuscript do not use FDR correction.

      (b) Percentile bootstrap confidence intervals are inaccurate when applied to means. Alternatively, use a bootstrap-t method, or use the pb in conjunction with a robust measure of central tendency, such as a trimmed mean.

      Rousselet, G. A., Pernet, C. R., & Wilcox, R. R. (2021). The Percentile Bootstrap: A Primer With Step-by-Step Instructions in R. Advances in Methods and Practices in Psychological Science, 4(1), 2515245920911881.

      Again, we are not sure what the reviewer’s is referring to. The confidence intervals are computed on latency estimates from bootstrapped waveforms. In the revised manuscript, these waveforms are obtained by averaging (i.e. mean) resampled trials, resampled channels, resampled participants. A trimmed mean could not be applied in this condition, except perhaps when averaging across trials. But then the trimmed mean would have to be applied separately at each time sample which would disturbed within-, or between-trial, variability.

      (c) Defining onsets based on an arbitrary "at least 30 ms" rule is not recommended:

      Piai, V., Dahlslätt, K., & Maris, E. (2015). Statistically comparing EEG/MEG waveforms through successive significant univariate tests: How bad can it be? Psychophysiology, 52(3), 440-443. https://doi.org/10.1111/psyp.12335

      The rule of contiguous significant points is a heuristic that many researchers have used successfully to avoid spurious detection due to temporal autocorrelation. While we are aware that more sophisticated methods exist to correct for autocorrelation, such as cluster-based approaches, it is not directly usable since it requires comparing 2 conditions. The approach described in Piai et al., 2015 is interesting but incorrect as well since it relies on split-half simulations, which reduced signal-to-noise ratio, resulting in over estimated correction to be applied. In our revised manuscript, we rely on multiple methods to estimate onset latency, some of which not relying on this heuristic. Moreover, we apply plausible physiological constrains to our latency estimates, such as rejecting any onset before 40 ms after stimulus onset.

      (8) Figure 5 and matching analyses: There are much better tools than correlations to estimate connectivity and directionality. See for instance:

      Ince, R. A. A., Giordano, B. L., Kayser, C., Rousselet, G. A., Gross, J., & Schyns, P. G. (2017). A statistical framework for neuroimaging data analysis based on mutual information estimated via a Gaussian copula. Human Brain Mapping, 38(3), 1541-1573. https://doi.org/10.1002/hbm.23471

      (9) Pearson correlation is sensitive to other features of the data than an association, and is maximally sensitive to linear associations. Interpretation is difficult without seeing matching scatterplots and getting confirmation from alternative robust methods.

      We rely on Pearson correlation because this replicates the method used in Kadipasaoglu et al., 2017. It is also a widely accepted measure of (linear) relationship (in our situation we did expect linear or near linear relationships) in the literature. To address the reviewers concern, in the revised manuscript we nevertheless report, as supplementary material (Figure S10), the same functional connectivity analyses performed using the methodology and code provided in Ince et al. (2017). The results of this analyses are extremely similar to the results using Pearson’s coefficients.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 6, the response onset latencies are rendered in a smoothed manner on the brain surface. However, with this smoothing, variability between electrodes cannot be seen, and it would be better visualized in color, rendered in each electrode.

      Latency estimates computed at individual channels are noisy, which is why we do not report individual channels latencies but rather rely on averaging signals across contiguous channels, either across whole regions (Figure 4) or across smaller volumes as in Figure 6.

      (2) Onset latencies of 60 seconds seem extremely early compared to literature typically citing evoked responses with a latency of ~170ms. It would help if some additional sanity checks were shown, such as showing the latency of early visual responses. This would help with relative comparisons.

      In the revised manuscript the earliest median latency is 95 ms, which is in line with previous intracranial electrophysiology literature (e.g. Jacques et al., 2016; Jacques et al., 2022 ; https://pubmed.ncbi.nlm.nih.gov/26212070/; https://pubmed.ncbi.nlm.nih.gov/36074548/). The 99% confidence intervals can result in earlier latencies both due to some participants showing early responses and noise in latency estimates. Also please keep in mind that latency estimates are usually earlier when combining data across channels/participants compared to individual channels simply due to differences in SNR or across participants (see e.g. Kadipasaoglou et al., 2017).

      The reviewer indicates “…to literature typically citing evoked responses with a latency of ~170ms.”. We are assuming that they refer to the face-selective N170 ERP component measured on the scalp in EEG. Even with this ERP component, the face-selective response usually starts around 120-130 ms after stimulus onset (e.g. Rousselet et al., 2008; Jacques, Retter and Rossion, 2016; https://pubmed.ncbi.nlm.nih.gov/18831616/; https://pubmed.ncbi.nlm.nih.gov/27138205/) at scalp level. With the same highly sensitive paradigm as used here in EEG, we have systematically shown latency onsets of face-selective activity shortly after 100 ms (e.g., Retter et al., 2020; also Quek & Rossion, 2017) not accountable for by low-level visual cues (i.e., not present for phase-scrambled stimuli; Rossion et al., 2015; Or et al., 2019). Our latency onsets are also in line with spiking activity recorded with the same approach in the LatFG (Laurent et al., 2026) https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5955677

      (3) Figure 6B shows that the variability of latencies in the ATL is larger than the variability of latencies in the PTL. It would be helpful to evaluate whether, rather than in mean onset latencies, there is a change in variability in onset latency along the VOTC.

      This is an interesting point. However, it is difficult to evaluate since SNR is reduced in the ATL compared to OCC or PTL (Jacques et al., 2022). As a result, any measured modulations in the variability of onset latencies along VOTC may simply reflect changes in the precision of latency estimation driven by SNR variability.

      (4) Line 415 typo: 'there appears to be no delay' instead of 'there appear to be no delay'.

      We thank the reviewer for their careful reading of our manuscript. This has been corrected.

      (5) The discussion states in lines 455-457 that "a large proportion of neuronal populations in anterior VOTC regions exhibiting similar activity to different face images independently of the context in which they appear". However, this claim about similar activity to different faces should be evaluated and tested at the single-trial level. In addition, it is not clear how context was varied in the experimental design.

      Context is variable because each face image appears directly after a different object image (or object images) in the sequence. We have shown also in previous studies with this paradigm in EEG that the time-course of face-selective responses is similar across base frequencies (3-15 Hz) unless the rate is too fast, and whether an orthogonal or explicit face categorization task is used (Retter et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32119982/). Note that we do not claim that activity is identical across images but similar – if it was not (largely) similar, the averaged response would be jittered and low.

      (6) Line 516-517: DTI does not provide evidence for whether connectivity is direct or not, and what the directionality of connectivity is between two areas. This sentence should therefore state "..., suggest independent connections between early visual cortex and face-selective regions...".

      This has been rephrased.

      (7) Line 555 in the discussion, the definition of low-level visual should be expanded to include other early visual areas that have been demonstrated to respond earlier than VOTC (e.g. Martin et al., 2019, JNeurosci https://doi.org/10.1523/JNEUROSCI.1889-18.2018), to avoid the suggestion that V1 directly projects synaptically to all of VOTC (e.g. Markov et al., 2014, Cerebral Cortex, https://doi.org/10.1093/cercor/bhs270).

      We are not proposing that V1 directly projects directly/synaptically to all of VOTC, i.e., without other low-level retinotoptic areas involved; only that face-selectivity in the association cortex is not organized hierarchically. We have revised this sentence.

      (8) Line 585, for the sentence: "with temporal synchrony strengthening their connections", evidence or citations should be provided.

      Citations have been provided.

      (9) It is not clear what is meant in the paragraph starting in line 571: do the authors suggest that top-down signals are not necessary for fast recognition of clear views of faces, or additionally argue that these top-down signals are not necessary for detecting ambiguous or degraded inputs as faces?

      Exactly: That top-down (i.e., descending) signals may contribute but would not be necessary for fast recognition of clear views of faces AND for detecting ambiguous or degraded inputs as faces.

      (10) No statement was provided on data or code availability.

      Data will be made available on a repository upon publication (see https://osf.io/2qzym).

      Reviewer #2 (Recommendations for the authors):

      (1) FDR correction: which one? Please provide a reference.

      We now provide a reference, both in the results and methods: Benjamini and Hochberg, 1995.

      (2) In the introduction, this statement is too strong: "arguably the most familiar and ecologically valid stimulus". It is unclear how static 2D representations of faces are the most familiar and valid stimuli. Could you rephrase this? What about other very familiar stimuli like letters, words and biological motion?

      This statement is not about static 2D images of faces, but faces in general (in their natural environment). We do consider human faces (in general, not restricted to laboratory context) to be indeed the most familiar and ecologically important stimulus, both from an ontogenetic and phylogenetic perspective, unlike written material.

      (3) About the questioning of a strict temporal hierarchy, this EEG reference comes to mind: Foxe, J. J., & Simpson, G. V. (2002). Flow of activation from V1 to the frontal cortex in humans. Experimental Brain Research, 142(1), 139-150. https://doi.org/10.1007/s00221-001-0906-7

      As confirmed by Foxe et al. ’s (2002) paper to which the reviewer is referring to, there is indeed ample evidence that areas in the dorsal stream or frontal cortex (e.g. FEF) are activated very soon after V1 and before many ventral stream regions (e.g. Lamme and Roelfsema, 2000; https://pubmed.ncbi.nlm.nih.gov/11074267/). While Foxe et al.’s 2002 is highly valuable, it can hardly be compared with our current study which looks specifically into ventral stream areas which are largely indistinguishable using scalp EEG as in Foxe et al. ’s paper.

      (4) Regarding statistical significance, there is no such thing as a "trend". The threshold for a trend should have been pre-registered and applied to both sides of the magical boundary, for instance, with matching conclusions for a "trend toward non-significance (p=0.04)". P values near 0.05 provide weak support against the null. I would suggest leaving it at that. Nothing special happens at 0.05.

      This no longer appears in the revised manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides evidence for our understanding of HIV transmission dynamics by age and sex in Zambia during the PopART trial; by combining phylogenetic and individual-based mathematical modelling (IBM), it adds depth to the epidemiological literature and may inform more strategic allocation of HIV prevention resources in sub-Saharan Africa. The authors employ two complementary and well-established methodologies (phylogenetics and IBM), and this dual approach is a notable strength. However, the evidence supporting key conclusions is incomplete, with several claims insufficiently substantiated by the data presented. Improvements in data presentation (e.g., quantification of qualitative statements, statistical estimates, and clearer description of results) would substantially strengthen the paper.

      We thank the editor and reviewers for their positive comments. We have revised the manuscript in response to the points raised, as described below.

      First of all, we would like to summarise what we have changed regarding the presentation of summary statistics throughout the text. We agree that many of the statements in the original submission tended towards being qualitative. This was the result of shying away from presenting two separate estimates, with different ways of quantifying uncertainty, in the text. The phylogenetics could be presented as mean and confidence interval, while the IBM would need some measure of centrality (mean or median) and the highest density interval for a summary statistic (e.g. the mean age gap) as it varies over the posterior. These are not directly comparable. We have now changed this to present both where appropriate, with cautionary note about the difference between the CIs and HDIs (lines 257-260).

      We also were somewhat arbitrary regarding where we chose to summarise the posterior in the IBM or look at the best-fitting single simulation, and where we presented the mean as opposed to the median. We have done a considerable overhaul of what is presented in this revision:

      (1) We always present the posterior summary unless the level of detail is such that summarising uncertainty over the posterior is not feasible (e.g. in figures 3, 4 and 5). In the latter case we still use the best-fitting IBM replicate.

      (2) In the main text we always present the mean. For the phylogenetics the summary statistics are mean and confidence interval. For the IBM this is the posterior mean, and 95% HDI, of the mean of a particular statistic as calculated in each of the 1000 IBM replicates. For example, each replicate will have its own distribution of male source ages which have a mean value. These means also vary over the posterior, and a mean of them is calculated, as well as the HDI interval to represent posterior uncertainty. This “mean of means” may be a slightly confusing piece of terminology at first glance, but it allows us to properly capture posterior uncertainty in a way we mostly avoided in the first submission.

      One result of 1) above is a change to figure 6. It is now summarised over the posterior, with the result that time trends that were previously not evident become clear. This changes our conclusions slightly (lines 529-537) but it should be noted that the magnitudes of the trends remain small.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of phylogenetic and epidemiological modeling of the PopART community cohorts in Zambia. The current manuscript draft is methodologically strong, but needs revision to strengthen the take-home messages. As written, there are many possible take-away conclusions. For example, the agreement between IBM and phylogenetic analysis is noteworthy and provides a methodological focus. The revealed age patterns of transmission could be a focus. The effects of the PopART intervention and the consequences of a 1-year disruption could be a focus. It is important, though, that any main messages summarized by the authors are substantiated by the evidence provided and do not extrapolate beyond the data that have been generated. I recommend that the authors think deeply about what the most important, well-supported messages are and reframe the discussion and abstract accordingly.

      We have rewritten the abstract, and also made changes to the discussion in order to centre our message around the contribution of particular of demographic groups to transmission, and how, with that contribution revealed, such groups can be selected for specialised interventions.

      Strengths/weaknesses by section:

      (1) ABSTRACT

      The Abstract summarizes qualitative findings nicely, but the authors should incorporate quantitative results for all of the qualitative findings statements.

      The abstract in the revision is extensively revised, and contains quantitative estimates throughout, from both methodologies where appropriate.

      The ending claim is not substantiated by the modeling scenarios that have been run: "targeted interventions for demographic groups such as under-35 men may be the key to finally ending HIV." It is straightforward to run this specific scenario in the model to determine whether or not this is true.

      Our modelling framework is not set up to model the “last mile” of HIV elimination, notably as it has no component for MSM or FSW transmission, and we do not feel that we could confidently present results regarding it. As a result, this statement has been greatly softened in the new abstract (lines 75-78).

      The authors should add confidence intervals to the quantitative metrics, such as the 93.8% and 62.1% incidence reduction.

      These have been added.

      (2) RESULTS

      The authors should check the Results section for any qualitative claims not substantiated by the analyses performed, and ensure the corresponding analyses are presented to support the claims.

      The Results and Methods describe the model's implementation of the PopART intervention differently. The Methods describes it as including VMMC, TB, and STI services, while the Results only mentions intensified HIV testing and linkage.

      This is a slight misreading of the text. That paragraph in the Methods is describing the trial itself, not the modelling framework.

      A limitation of the model is that HIV disease progression is based on the ATHENA cohort in the Netherlands, which is a different HIV subtype (B) than the one in the research setting (C). The model should be configured using subtype C progression data, which have been published, or at least a sensitivity analysis should be conducted with respect to disease progression assumptions.

      The available literature does not suggest a significant difference in progression between subtypes B and C, and we have added text and citations to this effect (lines 699-701).

      In Table 2, the authors should consider adding a p-value to establish whether or not IBM and phylogenetics estimates are different.

      We have done this; the appropriate test was a posterior predictive check. See lines 261-263, 575-579 and 805-814.

      (3) DISCUSSION

      The literature review and comparison of study results to previously published phylogenetic studies is very nice. The authors could strengthen this by providing quantitative estimates with CIs for a more scientific comparison of the study results vs. prior studies, perhaps as a table or figure.

      We have expanded the discussion on this point (lines 504-527). We considered adding a table, but the existing literature that directly answers the questions we ask is quite limited and fragmentary. For example, Monod et al do not present a complete treatment of age gaps. The literature using regression analyses to identify predictors of HIV prevalence or incidence related to partner age is extensive, but those results are not directly comparable to ours.

      The authors state that due to "the narrow geographical catchment area... The results should not be automatically extrapolated to apply to other SSA settings." The authors should exercise this caution when comparing the results to studies in South Africa and elsewhere.

      We have made more explicit acknowledgements of these limitations (lines 598-600).

      There are many other limitations to the analysis, including some mentioned above, that are not acknowledged. The authors should think carefully about what the most important limitations are and acknowledge them honestly at the end of the Discussion section.

      The limitations paragraph has been revised (lines 598-605).

      Reviewer #2 (Public review):

      Summary:

      The authors analyzed PopART data to better characterize the age and sex-specific heterosexual HIV transmission dynamics in Zambia, with the goal of allocating resources.

      Strengths:

      Important analysis to hone in on the key driver of HIV transmission in Zambia, which hopefully can be used to tune prevention efforts to maximize effect while limiting required resources. Two analytic approaches were used, and while the phylogenetic data were markedly more limited, they mirrored the simulated epidemic. The authors did a nice job reviewing the limitations of the data and the analyses. The authors did a nice job of providing analyses to support their goals and hypothesis, and this work may have more impact now that resources in SSA for HIV prevention and treatment may become more scarce

      Weaknesses:

      To increase the impact and utility of this work, it would be helpful to parse the analysis just a bit further to estimate the roles of undiagnosed vs diagnosed and untreated subpopulations on this transmission. PopART is a multifaceted intervention, but the cost, effort, and approach to reengagement in care vs testing/treatment can be quite different.

      We have now provided stratified results by diagnosed and non-diagnosed status of the source, as well as an overall summary of the proportion of undiagnosed sources by age and sex. See lines 305-310, 539-547, and table 3.

      Recommendations for the authors:

      Reviewing Editor:

      We commend you for conducting a rigorous and comprehensive study titled "The age and sex dynamics of heterosexual HIV transmission in Zambia: an HPTN 071 (PopART) phylogenetic and modelling study" that significantly advances the understanding of HIV transmission dynamics in sub-Saharan Africa. The study utilizes an innovative dual-methodology approach integrating individual-based mathematical modelling (IBM) and pathogen phylogenetics to characterize heterosexual HIV transmission patterns by age and sex during the PopART trial in Zambia.

      This manuscript reports on HIV transmission dynamics in Zambia using data from the PopART study, combining individual-based modelling and phylogenetic analysis. The use of two independent methodologies enhances confidence in the consistency of the findings and enables robust cross-validation. The work addresses an important topic in HIV prevention, particularly in settings where resources may become more constrained, and offers insight into potential demographic targets for intervention.

      However, several aspects of the manuscript limit its current impact. The main take-home messages are diffuse and not clearly presented. Some conclusions in the abstract and discussion appear to go beyond the scope of the presented data. For instance, the claim that targeting under-35 men may be key to ending HIV is not directly tested in the modelling scenarios and should be reframed or removed unless supported by new analyses. Furthermore, important quantitative details, such as confidence intervals, p-values, and precise age group estimates, are lacking in key sections (e.g., the Abstract and Results).

      The authors are encouraged to clearly identify and communicate their central findings, ensure all claims are fully supported by their analyses, and make the data more accessible to readers by adding detailed, quantitative summaries where needed.

      The following are our recommendations to the Authors:

      (1) Clarify Study Objectives and Central Messages

      Reframe the abstract and discussion to highlight a clear, well-supported set of main findings.

      Avoid overgeneralized or unsubstantiated claims, especially those not directly tested by your model (e.g., the effectiveness of targeting under-35 men).

      As stated above, we have revised this text accordingly.

      (2) Support Qualitative Claims with Quantitative Data

      Provide numerical results, including effect sizes and confidence intervals, wherever qualitative trends are mentioned.

      For example, restate: "The largest gaps for female recipients were among the youngest" as "... in the age group XX-YY with OR = Z.Z (95% CI: A.A-B. B)."

      As mentioned at the top of the review, we have overhauled the treatment of summary statistics extensively, and now give confidence or highest density intervals throughout the text.

      (3) Improve the Results Section

      Check that all claims are supported by the analyses, and ensure figure references are accurate.

      The statements that went beyond what was supported, notably about ending the epidemic by targeting young men, have been removed. The typo in table references has been fixed.

      Annotate Figure 6 with trendline coefficients and p-values where applicable.

      The takeaway message of figure 6 has now changed and we no longer see no trend, just a minor one.

      Revise Figure 4 for clarity or consider replacing it with a tabular format.

      We would prefer to keep the current figure 4, as we have not found any clearer way to illustrate the patterns, which are the consequence of the phenomenon observed in figure 5. We have put more explicit descriptive text in the discussion, linking the two figures (lines 470-476).

      (4) Address Potential Bias and Model Assumptions More Rigorously

      Explain sampling bias in IBM and phylogenetics (e.g., how the 355 high-confidence phylogenetic pairs were selected).

      The reviewer comment regarding the 355 pairs was based on a misapprehension; we used all the pairs we found using the phyloscanner pipeline. There are no sampling bias issues involved in the IBM as every individual in the simulations is considered. Appendix 2 includes some sensitivity analysis results if the procedure used to find the 355 is changed.

      Discuss how the use of subtype B disease progression data from the ATHENA cohort may impact results in a subtype C setting. A sensitivity analysis would strengthen this.

      Subtype B progression data was used in the absence of any appropriate data from subtype C, but the literature does not suggest any major difference between the two (lines 699-701).

      (5) Include More Detail on Undiagnosed Populations and ART Effects

      Estimate the roles of undiagnosed and untreated subpopulations in driving transmission.

      As mentioned above, this analysis has been added.

      Clarify mechanistically how ART might influence age gaps in transmission dynamics.

      This now is clarified in the introduction (lines 127-129).

      (6) General Improvements

      Provide p-values where comparisons are made (e.g., in Table 2).

      Use consistent terminology and definitions across Methods and Results.

      Add more discussion on limitations, especially regarding generalizability to other SSA settings.

      All of these have been inserted as previously mentioned.

      By addressing these points, the manuscript would present a more coherent narrative and a stronger, evidence-based contribution to the field. We appreciate you all for your fantastic effort and hope you will reflect the feedback in your final paper.

      Reviewer #1 (Recommendations for the authors):

      Thank you for the opportunity to review this interesting manuscript.

      In the public review, I have recommended that the authors should incorporate quantitative results for all of the qualitative findings statements. As one example, I would recommend that "We found the largest gaps for female recipients were among the youngest of those recipients" is re-written as "The largest gaps for female recipients were in the age group XXX-YYY with OR=ZZZ (XXX-YYY)." such as odds ratios, and specific outcome definitions including ages. To give one more example: "immediate increase in the average age at transmission of both sources and recipients" could be rephrased as "increase in the average age at transmission by XXX (YYY-ZZZ) years for sources and XXX (YYY-ZZZ) for recipients over [TIME PERIOD]."

      We hope the revisions we have made to the statistical presentation are satisfactory as a response to this request.

      Again in the public review, I recommended checking the Results section for any qualitative claims not substantiated by the analyses performed, and ensuring the corresponding analyses are presented to support the claims. An example is: "Trends are minor or non-existent in the former two variables." - please annotate Figure 6 (assuming the authors meant to reference Figure 6 and not 7 here?) to show over what period trendlines were fit and provide the coefficient and CI. To support the stated claim even more strongly, a p-value might be apt with a null hypothesis of a slope of zero.

      Please check the numbering on all figure references in the text, as some appear to be misnumbered. E.g., where the text refers to Figure 7, I believe the authors meant to reference Figure 6.

      The change to how we handled the statistics has changed the message of figure 6 (which is now figure 7) and rendered this somewhat moot. We have checked that all figure and table references are now correct.

      Figure 3 is very nice, but if the axes were flipped on one panel, it would make them easier to compare, and then adding some statistics to assess whether the patterns are the same or different when a man vs woman is the source.

      We have flipped the axes here.

      Figure 4 was too complicated for me. I could not follow the Sankey flows because there is too much going on and overlapping. Consider revising to make it easier to digest... perhaps to table format?

      As mentioned above, we would prefer to keep this figure, but we have situated it better in the text.

      Reviewer #2 (Recommendations for the authors):

      A few points that would improve the clarity and the strength of the manuscript

      (1) There is a need to clarify more about how the IBM and phylogenetic data does not suffer from sampling bias. For e.g.,

      Line 205: What proportion of the transmissions modeled in the IBM from Zambia?

      All of them. We confined the analysis of the IBM to the Zambian communities from which phylogenetic data was acquired (lines 755-758).

      Line 217: What proportion of the phylogenetic pairs (cherries) suggesting transmission were the 355 that had high confidence in directionality. How do these pairs compare to the others

      There was no identification of “cherries” involved in picking these pairs; the phyloscanner procedure does not use that step. We confined our analysis solely to the pairs for which we did identify a direction of transmission; that is the 355. Appendix 2 includes a sensitivity analysis involving varying the parameters by which these were identified.

      (2) I appreciate the authors noting that MSM transmissions are unlikely to be playing a role in this cohort, as noted in previous work by the group. However, systematic undersampling of men is common in other study cohorts of HIV. While the MSM and heterosexual networks may be relatively distinct, undersampled men who are bridging the networks could impact the estimates. Can the authors use the time to diagnosis analysis (HIV phyloTSI) to estimate rates of undiagnosed men and women?

      We feel that this is beyond the scope of this work. The phylogenetics dataset in its totality could be used for this purpose (although it is probably highly biased towards undiagnosed individuals due to the considerable majority of samples coming from the healthcare facilities). However, we concentrate here solely on the subset involved in our probable transmission pairs, which is fairly small. Extending the scope to an exploration of the full dataset would seem like a separate study, which we do have plans to do.

      We have used the IBM for this question instead (lines 303-321), however, as MSM transmission was not modelled, it is also not ideal for answering this question. Ultimately we feel that the way these studies were implemented makes it an unsatisfactory tool for answering the MSM question, important as it is.

      (3) Expanding on the point above, in other settings, transmission to young men has been associated with partnerships with older men, and if these young men then transmitted to young women, would we see a similar effect as noted in these models (assuming the young men were less well sampled).

      Our previous work (Hall et al., 2024) suggested no excess of identified male-male pairs in the phylogenetics dataset which might suggest cryptic male-to-male transmission. The age disparities would be worth exploring had this been found, but is curtailed by the lack of it.

      (4) Related to the point above, is there an estimate of the populations (age and sex) that are undiagnosed in the IBM model? Can this be teased out... is transmission from men to women more likely 2/2 lack of diagnosis... or lack of engagement in care?

      We have explored results by diagnostic status as it pertains to age and sex, but we feel that moving on to a more general exploration of the role of diagnosis and lack of engagement in care is again going beyond the scope of what is already a long paper.

      (5) I'm still not fully clear as to why ART might affect age gaps. Can this be explained in more detail?

      See lines 127-129.

    1. Author response:

      We are pleased that the reviewers viewed the core demonstration (that Dscam mutually exclusive splicing is preserved in a vector and that exon alternates can be replaced with genes of interest) as a solid foundation for the system. We agree that the manuscript would be strengthened by clearer quantitative characterization of expression, additional controls for fluorophore imaging, improved presentation of the figures, and more precise wording about the current scope of evidence. In a revised manuscript, we plan to address these points by adding or clarifying quantitative expression analyses, including S2 cell validation data, adding appropriate imaging controls where available, revising claims about 12-transgene expression to distinguish design capacity from direct experimental demonstration, and improving figure labels and legends throughout.

      We also plan to expand the discussion of PXGS limitations, including cell-type dependence on Dscam splicing machinery, possible position effects, and gene size considerations. Finally, we will improve Methods reporting by adding resource identifiers, cell culture quality control information, statistical design details, and data/code availability statements where appropriate.

      We appreciate the opportunity to revise the manuscript and believe these changes will make the strengths and limitations of PXGS clearer to readers.

    1. Author response:

      Reviewer #1 (Public review):

      Strengths:

      Overall, the computational model tested in the paper is novel and interesting.

      The demixing framework represents an appealing hypothesis that deserves further investigation.

      The current paper provides new empirical data showing that the target stimuli with the same absolute noise level can be either repelled from or attracted to non-target items, depending on the relative noise levels. The observation that biases depend on the relative noise levels is by itself an interesting one, and is consistent with the prediction of the demixing model.

      We are grateful for the positive evaluation of the model and the empirical observations.

      Weaknesses:

      While this manuscript contains interesting new experimental observations and theoretical ideas, it has several substantial problems in its current form, which limit the conclusions that can be drawn. The description of the computational model is too brief. The key modeling assumptions need to be better motivated and explained. As the computational models generate different predictions in different regimes, it is a bit difficult to evaluate how well the experimental data support the model at a more quantitative level. Also, the results focused on studying the biases in the behavior; it is unclear whether the model can fully explain the behavior data (such as error distributions or behavioral precision).

      We agree that the model description should be expanded and that quantitative agreement with the data should be assessed more thoroughly, and we plan to address this during the revision. In the initial version of the manuscript, we aimed to highlight the qualitative agreement of the data with the novel and counterintuitive predictions by the model. While the reviewer is correct that the model "generates different predictions in different regimes," the particular predictions we test (the interaction between the target noise level and the parity of the target and non-target noise levels in Experiments 1-3, and the effects of non-target noise when the target noise is held constant in Experiment 4) hold across regimes (Figure S1 shows this for the former prediction). We aim to further expand on this point in the revision.

      Major concerns:

      (1) Concerns/suggestions regarding the computational modeling

      The current paper seeks to test the predictions of the demixing-based computational model proposed in reference 22. There are several problems with the modeling component in the current paper.

      (1a) The description of the model is too brief and difficult to understand. Although the model was proposed in reference 22, it would still be beneficial to provide more details of the model so that readers can understand and appreciate the strengths/limitations of the model.

      The generative model and the inference procedure could be better explained to better link the model to the behavior. In particular, how was the observer's behavioral report in each trial modeled? This requires more explanation because currently the demixing procedure estimates four parameters for a given trial, yet for a given trial, only one behavioral report was produced (e.g., current Experiment 1), or two reports were produced sequentially (e.g., current Experiment 2).

      We will provide more details about the model and how it was fitted to the data. Please note that the model parameters were fitted per subject and condition, not per trial: 2 hue noise parameters,  and , corresponding to the noise of the target and non-target item across 4 noise combination conditions, plus a shared identifiability noise, , across conditions, determining the discriminability of the items along the identifying dimension. This strongly limits model flexibility as only 3 parameters (including the shared  across conditions) are used to create the bias curve for each subject in each condition.

      (1b) Key modeling assumptions need better justification.

      One such key assumption is that on a given trial, each stimulus triggers many samples (or approximately, an entire response distribution), rather than a single sample. This assumption deviates substantially from prior work on ideal observer models. It was not clear whether this assumption is realistic. For the type of stimuli used in the current experiments, perhaps one can argue that each pixel corresponds to one sample of brain activity, thus collectively each stimulus should trigger many samples of activity in the brain. If this were to be the case, it would have two implications. First, the noise parameter in the model should be directly related to the magnitude of the stimulus noise. Thus, one should be able to plug these experimentally-controlled parameter values into the model to directly generate predictions about the biases. Second, when using stimuli with no stimulus variability (e.g., simple grating stimuli), the predicted biases should change. However, it wasn't clear whether this would hold experimentally, i.e., using gratings would lead to different biases or no biases.

      If the variability of the samples for a given stimulus involves neural noise, it would be useful to justify why it is reasonable to consider that many samples were generated per stimulus.

      We are grateful to the reviewer for raising this point, and we will provide more details on it in the revision. In brief, we believe that it is the standard ideal observer assumption of one sample per trial that is unrealistic and works only in cases when there is a single signal source, so that the samples can be simplified to a single average. Consider that determining a stimulus value is a similar problem for an ideal observer to the one that a researcher who aims to decode neural data from populations of neurons (or fMRI voxels) has to solve. Different populations of neurons would provide responses that match different stimuli – in essence, creating different samples in an ideal observer framework. Thus, even without external noise, the demixing problem would be present when there is more than one stimulus, but internal noise is much more difficult to control, so in our experiments, we used multi-colored stimuli.

      (1c) As mentioned in (1b), the model assumes that on each trial, a large number of samples was generated. It would be useful to study and report how the prediction would change when the number of samples generated per stimulus is small. In particular, what happens when each stimulus only generates one measurement? This might be useful for interpreting previous experiment results with grating stimuli.

      This is an interesting point that we aim to address in the revision.

      (1d) Reference 22 studies how the predicted biases depend on the d-prime of the identifying dimension and found that the pattern of the biases varies substantially depending on the information available for the identifying dimension. However, the current paper didn't really discuss this important point. It is also unclear what parameters the authors used for the d-prime of the identifying dimension. Was it fitted directly to the data? The Methods section has some description on the "identifiability dimension", but it was a bit obscure.

      Intuitively, when the d-prime of the identifying dimension is very large, the demixing problem becomes irrelevant. In this case, there should not be any biases induced by demixing. In the case of the d-prime for the identifying dimension is 0, the problem should reduce to the simplified 1-d problem studied in reference 22. If my reading of reference 22 was correct, they reported different conclusions. It would be useful to clarify these points.

      We are grateful for the suggestion to expand the discussion of this point and will do so in the revision. The reviewer is correct that for very large d-prime in the identifying dimension, the demixing problem solution is trivial. However, the 2D case does not resolve to the 1D case when d-prime reaches zero. This is because the identifying dimension is still used to identify which item to report—unlike in the 1D case, when the reported dimension is the same as the identifying one. Consider what happens if the observer in our task does not remember at all which stimulus was left and which was right. It would report the other item in 50% of cases, leading to a strong attractive bias.

      In any case, the d-prime of the identifying dimension appears to be a key parameter. It would be great to constrain this parameter using the empirical data. When the d-prime of the identifying parameter is small, the observer would easily confuse the probed stimulus with the other stimulus in a given trial. This should lead to poor task performance. Thus, it may be possible to directly estimate the value of the d-prime of the identifying dimension based on the observer's performance, and then use this parameter to generate model predictions accordingly.

      We apologize for the confusion. We constrain the discriminability of items in the "identifying" dimension using the  parameter that determines the noise in that dimension for both items. The means in this dimension are fixed at an arbitrary value, as means and noise are interchangeable when considering discriminability. We will revise the description of the fitting procedure accordingly. Regarding the use of the same values in predictions, while possible, we prefer to keep predictions separate from fitting to avoid them becoming postdictions. The curves for the fitted model in Figure 2 already illustrate what the model predicts under the fitted parameter values.

      (1e) The current model assumes that a large number of samples are generated per stimulus and the brain can manipulate this information to perform the demixing task. It was well documented that visual working memory has a capacity limit (i.e., it can only hold information about a few items); this discrepancy needs to be clarified or addressed.

      We are grateful to the reviewer for raising this point, which we will address in the revised discussion. Briefly, we believe that the number of samples in the ideal observer model does not correspond directly to the working memory “slots”.

      (2) How well the computational model can explain the experimental data remains not entirely clear

      The authors show that there exists a parameter regime that can qualitatively explain the experimental finding. They also show that it is possible to fit the model to the data to explain the bias patterns. However, given that the model is flexible, it would be stronger if the authors could show that the same parameters that explain the biases could also explain other aspects of the behavior, for example, the magnitude of the errors.

      It would also help if the authors could report the best-fitted parameters from the experimental data. From these parameters, one can simulate synthetic data and apply the demixing model to see if the error distribution of the simulated observers is indeed similar to the experimentally measured error distribution. That way, one can check whether the fitted parameter explains the observer's behavioral performance beyond the biases.

      We are grateful to the reviewer for raising this point. We both agree and disagree with the reviewer here. The predictions reported come from an earlier paper describing the model (ref. 22). In our opinion, this represents a pure hypothesis-driven approach, where a prediction is formulated first and then tested with subsequently collected data. The model we test is normative, not descriptive; its goal is not to fit the data as closely as possible, but rather to make predictions about internal brain mechanisms. We do not suggest, for example, that demixing is the sole source of biases, so the resulting bias pattern might differ significantly from the predictions. That the model fits the data is, therefore, an additional bonus. At the same time, we agree that it is interesting to test whether the model can explain other parameters of the data. Note that our current fitting procedure was not geared toward this; we optimized the model to explain only the bias curve. In the revision, we aim to test whether the model can also explain the error variability.

      In other words, the model is not well constrained in the way it was tested in the paper. But it should be possible to improve it. First, if the noise parameter in the model is determined by the stimulus variability, one can determine it directly based on the external noise in the stimuli (discussed also in 1b) and see what prediction it leads to. Second, from the behavioral data, it may be possible to estimate the noise for the identifying dimension. Doing so will help better constrain the model.

      External noise accounts for only a portion of the total noise, as evidenced by behavioral errors. Even for a single item, the total noise consists of the amount of information the observer samples from the stimulus, the variability of these samples (external noise), and early (applied to each sample) and late (applied after integration) internal noise. Therefore, external noise alone might not constrain the model in the right regime. Regarding the identifying-dimension noise, as noted above, we do constrain it with the data. However, we aim to explore these points further in the revision.

      Other comments:

      (1) How does the model account for the swap errors? I am not sure I understood the way how the swap errors were treated in the paper. To me, substantial swap errors seem to be a consequence of having low d-prime values for the identifying dimension; that is, if there is only little information to discriminate the identity of the two stimuli, swap errors would be large. However, this possibility didn't seem to be mentioned in the paper.

      We apologize for the confusion. We will further clarify and perhaps reassess the treatment of swap errors in the revision. The model itself produces swap errors when the stimuli sources are misidentified.

      (2) Since the solution of the demixing problem was obtained using a numerical procedure based on EM. It would be useful to check whether the initialization has affected the biases obtained.

      Indeed, this is a valid point, and it's why we use a multi-initialization strategy. For each simulation of a single trial sample set (e.g., 100 random samples), we use a large number of initial points (50 in the initial submitted manuscript) to ensure the obtained EM solution is truly optimal. Additionally, we conduct a large number of simulated trials (10,000 for each parameter combination) to ensure the accuracy of the bias distribution we obtain.

      Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the origins of inter-item biases in visual working memory. The authors proposed a computational model where overlapping memory signals are disentangled, inducing memory biases that depend on relative noise levels across items. The key theoretical advance is the prediction that bias direction depends not only on absolute memory noise but on the relative noise levels of target and non-target representations. Using four experiments with color mosaics whose color variability manipulates memory precision, the authors report that biases reverse as a function of relative noise in a manner predicted by the model.

      Strengths:

      The manuscript is clearly written and theoretically motivated. The experiments are well designed and provide converging evidence for a distinctive and non-intuitive prediction of the proposed model. I found the central result compelling: independently manipulating target and non-target noise leads to qualitatively different bias patterns, consistent with the model's prediction that relative noise is a key determinant of bias direction.

      We are grateful for the positive evaluation of the model and the empirical observations.

      Weaknesses:

      The main limitation is that the evidence establishes consistency of the data with the proposed Demixing Model, but does not demonstrate that the model provides a unique explanation of the data. Although the manuscript argues that dominant theories struggle to account for the observed reversals, no formal comparison with alternative computational frameworks is presented. In addition, model fitting results are reported only briefly, making it difficult to evaluate fit quality at the level of individual observers.

      We agree and we aim to provide a comparison with alternative models and an expanded description of the fitting results in the revision. Note, however, that the majority of existing models are descriptive, while we believe that as a normative model, the Demixing Model should be compared with other normative models, thus limiting the selection of competitors significantly.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      In this paper, the authors provide a systematic investigation of structural brain differences associated with congenital aphantasia (self-reported lifelong absence of voluntary visual imagery). Specifically, the authors analysed a structural neuroimaging dataset involving 18 individuals with aphantasia and 18 visualizers to test two competing hypotheses: (1) that aphantasia reflects alterations in visual pathways and early visual cortex, and (2) that it instead reflects differences in higher-order frontotemporal and cingulate systems. To test these hypotheses, the authors employed multiple analysis approaches (e.g., cortical morphometry, tractometry, graph-theoretic network analysis).

      They report structural differences between the two groups in frontotemporal and cingulate systems. In contrast, they found no reliable group differences in early visual cortex or major visual tracts. On this basis, they propose that aphantasia is primarily associated with differences in higher-order systems supporting integration and conscious access to internally generated representations, rather than with deficits in sensory visual representations themselves.

      Strengths:

      (1) The present work addresses an important gap in the mental imagery literature, providing a systematic investigation of structural neuroimaging differences in congenital aphantasia. By showing that structural differences between aphantasics and visualizers are mainly concentrated in frontotemporal and cingulate systems (rather than in visual cortex), it makes an important step toward a better understanding of individual differences in mental imagery and provides a set of candidate regions for future mechanistic work.

      (2) A key strength of the study is the multimodal approach employed to address the main research question, integrating tractometry, functional region-of-interest (fROI)-based tractography, graph-theoretic network analysis, and surface-based cortical morphometry, which provide a converging assessment of structural differences between aphantasics and visualizers.

      (3) The complementary use of Bayesian analyses alongside NHST to assess evidence for null results is a further strength of this work.

      Weaknesses:

      (1) A weakness of this work is related to aspects of the framing and, in particular, what can be confidently inferred from the results. The framing of existing accounts of aphantasia in the Introduction appears limited in that it reduces the views on aphantasia to two options (sensory strength account versus conscious access account) without acknowledging a third distinct position, namely that aphantasia reflects a specific deficit in the voluntary generation of imagery (Milton et al., 2021; Zeman et al., 2015, 2020; Whiteley, 2021; Cavedon-Taylor, 2022). Like the conscious access account, the view that aphantasia involves a deficit in the generation of sensory representation also speaks against the hypothesis of reduced sensory strength of internally generated representations. This third view could be acknowledged/discussed as it also maps quite well onto the presented results.

      (2) Relatedly, I think the main weakness of the paper concerns the interpretation of results being restricted to a lack of "conscious access". The paper frames its findings as mainly evidence for a conscious access failure, the view that visual representations are generated by aphantasics but cannot be consciously accessed. However, the structural findings are equally consistent with a voluntary generation failure, especially since the same higher-order regions examined can also be implicated in the top-down generation and control of imagery. The authors themselves initially define aphantasia as "lifelong absence of voluntary visual imagery". Given the nature of structural imaging data (as opposed to functional data), it is not possible with the present study to distinguish between a lack of generation versus a lack of conscious access. As such, examining this alternative interpretation appears appropriate, and it would considerably strengthen the paper. Structural MRI alone is not sufficient to dissociate imagery generation from conscious access, as these are fundamentally functional questions.

      (3) Some inconsistency and lack of clarity around the specific choice of regions/networks, which could be better motivated and explained. E.g., the "core imagery network" analysed in the white-matter connections analysis was derived from a previous 7T study (with which the sample partially overlaps) and is not necessarily the network most commonly associated with visual imagery in the literature (e.g., see Dijkstra et al., 2019; Pearson, 2019). It is, for instance, unclear why V1 was examined in the cortical thickness analysis but not in the previous one, given that both analyses are related to the visual pathway hypothesis. Related to this, in the graph-theoretic analysis, the rationale for network selection is inconsistently established in the Introduction. The attention and salience networks do have some grounding in the Introduction through the mention of specific regions such as FEF and anterior insula, though these are discussed as individual regions rather than as networks. However, the default mode network receives no motivation in the Introduction. More explicit elaboration on these choices would be appropriate.

      (4) The interpretation provided in the Discussion tends to oversimplify what is in fact a heterogeneous and rich set of structural findings into a relatively coherent mechanistic account. The observed differences are spatially and directionally variable across tracts, cortical regions, and metrics: e.g., FA is reduced in the UF and posterior interparietal corpus callosum but increased in the dorsal cingulum; cortical thickness is reduced in aPFC but increased in medial temporal regions, and so forth. The Discussion acknowledges this in part (e.g., proposing increased dorsal cingulum FA as potentially compensatory) but does not address the directional heterogeneity systematically. The authors could discuss more explicitly what the opposing directions of effects mean for their overall interpretation. Relatedly, some parts of the Discussion link specific structural findings to specific imagery processes in ways that go beyond what the current data can support. The authors could more clearly distinguish between what the structural data show and what functional interpretations are taken from prior work.

      We will add two recent in-press Cortex papers to the Discussion. One provides lesion-based double-dissociation evidence against V1 as a necessary causal substrate of visual imagery. The other shows that aphantasic individuals can display visualizer-like oculomotor patterns during mental map exploration despite reporting little or no imagery vividness. Together, these studies help clarify our interpretation of our null V1 findings and structural effects in higher-order brain regions, which are consistent with aphantasia involving altered integration or access rather than a primary V1-dependent imagery deficit.

      Reviewer #2 (Public review):

      Summary:

      This paper addresses whether congenital aphantasia reflects an alteration of visual representations themselves, or rather of the systems that allow internally generated representations to reach conscious experience.

      Strengths:

      The study is novel and ambitious. The authors combine several complementary structural MRI approaches in a rare and well-characterised population, and the convergence of the findings toward frontotemporal and cingulate systems, with relative sparing of early visual cortex and major visual pathways, is particularly interesting because it could affect the way visual imagery is modelled and tested experimentally and clinically.

      Weaknesses:

      Overall, I found the manuscript conceptually and methodologically strong. My main concern regards the interpretation of the anatomical findings, rather than the findings per se. The authors discuss their results within a rich cognitive framework. However, the current dataset does not appear to include independent behavioural or neuropsychological measures that would allow the proposed cognitive interpretation to be tested in the same participants. As a result, the manuscript sometimes moves quite rapidly from 'these structural differences involve systems associated with higher-order control, salience, conscious access' to 'these structural differences may explain the cognitive mechanisms of aphantasia'. I agree that this is the most interesting interpretation, and probably the right one to explore. Although plausible, it remains indirect. The authors already acknowledge this point when discussing memory, affective control, and semantic processing. However, the same logic should be extended to the interpretation of the full set of findings. For example, if the salience/anterior insula findings are interpreted in relation to access to internally generated representations, it would be useful to know whether aphantasic participants also differ behaviourally on tasks tapping interoception or related aspects of internal monitoring. I appreciate that collecting additional behavioural data may not be feasible at this stage, especially given the difficulty of recruiting participants with such a specific manifestation. However, I think it should be acknowledged more explicitly in a dedicated limitation paragraph.

      We thank the reviewer for this thoughtful and constructive comment. Lack of introspective report of voluntary imagery is arguably the defining signature of aphantasia. This motivated us to primarily interpret our anatomical findings in a broader cognitive context of higher-order control, internal monitoring, and conscious access in aphantasia. We expect that a reliable behavioural test measuring imagery sensitivity and accessibility would allow us to direct link these findings to individual imagery ability. Nevertheless, to our best knowledge, this kind of test on imagery is still missing. Instead, our findings point to some plausible structural signature or brain regions that may be related to conscious imagery, which motivate future studies to examine their direct or causal roles. We agree with the reviewer, future studies should test the relationship between these anatomical structures and the accessibility to internal representation, together with related aspects of internal monitoring. We will therefore add a dedicated paragraph to discuss the plausible cognitive mechanisms during the revision.

      Reviewer #3 (Public review):

      Summary:

      The authors investigate the structural brain basis of congenital aphantasia, a condition characterised by a lifelong absence of voluntary mental imagery. They test two competing accounts: one predicting structural differences in early visual pathways, the other predicting differences in higher-order frontotemporal and cingulate systems. To do this, they combine four complementary structural imaging approaches: white-matter microstructure profiling along anatomically defined tracts, tractography seeded from functional regions of interest, whole-brain structural network analysis, and cortical thickness mapping. The main finding is that white-matter differences are selective for frontotemporal and cingulate pathways and absent in early visual pathways, which the authors interpret as support for the higher-order account.

      Strengths:

      The multi-modal design is a genuine strength: running four independent analyses increases the chance of detecting real effects and of identifying false positives that appear in only one stream. The statistical choices within each analysis are appropriate. Permutation-based correction with a threshold-free method is well-suited to the tract-level comparisons. The use of Bayes factors to quantify evidence for null results, rather than simply reporting non-significant tests, is particularly valuable here, since the absence of visual pathway differences is central to the argument. The robustness checks across multiple brain parcellations for the network analysis strengthen confidence in those findings.

      Weaknesses:

      The main limitation concerns the relationship between two of the analysis streams. The measure used to weight structural connections in the network analysis is calibrated to match fiber density estimates derived from the same diffusion signal that drives the white-matter microstructure differences. If the two groups differ in tissue organisation in certain pathways (which the microstructure analysis suggests they do), that difference will feed into both measures. The authors should acknowledge this dependency when discussing convergence across analyses.

      More broadly, the imaging metrics used throughout (measures of fiber organisation and weighted connection counts) reflect what the diffusion model captures from the tissue and cannot be directly read as measures of axon number or connection strength. This is a known limitation of the field, but it is relevant to the strength of structural claims made in this paper.

      The network analysis is presented without comparison to a null network. Without this, it is hard to know whether the node-level differences reflect specific network topology or simply follow from overall differences in connectivity weight or density between groups.

      The study runs four separate discovery analyses on the same 36 participants, each corrected within itself but with no control across analysis streams. At 18 participants per group, this is exploratory work. Some of the language used in the abstract and discussion, like "first comprehensive characterization" and "selective structural phenotype", reads as more definitive than the data support at this sample size. Framing the results as hypotheses to be replicated would make the paper stronger.

      The paper frames the results as distinguishing between two competing accounts. The positive evidence for the higher-order account is clear. The absence of differences in visual pathways is a different kind of result: it means such differences were not detected in this sample, not that visual pathways are uninvolved. The discussion at times moves toward that stronger conclusion, which the data do not support.

      The cortical thickness analysis finds one cluster in the predicted direction, while the other analyses each return multiple effects. One cluster in a whole-brain search with 18 participants per group is not strong evidence and should not be presented as equivalent to the other results.

      Effect sizes are reported without confidence intervals throughout. With 18 participants per group, the uncertainty around those estimates is large, and confidence intervals would give readers a more accurate sense of what can be concluded.

      We are grateful to the Reviewer for the constructive and thoughtful assessment of our manuscript. In response to the reviewer’s comments, we will revise the manuscript to clarify the dependency between diffusion-derived analysis streams, to state more explicitly the biological limits of diffusion MRI metrics, to add a null-network sensitivity analysis for the clustering coefficient findings, to include confidence intervals for reported effect sizes, and to temper the interpretation of the cortical thickness result. We will also revise the Abstract and Discussion to better reflect the exploratory nature of the study and to frame the findings as hypotheses requiring replication in larger independent samples. We believe that these revisions will make the manuscript more balanced, transparent, and appropriately cautious, while preserving the central conclusion that congenital aphantasia is associated with structural differences centered on higher-order frontotemporal and cingulate systems.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      In their manuscript, Arjun et al. investigate the role of the histone acetyltransferase Gcn5 in the control of drosophila blood cell homeostasis in the larval lymph gland. They use gcn5 zygotic mutants as well as targeted knock-down and over-expression of Gcn5 in various lymph gland populations to show that these modulations impact (in a rather haphazard manner) niche cell number, blood cell progenitor maintenance, plasmatocyte differentiation, crystal cell differentiation or DNA damage accumulation. Their results suggest that Gcn5 controls autophagy and they show that decreasing the expression of the autophagy machinery increases blood cell differentiation. Using drugs to modulate the mTOR pathway, they conclude that Gcn5 levels are regulated by mTOR but that the impact of this pathway on blood cell homeostasis can override Gcn5 function.

      While the authors did a lot of experiments and good quantifications of the blood cell phenotypes, many results do not make much sense or do not bring valuable information about Gcn5 mode of action. Several conclusions of the manuscripts are not backed by solid data (e.g. that Gcn5 action is mediated by TFEB and the autophagy machinery) and different aspects of the literature are not well taken into consideration. Some results (such as the validation of the knockdown and overexpression of Gcn5) seem flawed. There are some concerns about the results obtained with gcn5 zygotic mutants and an interpretation of the phenotypes observed upon manipulation of Gcn5 expression in different cell types is missing.

      We have now performed several experiments to address the comments raised by the reviewer and have also provided possible explanation of the phenotypes in cases where it was lacking.

      Important revisions are needed to improve the quality of the manuscript and confirm the authors' findings.

      Reviewer #2 (Public Review):

      Summary:

      Drosophila hematopoiesis has been shown to be governed by a number of signaling pathways such as JAK/STAT and Dpp. This important study shows the role of nutrient sensing and autophagy in determining blood cell differentiation. The authors show that General control non-derepressible 5 (Gcn5), a histone acetyltransferase affects blood cell differentiation. Gcn5 also negatively regulates autophagy through its effector TFEB which directly regulates autophagy genes. The authors also show that mTORC1 modulates Gcn5 levels and through it, TFEB activity thus acting as a fine-tuning mechanism that maintains optimal levels of autophagy.

      Strengths:

      The main strength of the work lies in the interesting finding that cellular metabolic processes such as autophagy have a direct role in blood cell differentiation and has the potential to be of interest to those working on vertebrate haematopoiesis as well. The report has generated intriguing data, using promoters specific for sub-sections of the lymph gland, that different cellular subsets of the lymph gland contribute differently towards haematopoiesis, but this is not followed up in detail and the final conclusions are derived from a combination of whole lymph gland perturbations as well as those from specific promoters.

      Weaknesses:

      (1) Gc5 seems to be expressed throughout the lymph gland but modulating it in the subsections does not have the same result. It is very striking that the knockdown of Gcn5 in the prohemocyte population does not have an effect on differentiation whereas overexpression does. The modulations of Gcn5 in PSC also have variable effects across hemocyte subpopulations which is not explored in the manuscript.

      We have now explained and discuss why Gcn5 modulation could be affecting the PSC size. Please check Discussion section Paragraph 1 line 10 onwards.

      Interestingly, also the domain deletion constructs show a differential effect on blood cell differentiation when altered solely in the prohemocytes which is not explained.

      Currently, with our observations all that we can comment about that data is that expression of domain deletion mutants causes aberrant hematopoiesis indicating a dominant negative phenotype since they are expressed in the wild type genetic background. Beyond this, we will be exploring mechanistically how these domains are functioning during hematopoiesis in future studies. We have already described the dominant negative effect in the text: Discussion Section Paragraph 3.

      While Gcn5 can be seen in all sections of the lymph gland in the first figure, under the HHLT-Gal4 and Hml-Gal4, Gcn5 looks cytoplasmic and almost completely excluded from the nucleus strikingly unlike Gcn5 expression under the Collier-Gal4 and Dome-Gal4.

      We have now revised Figure 1 and have only included the images with Collier-Gal4 and Dome-Gal4 which clearly shows both the niche cells, Dome-positive progenitors and Dome-negative cells of the primary LG lobe essentially showing that Gcn5 is expressed throughout the primary LG lobe. In Fig. 1C-F’, Gcn5 expression is both in the nucleus and cytoplasm as this molecule shuttles between cytoplasm and nucleus. The staining pattern with the other Gal4 could be due to problems in the immunofluorescence protocol and acquisition parameters. We have now removed those images from Figure 1. Please check revised Figure 1.

      The rest of the experiments in the manuscript are done with multiple promoters, with autophagy flux measured by modulating Gcn5 with a pan hemocyte promoter, but the mTORC1-Gcn5 axis is explored using chemical modulators which affect the whole of the lymph gland (Fig7) or using two pro-hemocyte promoters (Fig8).

      We have used a pan-hemocyte promoter for the autophagy analysis to investigate if Gcn5 regulation over autophagy is a hemocyte specific effect which we indeed see. We have removed the western blot data now in the revised manuscript where we looked at Atg8 and p62 levels in whole larval lysates when Gcn5 was perturbed using hemocyte driver as the results were puzzling and difficult to comprehend given the complete absence of a p62 band in Gcn5 knockdown conditions. Also, it’s worth noting that Hml-Gal4 is also active in the LG hemocytes. We did 2 alternate promoters for prohemocytes to cross-validate some of our results and the chemical modulators experiment was done since effects like mTOR inhibition/nutrient sensing effects are systemic and hence such modalities were employed.

      (2) The knockdown of Gcn5 seems to affect the gland size (A compared to B and C). Since mTORC1 is a central regulator of cell size, it is possible that some of the effects seen in these knockdowns are potentially through mTORC1 affecting size suggesting that the signalling axis between mTORC1 and Gcn5 might not be a one-way axis as suggested in Figure 9. Also, this would mean that in experiments where absolute cell counts of crystal cells or niche cells are used to assess blood cell differentiation, further analysis to consider total cell numbers in the lymph gland would strengthen the manuscript.

      It is a possibility that Gcn5 perturbation could be affecting lymph gland size although we have not seen any consistent trend that would point towards this phenotype either upon knockdown or over-expression. We believe Gcn5 controls blood cell differentiation phenotypes strongly via mTORC1. But in order to answer reviewer’s comment we have now re-analyzed our crystal cell differentiation data particularly and quantitated it and represented it as crystal cell differentiation index for dome-Gal4 specific Gcn5 modulation and for the data with genetic modulation of mTORC1 pathway. Please see Fig 3P and S10J for the revised analysis.

      (3) A genetic manipulation of mTORC1 specifically in the pro hemocytes would strengthen the role of mTORC1 in the pathway rather than the chemical modulation which affects the whole of the lymph gland.

      We thank the reviewer for their useful critique. We have now addressed this concern and we have genetically perturbed the mTORC1 pathway in the progenitors using both abrogation of TORC1 via depletion of Tor or Raptor or by activation using over-expression of Rheb. We have now included this data as Supplementary figures – Fig S10 and S11 and have described it in the results section. Please see results section “Chemical or genetic modulation of mTORC1 activity controls blood cell differentiation” in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      The abstract could clearly be improved. It does not make a clear presentation of what is new in the manuscript. The conclusions that Gcn5 function in the lymph gland is mediated by the autophagy machinery and the acetylation of its non-histone target TFEB are not grounded and purely circumstantial. The implication of mTOR and nutrition in drosophila larval blood cell homeostasis has already been studied but not mentioned here. Most of the time the authors do not provide any possible explanation about the phenotypes they observe and how they fit with the current literature. Several pieces of results are of serious concern.

      We would like to thank the reviewer for their feedback. We have revised the abstract and have incorporated the insights obtained from our study. We have now included relevant literature that talks about the implication of mTOR and nutrition in Drosophila larval blood cell homeostasis (Please see Introduction section Paragraph 2 in the manuscript). We have also noted the input of the reviewer on many phenotypes lacking any description of a possible explanation. We have worked on results section to provide possible explanation and speculation wherever relevant.

      In the introduction, the authors do not provide an up-to-date and accurate presentation of the field. For example, they could use much more recent and comprehensive reviews since Evans et al. 2003. (eg. MID: 30733377 or 35887113). Their choice for signaling pathways involved in Drosophila blood cell progenitors seems very much biased for lead author self-citation rather than more directly related citations. It is surprising too that the authors failed to mention a series of publications on Akt/mTOR and nutrient sensing impact on drosophila larval blood cells (PMID: 22951642 ; 22911822 ; 22407365 ; 22510984). Along the same line, there are already several reports on autophagy genes implicated in Drosophila hematopoiesis and blood cell functions (PMID: 23406899; : 33560224 ; 20498061 ; 37623416). The introduction on GCN5 is a bit of a catalogue and should be streamlined- citing a recent review would be useful (PMID: 32735945). Again, the authors fail to cite publications showing that Gcn5 levels can be modulated by nutrition (PMID 27022023; 27874008) and they do not mention that amino acids starvation or mTOR inhibition leads to a decrease in GCN5 activity / TFEB acetylation (ref 40). Taking into account all the missing information, the novelty of the present manuscript is strongly decreased.

      We would like to thank the reviewer for the detailed suggestions on including the relevant literature that are appropriate and relevant to be mentioned in the context of the observations in our manuscript. We have now included these references and have cited them as per reviewer’s suggestions. Please check Introduction section paragraph 2.

      Results

      While there is little doubt that Gcn5 is expressed in the entire primary lobes based on Fig 1C-F, the quality of the staining in G-J (especially H, J) is really poor and essentially looks like non-specific background with no clear signal in the nuclei. Better images should be presented. The conclusion of the paragraph ("all cellular populations of the LG") and title of Fig.1 are not fully accurate as the authors do not provide evidence that Gcn5 is also expressed in posterior lobes.

      As per reviewer’s suggestions, since the images in Fig 1A-D’ clearly show that Gcn5 is expressed in the entire primary LG lobe in PSC cells, MZ and CZ; we have removed panels E-H’ which lacked clear nuclear signal. Fig1A-D’ clearly show the nuclear staining pattern of Gcn5. We have also modified the conclusion of the paragraph to say that Gcn5 is expressed in cellular populations of the primary lymph gland lobe accordingly.

      Concerning, Fig S1 and Fig 2, while the analysis seems technically sound, the results are puzzling. The lack of P1 differentiation in gcn5 null heterozygotes is very surprising. The authors should check that this stock does not carry a mutation in nimC1 (for details see: PMID: 23899817) and use other plasmatocyte differentiation markers to confirm their observation (also with the different allelic combinations). I'm also concerned by the levels of plasmatocyte differentiation and crystal cell number in the control line (notably in S1H), which seem very low (and quite variable for P1 as there is a notable difference between S1H and Fig 2H). Moreover, the analysis of the allelic combinations gives rather incoherent results: PCSC cell numbers are affected only in null/hypomorph, whereas differentiation (NimC1 and Hnt), as well as DNA damage, was only increased in hypomorph homozygotes. The authors propose no hypothesis to explain these observations.

      We have now repeated these experiments with the E333st null allele by placing it on a different balancer and we observe homozygotes that are alive till late third instar/early pupal stage as shown before by Carre et al., 2005. We have now included these revised results on the plasmatocyte differentiation status of the E333st heterozygotes and homozygotes (See Fig 2 and Fig S1). We do find P1 positive cells in the E333St heterozygotes unlike earlier. Plasmatocyte and crystal cell numbers in the control line always shows some level of heterogeneity. We have included the wild type control individually with those respective mutants during the experiment hence drawing a cross comparison across two different experiments would not be appropriate. We have now explained the observations obtained on PSC cell numbers (Discussion section paragraph 1). Experiments to check all hematopoietic aspects of the gcn5 null have been done after changing the balancer line and the null mutants overall show a decrease in PSC size and a widespread increase in hemocyte differentiation which could be due to a systemic effect due to various signalling pathways being affected which needs to be investigated and is beyond the scope of this study. This has also been discussed in the Discussion section Paragraph 1.

      Although a side-by-side comparison would have been better suited, it seems that the homozygotes or trans-heterozygotes do not have stronger phenotypes than the heterozygotes as far as crystal cell and DNA damage are concerned, which is rather unexpected. Besides the authors should introduce why they look at DNA damage.

      We agree with the reviewer that for the crystal cell and DNA damage phenotype the homozygotes or trans-heterozygotes do not have a stronger phenotype as compared to the heterozygotes alone but since these are whole animal mutants there could activation/inactivation of various signalling pathways and systemic effects that would be difficult to account for and comprehend here which needs to be investigated further. The only conclusion that we draw from these observations is that Gcn5 is required for maintaining blood cell homeostasis. Regarding DNA damage, we have now included the rationale and supporting literature for why we have studied DNA damage in the context of Gcn5. Please see result section 2 paragraph 1.

      Importantly too, the authors failed to obtain gcn5 E333st/E333st (null) larvae, whereas Carre et al. originally reported that E333st/E333st individuals are viable until the late third instar larvae. I suspect that the stock they use carries additional mutations that need to be eliminated by back-crossing it to control flies for several generations. Of note too, a recent report showed that a deletion of gcn5 (generated by CRISPR) does not prevent adult emergence, challenging the conclusion that gcn5 expression is absolutely required for fly development (PMID: 37545086).

      The reviewer is right in pointing out that E333st homozygotes survive until late third instar as reported by Carre et al.,2005. We have procured the null allele again and used another balancer to obtain homozygotes and we were able to get homozygotes that survived till late third instar as reported earlier. We have now included new data from these homozygotes for all hematopoietic aspects and heterozygotes particularly for plasmatocyte differentiation Please see Fig 2 and Fig S1 and corresponding results section 2 of the manuscript.

      Concerning the validation of Gcn5 knock-down and overexpression: the results are highly dubious. In Fig S2B (hml>Gcn5 RNAi), there is virtually no Gcn5 signal in the primary lobes but hml is normally expressed only in the cortical zone. How is it possible? Similarly, the western blot (which is really too much cropped around the bands of interest) does not show any signal in the hml>Gcn5 RNAi lane (not even some background. According to the Methods section, the western was performed on whole larvae extracts; hml-mediated knock-down can not wipe out its expression in all the tissues. As for the overexpression, flag immunostaining in hml>Gcn5-flag is mostly cytoplasmic (S2E), which doesn't make sense and does not fit with S2C (Gcn5 immunostaining).

      Hml-Gal4 is a pan hemocyte driver and its expression is not limited to the CZ of the primary lymph gland lobe (Banerjee et al., 2019) and recent single cell sequencing data corroborate this that Hml domain is not limited to the cortical zone (Yarikipati and Bergmann, 2026). GFP driven by Hml-Gal4 is spread out across the primary LG lobe which could explain the phenotype of no Gcn5 signal obtained in the immunofluorescence experiment. Regarding the western blotting experiment which was performed on whole larval extracts, we were also puzzled by lack of Gcn5 bands in these lysates upon depleting Gcn5 using Hml-Gal4. We need to systematically probe further to understand expression of Gcn5 in other tissues and organs. We have now removed the western blot data as the data obtained cannot be comprehended at the moment. Regarding the FLAG staining experiment – the staining gave us a cytoplasmic pattern and since Gcn5 is known to shuttle between the cytoplasm and nucleus it is possible that the anti-FLAG staining detected the Gcn5 localizing in the cytoplasm. It is difficult to draw a direct comparison here between the images S2C and S2E as both are different antibodies.

      The initial analysis of Gcn5 level modulation in the prohemocytes, PSC or Hml+ cells is mainly descriptive and the authors do not elaborate on possible explanations based on the current literature.

      We have added a possible explanation wherever required for these respective results on Gcn5 modulation in prohemocytes, PSC and Hml positive hemocytes. Please see result section 3 where we elaborate on possible explanation for the phenotypes observed.

      The structure/function analysis of Gcn5 is based on overexpression of truncated mutants in the prohemocytes using the tep4-GAL4 driver and monitoring PSC cell, prohemocyte maintenance, plasmatocyte and crystal cell differentiation as well as DNA damage. As the overexpression of the full-length protein was made with a different driver (Dome), it is difficult to interpret the data. Nevertheless, no clear message emerges from this analysis and the authors do not reach any conclusion. Thus, the interest of these experiments remains limited.

      The structure-function analysis was largely done to understand which of the domains of Gcn5 upon over-expression results in a dominant negative like phenotype and our analysis shows that expression of some of these domain mutants results in a dominant negative phenotype in the wild type genetic background which we have now stressed upon in the text. However, further mechanistic understanding and in-depth analysis of each of these domains of Gcn5 warrants further separate investigation and is beyond the scope of this study. Please see the end of result section 4 for conclusion and possible explanation.

      The authors then analyze autophagy markers (in hml>Gcn5 LOF or GOF). Contrary to their say, hml-GAL4 is not a pan-hemocyte marker. It would have been interesting to ensure that the effects observed on Atg8 and Ref(2)P in the lymph gland are cell-autonomous- as expected for a direct role of Gcn5 on this pathway. Again, it is very surprising that p62 is not detected in the western blot on whole larval extracts when Gcn5 is knocked down in Hml+ cells only (Fig 5D). Moreover, quantifications on multiple samples will be needed to validate the increase/decrease of p62 and Atg8 as detected by western blot. As for the RT-qPCR (Fig S5), according to the Methods sections, they were made on adult blood cells but this is not explicit in the result section.

      We have corrected the text and mentioned Hml-Gal4 as a hemocyte specific Gal4 shown earlier as Gal4 marking both embryonic and larval hemocyte population (Goto et al., 2003, Yarikipati and Bergmann, 2026). Regarding the Atg8 and Ref (2)P blots – yes, it is surprising to us too that the p62 is not detected in the larval lysates when Gcn5 is depleted using Hml-Gal4. However, this result was consistent over the replicates performed and needs to be further studied. Since this phenotype of complete absence of p62 in larval lysates upon Gcn5 depletion cannot be comprehended and explained, we have removed the western blot data from the figure and have just retained the immunofluorescence data and have also quantified the Atg8 and p62 puncta per cell and included this data in Figure 5, Graphs D and E. For the qRT-PCR we have now included a description in the corresponding results section. Please see result section – result 5 under “Autophagic flux in the Drosophila blood cells is negatively regulated by Gcn5”.

      The knock-down of TFEB or several autophagy genes in the prohemocytes (tep4-GAL4) leads to a rather convincing increase in plasmatocyte and crystal cell differentiation. It would have been interesting though to quantify prohemocyte maintenance, PSC cell number, and DNA damage. Also, the authors should have performed Gcn5 GOF/LOF experiments with the same driver (they present tep>Gcn5 RNAi in Fig 8 but without the proper controls).

      We have now included data for prohemocyte index (Figure S8M) upon knockdown of TFEB and other autophagy genes along with PSC cell number, DNA damage (Supple Fig S8) in the revised manuscript. Please see corresponding results section titled “Genetic and chemical ablation of autophagy boosts blood cell differentiation in the primary lymph gland lobe” for the description of the results.

      The use of chloroquine should be better described. How long was the treatment? Did the authors observe an effect on autophagy in the lymph gland? Chrorloquine also affects lysosomal pH, so it remains to be demonstrated that the effects observed here are only autophagy-related.

      We have now written a detailed protocol for the treatment in the methods section and also mentioned the treatment time which is 16 hours in the results. We have included data to validate the effect of Chloroquine on autophagy by p62 and Atg8 staining in the LG and have quantitated the data (Refer Supple Fig S9) and the corresponding results section titled “Genetic and chemical ablation of autophagy boosts blood cell differentiation in the primary lymph gland lobe”

      Similarly, the use of drugs to activate (3BDO) or inhibit (Rapamycin) mTOR should be better controlled. More generally, given the promiscuous roles of mTOR (and autophagy) in the larvae, tissue-specific manipulations would be better suited.

      We have now perturbed mTOR pathway genetically by activation and in-activation and have studied the effect on blood cell differentiation. Please see Figure S10 and the corresponding result section titled “Chemical or genetic modulation of mTORC1 activity controls blood cell differentiation” where we discuss the results of genetic perturbation of mTOR pathway.

      Actually, as pointed out above, it has already been shown that modulation of Akt/TOR in hemocytes or amino-acid deprivation affects blood cell homeostasis (see above). The authors should definitely discuss how their results fit with the literature on this subject.

      We have added relevant literature in the introduction section and have also discussed how Gcn5 could fit into this context of nutritional sensing and control of hematopoiesis. Please check revised Introduction section paragraph 2. Also, check discussion section in last paragraph where we have discussed role of Gcn5 in nutrient sensing.

      Again, Gcn5 levels need to be quantified using multiple samples (Fig 7M, N) before concluding.

      Sorry for not including the quantitation earlier but we have now included the quantitation for the blots presented in Fig. 7 M and N.

      Finally, the authors show that 3BDO still induces an increase in blood cell differentiation when gcn5 is knocked-down in tep4+ cells and that Rapamycin still represses differentiation when Gcn5 is overexpressed in Dome+ cells. They conclude that mTORC1 overrides the effect of Gcn5. This seems a far-reaching conclusion given the available evidence.

      We have now toned down the conclusion that we make to accommodate other possibilities which we have been unable to test here currently.

      In particular, in the conditions used, the authors do not necessarily assess the activity/requirement for Gcn5 and mTORC1 in the same cell population.

      Other comments and suggestions:

      The discovery of the SAGA complex is not Grant 1999 but 1997 (PMID: 9224714).

      Ref 30 is not appropriate -nothing to do with HAT.

      GCN5 not only acetylates TFEB but also Atg7 (PMID: 28594263) to limit autophagy.

      Thank you so much for these suggestions. We have made the necessary amendments in the references.

      In the results section, the first paragraph is largely a repetition of the introduction. The same is true for most paragraphs in this section. A shorter (hypothesis-driven) introductory sentence would be more adequate.

      We have now taken the suggestion into consideration and made the necessary change in the results section throughout the manuscript.

      Fig 1: it seems that there is a higher accumulation of Gcn5 in a few cells in the cortical zone. This may correspond to crystal cells and could be easily confirmed.

      We have now checked this aspect. Please see supple fig S5 where we co-stain lozenge-GFP cells containing LG with Gcn5 to check for the accumulation. However, we do not see any accumulation in the Lozenge-positive crystal cells.

      Figure 3: the authors should also quantify the proportion of progenitors (dome>GFP+) in the different conditions.

      We have now done this and added it to the Figure. Please see panel N in Figure 3 and Figure S8M.

      Figure S3: how do the authors explain that Gcn5 knockdown in the PSC reduces plasmatocytes differentiation (but does not affect PSC cell number or crystal cell differentiation)? What could be the origin of the increase in DNA damage (essentially in CZ)? How do they explain that Gcn5 over-expression increases PSC size but does not affect (reduce?) blood cell differentiation?

      These observations need to be investigated further. We currently have no answer to these comments. The signals that are produced by the PSC could be affected due to which we observe these phenotypes like an effect on plasmatocyte differentiation and an increase in DNA damage whereas no effect on PSC cell numbers or crystal cell numbers which needs to be studied further. Also, in the case of Gcn5 over-expression in PSC we do not know how the increased size of PSC controls differentiation. This would need further experimentation and since this paper is not about the role of Gcn5 in PSC exclusively, we will look into this in our future studies. These aspects will be studied in our future follow-up studies as it is beyond the scope of the current manuscript.

      Figure S4: how do the authors explain the non-cell autonomous increase in PSC cell number upon Gcn5 KD/GOF in hml+ cells? How do they explain the increase in crystal cell number in Gcn5 GOF? Is it really cell-autonomous (i.e. all the Hnt+ cells are Hml+?)?

      We have discussed how Gcn5 depletion or over-expression in HmlΔ cells could affect PSC cell numbers. Please see discussion section, paragraph 1. Regarding the crystal cell phenotype - We have now tested if the increase in crystal cell numbers is cell autonomous by driving Gcn5 over-expression using a crystal cell specific driver and we find that the increase is cell-autonomous. Please refer to Supple Fig S5.

      The discussion is lengthy and should be reduced. It does not appropriately consider the current literature.

      We have tried to reduce the length of the discussion and have also added relevant references as per recommendations of the reviewer.

      Reviewer #2 (Recommendations For The Authors):

      (1) In general, it is not clear why in some of the experiments Tep-Gal4 is used to modulate proteins in prohemocytes while in others Dome-Gal4 is used.

      There is no particular reason. These Gal4’s have been used interchangeably as both label the hematopoietic progenitor population. Although recent single cell sequencing data has identified subsets within the progenitors namely core progenitors marked by tep4 largely and dome being a distal progenitor marker (Cho et al.,2020, Girard et al.,2021), in our study perturbations in Gcn5 using either of the Gal4’s results in a similar phenotype.

      (2) Considering alteration in lymph gland size (Figure 2), the number of positive cells should be analysed in relation to total cell numbers or s4ize.

      Although we do not find any visible differences or defects in the overall LG size in various genetic conditions discussed in this manuscript, we have done so for the plasmatocyte differentiation where we have represented it as plasmatocyte differentiation index (relative to the size of primary LG lobe) throughout the manuscript. We have now done this for crystal cell numbers too for critical genotypes in this manuscript and have represented it is as crystal cell index for example please see Figure 2O, 3P, S5G, S10J where these graphs have now been added.

      (3) Figure 1A G-I' does not look like mCD8 GFP expression, but rather cytoplasmic GFP.

      We have made the change in the figure and the corresponding text accordingly.

      (4) One of the main conclusions in the manuscript is that Gcn5 affects autophagy (Figure 5). Here, the puncta need to be quantified (relative to total cell numbers).

      Thank you for the suggestion. We have now quantitated the p62 and Atg8 positive puncta per cell and have represented it as panel D and E in Figure 5.

      (5) Figure 5 D and E show p62 and Atg8 total protein levels in the larvae when Gcn5 is modulated only in the hemocytes. It is surprising that there is a complete reduction in p62 levels across the whole larvae when Hml gal4 is used for the knockdown.

      Yes, we observe a complete absence of p62 in whole larval lysates when Gcn5 is depleted using Hml-Gal4 and we see this across replicates. This result is indeed puzzling to us and difficult to comprehend as to why a hemocyte specific driver would result in such a dramatic change hence we have decided to remove the western blot data as it is difficult to draw a solid conclusion from. We have retained the immunofluorescence data which shows a consistent alteration in autophagy upon Gcn5 perturbation using Hml-Gal4 and we have now included the quantification for the number of p62 and Atg8 positive puncta per cell for the IF data.

      (6) The beta-actin levels in the western blots in Figure 5 are highly oversaturated and do not represent loading control adequately. Also, it looks like there is substantially more total protein in 5D 3rd lane where Gcn5 is overexpressed.

      Thank you for pointing this out. We have loaded equal amount of protein in all the wells so we are unsure why the actin bands look over-saturated. We have now removed the western blot data from this figure as the data is puzzling and difficult to comprehend given a total absence of p62 in whole larval lysates in Gcn5 depletion conditions using Hml-Gal4. Hence, we are just retaining the immunofluorescence data.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It is unclear to what extent the model's success relies on the way non-decision time is formalised in the model. In the proposed PDG model, non-decision time is decomposed into separate visual encoding, saccadic execution, and manual execution components. Several values (assumed or recovered) do not match known physiological or behavioural ranges. This is a common issue in the literature, and the authors may want to address it in light of broader work discussing what non-decision time consists of in both manual and saccadic actions (e.g., Bompas et al., 2024, Non decision time: the Higgs boson of decision, Psychological Review).

      In particular, the "saccadic execution" parameter appears far too long and too variable to reflect merely execution; instead, it likely includes decisional components. This would make more sense since manual and saccadic planning essentially rely on distinct brain areas, hence it seems unrealistic that crossing a single threshold would trigger both manual and saccadic execution. Similarly, recovered manual non-decision times are substantially longer (though not more variable) than expected motor execution durations for button presses. These patterns suggest that parts of what the model treats as non-decision time are likely decisional in nature, although perhaps related to "action decision" rather than the "value-based decision" of interest to the authors. To what extent these two processes neatly follow each other or overlap could be usefully considered.

      We have added a paragraph to the Discussion explaining how our model’s estimates of sensory and motor latencies relate to corresponding values inferred from physiology or behavioral manipulations (e.g., Bompas et al., 2024). Specifically, we write:

      “The key assumption of the PDG model is that there is a delay between the moment a choice is internally committed and the moment it is externally reported with a key press. Because eye movements are typically faster than manual responses (𝜏<sub>e</sub> < 𝜏<sub>m</sub> in our simulations), this delay creates a window during which gaze can already be directed toward the covertly chosen item before the response is formally registered. We do not interpret these non-decision latencies as irreducible physiological minima for moving the eyes or pressing a button (Bompas et al., 2025). Rather, they are inferred indirectly by fitting an additive non-decision-time parameter to the behavioral data, which we decompose into a sensory delay (𝜏<sub>s</sub>) and a manual execution delay (𝜏<sub>m</sub>). Values of 𝜏<sub>e</sub> are then chosen so that the model reproduces the observed magnitude of the behavioral effects. This estimation procedure has important limitations. Some participants show relatively “flat” chronometric functions: response times vary little with value despite otherwise normal psychometric performance. Such patterns likely reflect processes not explicitly represented in the model, including procrastination, reduced motivation, task-unrelated thought, or noise in item ratings. Within a drift-diffusion framework, however, these cases are accommodated by assigning a long non-decision time together with a short evidence-accumulation period (Table S1). Consequently, some estimated non-decision times are substantially longer than would be expected if they represented only sensory and motor delays. A further limitation is conceptual. We model non-decision time as occurring either before or after evidence accumulation, whereas in reality decisional and non-decisional components are likely temporally interleaved (Graziano et al., 2011). This simplification may also inflate the recovered latency estimates. With these caveats in mind, sensory and oculomotor delays on the order of 300 ms remain broadly plausible, although they likely lie near the upper end of a realistic range. The estimated eye-movement latency is especially long. For instance, in monkeys trained to report simple perceptual decisions with a saccade, roughly 100 ms elapses between the threshold-crossing signal in parietal cortex (or the superior colliculus) and the executed eye movement (Roitman and Shadlen, 2002; Stine et al., 2023). Crucially, however, varying the assumed non-decision latencies across a reasonable range does not alter the qualitative predictions of the model (Fig. 8).”

      Further, we have added a parameter sensitivity analysis. Importantly, although the magnitude of the predicted effects depend on the non-decision latencies, the qualitative aspect of these predictions do not (new Figure 8). Specifically, (i) the increasing tendency to look at the ultimately chosen item as time elapses (new Fig. 8A), (ii) the lack of an interaction between the last-fixation bias and overall value (Fig. 8B), and (iii) the absence of an effect of choice consistency on Δdwell (Fig. 8C) are all findings that are independent of 𝜏<sub>e</sub>.

      Reviewer #2 (Public review):

      The paper focuses on analyzing the Krajbich 2010 data, but shows that the second effect replicates in many other datasets. A more principled approach, in which both effects are analyzed and presented for all datasets, would be more convincing. The results should then be shown together for clarity/readability.

      Following this suggestion (and the reviewer’s elaboration in the private comments to the authors), we have substantially restructured the manuscript. Both aDDM predictions are now presented together (new Fig. 2), and Figs. 3–4 test these predictions across multiple food-choice datasets. In doing so, we no longer treat the data from Krajbich et al. (2010) separately, and we extend the analysis of the last-fixation–choice association (MELFB) to additional datasets. We note that the same datasets could not be used in both Figs. 3 and 4, as some lack information on the final fixation required for the MELFB analysis. Nevertheless, results are highly consistent across datasets and align with findings from a recent study by Ting & Gluth (2025), which independently identified and examined one of our key predictions; this work is now cited in the revised manuscript. Finally, to reduce redundancy, we have consolidated all aDDM variants and optimal models into a single figure (new Fig. 10).

      Similarly, it would be nice to show to what extent the models' predictions depend (not depend) on using the best-fitting parameter values (are there any parameter settings under which the two effects are not predicted?)

      The key predictions of the model depend on the difference between the manual (𝜏<sub>m</sub>) and eye-movement-related (𝜏<sub>e</sub>) latencies. We have now added a parameter-sensitivity analysis to show how the model predictions depend on this difference. The new analysis shows that while the quantitative predictions do depend on the precise latency values, the results are qualitatively similar across values of 𝜏<sub>e</sub> (new Figure 8).

      Reviewer #3 (Public review):

      There was limited discussion about why one might allocate attention post-decision. I would have appreciated more discussion on the potential functional consequences or implications of post-decision gaze.

      Thank you for this suggestion. We added a new paragraph to the discussion (paragraph #2), where we argue that it is sensible for a decision maker to direct the gaze to the chosen item once a covert choice commitment has been made, as the benefits of attending to a stimulus do not end with the decision itself. Specifically we now write:

      “Instead, these observations are better explained by a post-decision account of the gaze-choice association that is, one in which gaze shifts to the selected item after a covert commitment to a choice. We argue that directing gaze to the chosen item after a covert choice commitment is sensible, as the benefits of attending to a stimulus do not end with the decision itself. In naturalistic settings, for instance, selecting a food item is typically followed by the action of reaching toward it, where visual attention supports spatial localization and motor planning for the upcoming action. Although participants in our computerized task did not physically act on their choices, these sensorimotor processes are likely highly automatized and may still be engaged by default, even when not strictly required. Beyond motor preparation, post-decisional attention may also serve additional functions, such as facilitating sensory anticipation of the reward, supporting metacognitive evaluation of the decision, and contributing to value updating for future choices. From this perspective, a degree of attentional “stickiness” whereby the chosen item remains preferentially attended after commitment could emerge as an effectively optimal policy once these post-decisional processes are taken into account. Moreover, a specific feature of the task design may further reinforce this tendency: in the snacks paradigm, the unchosen item typically disappears from the screen immediately after a response is registered. It is therefore plausible that directing gaze to the chosen item after commitment partly reflects anticipation of the imminent disappearance of the unchosen option. To disentangle these mechanisms, it would be interesting for future work to test whether this attentional bias persists when the chosen item, rather than the unchosen one, is the stimulus that disappears upon response.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments:

      (1) Framing of the modelling approach

      The manuscript would benefit from acknowledging the known limitations of DDM-based frameworks, especially given that the entire study is conducted within these constraints. The introduction highlights successes of the DDM, but the manuscript does not mention any of its conceptual or empirical limitations.

      We are unsure about what specific limitations the reviewer has in mind, but we have added a paragraph to discussion mentioning some limitations, like the inflation of the non-decision times and the difficulty of interpreting the fit parameters (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (2) Dependence on non-decision time assumptions

      The alternative model's explanatory power appears to rely heavily on assumptions regarding the decomposition of non-decision time: fixed visual encoding (𝜏<sub>s</sub>= 0.3 s), manual non-decision time (𝜏<sub>m</sub>; two free parameters), and saccadic execution (𝜏<sub>e</sub>; fixed parameters μ<sub>e</sub> = 0.35, σ<sub>e</sub> = 0.11).

      - 𝜏<sub>e</sub> is substantially longer and more variable than typical saccadic execution times, suggesting it likely incorporates decisional components.

      - Estimated 𝜏<sub>m</sub> values are approximately twice as long as known manual execution durations.

      - σnd is more plausible, implying that variability is captured correctly but mean durations are not.

      Together, these points raise the possibility that portions of what the model treats as non-decision time are in fact part of a (action) decision process. Only then does it make sense to assume that Tm is usually larger than Te. If Tm and Te were truly execution delays, then Tm would always be larger than Te.

      You may find it helpful to consider the framework in Bompas et al. Psych Review (2024), which discusses in detail what non-decision time is likely to comprise across effectors.

      Thank you we have added (i) a sensitivity analysis showing that our results are robust to changes in the specific value used for the eye movement related latencies (new Fig. 8), and (ii) a new paragraph in Discussion addressing the issue of the mismatch between our parameter estimates and the manual and saccadic execution times (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (3) Code availability.

      The authors should consider sharing all relevant code and data publicly.

      We agree, we now share the code and data on GitHub and indicate so in the revised manuscript.

      Minor Comments:

      (1) Lines 74-77. These are not worded as predictions but as questions; one tests predictions, but answers questions. I feel it would be clearer to stick to predictions (like in the abstract), and the introduction could benefit from explaining these predictions in a bit more detail (I found it difficult to get my head around these predictions from the intro text only).

      We rewrote the section in the introduction where we provide a gist of the model predictions (last paragraph of Introduction). We agree with the reviewer that the previous explanation was not clear.

      (2) It is confusing that panel B appears to the left of panel A in Figure 2.

      We agree. We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed and the panels follow a more logical order.

      (3) Figure 3C - remove MATLAB toggles.

      Yes, thanks.

      (4) Figure 5A shows the proportion of left choices, but the text and legend refer to right choices.

      Good catch, thank you.

      Reviewer #2 (Recommendations for the authors):

      This may appear self-serving, but the authors seem to be unaware of some highly relevant work from our group. Most importantly, in a recent publication (Ting & Gluth, 2024, JEP General), we have already looked at the dependency of the last- (or final-) fixation bias on overall value in value-based (VB) and perceptual (P) decisions. In VB, we found a negative effect; in P we did not find a significant effect. This is largely consistent with the current results, showing a negative but not significant trend. Another relevant work is Gluth et al. (2020, Nat Hum Behav), where we extended the aDDM by assuming that the probability to fixate on an option is a function of the accumulated evidence for that option. It would be interesting to know whether this assumption changes the predictions of the aDDM. Finally, we just published a new theory on how people search for information to make efficient value-based decisions (Gluth et al., in press, Psychol Rev; https://osf.io/preprints/psyarxiv/3qzak_v2). Although this theory focuses on multi-attribute choices, it can be applied to "simple" choices, too (by assuming that there is only one attribute = value). Interestingly, while the model also mispredicts a (slight) increase of the last-fixation bias with overall value, it correctly predicts the independency of the dwell-time advantage effect on choice consistency as well as the small increase of the effect with RT (attached here is a figure to show this: [https://elife-rp.msubmit.net/elife-rp_files/2026/01/22/00149589/00/149589_0_attach_9_477122. pdf], and the match with the empirical data shown in Figure 3B and 12 is striking). In general, the model shares many features of the Callaway and Jang models, but does not need to assume a biased value prior, which the authors suggest is responsible for the misprediction of the second effect. I leave it up to the authors to discuss this new theory, but I wanted to point this out.

      Thank you for pointing this out; these are all relevant points and studies.

      We now note that the first of our predictions has recently been identified and tested by Ting and Gluth (2025).

      We also considered extending the manuscript with a variant of the model proposed by Gluth et al. (Psychological Review, 2026). In fact, we attempted to fit this model to the Krajbich et al. (2010) dataset under the assumption that the duration of each sampling epoch is a free parameter. We find this model very interesting. However, in our current implementation it appears to make the same qualitative prediction as the aDDM, namely that ΔDwell depends on choice consistency (see Author response image 1).

      Given this, we have decided not to include these results in the manuscript. It remains possible that with further development particularly with a more realistic specification of fixation durations (e.g., allowing them to depend on value) the model could account for the full set of observed effects. We think this would be best addressed in a separate study.

      That said, we do find the model promising, as it provides a better account than most of the alternative models we explored for the patterns shown in panels D, H, and I.

      Author response image 1.

      Fits of a variant of the MACS model (Gluth et al. 2026) to the data of Krajbich et al. (2010).

      The paper would benefit substantially from restructuring. The aDDM's predictions are provided first, together with the empirical data, and then the optimal models are discussed. But Figure 2 shows all of this together. Later, the new (PDG) model is elaborated, and its predictions are shown. Towards the end of the results, variations of the aDDM and combinations of aDDM and PDG are shown in a series of figures (8-11), followed by a last figure showing one of the tested effects in other datasets. All of this feels pretty much thrown together without a clear structure. For instance, the aDDM and the optimal models could be described together (or the optimal models get a separate figure). The additive variants could be described earlier. And some figures could be put into the supplement. And the empirical results of the different studies could be shown together.

      We fully agree with this suggestion. We have now restructured the manuscript along the lines proposed by the reviewer (see the more detailed explanation of the restructuring in our response to the public comments).

      I strongly suggest avoiding the term "influence" in the y-axis of Figure 2, upper row, as it implies causality. Similarly, in line 182, the term "causal influence" is used in the context of the Callaway model, but as far as I know, this is not what the model assumes.

      We replaced the y-axis label with “Association of last dwell with choice (β)”

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 2 - Panel labels for A and B are reversed?

      We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed.

      (2) Does 3C include a .pdf screenshot?

      Thank you, it’s a Matlab bug on Mac. I guess they want us to switch to Python -:)

      (3) Figure 4 - It would be helpful if the green line were defined in the figure legend.

      Added

      (4) The effect size in 5B looks much more dramatic than in 2B(A?) - Is this for one example subject as opposed to all subjects? Please clarify what is different about the data.

      We are no longer showing the psychometric functions in Figure 2.

      (5) Line 252 - they say they compared the probability of choosing the right item (Fig. 5B) by the y-labels of that figure, which are all p(choose left).

      Yes, corrected now.

      (6) In general, they reference the subpanels of Figure 5 out of order, which causes the reader to jump around. They might consider reordering the panels of the figure so they follow the ordering of descriptions in the text.

      We agree, we have rearranged the figure panels to follow the ordering of the descriptions in the text.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      (1) Alternative mechanisms for performance differences.

      The authors assume that the difference in performance between the low-switch (LS) and high-switch (HS) frequency conditions is explained by a change in the "leakiness" of integration. However, several other mechanisms could potentially explain this effect:

      (1) Temporal Uncertainty: Integration might start later in the HS condition, leading to lower performance.

      (2) Reduced Efficiency: Integration could be less efficient in the HS condition (i.e., lower signal-to-noise ratio) without a change in the leak parameter itself.

      (3) Evidence Contamination: Motion information from the adapting stimulus in the HS condition may be integrated rather than ignored, which might be the case since the transition from the adapting to the test stimulus is not externally cued.

      To distinguish between these alternatives, I suggest two possible analyses. First, a formal model comparison could be performed, though I acknowledge this may be inconclusive in the absence of response-time data. Second, an analysis of motion energy kernels could be revealing; the leak hypothesis makes the specific prediction that for long test stimuli, early samples should contribute more to the choice in the LS condition than in the HS condition, relative to late samples.

      We thank the reviewer for raising these important points. We agree that we cannot definitively identify the algorithmic underpinnings of the behavioral effects we report and have made substantial revisions to the manuscript to be clearer about what is supported and what is speculative in our claims. Most importantly, we agree that we do not know if the context-dependent differences in how accuracy depends on viewing time are based on adjustments to a leak or to something else (e.g., a saturating non-linearity, as we identified in Glaze et al, 2015, that is separate from the leak itself), which we cannot resolve with this dataset, even with more formal model comparisons. We therefore:

      Changed the wording throughout the manuscript to refer to changes in leakiness as just one of several possible sources of the behavioral differences. We also added this point to the list of “limitations” (and possible future directions, including using motion-energy kernels, which would require us to use lower-coherence test stimuli) in the Discussion (L487-493).

      Added a new figure panel (Fig. 2D), a new Extended Data figure (Extended Data Fig. 3), and additional explanatory text (L168-175) that collectively describe the behavior in more detail, including quantifying a “crossover” dynamic similar to what we reported previously (Glaze et al, 2015).

      Added new explanations (L152-163) and analyses (Extended Data Fig. 9) indicating that the monkeys used some information from the end of the adapting stimulus to inform their decisions, which accounts for the patterns of choices at the shortest viewing durations.

      Indicate that the context-dependent differences in the slopes of the psychometric functions (and complementary analyses based on “raw” accuracy measures as a function of binned viewing duration) rule out the temporal uncertainty and evidence contamination explanations, but are consistent with effects on the temporal dynamics of the decision process (L175-179).

      (2) Independence of neural and pupil-linked signals.

      The authors take the lack of session-wise correlation between context-dependent contributions from neural and pupil terms as evidence that these two signals provide independent contributions to the behavioral effect. However, could this lack of correlation simply be a result of high variability or noise in these estimates? The data shown in Figure 7B suggests that measurements are very noisy, which might obscure a potential relationship.

      We agree that the lack of session-wise correlation between neural and pupil terms cannot be taken as definitive evidence of independence. We have both softened the language around the claim (L368) and added a sentence to the Discussion (L464-468) acknowledging that this lack of correlation may reflect underlying noise and/or variability rather than true independence of the underlying mechanisms.

      Reviewer #1 (Recommendations for the authors):

      (3) The neural data analyses rely fundamentally on "switch" trials (Figures 3-5). It might be informative to also examine "non-switch" trials to see if there are specific neural markers indicating the exact moment the motion stimulus becomes behaviorally relevant. Given that this may fall outside the primary focus of the paper, it is up to the authors whether to pursue this line of inquiry.

      We thank the reviewer for this suggestion. We agree and have added new analyses of data from non-switch trials (Extended Data Fig. 9), which show some effects of stimulus information from the adapting epoch on the monkeys’ choices, as we detail below in response to related comments from the other reviewers.

      Reviewer #2 (Public review):

      Aspects of the behavioral analysis would benefit from a tighter connection between theoretical claims about evidence accumulation and the empirical features of the psychometric functions. For example, the rightward shifts observed across adapting conditions are interpreted as consistent with a reset of accumulation on switch trials, but similar patterns could also arise from failures to detect the test stimulus on a subset of trials, leading responses to default to the final adaptor direction. Likewise, changes in psychometric slope and asymptote are attributed to differences in evidence accumulation without explicit modelling or consideration of alternative explanations.

      Clarifying how specific features of the psychometric functions map onto distinct components of the decision process will strengthen the link between the theoretical framework and the behavioral data.

      We agree and have made substantial revisions to address these important points. Specifically, we added a new figure panel (Fig. 2D), new Extended Data Figures (3 and 9), and several lines of explanatory text (L152-179) that collectively describe the behavior in more detail, including clarifying that: 1) for the shortest viewing durations, the monkeys’ decisions were informed by information from the adapting stimulus, which accounts for generally lower accuracy on LSF (longer exposure to the final adapting direction, thus more accumulated evidence for that direction before processing the switch) vs. HSF (shorter exposure to the final adapting direction, thus less accumulated evidence for that direction before processing the switch) switch trials; and 2) as viewing duration increased, the rate of rise of accuracy versus viewing duration was higher for LSF vs. HSF trials, implying differences in the process of evidence accumulation. As detailed in our response to a similar comment from Reviewer 1, above, we are now careful to temper our claims about the specific computational basis (e.g., a leak or other form of nonlinearity) for these differences.

      We also de-emphasized our treatment of the asymptotes of the psychometric functions. In principle, these regimes could give insights into leakiness (which can limit the total amount of information that can be accumulated) and lapses (which are measured at the asymptotes). In practice, however, the long-duration trials that constitute the asymptotes were relatively under sampled (to promote the unpredictability of the offset of the stimulus, which we believed was the more important consideration when designing the experiment), yielding unreliable estimates.

      A slight concern is the lack of a consistent analytical approach for relating behavioral changes to neural and pupil-linked measures. Different sections of the manuscript rely on different behavioral metrics-such as differences in accuracy within a selected stimulus-duration range (e.g., Figure 5C) or psychometric slope differences (Figure 6C) without clear justification for these choices. The analytical approach likewise varies between simple correlational analyses (Figure 5C, Figure 6C), pseudo-experimental group comparisons (Figures 5D, E), and the inclusion of neural or pupil terms in the behavioral psychometric regression model (Figure 7B). While each metric and approach may be defensible in isolation, adopting a more consistent framework will help convince readers that the reported effects are robust and not contingent on the selective choice of metric or analysis.

      We thank the reviewer for this thoughtful critique and agree that the rationale for our choice of behavioral metrics and analytical approaches could be stated more clearly. We have added text to the relevant sections of the Results (L247-251) clarifying these choices. In particular:

      The neural analyses (Figures 3D-E, Figure 4, Figure 5D-E) focused on preferred-motion switch trials, because: 1) low switch-frequency non-switch trials provide an additional 800 ms of exposure to the final adapting-stimulus motion direction relative to high switch-frequency non-switch trials, which confounds comparisons of context-dependent evidence encoding between conditions, and 2) MT neurons exhibit minimal responses to null motion (although note that we also included analyses based on ROC area, which is computed from both preferred- and null-motion switch trials, to account for possible contributions of null-motion responses; Figure 5A-C). Thus, to ensure a meaningful comparison between neural and behavioral measures, we used behavioral accuracy on switch trials as the relevant metric in Figure 5C-E, rather than psychometric slope, which is estimated across both switch and non-switch trials.

      The pupil analyses (Figure 6) focused on a time window preceding test-stimulus onset, representing the arousal state around when the decision process started, and included both switch and non-switch trials. Thus, for these analyses we used psychometric slope, which is estimated across both switch and non-switch trials.

      We used several different analyses to compare and contrast the neural-behavioral and pupil-behavioral relationships because they provide complementary and useful insights. The correlational analyses in Figures 5C and 6C characterize session-level relationships between neural/pupil signals and behavior. The group comparisons in Figures 5D–E provide a complementary visualization of the same relationship. The model-based approach in Figure 7 then allows direct quantification of the trial-wise contributions of each signal to behavior within a common framework. Importantly, the conclusions drawn from each approach converge on the same interpretation, which we believe speaks to the robustness of the reported effects.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 2 legend. Description of 'running average (5-trial window)' is unclear - presumably this is a running average in stimulus space rather than across trials.

      We thank the reviewer for flagging this ambiguity. We have updated the legend (L136-137) to clarify that the running average is computed across trials sorted by test-stimulus duration.

      (2) L158. Difficult to establish an asymptotic performance level for HSF conditions within the stimulus duration range tested.

      We have removed the reference to asymptotic performance and replaced it with a discussion of performance on longer-duration switch trials in the context of the newly added Figure 2D.

      (3) L515 Equation 1. While this is a standard formulation of lapse rate in psychometric functions, the construction here in terms of switch probability is not standard. Given the task and training, it seems more likely that on lapse trials, the animal will respond according to the last adapted direction (rather than randomly switch/stay with equal probability).

      We thank the reviewer for this point. We agree that it is possible that on at least some of the “lapse” trials the monkeys may respond according to the final adapting-stimulus direction rather than choosing randomly. However, we cannot distinguish those alternatives using this task design. We include a statement to this effect in Methods (L569-571).

      To explore the idea further, we refit the behavioral data using separate upper and lower asymptotes corresponding to lapse rates on switch and non-switch trials, respectively. Across monkeys, there were no significant differences between upper and lower lapse rates for either low (Wilcoxon signed-rank test for equal medians: p = 0.15, Cohen's d = -0.13) or high switchfrequency (p = 0.07, Cohen's d = -0.16) conditions. So, at the very least, there was no evidence for lapse-like errors driven by switch- (or non-switch-) specific defaults to the final adapting direction.

      (4) L256. Statistical significance of attenuation is not directly tested here.

      We have replaced "were attenuated" with "we did not identify any reliable context-stability differences" (L297) to accurately reflect what was directly tested without implying a statistical comparison between groups of sessions that was not performed.

      (5) L429. Does the increase in explanatory power warrant the increased complexity of the model here?

      We thank the reviewer for raising this important point. We used Tjur's pseudo-R<sup>2</sup> because it does not increase by default with added model complexity, making it more conservative than other R<sup>2</sup> measures in this respect. Tjur's pseudo-R<sup>2</sup> is a coefficient of discrimination, and as such its value increases only when additional terms improve the model's ability to separate predicted probabilities across response outcomes. Thus, the observed increases in explanatory power when adding neural or pupil terms reflect real improvements in discriminability rather than an artifact of model complexity. We have added a brief clarification of this point to the Methods (L662-664).

      Reviewer #3 (Public review):

      The task design may not be optimal. While the amount of time the monkey is exposed to each motion direction during the adapting stimulus is matched, it's hard to know if the reduced MT responses to the test stimulus are truly due to the greater frequency of switches during the HSF adapting stimulus or because the monkeys have been exposed to more repetitions of the stimulus. It's increased sensory adaptation in either case, but it makes it problematic to interpret this as temporal context-dependent adaptation specifically. I think this could potentially be partially addressed by an analysis that is in the paper, but could potentially be emphasized/fleshed out more, specifically the results shown in Figure 4D that seem to show that most of the reduction in neural response for adapting units occurs between the first and second stimuli.

      The reviewer raises an important point. The number of stimulus repetitions and switch frequency are confounded in the experimental design, making it difficult to attribute context-dependent differences in MT responses to the temporal pattern of switches rather than to accumulated repetitions. We also note, as the reviewer acknowledges, the observed differences reflect sensory adaptation either way. Figure 4D does offer relevant evidence, suggesting that a majority of the change in neural response occurred with just one stimulus repetition. This finding complicates an interpretation where adaptation scales with the number of stimulus repetitions. We have added several lines to the Results about these points (L231-233).

      The pupillometric analysis seems to be an indirect way of assessing whether the accumulator itself might be modulated by temporal context, but the link could be made clearer. The authors show that context-dependent behavior is related to pupil size, which is related to arousal/neuromodulation, but it would be helpful to have some idea of what neural mechanisms underlying adaptive decision-making are actually impacted by this neuromodulation. Lacking neural data to address this question (e.g., from a brain region proposed to be involved in the accumulation process), at least more discussion of this would be helpful. Essentially, I'm unsure of how to interpret the pupil results: the argument that temporal context affects instantaneous evidence encoding in MT that then drives the accumulator is very clear, but I am a bit confused about what, mechanistically, I should think about the effect of neuromodulation doing.

      We thank the reviewer for this thoughtful comment and agree that the mechanistic interpretation of the pupil results could be made clearer. We acknowledge that we cannot directly identify the neural mechanisms underlying the arousal-related contributions to adaptive evidence accumulation from pupil data alone, given that pupil size is an indirect and imperfect proxy for neural (e.g., LC-NE system) activity. However, we can offer some informed conjecture and have added to the Discussion (L469-482) in an effort to elaborate on possible mechanisms.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract could be retooled - does not emphasize the pupillometry/arousal results very much, and they are presented more as a control than an independent result.

      We agree and have revised the Abstract accordingly.

      (2) Do all neural/pupil analyses use only switch trials? Sometimes the figure captions do specify only switch trials, but not everywhere. It would be helpful to specify either in the Methods or at the beginning of each figure caption that all subplots show switch trial results. Also, if you do always use switch trials, it would be useful to see in the Supplement how the non-switch trial results differ from switch trials. It seems like they may in interesting ways based on the behavioral results (supporting a reset of evidence accumulation on switch but not non-switch trials).

      We thank the reviewer for flagging these important points. We have added a justification for switch trials (L186-190) as well as clarification about which trial types were used for which analyses (L246-249) and information about trial types to relevant figure captions. We have also added a new Extended Data figure (Extended Data Fig. 9) examining relationships between neural activity and behavior on non-switch trials. As inferred by the reviewer, behavior on non-switch trials is consistent with the use of information from the adapting stimulus.

      (3) In Figure 3C, 5B, etc, when computing firing rate for the test stimulus (50-500 ms), are differently sized windows used to compute the rate for different test stimulus durations (since some will be <500 ms)? Or are only trials where the test stimulus duration is > 500 ms used for this analysis?

      We thank the reviewer for raising this point. To clarify, the 50–500 ms window does not reflect a fixed window applicable for all trials. Rather, neural activity from 50 ms after test-stimulus onset through test-stimulus offset was included for each trial, with 500 ms serving as the upper bound for trials with longer durations (> 500 ms). We have clarified this in the Methods (L607-610) to avoid ambiguity.

      (4) I think it might be better to be consistent with the time windows used for analysis; specifically, to choose either the 50-500 ms window used in Figures 3, 4, and 5B, or the 200- 400 ms window used for the remaining analyses in Figure 5.

      We agree that using the same window for all of the analyses would improve consistency, but not doing so provides advantages that we believe take precedent and now describe in more detail. The broader 50–500 ms window used for Figures 3, 4, and 5B was chosen to characterize MT neural activity over a relatively large a time window, ensuring that every trial contributes to each estimate. Because test-stimulus durations were drawn from a truncated exponential distribution (100–1200 ms), restricting these analyses to the 200–400 ms window would have excluded the substantial proportion of trials with durations <200 ms (but would yield similar figures and conclusions). The narrower window used in subsequent analyses allows us to focus on the conditions that exhibited the biggest modulations of neural activity when comparing them to behavior.

      (5) Similarly, provide justification for using only trials ending 375-600 ms after test stimulus onset for the behavioral correlations. It seems reasonable to choose a subset of test stimulus durations where the monkeys' behavior is greater than chance but less than ceiling, but it would be good to specify this so that it doesn't seem arbitrary.

      We agree and have added text to make this important point (L249-251).

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Okabe and colleagues build on a super-resolution-based technique they have previously developed in cultured hippocampal neurons, improving the pipeline and using it to analyze spine nanostructure differences across 8 different mouse lines with mutations in autism or schizophrenia (Sz) risk genes/pathways. It is a worthy goal to try to use multiple models to examine potential convergent (or not) phenotypes, and the authors have made a good selection of models. They identify some key differences between the autism versus the Sz risk gene models, primarily that dendritic spines are smaller in Sz models and (mostly) larger in autism risk gene models. They then focus on three models (2 Sz - 22q11.2 deletion, Setd1a; 1 ASD - Nlgn3) for time-lapse imaging of spine dynamics, and together with computational modelling provide a mechanistic rationale for the smaller spines in Sz risk models. Bulk RNA sequencing of all 8 model cultures identifies several differentially expressed genes which they go on to test in cultures, finding that ecgr4 is upregulated in several Sz models and its misexpression recapitulates spine dynamics changes seen in the Sz mutants, while knockdown rescues spine dynamics changes in the Sz mutants. Overall, these have the potential to be very interesting findings and useful for the field. My major concerns from the initial manuscript, especially regarding cherry picking and circularity have been addressed with revised analytical approaches. I have some remaining minor comments.

      (1) The comparison between two wild-type samples versus wild-type-mutant samples is helpful - I think this could be added to the manuscript.

      As suggested, we added the figure comparing two wild-type samples against wild-type mutant samples as Supplementary Figure 2. 

      (2) For results of time-lapse imaging - please spell out in the results section the direction of change (lines 270 - 277).

      As suggested, we added the direction of change (an increase in the turnover rate) to the text (page 12, lines 270-271).

      (3) Using linear mixed effect models for statistical analysis is a significant improvement. While a sample size (n) of mice = 3 is not ideal, I think given the multiple different mouse lines used and intensity of analysis, this is probably the best that can be done, although further validation in larger samples eventually is to be hoped for.

      We appreciate the reviewer for recognizing the effort required to collect data across multiple mouse lines.

      (4) The revised text is much improved, but I still think the authors should be upfront somewhere in the text that the schizophrenia-associated genes can only confer biased risk for schizophrenia (and that the clinical phenotype can also include autism). As I said before, I think this is the best we can do and I agree with their choices, but it is important not to overstate the link. The differences they see make it clear that these are still relevant distinctions.

      As suggested by the reviewer, we further modified the discussion related to the comparison between ASD- and schizophrenia-associated mouse models (pages 23-24, lines 508-522).

      “The nanoscale features of dendritic spines in mouse models of Nlgn3<sup>R451C/(y or R451C)</sup>, Syngap1<sup>+/−</sup>, POGZ<sup>Q1038R/+</sup>, and 15q11-13<sup>dup/+</sup>, which we classified as being related to ASD, are highly heterogeneous. This heterogeneity may reflect the broad clinical spectrum of ASD, which ranges from mild impairments in social skills to severe intellectual disability. Accordingly, these four mouse models may represent distinct subgroups characterized by different degrees or forms of hippocampal dysfunction. Notably, among the ASD-related models, 15q11-13<sup>dup/+</sup> showed population-level spine properties closer to those found in the 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> mouse models. Although we classified 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> as schizophrenia-related models, both 22q11.2 deletion syndrome and Setd1a haploinsufficiency in humans are also associated with ASD, suggesting substantial overlap in the genetic risk factors underlying ASD and schizophrenia. Further systematic analyses linking rare genetic variants to synaptic phenotypes in mouse models may provide important insights into the mechanisms underlying both shared and disorder-specific synaptic alterations in neurodevelopmental and psychiatric disorders.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) I would suggest that it might be preferable to use the word 'neuropsychiatric' rather than 'mental' in the title.

      As suggested, we modified the manuscript title.

      (2) I think it would be clearer to say that DEGs are listed if present 'in three or more models' rather than >2 (I appreciate the latter is mathematically clear, but can easily be read as 2 or more if reading fast). This is changed in the figure legend, but I suggest it is also changed in the main text (line 352-3)

      As suggested, we changed the main text to incorporate "in three or more models" (page 16, line 352).

      (3) Please add to Methods (line 557) that 'control cultures were prepared from littermate embryos....'

      As suggested, we added the phrase "control cultures were prepared from littermate embryos" (page 26, line 559).

      (4) Sorry to add something, but please could the authors add a definition of how they calculate spine turnover (and add units to the y axis of Figure 5A-C)?

      As suggested, we modified the y-axis of Figure 5A-C (% as unit) and added the method of calculating spine turnover rate in the text (page 36, lines 808-811).

    1. Author response:

      We would like to thank the editors for their interest in our work and the three referees for their time and careful reading of the manuscript. The reviewers have provided a series of helpful suggestions that we discuss in this provisional reply and will seek to address in the revised version of the manuscript.

      The main concern raised is that the bacterial community we refer to as a biofilm may instead correspond to a cell aggregate. Following the passing of Prof. Kevin Wood, in whose lab the experimental work was carried out, our ability to perform additional experiments is limited. Nevertheless, we plan to wash and fluorescently stain the extracellular matrix before imaging to measure the extent to which the observed bacterial community is an attached biofilm. In the meantime, we would like to highlight the work of Wen Yu et al. [1], in which E. faecalis biofilms were grown in 96-well plates under antibiotic stress. In particular, one of the strains of E. faecalis used in this article was OG1RF, the same strain used in our study. Crystal violet staining was used to quantify biofilm biomass, providing evidence for biofilm formation under those conditions. While we recognize that the experimental setup differs from ours and that the OG1RF sample used did not contain fluorescent and resistance plasmids, these results nevertheless support the expectation that OG1RF will readily form biofilms.

      Reviewer #1 (Public review):

      The mechanistic interpretation could, however, be clarified further by more explicitly emphasizing the competing timescales associated with detoxification, growth, and resource limitation. The current results suggest that when resistant cells are initially abundant, detoxification occurs rapidly relative to growth, allowing the population to approach carrying capacity after relatively few doublings, whereas slower detoxification at lower resistant fractions may permit greater expansion of sensitive cells once antibiotic concentrations decline. Additional direct measurements of antibiotic concentrations over time would also strengthen the connection between the experimental system and the modeling framework by testing whether the detoxification dynamics assumed in the model are quantitatively appropriate, although this seems very plausible.

      The timescale of drug degradation is an important system metric. We appreciate the referee’s suggestion to quantify antibiotic concentration over time. We plan to perform experiments in which samples are collected from the culture at fixed time intervals. After removing the bacteria from the samples via centrifugation, serial dilutions of the supernatant will then be spotted on a lawn of sensitive cells to measure the antibiotic efficacy at each time point.

      Reviewer #2 (Public review):

      One clarification the author should make is on the biofilm growth process. Specifically, could staining experiments be performed to demonstrate the secretion of the extracellular matrix? Just by looking at Figure 1b, it is hard to say. It remains a question whether the biofilm culture simply contains unstructured clusters rather than real biofilms (that are usually structured).

      We agree with the referee that additional evidence would strengthen our study. As noted above, we will perform additional experiments to demonstrate the presence of an attached biofilm.

      Reviewer #3 (Public review):

      The observed results are tied very closely to the experimental setup of adding antibiotics very close to the time of inoculation, but this connection is not discussed. [...] The mathematical model is used to confirm the result that no spatial components are needed to describe the results; however, this is mostly linked to the initial setup of the experiment, where antibiotics are added at the time of inoculation, and no biofilm could form before the outcome of the antibiotic-cell interactions was concluded.

      The experiment was designed to address how coupled planktonic and biofilm populations develop in the presence of antibiotics, which we will more explicitly discuss in the revised manuscript. We do agree that investigating how mature biofilms and their planktonic populations respond to antibiotic stress is an exciting direction for future studies. However, we believe that is beyond the scope of the our study on the development of coupled populations. We will be sure to explicitly identify this limitation in our revisions.

      The described ‘population inversion’ effect is better described as frequency-dependent selection for resistant cells, but frequency-dependent selection is not discussed.

      The reviewer is correct that this ‘population inversion’ is a frequency-dependent (perhaps also density-dependent) effect, and we should have situated it within that broader ecological framework. We will use this terminology in our revisions. We do want to acknowledge that the late Dr. Kevin Wood was fond of this phrasing to describe the reversal of the dominant strain, which is not necessarily true for frequency-dependent effects. Although we do not know for certain, we suspect this was a play on the ‘population inversion’ term used in quantum physics, used to describe a system in which its excited state (high energy) population unexpectedly outnumbers its ground state (low energy) population.

      The authors claim that biofilm and planktonic bacteria are protected equally by the presence of resistant bacteria; however, Figure 1a and b seem to clearly show that the proportion of sensitive cells is higher in the planktonic cells compared to biofilm cells when started from an equal frequency inoculum, meaning this is not always the case.

      If the reviewer is indeed discussing Figures 1a and 1b, these are not comparable as the starting fractions differ. On the other hand, if the reviewer was talking about Figures 2a and 2b (which is more clearly discussed by looking at Figures 2c and 2f), we agree that it appears that planktonic communities tend to have a slightly greater frequency of sensitive cells than the biofilms. We will be sure to highlight this observation and possible explanations in our revisions. However, given the uncertainty in these observations, we do not believe the differences are sufficient to alter our overall conclusion that final resistant fractions in biofilm and planktonic populations are quantitatively similar. Furthermore, the no drug treatment shows the same trend, which suggests it’s an effect of different growth dynamics of these two strains at high density rather than driven by the protective effects of resistance cells.

      Confocal microscopy was used to quantify the relative proportion of antibiotic-resistant and sensitive cells in the biofilm; however, it is unclear if the entirety of the Z stacks was used to determine these proportions. This is also the case for the analysis of whether the sensitive/resistant cells are non-randomly distributed in the biofilm: it is unclear whether the vertical distance between cells was taken into account.

      The entirety of the Z stack was used to measure the final resistant fraction in the biofilm. On the other hand, we used only the densest slice of the Z stack to calculate the correlations. The correlations follow the same trend when calculated over less dense slices, but as density decreases, noise increases, so such plots did not bring more clarity to our conclusions and were not included in the manuscript. Additionally, only horizontal correlations (over a slice) were calculated because consecutive Z-stack slices were imaged with a 2.5 µm spacing. Given that the average cell diameter is approximately 1 µm, calculating vertical correlations may miss neighboring cells located between imaged slices, making such measurements unreliable. We will clarify the points raised by the reviewer in the results section and add more detail to the imaging methods section in the revised manuscript.

      References

      (1) Wen Yu, Kelsey M. Hallinen, and Kevin B. Wood. “Interplay between Antibiotic Efficacy and Drug-Induced Lysis Underlies Enhanced Biofilm Formation at Subinhibitory Drug Concentrations”. In: Antimicrobial Agents and Chemotherapy 62.1 (Dec. 2017), 10.1128/aac.01603–17. doi: 10.1128/aac.01603-17. url: https://journals.asm.org/doi/10.1128/aac.0160317 (visited on 01/11/2026).

    1. Author response:

      We thank the reviewers for their careful reading of our manuscript and for providing positive, constructive feedback. In particular, we thank the reviewers highlighting the several strengths of our study.

      To address the reviewers’ major concerns, we will revise the presentation of our main findings (specifically data/animal vs data/ROI), provide more clarity in the Results, Methods, and Discussion sections, and modify the title to better reflect these nuances.

      Additionally, we will perform the following new experiments:

      (1) Astrocytic Marker Validation: To further confirm comparable astrocyte cell counts between the CTRL and Upf2-cKO conditions, we will perform immunostainings using Aldh1L1, Sox9, or S100b instead of GFAP.

      (2) NMD Candidate Validation: To validate top candidate NMD target transcripts, we will perform immunostainings or qRT-PCR for Gabbr2, Adora1, S100b, or Cldn9.

      (3) Sample Size Expansion: To strengthen the morphological and PSD-95 quantifications, we will increase the sample size (N) by incorporating additional animals.

      (4) Mechanistic Timeline & Phenotype Linkage: We value the reviewer’s comment regarding the timeline of morphological and Ca<sup>2+</sup> phenotypes. To gai insight into whether these phenotypes are independent or linked, we will perform 3D reconstructions in CalEx conditions to assess whether Ca<sup>2+</sup> restoration rescues astrocyte morphology in Upf2-cKO mice. This will allow us to determine if increased Ca<sup>2+</sup> activity is upstream of the morphological alterations. Taken together, we believe that incorporating these manuscript revisions will strengthen the clarity and conclusions of our work. We thank the reviewers for their time and careful evaluation of our study.

    1. Author response:

      We sincerely thank the editors and reviewers for their overall positive assessment and constructive feedback on our manuscript detailing the nanoscale organisation of βII-spectrin of the membrane-associated periodic skeleton (MPS) in mouse sciatic nerve axons. Their perspective and comments will help refining the manuscript.

      A common comment by the reviewers relates to the description of the characteristic longitudinal periodicity of the MPS. We value these comments, which we believe are motivated by the fact that the longitudinal periodicity of the MPS is undoubtedly the most studied and prominent feature of the MPS in cultured neurons. However, the main goal of the present project was to describe how βII-spectrin is organised in the transverse axis of individual segments of the MPS in nerve tissue. This is why we utilised cross-sections of the sciatic nerve, hence achieving the best resolution possible in that plane, at the expense of the resolution in the axial axis. Furthermore, this study clearly shows that the transverse morphology of axons, and thus of the MPS, of neurons in the tissue is highly irregular, in comparison to cultured neurons. This imposes an extra challenge to observe correlated longitudinal structures when the observation length is limited, as in our studies. Nonetheless, to improve this aspect of the manuscript, we will revise our data and previous evidence, clarify the methodological trade-offs made, and make our interpretations more accurate.

      Additionally, we will clarify several imaging- and definition-related inquiries, including tests for insufficient staining, the interpretation of βIII-tubulin staining, the assessment of axon–glia boundaries, and consistency in the use of terms like ‘clusters’ and ‘periodicity’, among others.

      We believe these and other revisions will substantially strengthen the manuscript and comprehensively address the reviewers' feedback.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In the wild, bacteria can be found in a wide range of metabolic states, including states in which they are resource-limited. Because phages heavily rely on the infected cell's molecular machinery to replicate, it is natural to wonder how phage-bacteria interactions depend on the metabolic state of the cell. In this work, Marantos et al. investigate specifically how the rate of infection of 5 different phages changes between cells grown in energy-rich conditions and cells grown in energy-depleted conditions. Their results clearly show that 4 out of the 5 phages studied display a significant reduction in infection rate in cells that are energetically depleted and provide a potential explanation for this observation by looking into the mechanisms that these phages use to irreversibly infect their host cells.

      The work also tries to explain the observation using a mathematical/mechanistic model that describes infection as the sequence of two steps, where a phage first needs to bind to a cell receptor, from which it can potentially unbind, and then irreversibly infects by injecting its genome. While the model is sensible from a mechanistic perspective, the experimental evidence that supports how each model's rate is affected by the cell metabolic state is weak, as only ratios of these rates can be inferred from the data.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate the dependence of phage adsorption rates on host metabolic state, using 5 coliphages that differ in their infection cycles and host receptors. They find that four of the 5 phages showed significantly reduced infection under low metabolic states, with phages that generally have weaker adsorption being more strongly affected by low metabolism. The authors complement their findings with a 2-step infection model where phages can disengage from their hosts after initial adsorption. The paper illustrates the power of standardized experimental protocols for quantitative trait comparisons and highlights the dependence of phage infection success on host physiology.

      Strengths:

      The paper is well written and clearly structured.

      The experiments are well-designed, and particularly commendable is the diligent use of control scenarios to allow for quantitative comparison between phages. This standardized protocol will be valuable for the entire phage community.

      The authors convincingly show the impact of host physiology on phage adsorption success. This dependence has so far mainly been considered for intracellular phage replication, and the paper shows that host physiology has to be taken into account at all steps of phage infection.

      Weaknesses:

      There are some concerns about the experimental setup and which conclusions can be drawn from it:

      Before phage infection, bacterial cultures are grown to exponential growth, washed, and then resuspended with glucose or arsenate-azide for 10min. It is however, questionable that 10 minutes is enough to simulate high and low metabolic states realistically. 10 minutes seems to be quite short to go from exponential growth to a low metabolic state, given the transcriptional memory of previous environments. It seems more likely that the population will be quite heterogeneous, with cells in various states of transition towards low metabolic states.

      While we agree with the reviewer that during metabolic transitions there may be a period in which the population is heterogeneous, with cells in different stages of transition toward a low metabolic state, the 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyper diffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027). We have also corrected the DOI for this reference in the manuscript. Furthermore, the ATP pool of log-phase E. coli turns over several times per second (Holms et al., Arch. Mikrobiol. 1972, http://dx.doi.org/10.1007/BF00425016). We therefore assumed the bacteria were energy depleted after 10 minutes. We have clarified this point in the revised manuscript.

      Given that arsenate and azide inhibit cellular metabolism, i.e., have antimicrobial effects, cells might not just downregulate metabolism but also activate the stress response, and this causes some of the observed effects on phage adsorption. Therefore, the 'low metabolic state' of the cells in this paper could mean that cells are starved or that they are stressed or both.

      The reviewer is correct. We don’t exclude indirect effects. However, as nutrients were removed from the bacteria by washing and energy metabolism was inhibited by the addition of arsenate and azide, we assumed a stress response requiring biosynthesis would be unlikely to occur.

      The abundance of receptors could change between the high and low metabolic media conditions and contribute to the observed differences in adsorption, while the authors seem to assume in their model that the initial adsorption rate always remains the same.

      We do not think that the observed differences in adsorption are explained by a change in receptor abundance. In a previous study using the same experimental protocol as in the present work, phage λ was compared to the metabolically insensitive mutant λh (Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119). If the lower adsorption in the low-metabolic condition were caused by a reduced number of receptors, then λh should also have shown a lower adsorption rate under the same condition. Instead, no measurable effect on λh adsorption rate was observed. We therefore conclude that the effect is not explained by changes in receptor number on the timescale of the experiment. We have clarified this point in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      Marantos et al. showed that for some coliphages, the energetic state of the bacterial host cell has a strong impact on whether phage infection is initiated. The authors drew this conclusion from the observation that there are more free phages remaining in the medium after infection of arsenate-azide-treated cells as compared to after infection of untreated cells. These data were analyzed and reported both as ratios of the treated vs. untreated conditions and using a mass-action kinetic model of phage-cell collision in the infection mixture. The data supported the findings that for four phages infecting Escherichia coli bacteria, namely, phages λ, ɸ80, m13, and T6, the phages are less likely to initiate infection if the host bacteria are energy-depleted. However, for phage T5, the authors found that their infection propensity is not impacted.

      Strengths:

      The data presented by the authors clearly supported the principal conclusion of the study ("Viral commitment to infection depends on host metabolism"). The five phages chosen by the authors represent different viral lifestyles and infection mechanisms, highlighting the potential applicability to other Escherichia coli phages. Finally, the authors successfully used a classic mass-action model of phage-cell collision to interpret their data. The simplicity of their experimental assay, combined with the use of this mathematical model, offers other investigators who study phage-bacterial interactions in other contexts a potentially useful toolkit to examine infection in general, and specifically, the dependence of phage infection on the host's metabolic state.

      Weaknesses:

      (1) The authors isolated and measured the numbers of free phages in the medium after infection of bacteria under different treatments. These measurements were analyzed in two different ways: (1) simply as ratios (corrected/normalized using different controls), and (2) fitted using a simple mathematical model. I have concerns regarding both analyses.

      (1.1) For the first method, having different time points at which the sample of each phage is collected critically complicates data interpretation. As one incubates the phage-bacteria mixture for a longer time, more infection occurs, and the number of phages collected from the mixture decreases. Therefore, the different incubation time forfeits the goal of "a systematic and quantitative comparison across different phages [...]", just as the authors self-criticized. Conceivably, the authors could have used the shortest measurement time for all phages (i.e., 10 minutes, as for phage λ). Alternatively, the authors could have applied a systematic criterion such as half (or any other fraction) of the latent period of each phage, which would still "maximize the incubation period while ensuring that manipulations were completed before the first infection cycle concluded". In my view, the seemingly arbitrary measurement time for each phage renders the entire first analysis very challenging to interpret. It also goes against the author's proposition that the protocol was "standardized" or "consistent". It is not clear what the readers are supposed to take away from this first analysis, or rather, which evidence, finding, or conclusion the manuscript would lose if the authors only presented the modeling-based analysis.

      (1.2) The second method of analysis sought to remove the dependence of the measurements on time. I completely agree with this goal, and the findings extracted from this analysis significantly contributed to the merits of this manuscript. However, the authors achieved this goal using a single time point for each phage to calculate the infection rate (η). As shown in Figure S3, each of the phage depletion curves is anchored by only one data point (note that the P(t)/P(0) = 1 at t = 0 is assumed, not measured). This goes against the typical way this collision model is used in the literature, where a time series is measured and used to fit the model (e.g., DOI 10.1007/978-1-60327-164-6 18, or more recently, PMID 39700139). This practice in the current manuscript reduced the robustness of the inferred η values. This problem is exacerbated by assumptions used by the authors in formulating this model. For instance, the authors used a constant value for the bacterial concentration, B, because "bacterial growth and lysis were negligible" (lines 135-136). However, considering that the bacteria were cultured at 37oC in a very rich medium (first in YT broth, then in 2% glucose), the measurement times of 20, 30, and 55 minutes are most likely one or a few generations of bacterial growth and division.

      Related note: I suggest that one of the panels in Figure S3 should be moved to the main text, since it is critical to the second method of analysis.

      We would like to clarify that the manuscript does not present two separate methods, but rather one method presented in two steps: a first step with results that are directly tied to the experimental measurements and show whether the effect is present for each phage, followed by a second, analytical step that makes the results comparable across phages.

      The first step presents the ratios because they directly reflect the measurements performed in the experiment and allow the reader to see the effect of the metabolic state for each phage in contrast to its control. We agree that these ratios are time-dependent and therefore not suitable for quantitative comparison between phages. Their purpose is to illustrate the experimental outcome and to show that the effect is present (or absent) on a per-phage basis not to compare magnitudes across phages.

      We then follow this with the second step, allowing the reader to follow the logic of the analysis. The analytical step that follows does not represent a second method, but a continuation of the same analysis. Here, we remove the time-dependence specifically in order to make comparison of the effect across phages possible, by connecting our results to standard measures such as the adsorption rate η. Importantly, P(0) is measured for every phage in every experiment. The only modeling assumption used (a standard one in the field) is the exponential form for the decay in free phage number, which naturally yields P(t)/P(0) = 1 at t = 0.

      Regarding the reviewer’s concern that bacterial growth may not have been negligible over the relevant time window, we note that recent work on rich-to-minimal growth lags in E. coli reports substantial delays before growth resumes after nutrient downshift. One 2023 study (Wu et al., Nature Microbiology 2023, https://doi.org/10.1038/s41564-022-01310-w) considering wild-type E. coli shows in Fig. 2c a lag of up to about 2 hours after a shift from MOPS minimal medium with 0.2% glucose plus 18 amino acids to the same medium without amino acids. Another 2023 study (Zhu and Dai, Nature Communications 2023, https://doi.org/10.1038/s41467-023-36254-0) examining both rel+ and rel− strains reports a growth lag of about 49 minutes for rel+ and more than 5 hours for the relA deletion strain. While these conditions are not identical to ours, they support the general point that growth does not immediately resume after such shifts. We therefore think it is unlikely that, following transfer from YT, the cells underwent one or a few full generations during the time window of our adsorption measurements.

      On the related note: Following the comments of all reviewers on Figure S3, we have decided to remove it to avoid confusion.

      (2) The data were able to distinguish phages that successfully infected bacteria and those that remained free in the medium, and the authors appropriately interpreted the data as such throughout the Results section. However, in the Discussion (starting from the very first sentence, line 172), the authors used terms that include "adsorption" and "entry" more interchangeably (for example, see the three sentences in lines 310-313, for "viral entry efficiency is shaped by [...]", then "adsorption kinetics modeling"). I do not see how the authors' data could distinguish between adsorption (the phage particles attaching to the outside of the cell) and entry (the phage DNA being injected into the cell). Conceivably, any phage particles that irreversibly attach to a cell but do not yet inject their genome into the cell would still be removed from the medium and therefore not quantified. Another example: in lines 189-191, the authors interpreted that "[...] when the bacterium is in a low metabolic state, the phage does not bind irreversibly to the host", but how do the authors eliminate the case of no phage binding (i.e., the reversible step) to begin with?

      We agree with the reviewer that our use of the terms adsorption, entry, and infection should have been more careful. Our experiment can only identify the irreversible commitment of phage to a host cell. We have therefore revised the text to refer consistently to phage commitment.

      Similarly, in lines 283-293, how do the authors delineate whether energy depletion would increase the k_off term or decrease the k_inj term, because either would result in more free phages in the medium as observed in the data? I believe that the writing of the Discussion, as it stands now, is doing a disservice to the conclusions presented in the Results section.

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: if energy depletion leads to reduced commitment, this can arise either because k_off increases, because k_inj decreases, or because both change, as long as k_off/k_inj becomes larger in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also leads to the trade-off now discussed in the revised manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become significantly larger in the inactive case. Depending on whether this is achieved through changes in k_off or k_inj, the cost of discrimination appears either as slower commitment or as additional energy dissipation. We agree that the previous wording overstated the mechanistic interpretation, and we have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (3) The authors presented an argument that performing infection of all five phages in the same condition is an advantage, allowing for comparison across different phages. While this goal is a completely valid one, it is difficult to reconcile that with the fact that different phages require different optimal conditions for successful infection. For instance, phage T5 famously requires Ca2+ for successful infection into the host bacterium (and later successful replication); see PMID 13174489. However, all infections were performed in TMG, which lacks Ca2+. Perhaps the absence of T5 dependence on the host metabolism is because the infection condition used by the authors was not optimal for T5 to begin with? Similar arguments could be made for other phages.

      Our study alone cannot eliminate that possibility. However, we have cited multiple previous studies, for example references citing Braun et al., showing that T5 remains insensitive to the host metabolic state under different buffer conditions. We therefore believe it is unlikely that the lack of metabolic dependence we observe for T5 is simply due to suboptimal infection conditions.

      (4) Whereas the manuscript examined five coliphages, only phage T5 and phage λ were discussed extensively. I believe some discussion points for these two phages need clarification.

      We focused our discussion on the phages T5, λ and φ80 because these are the phages for which similar effects have been reported previously in the literature. This allowed us to connect our findings directly to existing work and to discuss mechanistic hypotheses in a meaningful comparative framework. For the remaining phages, to our knowledge no prior studies have examined their behavior under comparable metabolic conditions, and therefore a similarly detailed discussion would have been speculative. Nevertheless, all five phages are treated equally in the presentation of the experimental results and in the quantitative comparison of adsorption rates.

      (4.1) Phage T5: The data obtained by the authors show that the infection rate of phage T5 is not impacted by the metabolic state of the host cell. Considering that the authors used the terms "infection", "adsorption", and "entry" interchangeably to refer to the irreversible commitment of a phage to a host cell (see point 2), this discussion regarding phage T5 lacks one critical literature context: DNA entry of phage T5 is known to occur in two phases (first-step transfer and second-step transfer). Critically, the second step can only occur if phage proteins encoded by the phage DNA transferred in the first step are expressed (see PMID 10577483 and the cited papers therein). In that context, metabolic poisoning of the host bacteria should have impeded T5 infection. The authors should comment on this point.

      As the reviewer pointed out, our usage of the terms infection, adsorption, and entry should have been more careful. Our experiment can only identify irreversible commitment of phage to a host cell. For T5, we expect that this irreversible commitment already occurs upon first-step transfer of phage DNA. As a result, even if second-step transfer is impeded under metabolic poisoning, our method would not resolve that effect. We have added this clarification to the revised manuscript.

      (4.2) Phage λ: The experiment using phage λ in this current study shares many resemblances to that in Brown et al. 2022. That feature alone is not a problem, but at many places in the text, the writing is ambiguous as to whether it is discussing the results in Brown et al. 2022 or in the current manuscript. I am giving three examples below, but this is not exhaustive: (i) Lines 67-69, there is no Brown et al. 2022 reference immediately after "a mutant phage variant (λh) could bypass this dependency [...]" (not just in the previous sentence); (ii) Line 228 should clearly say "Our previous findings suggested that phage λ is capable of [...]", since it concerns Brown et al., 2022, not the current study; and (iii) Lines 245-246, there is no Brown et al., 2022 reference immediately after "we observed that a mutant variant [...] even energy-depleted host" (without a reference, it reads like the authors "observed" that finding in this current manuscript).

      The reviewer is right. In those places, the text was ambiguous as to whether it referred to the present study or to Brown et al. (2022). We have now inserted the reference at the relevant points and revised the wording where needed to make this distinction explicit.

      Also, regarding phage λ: The discussion between line 230 and line 249 is very interesting, but since it concerns the differences between λ PaPa and Ur-λ, the authors should consider mentioning and discussing a very relevant recent study, PMCID: PMC6312755.

      We agree that the study by Guan et al. is very relevant and interesting. However, our point in this part of the Discussion is only to clarify that we used λ PaPa and not the originally isolated λ strain. We have therefore limited the discussion here to that distinction.

      (5) Control experiments, or references to prior studies, are needed to support that the As/Az treatment at this concentration and duration (at least 10 minutes) is sufficient to deplete the metabolic state of the cell. For instance, this can be shown by impeded or null cell growth, arrested motility (using a standard swimming assay), or a fluorescent reporter for the energetic state of the cell.

      The 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyperdiffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027) where the effect was assessed by monitoring the rate of movement of the λ receptor on the bacterial surface. We have clarified this point in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As mentioned earlier, I found the paper interesting and addressed an important and significant knowledge gap.

      My biggest concern is about the interpretation of the experimental data in light of the two-step model. In particular, around line 286, it is stated "k_inj is more sensitive to metabolic state than k_off". Assuming k does not depend on metabolic state, which is a fair assumption, the equation for eta only depends on the ratio between k_inj and k_off and not on the individual parameters separately. Consequently, there is no way of saying which one of the two is more affected by metabolic state, unless the model already assumes that k_off is not influenced by metabolic state. The results could equally be explained by k_inj decreasing in metabolically depleted cells, or k_off increasing in such cells. If this is an assumption of the model, this should be clearly stated and not reported as a consequence of the data, as it is at the moment. Also, how does this mathematical model connect to the fitting function used in Figure 2b?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: discrimination requires that k_off/k_inj be larger for inactive hosts than for active hosts, such that commitment is specifically reduced in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This introduces a trade-off: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination requires this ratio to become significantly larger for inactive hosts. If this is achieved through changes in k_off, discrimination comes at the cost of slower commitment by allowing more time to leave; if it is achieved through changes in k_inj, it can preserve fast commitment to active hosts but requires additional energy dissipation in order to actively modulate commitment. We have therefore revised the text accordingly to frame the argument in terms of this trade-off, rather than attributing the effect specifically to k_inj. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      I have a related experimental criticism. The kinetic model presented assumes an exponential decay of free phage, which is a commonly used assumption in the phage literature. Given that the phage types used in this study lyse relatively slowly, it would be good to actually see adsorption curves, in which free phage is measured at different time points between inoculation and lysis. This data would not only provide useful evidence for the kinetic model, but it should also replace what is now in Figure S3, which consists of fitting one experimental point with one line. As it currently stands, Figure S3 is not useful actually misleading.

      We appreciate the reviewer’s point. We agree that adsorption curves, in which free phage is measured at different time points between inoculation and lysis, would provide a stronger basis for evaluating the kinetic model. However, we do not have the resources to perform these additional experiments within the scope of the present study. Following the comments of all reviewers on this point, we have therefore decided to remove Figure S3 to avoid confusion.

      Finally, it is not clear to me why the quantity "Ratio" has been chosen to be presented in Figure 1, rather than the ratio of estimated adsorption rates eta'/eta, which is much more intuitive for a phage study and contains the same information. I would recommend switching to this choice, unless there is a clear rationale for why the quantity "Ratio" is more useful/effective. Showing eta'/eta would also increase the readability of Figure 1, as it would move the y-axis to a logarithmic scale and better visualize values around 1.

      We used “Ratio” in Figure 1 to illustrate the experimental design, controls, and measured quantities directly, as it more transparently reflects the data collected. In the second part of the analysis, where we compare time-independent adsorption rate estimates, we have presented the corresponding values of η′/η as suggested.

      Minor comments:

      (1) Introduction

      Line 31: "... such as nutrient limitation, fluctuating temperatures, and variable energy availability" - if drawing a distinction between energy availability and nutrient limitation, please make explicit what this distinction is. Energy availability seems like a natural consequence of nutrient availability.

      While energy and nutrient availability are often linked in E. coli, they represent distinct physiological constraints. Nutrient limitation refers to the lack of essential biosynthetic precursors such as nitrogen, phosphorus, or amino acids. Energy availability, in contrast, reflects the cell’s ability to generate ATP and reducing equivalents through metabolic processes. For example, under anaerobic conditions, E. coli may have ample nutrients but limited energy production due to the lower efficiency of fermentation compared to aerobic respiration. Thus, energy limitation can occur independently of nutrient limitation.

      (2) Results

      (a) Whole Section: Please label equations.

      All equations have now been labelled in the revised manuscript.

      (b) Lines 105 to 114: As stated in Major Comments, I think the clarity of the paper would be improved by introducing the relative adsorption rate here and dropping the concept of Ratio entirely. However, if the authors wish to use Ratio, I would recommend the following:

      Lines 105 to 109 are confusing to read because of the number of connectives: "... ratio of free viruses from permissive AND resistant hosts respectively TO the free viruses in buffer under energy-depleted AND energy competent conditions". This would be clearer if each quantity were given an algebraic symbol, and RP, RR, and Ratio were defined through formal algebra, rather than mixed mathematical and sentence notation.

      This section has been rewritten for clarity. We now introduce explicit algebraic symbols and define the quantities formally, which removes the ambiguity present in the sentence-only description while retaining the intended meaning.

      The chemical names "arsenate" and "azide" should appear in the body of the text before they appear abbreviated in an equation. Please state at this point that these are both metabolic inhibitors, as it is not immediately clear what role they play or why you are using them.

      The text has been updated to introduce arsenate and azide by name before the abbreviations are used, and we now explicitly note that they act as metabolic inhibitors.

      On line 114, the authors helpfully provide an interpretation of Ratio = 1. It would be useful to provide at the same time interpretations of Ratio >1 and <1, perhaps 2 and 0.5 specifically?

      We have added brief explanations illustrating the interpretation of Ratio values greater than and less than 1, including examples of 2 and 0.5.

      I would consider giving this quantity a more interpretable name than Ratio. This quantity represents how much a bacteriophage preferentially adsorbs to metabolically active cells, so perhaps "Selectivity" or "Adsorption Bias"?

      We intentionally retained the generic term “Ratio”, as this quantity reflects an intermediate experimental measure used to describe the process rather than a newly defined metric. Its purpose is to bridge the experimental observations and the subsequent quantification of effects on the adsorption rate (η).

      (c) Lines 117 to 122: the authors sometimes refer to ratios explicitly, "average ratio of around 1.6" and other times say e.g., "a greater than 3 times increase in viral particles". Using more consistent language (saying "Ratio" every time) would be clearer.

      We have standardized the terminology in this section and now refer to all fold-changes consistently using “Ratio” to avoid ambiguity.

      (d) Figure 1

      Phages λ and T6 look like they have ratios less than 1 for resistant cells? If this is true / if the ratio is statistically significantly below 1, please comment.

      The ratios for λ and T6 are not statistically different from 1. The apparent deviation is within the standard error of the mean. To make this clearer, we have added the corresponding p-values to Table S2 in the Supplementary Information.

      Ratios near 1 are difficult to distinguish from 1, especially in panels A and D. Using a logarithmic scale on the y-axis would make the plots more readable.

      Because the values in these panels are not statistically different from 1, changing to a logarithmic scale would not alter the interpretation. We therefore retained the current axis scaling to reflect that there is no meaningful deviation from 1 in these cases.

      The data corresponding to individual experiments have no error bars. Given that the number of free virions was determined by plaque assay, which carries an intrinsic sampling error, this uncertainty should be reflected in the plots.

      We thank the reviewer for this important comment. Because plaque assays have compound sources of stochastic variation, assigning a per-measurement error bar would risk implying false precision. For this reason, we present the values from each biological replicate directly, and the uncertainty is represented in the statistical summary across replicates. Specifically, for each phage and condition we show the three independent experimental measurements and report the mean along with the standard error of the mean. This approach allows us to represent biological variability without implying a precision that cannot be accurately quantified at the level of single plaque counts.

      Similarly, the average value does show error bars, but it is not stated what these error bars correspond to: standard error in the mean, standard deviation of the sample, or combined uncertainty?

      The caption has been updated to state that the error bars represent the standard error of the mean.

      The resistant bacteria seemed to have ratios close to 1 in all cases. Is this because very few virions adsorbed under both energy conditions?

      Resistance is commonly associated with a lack of a surface receptor for the phage (or generally an entry pathway). We use the resistant bacteria as a control group for the effect of the conditions on adsorption. For resistant bacteria, the Ratio should be 1 since virions do not adsorb under both energy conditions. Any slight variations from 1 should come from sampling errors or small heterogeneity in the population.

      (e) Figure 2

      Please comment on what the error bars here represent. Error bars in Figure 2 A seem to permit negative (or at least zero) values of relative adsorption rate for phages m13 and T6, possibly implying an overestimate of the error? If it is the case that multiple values used to calculate the mean are far apart, possibly showing the values individually through a superimposed swarm plot would be clearer.

      This point is now addressed in the Supplementary Information, where we clarify how the error bars were calculated.

      (3) Discussion

      (a) Line 189: "high metabolic state" is imprecise. Say "energy-competent" to be consistent with earlier language.

      To maintain continuity with earlier terminology, we now include “energy-competent” in parentheses alongside “high metabolic state,” while retaining the original phrasing for readability.

      (b) Figure 3, population level

      Show adsorbed virions physically attached to bacteria, rather than removing them completely from the image, as currently, the implication is that at a high metabolic state, there are fewer virions total, not fewer virions remaining in solution because more are adsorbed. You could go as far as to add a third "after centrifuging" row, showing the adsorbed phages stuck in the pellet and the unadsorbed phages remaining in solution.

      Thank you for this suggestion. Figure 3 has been updated to depict adsorbed virions attached to bacterial cells, clarifying that the decrease represents adsorption rather than loss of total particles. This change improves the accuracy and interpretability of the schematic.

      (4) Methods and Materials

      (a) Figure 5

      The step "estimate cell numbers from OD" appears to follow incubating plates overnight. If the cells you are counting come from the pellet produced by centrifuging 3 steps prior, you could add a fork into the black line connecting the steps, with one branch corresponding to the supernatant and phages, and the other to the pellet and cells?

      Thank you for pointing this out. The order in the figure has been corrected: cell numbers are estimated from OD before overnight incubation. This resolves the confusion without the need for branching in the workflow diagram.

      (a) Line 332

      You allow as much time as possible for adsorption without the possibility of lysis. Did you determine the lysis times / latent periods of these phages through one-step-growth-curves, or use published results, in which case please cite? Having obtained the lysis time by either method, what fraction of the lysis time did you allow for adsorption? Also, please add supplementary tables with lysis times used for the different phages.

      We thank the reviewer for this comment. We used published latent-period values as guides and verified compatibility with our own system when selecting incubation times. We have clarified this in the text and added the relevant citations. We did not use a common fixed fraction of the lysis time for all phages; instead, incubation times were chosen to allow sufficient time for adsorption but not for completion of the first lytic cycle. For λ, productive lytic development was blocked in the host background used, as in Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119. For ϕ80 and T5, we used published latent-period values as guides and verified their compatibility with our own system (De Paepe and Taddei, PLoS Biology 2006, http://dx.doi.org/10.1371/journal.pbio.0040193). M13 is a chronic filamentous phage and therefore does not have a standard lytic latent period; in our host–phage combination, it required more than 1 h before phage release. For T6, we relied primarily on the kinetics observed in our own system, since adsorption was unusually slow for this phage–host pair under our assay conditions. Although literature reports describe shorter T6 latent periods under specific assay conditions (Foster and Johnson, Journal of General Physiology 1951, http://dx.doi.org/10.1085/jgp.34.5.529), this is consistent with published work showing that adsorption and infection kinetics can vary substantially with host background, surface structure, and experimental conditions (Heller and Braun, Journal of Bacteriology 1979, http://dx.doi.org/10.1128/jb.139.1.32-38.1979; Storms et al., Biochemical Engineering Journal 2012, http://dx.doi.org/10.1016/j.bej.2012.02.010).

      (5) Supplementary

      Figure S1

      This data is useful in understanding the main body of the paper, and I think this should form part of a main figure (possibly with the individual experimental data points superimposed over the bars). This could come before or as part of Figure 1?

      We thank the reviewer for this suggestion. We have explored including these data directly in the main figure but found that doing so substantially reduced the readability of the figure, as the underlying table is visually dense. For this reason, we chose to summarize the results in Figure 1 and present the detailed data separately in Figure S1 of the Supplementary Material, along with the Ratio analysis, which more effectively conveys the trends without overloading the main figure.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) L16-18: This sentence could be made more accessible as 'error correction' is not an intuitive term in the phage field.

      We have updated the overall theory section including the terminology. Instead of error correction, we now refer to it as a discrimination process.

      (2) L96-98: Does this potentially indicate a trade-off where evolution for stronger binding cannot evolve at the same time as responsiveness to metabolic activity?

      We agree that this sentence made a stronger evolutionary claim than our data support. Since we only tested four laboratory phages, we cannot conclude that there is an evolutionary trade-off between stronger binding and responsiveness to host metabolic activity. We have therefore removed this sentence to avoid making an unsupported evolutionary interpretation.

      (3) L102: What does 'post-cellular' mean?

      Postcellular supernatant is simply the liquid that remains after cells have been removed. During centrifugation, the cells pellet at the bottom, and the liquid above (which can contain viruses) is the postcellular supernatant.

      (4) L105-107: Worth splitting into two sentences as it is a bit unclear if ratios are built between permissible and resistant hosts or between buffers or both.

      Thank you for the suggestion. We have rewritten this section into two sentences to clarify how the ratios are constructed, and we hope the revised wording improves readability.

      (5) L110-122: Figures S1 and S2 could be referenced here.

      References to Figures S1 and S2 have now been added in this section.

      (6) L137: As P(0) is the viral concentration in buffer, I am assuming that the phage lysate has been diluted in buffer and phages have been added to cultures from the same dilution tube to guarantee equal starting numbers, but I couldn't find this in the methods.

      This clarification has been added to the Methods and Media section of the Supplementary Information.

      (7) L243: It would be worth defining what 'hyperdiffusion' means.

      We have added a brief definition of “hyperdiffusion”.

      (8) L253-256: I do not entirely follow this explanation.

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #3. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (9) L284: Why is k_inj necessarily more sensitive to the metabolic state than k_off? Could membrane changes under stress increase k_off?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: reduced commitment in inactive cells can arise through an increase in k_off, a decrease in k_inj, or both, as long as k_off/k_inj becomes larger in the inactive case. What matters is therefore not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also underlies the trade-off now discussed in the manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become much larger in the inactive case. We have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (10) Figure 1: There seems to be more variation between replicates in phage Lambda than in other phages. Is this caused by receptor number heterogeneity in the population?

      Unfortunately we do not have a way to compare receptor number heterogeneity across the different phage receptors in our experiments. We therefore cannot conclude that the larger variation observed for phage λ is caused by receptor number heterogeneity in the population.

      (11) Figure S1: There seems to be a significant difference between phage Lambda viability in the two buffers - do the authors have an idea where this comes from?

      There is no difference in λ viability between the two buffers. The apparent difference in the figure is due to sampling variability.

      (12) Figure S3: Last sentence of the legend probably shouldn't say 'upper'.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      Reviewer #3 (Recommendations for the authors):

      (1) The text reads as incomplete in some places. Can the authors please provide clarifications on the following points?

      (1.1) Lines 235-256: How do the authors draw a conclusion that "a phage can detect host metabolic status" from a study that used purified LamB receptors (i.e., no live cells with any metabolism) extracted from two different bacterial species (i.e., not a difference in metabolic states)?

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #2. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (1.2) Line 270, in the abstract, and in the caption of Figure 4: The authors described the model using terms such as "an error-correction mechanism" or "standard error correction", but there is little explanation. Can the authors clarify what kind of "error" is discussed here, and how it is "corrected"? In the "standard error correction" model, what determines which method of correction is "standard"? If "error correction" is a standard term in phage-bacterial interaction modeling, please provide references.

      We agree with the reviewer that our use of the term error correction was not appropriate in this context. The proper term is discrimination process rather than error correction. We have now corrected this terminology throughout the manuscript and clarified the underlying logic in the relevant sections.

      (1.3) Line 301: The authors speculated that phage T5 is "better suited to ecological niches", but I am not sure how that is consistent with their data showing T5 is more rampant, that they infect both energy-competent and energy-depleted cells, not just depleted cells. Why "niches", and why are T5 better suited to environments "where energy-limited cells dominate", not just any environment?

      We agree that this point was not stated clearly enough. What we intended to convey is that T5 would be at a net disadvantage in a niche containing a mixture of energy-competent and energy-deficient hosts. We have updated the main text accordingly.

      (1.4) Line 303, and related to point 6.3. above: Phage λ can also infect and replicate in "starved bacterial cells" (shown in Kourilsky 1974 and Geng et al. 2024, both of which were cited in this manuscript). How do the authors reconcile these reports with the discussion point in line 303, and their data that only phage T5, but not λ, shows insensitivity to the host metabolic state?

      Our data do not imply that phage λ is unable to infect starved bacteria. As shown in Kourilsky (1974) and Geng et al. (2024), λ can indeed infect and replicate in nutrient-limited cells. Our results specifically indicate that λ infection under starvation proceeds with a reduced adsorption rate, while T5 maintains the same adsorption rate even when the host is starved. Thus, our conclusion is that T5 is insensitive to the host metabolic state at the level of adsorption, whereas λ is not. We acknowledge that the wording in line 303 may have unintentionally led to confusion, and we have revised this part of the text to avoid that.

      (2) The following comments relate to the text and figures in the manuscript. There are many places in the manuscript that could use fine proofreading and copy-editing for clarity and consistency. For example:

      (2.1) If I understand it correctly, the equation in between lines 109 and 110 should be clarified using terms such as "Free viral particles after mixing with bacteria in Arsenate and Azide" and "Free viral particles in bacteria-free buffer with Arsenate and Azide". As it stands, it is not clear which terms correspond to conditions where bacteria are present.

      The equation has been updated to explicitly indicate which terms refer to mixtures containing bacteria and which refer to bacteria-free controls, so that the correspondence between conditions is now clear.

      (2.2) Equations in between line 276 and 283, and elsewhere: Some concentration terms are enclosed in brackets ("[BP]"), while most are not.

      This notation has been clarified. We now use “[PB]” specifically to denote the transient phage–bacterium complex, distinguishing it from the product P⋅B. All other concentration terms are written without brackets for consistency.

      (2.3) Figure 4 and in equations: "BP" or "PB"?

      The notation has been made consistent throughout; we now use “PB” exclusively to denote the phage–bacterium complex.

      (2.4) Line 284 and line 286: The "inj" in "k_inj" is sometimes italicized, sometimes not.

      The notation has been standardized so that k_inj is now formatted consistently throughout the manuscript, without italicizing “inj.” Also we have replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (2.5) Figure 5: Was the step "Estimate cell numbers from OD" really performed on the next day after the experiment (i.e., >12 hours after infection and phage plating), not immediately after cell washing?

      Thank you for pointing this out. The figure has been updated to reflect the correct order of steps: cell numbers are estimated from OD immediately after washing, followed by overnight incubation of the plates.

      (2.6) Figure S1: As it stands now, the x-axis of each panel can be read either as "Permissive, Resistant bacteria, Buffer" (missing "bacteria" for the first pair of bars), or "Permissive (bacteria), Resistant (bacteria), Buffer (bacteria)" (extra "bacteria" for the last pair of bars).

      The intended interpretation is the second one (permissive bacteria, resistant bacteria, buffer).

      (2.7) Figure S3: The panel letters "A" and "B" are missing in the figure. Also, it is not clear why the legend for the five phages and the legend for the measurement times are not combined.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      (2.8) Strain table in the Methods and Materials: Please write genotypes with italicization, and consistently indicate mutations and deletions with the minus sign superscript or the Δ prefix. Also, for the S3222 strain: Is it really the entire Mal regulon mutated ("Mal-"), or just lamB-? In Brown et al. 2022, it was only the latter.

      Genotypes have been reformatted with consistent notation. For S3222, the correct designation is Mal-, as in the SI of Brown et al. 2022. In this case, Mal- is intended as a phenotypic designation rather than a specific genotype, and we have therefore formatted it accordingly, i.e. neither italicized nor written in lower case.

    1. Author response:

      The following is the authors’ response to the original reviews.

      In revising the manuscript, we have focused on three main priorities raised during review: (1) improving precision around evidential claims, particularly concerning vector maintenance and P2A-mediated protein separation; (2) substantially improving figure quality, accessibility, and legend clarity; and (3) correcting inconsistencies and expanding methodological detail where requested.

      This study was intended as a foundational genetic toolkit and methodological framework for Blastocystis ST7-B, establishing practical workflows for DNA delivery, endogenous regulatory-element benchmarking, antibiotic-selected recovery, clonal propagation, and reporter-based analysis in a genetically challenging anaerobic microbial eukaryote. The central evidence presented is therefore functional in nature: reproducible transgene delivery, selectable recovery and propagation of colony-derived transgenic lines, and detectable reporter expression using multiple anaerobic-compatible reporter systems.

      We agree with the reviewers that several additional experiments, including Western blot analysis of P2A-containing constructs, outward-facing PCR, plasmid rescue assays, and selection-withdrawal experiments, would further strengthen the mechanistic interpretation of the system and help distinguish episomal persistence from genomic integration. We have therefore revised the manuscript throughout to clearly separate what is directly demonstrated from what remains a plausible working interpretation or important future direction.

      Importantly, the revised manuscript no longer presents episomal maintenance or complete P2A-mediated protein separation as demonstrated conclusions. Instead, these are now discussed explicitly as unresolved mechanistic questions requiring future molecular analysis. Nevertheless, the central methodological conclusion remains unchanged: stable selectable transgene expression, recovery of colony-derived transgenic lines, and reporter-positive Blastocystis ST7-B transformants can now be reproducibly obtained.

      Reviewer #1 (Public review):

      Summary:

      This paper presents a toolkit for the transformation of Blastocystis. The authors have screened a number of selectable agents, promoters and reporter genes and present their findings. This resource will be of immense use to those in the Blastocystis field, as well as those seeking to establish transformation tools in other species where such tools do not yet exist. Establishing new transformation tools is extremely challenging, and the authors have done an excellent job.

      Strengths:

      The authors have carried out a systematic screen of promoters, reporter genes and selectable agents. They have screened numerous for each, and all the data is presented. It is good to see when things did not work as well as when things did, so this data set is extremely useful indeed.

      Weaknesses:

      The findings are reported by reporter gene assay (microscopy). No evidence is given using genetics. The authors claim that the DNA is maintained episomally. However, could it be possible that there is integration? No PCRS/RT-PCRs are shown (although it can safely be assumed that the DNA/RNA is present where the transformation was successful), nor are any Western blots. These would have been useful to show that the P2A ribosomal skipping had occurred, and that proteins were expressed individually rather than as a polyprotein.

      We thank the reviewer for the positive assessment of the manuscript and for recognising both the technical difficulty and broader utility of establishing genetic tools in Blastocystis and other experimentally challenging microbial eukaryotes. We also appreciate the reviewer’s identification of the main evidential limitations in the original manuscript, particularly regarding vector maintenance and P2A-mediated protein separation.

      First, regarding the question of vector topology and the interpretation of episomal maintenance.

      We agree that the original manuscript presented episomal persistence too strongly relative to the evidence currently available. We have therefore revised the manuscript throughout to clarify that episomal maintenance should presently be regarded as a plausible working model rather than a directly demonstrated conclusion.

      The transfection system used here was adapted from Li et al. (2019), including use of the pXS2-P<sub>Legumain</sub>-derived plasmid framework. Importantly, the construct used in the present study does not contain the original Trypanosoma brucei tubulin-targeting region associated with homologous integration in the original pXS2 system. Complete plasmid sequencing confirmed that the constructs function here as heterologous expression plasmids carrying Blastocystis ST7-B regulatory elements and transgenes. While this does not demonstrate episomal persistence, it also means that genomic integration cannot be inferred from the historical pXS2 vector architecture alone.

      We further note that comparative genomic analyses by Gentekaki et al. (2017) suggest that Blastocystis lacks components of the canonical non-homologous end-joining (NHEJ) machinery, implying that homologous recombination is likely to represent the principal route for double-stranded DNA repair. Because the constructs used here did not contain Blastocystis homology arms, there is currently no obvious mechanism favouring targeted homologous integration. Nevertheless, we fully agree that genomic integration cannot presently be excluded.

      To reflect this appropriately, the revised manuscript now explicitly separates the demonstrated functional outcomes from unresolved mechanistic questions concerning vector maintenance. We also identify several future approaches that would help distinguish episomal persistence from genomic integration, including outward-facing PCR, plasmid rescue followed by full plasmid sequencing, Southern blotting, FISH, selection-withdrawal experiments, and long-read sequencing approaches.

      We have revised the manuscript throughout to remove statements implying demonstrated episomal maintenance and now present episomal persistence only as a plausible working interpretation.

      In the Methods section under Cloning, the following text has been added:

      Lines 202–206: “The constructs used in this study were derived from the pXS2-P<sub>Legumain</sub> vector described by Li et al. (2019), which adapted a heterologous expression-vector backbone for transient plasmid-based expression in Blastocystis ST7-B. Here, the same molecular backbone was used as a plasmid scaffold carrying Blastocystis-derived regulatory elements and transgenes.”

      In the Discussion, the following text has been added/edited:

      Lines 665–673: “The molecular maintenance state of the introduced constructs remains unresolved: episomal maintenance is a plausible working model, but genomic integration cannot be formally excluded. The constructs used here lack Blastocystis homology arms, and comparative genomic analyses suggest that Blastocystis lacks canonical non-homologous end-joining components (Gentekaki et al., 2017), making targeted integration by standard repair routes unlikely but not impossible. Direct assays such as outward-facing PCR, plasmid rescue followed by full plasmid sequencing, FISH, or selection-withdrawal experiments will be required to distinguish episomal persistence from integration.”

      Second, regarding P2A-mediated protein separation.

      We agree that Western blotting would provide the most direct biochemical assessment of P2A-mediated ribosomal skipping efficiency in Blastocystis ST7-B and would help determine the extent of any residual uncleaved fusion product. We have therefore revised the manuscript to avoid implying that complete protein-level separation was directly demonstrated.

      The revised manuscript now states only what is directly supported by the current data: that P2A-containing bicistronic constructs supported antibiotic-selected recovery of transgenic lines together with detectable downstream reporter expression. The microscopy data therefore support functional downstream reporter expression, but do not by themselves exclude residual uncleaved fusion products.

      We selected P2A because it is a compact and well-characterised peptide with high reported separation efficiency across multiple eukaryotic systems, including microbial eukaryotes. However, we agree that P2A performance can be context-dependent, and we now explicitly identify biochemical validation of P2A cleavage efficiency as an important future direction.

      Importantly, these revisions do not alter the central methodological conclusion of the study, namely that selectable transgene expression, propagation of reporter-positive lines, and recovery of colony-derived Blastocystis ST7-B transformants can now be reproducibly achieved.

      Text inserted in the Results:

      Lines 394–396: “The P2A peptide is expected to promote ribosomal skipping during translation, allowing two separate polypeptides to be produced from a single open reading frame.”

      Lines 403–404: “However, protein-level separation was not directly tested, and the extent of any residual uncleaved fusion product remains unresolved.”

      Text inserted in the Discussion:

      Lines 619–629: “P2A was selected because it is a well-characterised peptide with high reported separation efficiency in human cell lines, zebrafish embryos, and mice (Kim et al., 2011). It also has precedent across microbial eukaryotes, including the protest Dictyostelium discoideum (Zhu et al., 2023), the fungi Aspergillus niger (Schuetze and Meyer, 2017) and Ustilago maydis (Müntjes et al., 2020), and the apicomplexan parasites Toxoplasma gondii (Markus et al., 2019) and Plasmodium falciparum (Dans et al., 2024). However, P2A performance is context-dependent, and the evidence presented here is functional rather than biochemical. P2A-containing constructs support antibiotic-selected recovery and downstream reporter expression in Blastocystis ST7-B, but ribosomal skipping efficiency and any residual uncleaved product will require direct protein-level validation.”

      Reviewer #1 (Recommendations for the authors):

      (1) Please could you show a Western blot to confirm if P2A has worked? It could be that the proteins are being expressed as a polyprotein.

      We agree that Western blotting would provide the most direct biochemical assessment of P2A-mediated ribosomal skipping efficiency in Blastocystis ST7-B and would help determine the extent of any residual uncleaved fusion product. This is an important point, and we have revised the manuscript accordingly to avoid implying that complete protein-level separation was directly demonstrated.

      The current study was designed as a first-generation functional genetic toolkit for Blastocystis ST7-B, focused primarily on establishing reproducible workflows for selectable transgene expression, reporter recovery, and propagation of transgenic lines in this experimentally challenging anaerobic microbial eukaryote. The toolkit is therefore validated here through functional outcomes, including antibiotic-selected survival, stable propagation through extended passaging (>15 passages) and cryopreservation, and detectable reporter fluorescence above wild-type autofluorescence.

      P2A was selected because it is a compact and well-characterised peptide with high reported ribosomal skipping efficiency across multiple eukaryotic systems, including microbial eukaryotes, as discussed above. Nevertheless, we fully agree that direct biochemical validation would strengthen the mechanistic interpretation of the bicistronic system in Blastocystis ST7-B. We therefore now explicitly identify Western blot analysis, ideally using epitope-tagged upstream and downstream products, as an important future direction for quantitative assessment of P2A cleavage efficiency and any residual uncleaved fusion products.

      Relevant manuscript revisions are described above under the general response to Reviewer 1.

      (2) Something has gone wrong with figure formatting. Figure 2 is nearly illegible and I cannot read the text in section A. Sections B, C, and D have lost their labels and are fuzzy and surrounded by black. A similar issue affects Figure 3. Everything is just black with a few cells. It is illegible when printed.

      We thank the reviewer for highlighting these presentation issues and agree that the submitted figure quality significantly impaired readability and interpretation. The problems appear to have arisen primarily during manuscript compilation and export, particularly affecting image resolution, contrast, and panel labelling in the review PDF.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions. We have also improved typography, panel separation, colour scaling, and legend clarity throughout. In response to additional reviewer suggestions, individual data points have now been added to Figures 2B and 2C to improve transparency and interpretability of the underlying data distributions.

      Figures 2 and 3 have been replaced with fully revised high-resolution versions with improved panel labelling, accessibility, typography, and figure legends.

      (3) The data from Figure 2B would be better placed in Table 1 with a column for robust/moderate/intermediate/weak/very weak. This would be much easier for the reader.

      We thank the reviewer for this helpful suggestion. We believe the comment refers to the promoter activity data shown in Figure 2A rather than the voltage optimisation data in Figure 2B. To improve readability and accessibility of these data, we have revised Figure 2A extensively to make the promoter activity tiers more legible and easier to interpret directly from the heat map and accompanying box plots.

      We considered incorporating simplified activity classifications into Table 1. However, activity patterns were construct-specific rather than simply locus-specific. In several cases, multiple promoter fragments derived from the same locus produced substantially different reporter outputs, and activity did not scale monotonically with promoter fragment length. We therefore felt that assigning a single categorical activity label at the locus level would oversimplify the dataset and reduce the construct-level resolution that is central to the toolkit value of the study.

      Instead, we addressed the reviewer’s concern by substantially improving the presentation and readability of Figure 2A, allowing readers to identify robust, moderate, intermediate, weak, and very weak expression constructs more directly while preserving the underlying construct-specific information.

      Figure 2A has been revised to improve clarity, accessibility, and legibility of the promoter activity tiers, allowing construct-level expression classes to be interpreted more directly from the heat map and accompanying boxplots.

      (4) How do you know if the constructs are maintained as episomes? Have you done an outward-facing PCR?

      We agree that direct molecular evidence distinguishing episomal persistence from genomic integration is currently lacking, and we appreciate the reviewer highlighting this important limitation. We have therefore revised the manuscript throughout to avoid presenting episomal maintenance as a demonstrated conclusion and now describe it only as a plausible working interpretation based on the current evidence and vector design.

      We have not performed outward-facing PCR in the present study. As discussed in the general response above, we now explicitly identify outward-facing PCR, plasmid rescue followed by full plasmid sequencing, selection-withdrawal assays, FISH, and long-read sequencing approaches as important future directions for resolving the molecular maintenance state of the constructs.

      The revised manuscript now clearly separates the demonstrated functional outcomes, including selectable transgene expression, recovery of colony-derived transgenic lines, and stable reporter-positive propagation under selection, from the unresolved mechanistic question of vector topology.

      This issue has been addressed throughout the revised manuscript, including in the Methods and Discussion sections, where episomal maintenance is now presented as a plausible but unconfirmed interpretation rather than a demonstrated conclusion.

      Minor Comments

      Line 66: is this one to two billion individuals with Blastocystis, or one to two billion Blastocystis cells per gut?

      The intended meaning was colonised individuals globally. We agree that the original phrasing was ambiguous and have corrected it for clarity.

      Lines 66–67 revised to: “…microorganisms in the human gut, and is estimated to colonise approximately one to two billion people globally (Scanlan and Stensvold, 2013).”

      Line 148: Supplier of IMDM?

      The supplier information was already present in the original manuscript as IMDM L0191 (Biowest).

      No additional manuscript change required.

      Line 157: Who annotated the dataset, the 2017 paper or the present study?

      The dataset annotation derives from Armengaud et al. (2017). We agree that the original wording was unclear and have revised this section substantially to improve clarity regarding the rationale and workflow used for promoter and terminator candidate selection.

      “The relevant Methods section has been extensively revised for clarity and expanded detail” (Lines 156–189).

      Line 166: Who predicted the 3′ UTR, the 2017 paper?

      This information derives from the NCBI annotation associated with the Blastocystis ST7-B genome based on Denoeud et al. (2011). This has now been clarified in the Methods section.

      Clarified in revised Methods section.

      Line 237: How long did it take in days?

      Approximately 2 days.

      Line 270 revised to: “…turned yellow without drug treatment, usually within 2 days post-transfection.”

      Line 325: Typo, missing gap between Figure and 1A.

      Corrected in revised manuscript.

      Reviewer #2 (Public review):

      This manuscript presents a substantial technical advance for the genetic manipulation of Blastocystis by establishing an integrated workflow for stable episomal transgenesis, antibiotic selection, clonal recovery, and reporter-based imaging in the ST7-B subtype. The study is particularly valuable because it combines multiple previously fragmented approaches into a coherent and practically applicable toolkit, including endogenous regulatory elements, optimized electroporation conditions, selectable markers, and anaerobic compatible fluorescent reporters. This methodological work greatly expands the molecular toolbox and future studies focused on both basic and infection biology can now build on the ability to express and localize proteins in fixed as well as live cells.

      The microscopy data are convincing and clearly demonstrate functional reporter expression and successful recovery of stable transgenic lines. Nevertheless, because this is primarily a methodological paper, the study would be further strengthened by the inclusion of Western blot validation of reporter expression and bicistronic constructs. In particular, biochemical analysis of the P2A-containing constructs would help assess the efficiency of ribosomal skipping and exclude the possible presence of uncleaved fusion proteins, thereby providing stronger support for the interpretation of the imaging data and the functionality of the expression system.

      We thank the reviewer for this thoughtful and positive assessment of the manuscript and for recognising the value of integrating previously fragmented approaches into a coherent and practically usable genetic toolkit for Blastocystis ST7-B. We particularly appreciate the reviewer’s recognition that the system expands the currently available molecular toolbox for both cell biological and infection-related studies in this experimentally challenging anaerobic microbial eukaryote.

      We also appreciate the reviewer’s comments regarding biochemical validation of the P2A-containing bicistronic constructs. We agree that Western blot analysis would strengthen the mechanistic interpretation of the reporter system by directly assessing ribosomal skipping efficiency and the possible presence of residual uncleaved fusion products. In response, we have revised the manuscript throughout to ensure that the conclusions remain appropriately evidence-based and do not imply that complete protein-level separation was directly demonstrated.

      The revised manuscript now explicitly distinguishes the demonstrated functional outcomes, including selectable transgene expression, stable propagation of reporter-positive lines, and detectable downstream reporter expression, from unresolved mechanistic questions concerning P2A cleavage efficiency and vector maintenance state. We now also identify biochemical validation of P2A-mediated protein separation as an important future direction for further refinement of the system.

      Relevant manuscript revisions addressing these points are described above under the response to Reviewer 1.

      Reviewer #2 (Recommendations for the authors):

      The quality of images could be better. The figures lacked resolution — possibly a conversion artefact.

      We agree that the figure quality in the submitted review PDF significantly reduced readability and visual interpretation. The issues appear to have arisen primarily during manuscript compilation and export, particularly affecting image resolution, typography, panel labelling, and contrast rendering.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions. We have also improved panel separation, typography, colour scaling, contrast settings, and figure legends to improve accessibility and interpretability both on screen and in print. In addition, the export workflow and file formatting have been updated to improve compatibility with journal production requirements and reduce the likelihood of compression-related rendering artefacts during manuscript compilation.

      Figures 2 and 3 have been replaced with revised high-resolution versions with improved typography, panel labelling, contrast settings, and accessibility.

      Reviewer #3 (Public review):

      Summary:

      The primary objective of this study was to establish a practical and functional framework for the propagation of stable transgenic cell lines of Blastocystis, a common animal gut microeukaryote. Although the work focused on Blastocystis ST7-B, a subtype with relatively low prevalence in humans, this choice is justified by its association with more frequent negative health effects. Beyond their relevance to the medical field, the methodological advances described here have the potential to also expand cell biology studies of this anaerobic organism, including its unusual mitochondria and redox metabolism.

      Strengths:

      Prior to this work, genetic tools for Blastocystis were very limited, relying on a single strong promoter-terminator combination. The authors successfully expanded the available promoter set across a range of expression strengths by testing two dozen variants in luciferase-based assays. Critically, they developed an integrated workflow from a modular transgenic construct design, to an expanded inventory of molecular components (promoters, reporters), optimized DNA delivery, stepwise antibiotic resistance-mediated clonal selection and propagation, and to reporter validation. The evaluation of several anaerobiosis-compatible labeling strategies for live (and fixed) cell optical imaging will be particularly useful, with the SNAP-tag system appearing especially promising for Blastocystis.

      Weaknesses:

      The presented data generally provide solid support for the conclusions that the work reached, but clarification of reasoning and several inconsistencies, as well as amendments to the visual presentation of the data, would be highly beneficial, as detailed below.

      (1) Episomal persistence of the construct:

      The manuscript repeatedly assumes, including in its title, that constructs persist in Blastocystis in their episomal form, but no direct evidence is provided. Although this interpretation is plausible, it should be identified more clearly as provisional. Nuclear genomic integration (e.g., via NHEJ) remains a possible explanation unless supporting evidence or rationale is provided to exclude it. Testing whether the phenotype persists without drug-mediated selection in the generated transgenic cell lines would help strengthen the case for episomal maintenance.

      We thank the reviewer for this important point and agree that the original manuscript presented episomal persistence too strongly relative to the currently available evidence. In particular, we agree that the title and several sections of the manuscript implied a level of mechanistic certainty that was not directly demonstrated.

      We have therefore revised the manuscript throughout to clarify that episomal maintenance should presently be regarded as a plausible working interpretation rather than a demonstrated conclusion. The revised text now explicitly distinguishes the demonstrated functional outcomes, including selectable transgene expression, recovery and propagation of colony-derived transgenic lines, and stable reporter-positive maintenance under selection, from the unresolved mechanistic question of vector topology.

      As discussed in our response to Reviewer 1, the constructs used here do not contain Blastocystis homology arms, and comparative genomic analyses suggest that Blastocystis lacks canonical non-homologous end-joining components, making targeted integration by standard repair routes less strongly supported mechanistically, although genomic integration cannot presently be excluded.

      We agree that selection-withdrawal experiments would provide useful additional evidence regarding construct persistence and have now explicitly identified such assays, together with outward-facing PCR, plasmid rescue, FISH, and long-read sequencing approaches, as important future directions for resolving the molecular maintenance state of the transgenes.

      The manuscript has been revised throughout to remove wording implying demonstrated episomal maintenance. Episomal persistence is now discussed only as a plausible working interpretation pending direct molecular validation.

      (2) Promoters and terminators:

      (2.1) There is a discrepancy between the claimed number of loci (14), from which promoters used to drive luciferase expression were derived, and those detailed as having been actually generated in Table 1 (11). This inconsistency should be corrected or explained, as it creates uncertainty around the accuracy of the dataset.

      We thank the reviewer for this careful reading and for identifying this inconsistency. We agree that the distinction between candidate loci and successfully generated promoter constructs was not sufficiently clear in the original manuscript and could create uncertainty regarding the dataset.

      The original candidate set comprised 14 loci selected for promoter and terminator discovery. However, only 11 loci yielded successfully cloned and experimentally tested promoter constructs. The remaining three loci were retained in Table 1 for completeness and transparency, as repeated cloning attempts were unsuccessful despite two independent efforts.

      We have revised the manuscript to make this distinction explicit and to clarify that the reported NanoLuc benchmarking experiments were ultimately performed using constructs derived from 11 successfully cloned loci.

      Lines 361–364: “To expand the available regulatory parts, we screened 23 NanoLuc reporter constructs containing putative endogenous promoter–terminator pairs from 11 of 14 candidate loci; three loci could not be cloned after two independent attempts and are indicated in Table 1.”

      (2.2) Based on the presented evidence, constructs benchmarked in bioluminescence assays differed only in their promoter composition. Although terminator selection is mentioned in the Methods section, no additional details are provided; for instance, Table 1 and Figure 2 only list 23 promoters in total. Figure 2A likewise shows only promoter-dependent variation. If the terminator was held constant (LeguP1?), this should be stated explicitly. The authors may then consider revising the wording of having tested “23 promoter-terminator pairs” to better reflect that only promoters varied.

      We thank the reviewer for the opportunity to clarify this point. We agree that the original presentation may have created the impression that promoter and terminator regions were independently varied and benchmarked, whereas the experimental design was primarily focused on construct-level comparison of endogenous regulatory modules.

      As described in the Methods, each construct contained a candidate endogenous upstream promoter region together with the corresponding endogenous downstream terminator region derived from the same locus. For consistency and to keep the cloning and screening strategy experimentally tractable, a fixed 500 bp downstream terminator fragment was used for each locus rather than systematically varying terminator length or independently testing terminator activity.

      We therefore retain the description “endogenous promoter–terminator pairs,” since each construct contains both endogenous upstream and downstream regulatory regions from the same genomic locus. However, we agree that the assay was not designed to independently dissect promoter versus terminator contributions to reporter output. We have revised the manuscript accordingly to make this distinction explicit and avoid ambiguity regarding the scope of the benchmarking analysis.

      Lines 365–368: “Each construct paired a candidate upstream promoter region with the corresponding downstream terminator region from the same locus, defined here as the native 500 bp sequence immediately downstream of the stop codon. Where multiple promoter lengths were tested for the same locus, the terminator fragment was kept constant (Table 1; Figure 1A).”

      This design allowed construct-level benchmarking of paired promoter–terminator modules but did not test promoter strength or terminator activity independently.

      (2.3) Promoter benchmarking was done with a plasmid lacking a selection marker, so it is unclear how the maintenance of the luciferase construct was ensured. Without selection, the observed reporter intensity could reflect differential or stochastic plasmid retention rather than promoter strength alone. The luminescence assay was performed 16-18 hours after transfection, but the rationale for this particular timeframe should be explained. In this context, the authors should explicitly state whether the experiments shown in Fig.2A represent biological triplicates or technical triplicates from a single transfection.

      We thank the reviewer for these important methodological points. We agree that the original manuscript did not sufficiently clarify the transient nature of the NanoLuc benchmarking assay or the rationale underlying the assay design and timing.

      The promoter benchmarking assay was designed as an early transient-expression screen adapted from the NanoLuc-based workflow of Li et al. (2019), with modifications, rather than as a stable-maintenance assay. No selectable marker was included because the objective was to compare relative early reporter output across constructs shortly after DNA delivery, before prolonged culture effects became dominant.

      The 16–18 h post-electroporation time point was selected based on the NanoLuc expression kinetics reported by Li et al. (2019) and empirical optimisation during assay development. This window allowed robust transient reporter detection while limiting confounding effects arising from prolonged plasmid loss, differential outgrowth, variable recovery, or later culture-level changes.

      We agree that, in the absence of selection, the observed NanoLuc signal cannot be interpreted as an absolute measure of promoter strength independent of DNA uptake efficiency, early plasmid retention, or post-transfection recovery dynamics. We have therefore revised the manuscript to clarify that Figure 2A reports relative transient reporter output under standardized early post-transfection conditions rather than isolated promoter activity alone.

      We now also explicitly state that the data shown in Figure 2A derive from three independent electroporation experiments per construct, each assayed in technical duplicate.

      Lines 241–248: “Promoter–terminator activity was assessed 16–18 h after electroporation using a transient NanoLuc assay adapted from Li et al. (2019), with modifications. This early time point was selected to capture reporter output within the transient-expression window after DNA delivery, before prolonged plasmid loss, differential outgrowth, or culture-level changes could dominate the readout. Because the constructs did not contain a selectable marker, the measured NanoLuc signal reflects early transient reporter output rather than promoter strength independent of DNA uptake, early plasmid retention, or post-transfection recovery.”

      Additional clarification added to Figure 2 legend stating that measurements derive from three independent electroporation experiments, each assayed in technical duplicate.

      (3) Figure 2:

      (3.1) Several aspects of the current design may lead to ambiguity for the reader. The boxplots are colour-coded, but it is unclear whether the colours carry meaning or are purely decorative. Because the data are already spatially separated into bins, additional random colouring is redundant and may suggest distinctions that are not intended. In addition, part A of Figure 2 is split into two panels, with the scale for the left panel shown in the right panel and some of the boxplot colours falling in the range of the scale, but not in line with their counterparts in the left panel. Because the colour use is not consistent, it is difficult to tell whether the same scale should be applied to both panels or how it should be interpreted.

      (3.2) The left panel of part A uses a diverging blue-white-red colour scheme, which is most appropriate when the midpoint represents a meaningful central value such as zero. Because the values shown in this graph are only positive, a non-diverging 2-colour scale or a colour palette such as 'viridis' would make the plot easier to interpret.

      (3.3) A black background should be avoided: 'B' and 'C' labels are invisible, and it draws attention to a distracting design feature rather than the data themselves.

      We thank the reviewer for these detailed comments regarding figure design and visual interpretation. We agree that the original presentation of Figure 2 introduced unnecessary visual ambiguity through inconsistent colour usage, the use of a diverging colour scale for strictly positive values, and poor readability associated with the dark background and low-resolution export.

      In response, Figure 2 has been extensively redesigned to improve clarity, accessibility, and interpretability. The previous blue–white–red diverging heatmap has been replaced with a sequential colour palette appropriate for positive-only expression data. Boxplot colouring has also been simplified and harmonised with the heatmap scheme to avoid implying unsupported categorical distinctions. In addition, panel organisation, typography, scaling, and legend structure have all been revised to improve readability and reduce ambiguity regarding interpretation of the plotted values.

      We also agree that the black background distracted from the data presentation and impaired visibility of panel labels and image boundaries. The revised figures therefore use white backgrounds together with clearer panel separation and improved label visibility throughout.

      Figure 2 has been completely reformatted using a sequential colour scale in panel A, simplified and harmonised boxplot colouring, larger typography, improved panel separation, revised legends, and white backgrounds throughout. Corrected high-resolution source figures have been provided.

      (4) Figure 3:

      (4.1) Individual snapshots should be separated more clearly, either by using a white background or by adding visible borders to make the overall composition clearer. As currently displayed, some boundaries between fluorescent channels resemble image artifacts rather than intentional panel divisions.

      We thank the reviewer for this helpful comment regarding figure composition and panel separation. We agree that the original presentation made it difficult to distinguish intentional panel boundaries from imaging artefacts, particularly in the low-resolution review PDF generated during manuscript compilation.

      To improve clarity, Figure 3 has been reformatted using white backgrounds, clearer panel spacing, and more explicit separation between individual snapshots and imaging channels. High-resolution source images have also been provided to ensure that fluorescence patterns, image boundaries, and panel organisation remain clearly interpretable both on screen and in print.

      Figure 3 has been reformatted with improved panel separation, white backgrounds, clearer image boundaries, and revised high-resolution source figures.

      (4.2) In parts B-D, the legend should explain more clearly what each image shows, and the figure itself would benefit from annotations. There seem to be three sub-panels in each 'condition' of part B (as well as C and D): while the middle and rightmost panel can be easily inferred to represent the fluorescent protein and bright-field image, what the leftmost panels represent is not specified. If DAPI was used to dye DNA, an explanation why mostly multiple labelled regions are visible should be provided.

      We thank the reviewer for these helpful suggestions regarding figure annotation and legend clarity. We agree that the original presentation did not sufficiently explain the composition of the imaging panels, particularly under the low-resolution conditions of the review PDF.

      To improve interpretability, the revised Figure 3 now includes clearer panel organisation, improved annotations, and expanded figure legends explicitly identifying the individual imaging channels and staining conditions shown in each subpanel. The leftmost panels in parts B–D are now more clearly identified in both the figure and legend, together with the corresponding fluorescence or staining conditions used in each experiment.

      As mentioned in the Methods sections we used Hoechst 33342 to visualise DNA; but we agree that the Hoechst 33342-labelled structures required additional clarification. The revised legend section now explains that multiple Hoechst 33342-positive regions are commonly observed because Blastocystis cells can contain multiple nuclei depending on cell stage and subtype-specific morphology.

      In addition, high-resolution source images have been provided to ensure that fluorescent signals, panel boundaries, and imaging features remain clearly interpretable both on screen and in print.

      Figure 3 legends and annotations have been revised to clarify imaging channels, staining conditions, and panel organisation. The figure caption was also edited to include: “DNA was visualised using Hoechst 33342. Most cells contained two nuclei, and smaller Hoechst 33342-positive signals consistent with mitochondrial DNA were also observed in some instances.”

      (4.3) Cell morphology and appearance differ markedly between UnaG/smURFP and SNAP-tag images, which should be explained. A microscope issue is mentioned in the main text, but if that was the cause, the authors should consider replacing the images, as the current distortions complicate interpretation.

      We thank the reviewer for this important observation and agree that the apparent morphological differences between the UnaG/smURFP and SNAP-tag panels required additional clarification.

      The images shown for the different reporter systems were acquired under different imaging conditions and microscope configurations following an instrument-related issue during part of the imaging workflow, as noted in the Methods section. As a result, direct visual comparison of cell morphology between reporter systems is not appropriate. The primary purpose of these panels is instead to demonstrate reporter detectability, live-cell labelling capability, and the characteristic fluorescence patterns obtained with the different anaerobiosis-compatible reporter systems.

      In particular, the SNAP-tag panels were included to demonstrate successful live-cell labelling without permeabilisation together with the expected increase in fluorescence signal at higher substrate concentrations, rather than to support quantitative comparison of cell morphology across imaging conditions.

      We considered replacing the affected images. However, equivalent replacement datasets acquired under directly comparable conditions are not currently available. We have therefore retained the original images but revised the figure legend to clarify the intended interpretation and limitations of these panels explicitly.

      Figure 3 legend revised to include:

      “Because images for the different reporter systems were acquired under different imaging conditions, they are presented to demonstrate reporter detectability and labelling pattern and should not be used for quantitative comparison of cell morphology across reporter systems.”

      Reviewer #3 (Recommendations for the authors):

      The reader may find the current order confusing starting with construct design before testing which drug to use for selection. The narrative would work better if it started with antibiotic selection as the first logical step for generating stable cell lines.

      We thank the reviewer for this thoughtful suggestion regarding narrative structure and agree that multiple organisational strategies are possible for presenting a methodological workflow of this type.

      We considered reorganising the Results section to begin with antibiotic selection and drug sensitivity profiling. However, we ultimately retained the overall structure because the manuscript is organised as a toolkit-development framework rather than as a strictly chronological experimental protocol. The Results therefore begin with regulatory-element discovery and construct design, which form the conceptual and experimental foundation of the toolkit, before progressing to DNA delivery optimisation, drug sensitivity profiling, clonal recovery, and reporter validation.

      We felt that this structure most clearly reflects the dependency relationships within the system: regulatory elements are required before constructs can be assembled, constructs are required before electroporation conditions can be evaluated, and selectable constructs are required before stable selection and clonal recovery can be meaningfully assessed.

      (2) The text states that the screen 'focused on the 1,000 most abundant proteins to establish a preliminary library capable of supporting varying levels of transcription.' Since the genome has ~6,000 protein-coding genes, the top 1,000 cover the most abundant proteins — not a wide expression range.

      We thank the reviewer for this important clarification. We agree that the original wording could incorrectly imply that the screen was intended to sample broadly across the full transcriptional range of the Blastocystis genome. This was not the case, and we have revised the manuscript accordingly.

      Our strategy was instead designed to enrich for candidate loci with a higher prior likelihood of supporting detectable transgene expression. Because no genome-wide promoter map, transcription start site dataset, or experimentally validated regulatory annotation was available for Blastocystis ST7-B at the inception of this work, we used the abundance-ranked Blastocystis ST4-WR1 proteomic dataset of Armengaud et al. (2017) as a practical starting point for candidate discovery.

      Importantly, the Blastocystis ST4-WR1 proteome is highly skewed, with 193 proteins contributing approximately 50% of the detected proteome and the 13 most abundant proteins contributing approximately 10% (Armengaud et al., 2017). We therefore selected the top 1,000 proteins not as a representation of the genome-wide expression range, but as a proteomics-guided enrichment strategy to identify loci more likely to contain active endogenous regulatory regions suitable for initial toolkit development.

      We have revised the relevant Methods section substantially to clarify both the rationale and the workflow used for candidate selection, homolog identification, and promoter/terminator definition.

      The Methods section (Lines 156–189) has been extensively revised to clarify the rationale underlying candidate regulatory-element selection. The revised text now explicitly states that the strategy was designed to enrich for likely active loci for toolkit development rather than to systematically survey the full range of promoter strengths across the Blastocystis genome.

      Additional methodological detail has also been added regarding:

      Use of the Armengaud et al. (2017) proteomic and proteogenomic datasets,

      Homolog identification in Blastocystis ST7-B,

      Locus selection criteria,

      Promoter boundary definition,

      And operational definition of candidate terminator regions.

      (3) The Methods contain an inconsistency: cells were left in 0.5 mL, then 1 mL was added, but then only 0.5 mL is apparently used for transfection. What happened to the 1 mL?

      We thank the reviewer for identifying this ambiguity in the transfection workflow description. The apparent inconsistency arose because the protocol description moved from bulk cell resuspension to preparation of individual electroporation reactions without explicitly stating how the intermediate suspension was used.

      After washing, approximately 0.5 mL of cytomix buffer remained above the pellet, and 1 mL of complete cytomix buffer was then added to generate an approximately 1.5 mL cell suspension. Cells were counted from this pooled suspension, after which the volume corresponding to 5 × 10<sup>7</sup> cells was transferred into each individual electroporation reaction. Following addition of DNA, each electroporation reaction was adjusted to a final volume of 500 µL with complete cytomix buffer. The remaining cell suspension was retained for additional transfections or control reactions.

      We agree that the original wording could be misinterpreted and have revised the Methods section to clarify the sequential handling steps more explicitly.

      Lines 225-229 revised to read: “The resulting approximately 1.5 mL pooled cell suspension was used for total viable cell counting using a hemacytometer.”

      “After counting, the volume corresponding to 5 x 10<sup>7</sup> cells was transferred to each electroporation reaction and combined with 25 µg of plasmid DNA. The total electroporation volume was adjusted to 500 µL with complete cytomix buffer.”

      (4) Figures 2 and 3 are too low-resolution for the font size used and for clearly viewing the microscopy images.

      We thank the reviewer for highlighting these readability issues. As noted in our responses above regarding Figures 2 and 3, the low-resolution appearance primarily resulted from manuscript compilation and PDF export artefacts affecting typography, image rendering, and panel clarity in the review version.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions featuring improved typography, panel labelling, contrast, accessibility, and image clarity for both on-screen viewing and print reproduction.

      Revised high-resolution versions of Figures 2 and 3 have been provided as described above. No additional manuscript changes were required beyond the figure revisions already outlined.

      (5) Figure 4 is confusing because the left and right panels appear inconsistent, with much higher concentrations required for growth inhibition in the culture-based assay than the resazurin assay indicated. The rationale for the resazurin assay should be explained, and the complete growth inhibition (CGI) concentration should be highlighted in the right panel.

      We thank the reviewer for highlighting this potential source of confusion. We agree that the distinction between the two assay endpoints was not sufficiently emphasised in the original figure presentation and legend.

      The apparent discrepancy arises because the two assays measure different biological endpoints under different assay conditions. The resazurin assay was used to estimate IC<sub>50</sub> values, corresponding to the concentration at which metabolic activity was reduced by approximately 50% under the assay conditions. In contrast, the small-culture assay was designed to determine complete growth inhibition (CGI), defined operationally as the concentration at which no detectable culture outgrowth occurred after incubation, using phenol red acidification as a culture-level readout.

      Because these assays measure partial metabolic inhibition versus complete suppression of detectable culture outgrowth, the corresponding concentration ranges are not expected to coincide directly. The higher concentrations observed in the right-hand panels therefore reflect the more stringent endpoint associated with complete growth inhibition rather than inconsistency between the assays.

      We agree that this distinction should have been explained more clearly in the original manuscript. We have therefore substantially revised the Figure 4 legend to clarify the rationale underlying both assays, explicitly distinguish IC<sub>50</sub> and CGI endpoints, and explain how the CGI values were used to guide subsequent antibiotic selection conditions for Blastocystis ST7-B transformants. The CGI transition range has also been made more visually explicit in the revised figure presentation.

      Figure 4 caption revised to: “Antibiotic potency and selection-window determination in Blastocystis ST7-B. Dose–response curves for puromycin, trimethoprim, and WR99210 were estimated from a resazurin-based viability assay (n = 3 independent replicates per drug per concentration). Points show mean ± SD, and the insets list the estimated IC50 values with R<sup>2</sup>-values > 0.75 for all fitted curves. The IC<sub>50</sub> estimates represent the drug concentrations that reduced resazurin-based metabolic activity by 50% under the assay conditions.”

      Right panels: “small-culture complete growth inhibition assay using 1 × 10<sup>7</sup> WT Blastocystis ST7-B cells per culture, assayed in triplicate across a wide range of concentrations. Cultures were incubated for 2 days, and outgrowth was assessed using phenol red acidification of the medium as a culture-level readout, with yellow indicating growth and red indicating no detectable growth. The yellow-to-red transition was used to estimate the concentration required for complete growth inhibition and to guide the subsequent antibiotic selection strategy for Blastocystis ST7-B transformants.”

      “IC<sub>50</sub> and CGI represent distinct assay endpoints: the former measures partial reduction in metabolic activity, whereas the latter identifies the concentration at which no detectable culture outgrowth occurs under the small-culture assay conditions.”

      (6) In Figure 3B, the unexpected UnaG fluorescence pattern could be due to protein sequestration because the protein is mildly toxic to the cell. This should be discussed in addition to the reasons already provided.

      We thank the reviewer for this thoughtful suggestion and agree that protein sequestration or reporter-associated cellular stress represent plausible alternative interpretations of the observed UnaG fluorescence pattern.

      We considered the possibility of UnaG-associated toxicity during interpretation of these data. However, under the conditions tested, we did not observe clear evidence of a substantial toxic effect: UnaG-expressing Blastocystis ST7-B cells could be recovered as stable lines, maintained under antibiotic selection, and propagated through continued culture. We therefore felt that direct attribution of the observed fluorescence pattern to reporter toxicity would currently remain speculative.

      At present, we consider the biochemical properties of the UnaG system itself to provide a more parsimonious explanation for the observed localisation pattern. In particular, unconjugated bilirubin is highly hydrophobic and would be expected to partition preferentially into lipid-rich cellular environments. This interpretation is consistent with the lipid-rich peripheral and intracellular structures previously reported in Blastocystis ST7-B (Liao et al., 2023).

      We have therefore revised the Discussion to acknowledge that the observed UnaG fluorescence pattern may reflect a combination of reporter-specific biochemical behaviour, bilirubin partitioning, local intracellular environment, or possible sequestration phenomena. At the same time, we avoid assigning toxicity as a demonstrated mechanism in the absence of direct measurements of cell fitness, reporter abundance, or bilirubin distribution. Such experiments would be required to evaluate this possibility rigorously.

      Lines 642-648: “Consistent with this, lipid-rich peripheral and intracellular structures have been reported in Blastocystis ST7-B, potentially providing favourable microenvironments for BR partitioning and contributing to the punctate UnaG fluorescence pattern (Liao et al., 2023). An alternative possibility is that the observed signal pattern reflects reporter sequestration or reporter-associated cellular stress. However, because UnaG-expressing lines were recovered, maintained under selection, and propagated through continued culture, toxicity remains a possible but untested explanation rather than a demonstrated mechanism.”

      Minor Comments

      Figure 2: Parts B and C should also show individual datapoints for better reader assessment.

      We agree that inclusion of individual data points improves transparency and interpretability of the underlying data distributions.

      Individual data points have now been overlaid on the boxplots in Figures 2B and 2C.

      Figure 3A: Separate channels (fluorescence, bright-field, merge) should be shown rather than only the merge. The current overlay is difficult to interpret, especially for colour-blind readers.

      We appreciate the reviewer’s concern regarding accessibility and interpretability. We considered separating the fluorescence, bright-field, and merged channels for Figure 3A. However, this panel was intended primarily as an overview demonstrating reporter detectability within the bicistronic construct context, while the detailed fluorescence distribution is explored more extensively in the subsequent UnaG panels. We therefore retained the merged presentation for Figure 3A. Importantly, the image is not dependent on red–green discrimination, as it combines a greyscale bright-field background with a high-contrast green/cyan fluorescence signal that remains distinguishable through brightness and contrast differences. In addition, colour-blind-friendly lookup tables (LUTs) were used throughout the revised figure set.

      To further improve accessibility, the original red annotation arrow has been replaced with a colour-blind-friendly annotation colour.

      Briefly define system components (P2A, UnaG, smURFP, SNAP-tag) and add an abbreviation list.

      We agree that brief contextual definitions improve accessibility for readers less familiar with these reporter systems. Rather than adding a separate abbreviation list, we have added short explanatory descriptions at the points where these components are first introduced in the manuscript.

      Lines 394–396: “The P2A peptide is expected to promote ribosomal skipping during translation, allowing two separate polypeptides to be produced from a single open reading frame.” Line 515–516: “UnaG, a bilirubin-binding fluorescent protein originally isolated from the muscle of the Japanese eel (Kumagai et al., 2013)…” Line 527: “smURFP (small ultra-red fluorescent protein)…”

      Abstract: “among the most prevalent microbial eukaryote” should be “eukaryotes”.

      Corrected in revised manuscript.

      Conclusion (2nd sentence): unclear what “endogenous regulatory part discovery” means.

      We agree that this phrase required clarification. The intended meaning was the identification and benchmarking of native Blastocystis ST7-B promoter and terminator elements for construct design and toolkit development. We have clarified this directly in the revised Conclusion section.

      Lines 682–683 revised to: “By bringing endogenous regulatory part discovery, namely the identification of native promoter and terminator elements, …”

      Author contributions: “critical advise” should be “advice”.

      Corrected in revised manuscript.

      Again, we thank the reviewers for their careful evaluation, constructive criticism, and thoughtful feedback on the manuscript. The review process has substantially strengthened the manuscript by helping us clarify the distinction between what is directly demonstrated experimentally and what remains mechanistically unresolved.

      The central methodological conclusions of the study remain unchanged: the toolkit enables selectable transgene expression, recovery of colony-derived lines, and propagation of reporter-positive transgenic Blastocystis ST7-B lines, extending genetic accessibility in this organism substantially beyond the previous transient transfection framework.

      At the same time, the revised manuscript now more explicitly acknowledges important unresolved mechanistic questions, including vector topology, P2A-mediated protein separation efficiency, and persistence in the absence of selection. These are now discussed transparently together with the future experimental approaches that will be required to address them directly.

      We believe the revised manuscript now presents a clearer, more rigorous, and more accessible description of a practical genetic toolkit for Blastocystis ST7-B and hope that the revisions and clarifications satisfactorily address the reviewers’ concerns.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Vasquez-Correa and colleagues describes the expression pattern of the ocelli (simple eye) gene regulatory network in ants. They correlate the expression pattern of these genes with the presence and absence of ocelli in different classes and species of ants. The presence of ocelli is a polyphenic trait in ants - understanding the molecular and developmental underpinnings of polyphenic traits is of significant interest to evolutionary biologists, developmental biologists, and ecologists. The authors propose that the presence of the latent expression of the ocellar network in classes of ants that do not display ocelli in the adults may underlie the re-evolution of ocelli within the ant lineage.

      Strengths:

      The strengths of the manuscript are that it is well written, the images are of the highest quality, and the data support the conclusions of the authors.

      We thank Reviewer 1 for their positive comments.

      Weaknesses:

      One improvement that could be made is to include imaginal discs of the queen ants as well as scanning electron images of the ocelli of the queen ant to match the pupal stage images of the worker and soldier ants. A second improvement is to attempt a gene knockdown using RNAi or similar methods to ensure that the genes that are being studied are, in fact, responsible for ocelli development in the ant.

      The reproductive caste in ants is typically composed of both winged males and winged queens. We agree with Reviewer 1 that the queen caste, which develop 3 fully functional ocelli, is an important point of comparison in our study to the wingless minor workers and soldiers. Unfortunately, however, laboratory colonies rarely produce reproductive queens, and in the field, queen production in colonies of C. floridanus occurs within a narrow seasonal window, making the collection of queen larvae particularly challenging for developmental work. In contrast, the winged males, which also develop 3 functional ocelli like the queens for help during mating flights, can be readily generated in the lab throughout the year. Therefore, we use males as a proxy for characterizing ocelli development and GRN in queens and the winged reproductive caste as a whole. Given the deeply conserved gene regulatory networks underlying this trait across insects, we believe this is a reasonable assumption.

      We also agree with Reviewer 1 that using RNAi to knock down genes in the ocelli GRN would improve the study. For completeness of the scientific record, we would like reviewers and readers to know that we actually did, in fact, invest significant effort trying to knock down otd-1 (ortholog of the Drosophila otd gene), which functions as key upstream regulator of ocellar development. In Drosophila, RNAi knockdown of otd disrupts the development of all three ocelli as well as fine morphological features on the anterior of the head. In C. floridanus, otd -1 is expressed in the head capsule and brain (see Author response image 1 in this response). Injection of dsRNA of otd-1into whole soldier-destined larvae, significantly reduced otd -1 expression in the brain relative to its control, while in the head capsule, otd -1 expression remained largely unchanged relative to its control (see Author response image 1 in this response). This indicates that in the same individual, the injected otd -1 dsRNA was able to penetrate and significantly reduce otd -1 expression in the brain, but, was unable to penetrate the head capsule, where otd -1 expression remained largely unchanged. No ocellar phenotypes could be observed in pupae or adults. Therefore, for technical (not biological) reasons, we were unable to knockdown genes in the ocelli GRN in the head capsule. We hope to solve this technical problem in the coming years to add a mechanistic explanation for the latent expression and maintenance of the ocelli GRN in workers that completely lack ocelli as adults.

      Reviewer #2 (Public review):

      Summary:

      The manuscript titled "Latent gene network expression underlies partial re-evolution of a polyphenic trait in the worker caste of ants" by Vasquez-Correa et al. aimed to study genetic mechanisms underlying developmental plasticity, especially binary polyphenism in queen vs worker ant castes. This is an interesting question regarding the extent to which phenotypic traits were altered, lost or regained, and how molecular pathways (upstream vs. downstream) can facilitate this process.

      In ants, reproductive castes (queens and males) develop wings as well as 3 ocelli for mating flights and other activities, while worker castes are wingless, and in some species, they have either no or a reduced number of ocelli. The phylogenetic analysis showed that in the Camponotini ant clade, the one-ocellus phenotype revolved in three species independently. The authors analyzed the conserved developmental pathways between Drosophila (well-established) and ants using HCR (a high-quality in situ hybridization technique). They found that although upstream genes for the development of ocelli (otd and hh) showed similar expression between castes, downstream genes (toy, eya, and so) had reduced or no expression in workers of C. floridanus, and this differential expression may lead to partial or complete loss of ocelli. Consistently, workers develop rudimentary tissues, suggesting that they initiate the ocellus developmental process but somehow stop it before adulthood.

      Strengths:

      Evo-devo approaches to reveal conserved molecular pathways of ocellus development. High-quality HCR provided convincing evidence of the expression of key genes in ocelli, eyes and antenna throughout larval development.

      Using HCR, the authors showed differential expression of downstream genes in males vs. soldiers vs. minor workers of C. floridanus, which might explain phenotypic differences between castes.

      We thank Reviewer 2 for their positive comments.

      Weaknesses:

      Although the molecular pathway is conserved, the mechanism underlying the lack of ocelli in workers remains unclear. In C. floridanus, it could be explained by the evidence of no expression of certain developmental genes, but in other species, e.g. Polyrachis rastellata, is their expression intact, or reduced? There is no control male.

      In addition, HCR in species with partial re-evolution (if their genomes have been sequenced) would be useful to understand the mechanism. For example, there might be differential spatial expression between medial and lateral ocelli.

      We agree with Reviewer 3 that investigating the mechanisms underlying the lack of specific ocelli in these and other species is the next step for this research. Here, our main focus was instead on trying to explain the mechanisms underlying partial reversion of ocelli through the persistence of ocelli GRN expression in adult workers lacking ocelli. We therefore focused on the latent expression of the ocelli GRN in Polyrachis rastellata, a species that completely lack ocelli in adult workers, and how it may have facilitated the partial reversion of a single ocellus in its congener Polyrachis bihamata. Therefore, although we did not reveal specific interruption points in the ocelli GRN in Polyrachis rastellata, our results showing that this species expresses three genes of the ocelli GRN, offers sufficient evidence that this network is conserved and likely facilitated the partial reversion to a single ocellus in P. bihamata.

      We also agree with Reviewer 3 regarding the male control in P. rastellata and obtaining the species in our study that have undergone partial re-evolution. Unfortunately, these ants occur in Southeast Asia and are very difficult to collect. For males in P. rastellata, our colony died before we could try to induce male development. However, given the deep conservation of the network in the males of a genus within the same subfamily (Camponotini), we feel it is reasonable to assume that the network would also be conserved in the males of P. rastellata, especially since the genes we sampled are conserved in workers that do not develop ocelli as adults. As am sure the Reviewer may know that this is a continual challenge of working with emerging models in evodevo.

      Reviewer #3 (Public review):

      Summary:

      This paper examines the loss and re-evolution of specific organs during the evolution of ants. The authors show that these organs, the ocelli, disappear and are re-evolved in different ant species and in different ant castes within these species. The authors show that this is linked to to a conserved GRN discovered in Drosophila, that appears to underlie the development of the ocelli, and demonstrate that this GRN appears to remain active in the developing heads of ants that have no ocelli- implying that it is the evolutionary latency of this GRN that allows loss and subsequent evolution.

      Strengths:

      This manuscript has outstanding imaging of a very difficult developing organ, and the key data, fluorescence in situ hybridisation, is done well and clearly shows what the authors wish to demonstrate. The methods are well described and underpin the whole work.

      The authors convincing demonstatrate that gene expression patterns imply the conservation of the ocellus gene regulatory network from Drosophila to ants. They further show that this network is present even in ants that don't produce an adult ocellus, but do show that in those species, loss of a developing nascent ocellus (which they identify) occurs at the same time as an interruption in the expression of the key genes in the GRN. All of this data is beautifully presented and explained.

      We thank Reviewer 3 for their positive comments.

      Weaknesses:

      There is one key weakness in that there are no functional students that indicate that the GRN actually does make the ocellus, though the expression patterns are convincing. This applies to loss of the ocellus as well. It would be nice to see that transient loss of the ocelli GRN might lead to loss of ocelli in ant species that have them. These are very difficult things to achieve, as the key genes have earlier developmental roles, such that CRISPR knockouts would not be interpretable, and transient RNAi in the head capsules of developing pupal ants would be challenging.

      We agree with Reviewer 3 that functional experiments in species where workers both have ocelli present and absent is a key next step in this research. Please see our response to Reviewer 1 on our failed attempts to achieve this. We are therefore grateful to Reviewer 3 for acknowledging the challenges in trying to establish RNAi and CRISPR in the head capsules of developing workers in these ants. Also, please see our response to Reviewer 2 on the difficulty of finding and collecting these ants, which occur mainly in Southeast Asia.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      One improvement that could be made is to include imaginal discs of the queen ants as well as scanning electron images of the ocelli of the queen ant to match the pupal stage images of the worker and soldier ants.

      A second improvement is to attempt a gene knockdown using RNAi or similar methods to ensure that the genes that are being studied are in fact responsible for ocelli development in the ant.

      Please see our response to Reviewer 1 above.

      Reviewer #2 (Recommendations for the authors):

      For the questions below, if there is no experimental evidence, consider addressing them in the Discussion.

      Do sizes of ocelli different between castes? For example, even workers have 1-3 ocelli, their sizes are smaller than those of males/queens, especially in workers with 1 ocellus. If so, might it be continuous (not binary) changes in downstream gene expression that control ocellus size, with no ocellus below threshold? Does this favor the hypothesis of threshold but not switch?

      We thank Reviewer 3 for highlighting an important point about the size and development of ocelli. Observations suggest that ocelli tend to be larger in queens and males than in workers and soldiers in species with ocelli. However, we lack quantitative data to test this conclusively. We now include a sentence on the Discussion stating that an important avenue of future work should investigate whether threshold or switch mechanisms influencing the presence/absence, as well as size, of ocelli between queens and workers.

      For the species whose workers have a single ocellus, are there variations, e.g. spanning from 0, 1 to 2? If 2, always one medial plus one of the two laterals? If always one, it would be a good control for staining to see up- vs down-regulation of downstream gene expression within the same individual.

      We agree with Reviewer 2 that this is a fascinating approach to our question. We have not observed natural wild-type variation in the number of developing ocelli in the same-sized individuals in the worker caste. However, in a distantly related leaf-cutting ant species (Atta cephalotes) belonging to different subfamily (the Myrmicinae) individuals with different head-to-body scaling within the same colony can vary in the number of ocelli. For example, soldiers of Atta cephalotes include individuals developing one, two, or three ocelli. These configurations can appear as only the median ocellus, only the two lateral ocelli, or even the median plus a single lateral ocellus. Interestingly, these correlations vary with changes in the size and head-to body scaling, suggesting that each ocellus can undergo different degrees of development, with one or more remaining vestigial or completely absent. On the other hand, workers in other species consistently develop a single ocellus, like in workers of Polyrachis bihamata, with no correlation to size or head-to-body scaling. These cases highlight how evolutionarily labile this trait is among workers of different ant species, which supports our proposal that the underlying gene regulatory network remains latent, thereby facilitating the emergence of novel trait combinations. We therefore agree on the importance of comparing the developmental mechanisms underlying these patterns temporally across larval stages and between individuals within a colony. We have now incorporated 2 sentences into the discussion, stating that this will be an important avenue for future work.

      Is there any function of a single ocellus in workers, or just a consequence of incomplete down-regulation of gene expression?

      Thank you again for highlighting these important points that help us to elaborate on the discussion of our study. The functional role of ocelli in species that develop these structures remains largely understudied. However, for some species particularly within the Formicinae clade the function of the three ocelli in workers has been investigated, revealing that they serve as a celestial compass that facilitates navigation. We reference these findings in our Introduction and Discussion to illustrate that the presence of three ocelli in workers can represent an adaptive trait. In contrast, the functional significance of a single ocellus or of partially developed ocelli remains an important question. This knowledge gap presents a promising avenue for future research to understand the adaptive value of reduced, partially suppressed ocellar development. We have now added a sentence in the discussion stating this.

      In previous studies, JH treatment can increase the number of ocelli in workers, consistent with its role in promoting reproductive development. In the ocellus developmental pathway, what causes the reduction of downstream gene expression in C. floridanus? Does JH directly regulate their expression?

      We thank Reviewer 2 for proposing yet another interesting question for future investigation, which we have added to the Discussion.

      The only current evidence available in C. floridanus is a recent study (MacMillan et al. 2025), in which minor workers and soldiers were treated with JH at different developmental stages. Unfortunately, no evidence of ocelli induction was observed in JH-treated individuals, suggesting that the mechanisms of ocelli development in C. floridanus might be highly canalized, especially in species that exhibit worker polymorphism (inter-individual variation in size and head-to-body scaling within the worker cate). However, more studies are required to understand why in Monomorium pharonis (no worker polymorphism) ocelli development can be readily induced by JH, while in another C. floridanus (with worker polymorphism) it appears quite difficult.

      "In D. melanogaster, the head develops from the eye-antenna disc" This statement is not correct. The brain does not belong to the eye-antennal disc.

      We thank Reviewer 2 for catching the misspelling. We have changed the name to eye-antenna disc in the sentence.

      Reviewer #3 (Recommendations for the authors):

      It is hard to see the developing ocelli in Figure 7 - could the authors increase the contrast to make them more visible?

      We have made the suggested changes to Figure 7 in the main article, and it has indeed improved the figure.

      Author response image 1.

      RNAi knockdowns in developing soldiers of Camponotus floridanus show a reduction of otd -1 expression in the brain, but no effect on otd -1 expression in the eye-antenna disc. A. HCR revealing otd -1 expression in the brain B. qPCR of otd -1 expression after RNAi knockdown shows significantly reduced otd -1 expression in the brain, C. HCR revealing otd -1 expression in the eye-antenna disc D. qPCR of otd -1 expression after RNAi knockdown shows no significant affect on otd -1 expression in the eye-antenna disc.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      Cisplatin, a platinum-based chemotherapeutic agent, induces intra- and interstrand crosslinks, thereby blocking DNA replication and transcription and triggering apoptosis. The authors aim to demonstrate that DNA polymerase κ (Polκ), traditionally seen as a translesion synthesis (TLS) polymerase, able to synthesize DNA through DNA lesions, plays a non-catalytic, structural role in stabilizing replication forks and protecting cells from cisplatin-induced cytotoxicity. A key finding of this work is the identification of two novel molecular axes: PCNA-Polκ-Polδ, which facilitates efficient DNA replication; PCNA-Polκ-USP18, which stabilizes DNA damage response proteins. These findings provide actionable therapeutic targets for overcoming head and neck squamous cell carcinoma chemoresistance, a cancer with rising incidence and limited treatment options.

      Strengths:

      The study relies on a robust experimental design, including Polk allegedly CRISPR-Cas9 knockout, siRNA knockdown, and rescue experiments with wild-type, catalytically dead, and PCNA interaction-deficient Polκ variants, supporting a non-catalytic role of Polκ. The work also reports a strong implication of Polk in cisplatin resistance, the identification of USP18 as a possible Polk partner and the consequences of Polk depletion on post-translational stabilisation of DNA damage response proteins.

      Thank you so much for appreciating our efforts to demonstrate role of Polκ mediated axes in cisplatin resistance in head and neck cancer cells.

      Weaknesses:

      The findings reported in this manuscript cannot be generalized to all cisplatin resistance mechanisms, as cells may develop multiple adaptive strategies to survive chemotherapy. Polκ's role varies across cancer types. For example, it is downregulated in stomach and colorectal cancers but upregulated in HNSCC, lung, and ovarian cancers. Thus, its use as a biomarker or drug target may be context-dependent.

      We completely agree with you, and the presented data only support Polκ's role in HNSCC as demonstrated in both acute cisplatin exposure as well as the cisplatin-resistant HNSCC models. Other cell and cancer types may adopt different strategies for cisplatin resistance.

      Acute cisplatin exposure is sufficient to trigger Polκ upregulation to levels similar to those in resistant cells. However, it remains unclear how long this upregulation persists and to what extent it contributes to survival. Further, the sensitivity of cisplatin-naïve H357 or SCC9 cells (H357-S and SCC9-S) to Polκ knockdown has not been addressed. This is a critical question, as acute cisplatin exposure induces Polκ expression to levels similar to those in resistant cells. This could argue against a direct role for Polκ in mediating resistance and instead suggest indirect mechanisms (like Polκ-dependent mutations during adaptation).

      Since H357-S and SCC9-S cells are highly sensitive to cisplatin, knocking down of Polκ unlikely will alter the phenotype, as other TLS DNA polymerases like Polκ and Polκ play critical role in such lesion bypass. Since no other DNA polymerase was upregulated in these cells upon cisplatin exposure and in the cisplatin-resistant cells, it was intriguing to demonstrate a direct role of Polκ in chemoresistance and that has been proven in this study. Since the catalytic activity of Polκ is not required to induce chemoresistant in these cells, we strongly believe that Polκ-dependent mutagenesis play minimal or no role in adapting cells to tolerate cisplatin. Nevertheless, we will knock down Polκ in these cells and determine cisplatin sensitivity

      The experimental design and results aimed at demonstrating the existence of a PCNA-Polκ-USP18 axis (Figure 9A) do not fully support the conclusion that these proteins form a stable complex. This set of experiments also lacks essential controls, such as the immunoprecipitated bait and the amount of immunoglobulins precipitated in all conditions. This also applies to the colocalization experiments in cells shown in Figure 9B. Images are poor and lack quantification. Further, Polk is seen mainly cytoplasmic in the upper panel, while it is nuclear in the lower panel. Discrepancies in Polk subcellular localization are also evident in the Supplementary data.

      We appreciate the Reviewer's critical and insightful comment. In our view, the interaction between Polκ and USP18 is very specific as USP2 and IgG alone do not pull down Polκ. Similarly, we also show that both Polκ and USP18 interact with PCNA. We agree with the reviewer that the existence of a stable complex of PCNA-Polκ-USP18 has not been fully demonstrated in the current version. We will perform additional experiments to strengthen our finding: a) Co-IP experiments with Polκ PIP mutants (wild-type vs. mutant) should be performed to determine whether USP18 loses its ability to bind PCNA in the absence of Polκ-PCNA interaction. b) Mapping the domain in Polκ that is involved in USP18 binding and their Co-IP experiment. Additionally, high resolution co-localisation images including quantified data will be provided.

      USP18 is known to deubiquitinate ISG15-modified proteins (not just ubiquitin). The study does not rule out ISGylation as a contributing mechanism.

      We find the point raised by the reviewer is very intriguing, however, as it will require a significant amount of time and effort to demonstrate ISGylation of DDR proteins and deISGylation by UPS18, and the insight that we may gain is unlikely to add to the central theme of this paper, we will expand this in our subsequent related study. Thank you for the suggestion.

      The experimental design involving analysis of DNA synthesis dynamics at a single-molecule level is not appropriate. Over interpretation of the data in several parts of the manuscript and lack of rigor in performing the experiments. Inappropriate consideration and absence of discussion of previously published literature directly related to the subject studied in this manuscript. Discrepancy with a previous report regarding the role of Polκ in Chk1 phosphorylation (Tonzi et al., eLife 2018). Synergic effect of T2AA inhibitor and Cisplatin have been already described in « naive » cancer cells (Inoue et al, 2014).

      Thank you very much for the suggestions. We will take care of the portions and modify as suggested. The necessary reference will be added as appropriate.

      Another critical point is that the proliferation rate of Polk-depleted cells is slower than that of wild-type cells. Hence, the colony formation assay shown in Figure 2B can be misleading, since the observed differences can be interpreted only as a proliferation problem.

      Thank you for pointing this out and we will modify the portion for better clarity.

      Reviewer #2 (Public review):

      Summary:

      Building on earlier studies, the authors report a role for pol kappa in mediated cisplatin resistance. Their data on dispensability of pol kappa catalytic activity for cisplatin resistance is consistent with previous reports. They further demonstrate that the PIP box of pol kappa is critical for cisplatin response. Based on these observations, the study concludes that targeting pol kappa and PCNA interaction can be a viable approach to overcome cisplatin resistance.

      Strengths:

      Indications that interaction between Pol kappa PIP box and PCNA can be targeted to overcome cisplatin resistance.

      Thank you for appreciating our finding that the PIP box of Polκ is critical for cisplatin response

      Weaknesses:

      (1) The study has used a model of cisplatin resistance and found that the phenotype is specifically reliant on upregulation of Pol kappa. They also observe that in this model of cisplatin resistance, there is rapid degradation of multiple repair proteins, including ATM, ATR, HR and NHEJ proteins upon knocking out Pol kappa. However, it is unclear how the resistant model was derived. Also, since the data and almost all experiments in this manuscript were performed with a single model of cisplatin resistance, the conclusions should be taken with caution.

      We are extremely sorry for the lack of clarity. Please note that two cisplatin-resistant models (H357 and SSC9) have been used and the results were very consistent in both cells. Fig. 1C clearly demonstrates about the generation of these resistant models and the original reference has been already cited.

      (2) There are also inconsistencies in findings. Increased G2 arrest and no change in origin firing are being observed despite a significant reduction in Chk1 protein levels.

      Thank you for pointing this out. In our view, the increased G2 arrest is due to more fork stalling or collapsed than the new origin firing. Also, in our assay we observed less than 10% of new origin fired DNA fibres, and that could be the reason of no significant change in new origin firing among various cells.

      Reviewer #3 (Public review):

      This manuscript investigates the role of PolK in cisplatin repair. While in general it is considered that polK is not involved in the repair of cisplatin-induced DNA damage, the authors show that in a very specific scenario, namely cisplatin-resistant head and neck cancer cells, loss of PolK causes cisplatin sensitization, implying a role in cisplatin repair by polK in these cells. It is also implied that these cells acquire cisplatin resistance by overexpressing polK, but this is not really investigated. The authors then go on to show that DNA replication in the presence of cisplatin is affected by the loss of polK in these cells and also identify USP18 as a potential polK interactor in these cells with a similar phenotype. They claim that polK and USP18 form a pathway that allows cisplatin tolerance in these cisplatin-resistant head and neck cancer cells. The findings are interesting and useful to the field; however, the manuscript, in its current form, has several issues. Most importantly, the mechanism of USP18 has not been investigated. In addition, the manuscript does not flow fluidly, and instead, various experiments are put together without a clear logic. Some of the claims are not substantiated by the data shown.

      Thank you very much for finding our study interesting and the pending concerns will be addressed as suggested.

      (1) The experiments in Figure 1 using a few cell lines from various types of cancers are not enough to conclude that polK expression is specifically induced by cisplatin in some types of cancers but not others. Since the focus of this study is head and neck cancer, the authors should show the expression of PolK after cisplatin treatment in more head and neck cancer cell lines, and not just the two investigated.

      In this study, we have explored eight different cell types (breast, brain, liver, head and neck, pancreatic, prostrate, lungs, and kidney) to check the expression of Polκ upon cisplatin exposure, and HNSCC cells only showed Polκ up-regulation. Therefore, we went ahead for further demonstration of the role of Polκ in cisplatin resistance in OSCC using four different cell models (H357-S, H357-R, SSC9-S, and SSC9-R). By adding more cell lines to study will unlikely change the central theme of the paper. Yes, by acquiring and analysing clinical samples from the cisplatin responder and non-responders would have strengthen our finding.

      (2) It is unclear to me why the authors include H357-S in their experiments. If the idea is that these cells acquire resistance because they overexpress polK, then the authors should investigate this by exogenously overexpressing PolK in H357-S cells and test if these cells are cisplatin resistant.

      It’s an interesting point and we will check whether overexpression of Polκ in H357-S cells could induce resistance to cisplatin and alters IC<sub>50</sub>. Thank you for the suggestion.

      (3) In addition, the authors should create the polK knockout in H357-S cells as well and include it as a control in their experiments.

      We appreciate your suggestion. As suggested by Reviewer #1 also, we will check the phenotype of Polκ knockdown H357-S cells.

      (4) Page 6, line 28: the comet assay does not measure DNA degradation, but rather DNA breaks.

      Thank you for the suggestion, we will modify the text accordingly.

      (5) Figure 4B: How does the overexpression of PolK mutants compare to endogenous PolK expression? It is important to assess if this expression is similar or of much higher magnitude.

      Please note that GFP-Polκ has been overexpressed in H357 Polκ knockout cells to nullify the effect of endogenous Polκ, otherwise we will not be able to test the role of various Polκ mutants.

      (6) Page 9, line 22: "For such a function, the catalytic domain of PolK becomes dispensable, whereas its interaction with PCNA is sufficient to drive efficient replication". I do not understand what data the authors used to make this claim. The interaction and colocalization studies should be performed with the PIP mutant. Similarly, this mutant should be used in the HU DNA fiber assays.

      We are extremely sorry for the lack of clarity. The inference has been derived from two sets of experiments as shown in Fig. 4C and Fig. 4D (and is with HU).

      (7) It is unclear how USP18 acts. What are its substrates? Chk1/2, BRCA1, BRCA2? This needs to be investigated. The impact of PolK on this activity needs to be assessed as well (is PolK needed for USP18-mediated de-ubiquitination of these DSBR proteins?). As it stands, the manuscript does not address the mechanism of USP18 in DNA repair, which is billed as the main finding of the paper.

      It has already been demonstrated in Fig. 9C where by knocking down USP18, the DDR proteins like Chk1, Chk2, CtIP, and Artemis can be recovered for ubiquitin-mediated proteasomal degradation. The same results are also obtained when its interacting partner Polκ is deleted. In our view, the presented results have sufficiently demonstrated the role of Polκ-Usp18 in the repair of cisplatin adducts through DDR proteins.

      (8) Do PolK and USP18 interact directly? Experiments using recombinant proteins would be useful to address this.

      We appreciate your suggestion. Since the Usp18 protein is not readily available, we will not be able to show; however, we believe the interaction is direct, and we will be able to map the binding site in Polκ.

    1. Author response:

      eLife Assessment

      This is a potentially important study comparing LTP mechanisms between primates and rodents. The experimental methods have some possible confounds, and the power (replicates) and design of the statistical methods could be strengthened, hence the support for the central claims of species differences is currently incomplete.

      We thank the Editor and the Reviewers for taking the time to carefully review our manuscript and for providing constructive comments and suggestions, as well as the opportunity to revise our work.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an important paper examining LTP induced by theta-burst stimulation in hippocampal slices from macaques and rats. While both species show theta-burst-late-LTP, only the non-human primate theta-burst-late-LTP showed synaptic tagging and capture that converts early-LTP into late-LTP in an independent synaptic pathway.

      Strengths:

      Synaptic tagging is a fundamental feature of repeated 100 Hz-tetanus-induced LTP, whereas theta-burst induction is arguably more physiologically relevant. Thus, synaptic tagging during theta-burst may differ in the two species, a distinction that may prove important in the mechanisms underlying the cognitive differences between the species.

      Weaknesses:

      Bursts repeated at the frequency (~5 Hz) of the endogenous theta rhythm induce strong LTP, primarily because this frequency disables feed-forward inhibition and allows sufficient postsynaptic depolarization to activate voltage-sensitive NMDA receptors. Therefore, the species differences may be due to differences in inhibition, rather than in molecular mechanisms of maintenance. One way to assess the relative strengths of this early induction mechanism in rats and macaques is to examine the "depolarization envelope" during the sequential bursts, which may be determined from the recordings already obtained. (Larson and Munkácsy, Theta-burst LTP, Brain Res 2015 Sep 24:1621:38-50. doi: 10.1016/j.brainres.2014.10.034)

      Another issue is that the PKMzeta-antisense oligodeoxynucleotides block the synthesis of the kinase. However, Mei F, Nagappan G, Ke Y, Sacktor TC, Lu B (2011), BDNF Facilitates L-LTP Maintenance in the Absence of Protein Synthesis through PKMzeta. PLoS ONE 6(6):e21568, provided evidence that BDNF and theta-burst stimulation can act to increase PKMzeta by a protein synthesis-independent mechanism, presumably through decreased degradation. Therefore, the absence of an effect of the PKMzeta-antisense does not exclude the possibility that persistently increased PKMzeta is the mechanism of theta-burst-late-LTP maintenance in mice or macaques. This issue is worth discussing.

      We sincerely thank the reviewer for the positive evaluation of our study and for highlighting the significance of examining synaptic tagging and capture following theta-burst stimulation (TBS) in rodents and non-human primates.

      We agree that TBS is a physiologically relevant induction paradigm and that differences in inhibitory circuit dynamics may also contribute to the species-specific effects observed in our study. As highlighted by Larson and Munkácsy (2015), repeated bursts delivered at theta frequency (~5 Hz) can transiently suppress feed-forward inhibition through GABAB receptor-mediated mechanisms, thereby enhancing postsynaptic depolarization and facilitating NMDA receptor activation. We therefore agree that species differences in inhibitory regulation and burst-evoked depolarization may contribute to the distinct expression of synaptic tagging and capture observed between rats and non-human primates.

      We further agree that analysis of the “depolarization envelope” during sequential bursts may provide additional insight into the relative strengths of early induction mechanisms. We will therefore perform these analyses using the existing recordings and compare the depolarization envelope between rodents and NHPs in the revised manuscript. Following the reviewer’s suggestion, we will expand the Discussion section to acknowledge the potential contribution of inhibitory circuit dynamics and depolarization envelope differences during sequential bursts.

      Importantly, however, we believe that differences in downstream molecular maintenance mechanisms also contribute to these species-specific effects. In support of this, our molecular analyses revealed enhanced recruitment of plasticity-related proteins and transcriptional pathways in NHP hippocampus following TBS, including increased expression of BDNF and PKCζ. These findings suggest that both induction-related network properties and downstream molecular stabilization mechanisms may collectively contribute to the enhanced associative plasticity observed in NHPs.

      We also thank the reviewer for the important point regarding PKMζ antisense experiments and the study by Mei et al. (2011). We agree that the absence of an effect of PKMζ antisense oligodeoxynucleotides does not necessarily exclude a role for persistently elevated PKMζ in the maintenance of theta-burst late-LTP. As demonstrated by Mei et al., BDNF together with theta-burst stimulation can maintain late-LTP in the absence of protein synthesis, potentially through stabilization of PKMζ protein levels by reducing degradation rather than through de novo synthesis. However, these findings are not directly comparable to our study, since our experiments involved theta-burst stimulation alone without exogenous BDNF application. Interestingly, our results suggest species-specific differences in the interaction between BDNF and PKMζ signaling pathways. In rats, TrkB/Fc-mediated blockade of BDNF impaired TBS-LTP maintenance, whereas PKMζ inhibition alone had no significant effect. In contrast, in NHP hippocampal slices, inhibition of either BDNF signaling or PKMζ alone failed to abolish late-LTP, whereas simultaneous inhibition of both pathways disrupted LTP maintenance.

      These findings suggest that endogenous BDNF signaling and PKMζ may operate through partially redundant or compensatory mechanisms, particularly in the primate hippocampus. Therefore, although our findings indicate that de novo PKMζ synthesis may not be strictly required under the present experimental conditions, we cannot fully exclude the possibility that protein synthesis-independent stabilization or maintenance of PKMζ contributes to theta-burst late-LTP maintenance in rodents or NHPs. We will now clarify this point in the revised Discussion section.

      Reviewer #2 (Public review):

      Summary:

      This study compares theta-burst stimulation (TBS)-induced synaptic plasticity in hippocampal CA1 slices from rats and non-human primates (Macaca fascicularis). The authors report that while TBS induces persistent LTP in both species, only primate hippocampal slices exhibit synaptic tagging and capture (STC) under these conditions. They further show increased BDNF and PKMζ expression following TBS in primates and propose that a redundant BDNF/PKMζ signaling architecture supports persistent plasticity in primates, whereas rodent TBS-LTP depends primarily on BDNF. The work aims to identify species-specific specializations in associative plasticity with implications for translational neuroscience.

      Strengths:

      The topic is potentially important because direct comparisons of hippocampal plasticity mechanisms between rodents and primates are rare.

      Weaknesses:

      (1) Limited biological replication in the primate experiments

      The manuscript's strongest claims rely on data obtained from 36 slices from 7 monkeys, qPCR analyses with n=3 biological replicates, and Western blot analyses with n=3 biological replicates. The effective sample size for species-level conclusions is therefore not large. The manuscript frequently treats slices as independent observations while drawing conclusions about species differences. This is particularly problematic for electrophysiological experiments because multiple slices appear to originate from the same animals. The statistical unit should be the animal, not the slice, unless nested analyses are performed.

      The authors should (1) report the number of animals contributing to each experiment, (2) provide animal-level analyses, (3) use mixed-effects or hierarchical models where appropriate, and (4) clarify whether multiple slices from the same monkey contributed to the same experimental condition. Without these analyses, the evidence for species-specific mechanisms remains weaker than presented.

      We thank the reviewer for this important and thoughtful comment regarding statistical interpretation and biological replication. We agree that, particularly for electrophysiological experiments where multiple slices may originate from the same animal, the effective sample size for species-level conclusions should be considered at the animal level rather than solely at the slice level.

      In the revised manuscript, we will clearly indicate the number of biological replicates (animals) together with the number of slices contributing to each electrophysiological experiment, as well as the biological replicates used for qPCR and Western blot analyses. We will also clarify whether multiple slices from the same NHP/rat contributed to the same experimental condition. These details will be incorporated into the figures and figure legends wherever appropriate.

      In addition, we will perform animal-level analyses by averaging slice responses within each animal prior to statistical comparison and, where appropriate, apply hierarchical or mixed-effects statistical models to account for the nested structure of slices within animals.

      We acknowledge that the number of non-human primates (NHPs) available for this study was inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with primate electrophysiology and tissue collection. Consequently, achieving sample sizes comparable to rodent studies is often not feasible in NHP research. Nevertheless, to further strengthen the biological robustness of the findings, we are currently in the process of obtaining additional NHP brain samples and plan to repeat key experiments in an additional 3-4 animals. We believe these revisions and additional experiments will substantially strengthen the statistical rigor and overall interpretation of the study.

      (2) The central STC conclusion requires stronger controls

      The most important result is that TBS supports STC in primates but not rats (Figures 1F-G). However, several alternative explanations are not excluded. For example, only a single interval (30 min) between TBS and WTET is examined. Classical STC studies characterize tag duration, PRP availability window, and temporal asymmetry. The current work does not determine whether primates exhibit longer tag persistence, increased PRP synthesis, altered capture efficiency, or merely a shifted temporal window. A temporal series (e.g., {plus minus}15, {plus minus}30, {plus minus}60, {plus minus}90 min) would substantially strengthen the mechanistic interpretation.

      We thank the reviewer for this insightful comment regarding the mechanistic interpretation of the STC findings. In the present study, we selected the 30 min interval based on well-established classical STC paradigms in rodents, where this interval reliably falls within the effective tagging and capture window. Using this experimentally validated interval allowed us to directly compare whether TBS is sufficient to support STC in primates versus rats under equivalent experimental conditions. Accordingly, the primary objective of this study was to determine whether TBS-induced STC varies across species, rather than to comprehensively define the temporal dynamics of the tagging window.

      We agree, however, that the current experiments do not distinguish whether the primate-specific effect reflects prolonged tag persistence, enhanced plasticity-related protein (PRP) synthesis, altered capture efficiency, or a shifted temporal window. Addressing these possibilities would indeed require systematic temporal interval analyses (e.g., ±15, ±30, ±60, and ±90 min), which represent important future directions. Such experiments are particularly challenging in non-human primates because the availability of primate tissue and experimental resources for large-scale electrophysiological studies remains limited and is currently beyond our experimental capacity due to substantial ethical, logistical, financial, and technical constraints.

      Nevertheless, we fully agree with the reviewer that these experiments are important for advancing the mechanistic interpretation of the findings. Similar temporal analyses have recently proven informative in our rodent studies (Chong YS, Ang SR, Sajikumar S. Commun Biol. 2025;8:553). Importantly, we are currently in the process of obtaining additional non-human primate samples and plan to extend the present work by examining an additional 60 min temporal interval to further characterize the temporal properties of synaptic tagging and capture in non-human primates.

      (3) Species differences may reflect tissue quality or preparation differences

      The manuscript compares 5-7 week-old rats with 5-7 year-old monkeys. These are very different developmental stages. Moreover, euthanasia methods, extraction procedures, and post-mortem handling are different. These factors can affect BDNF expression, protein synthesis, LTP magnitude, and transcriptional responses. The authors should discuss these caveats more explicitly.

      We thank the reviewer for raising this important and insightful point. We agree that differences in developmental stage between the experimental groups represent an important consideration when interpreting potential species-dependent effects. In the present study, rat experiments were performed in 5-7 week-old animals, whereas non-human primate (NHP) tissues were obtained from 5-7-year-old monkeys. This difference largely reflects the practical, ethical, and logistical constraints associated with NHP research and tissue availability. We acknowledge that these ages are not developmentally equivalent and that maturation state may influence BDNF signaling, protein synthesis capacity, synaptic plasticity thresholds, and transcriptional responses relevant to late-LTP and STC mechanisms.

      We also recognize that differences in euthanasia procedures, tissue extraction, slice preparation, and postmortem handling between rodent and primate tissues may influence tissue physiology and electrophysiological properties. Although extensive care was taken to optimize tissue viability and maintain stable recordings within each species, these variables cannot be completely excluded as contributing factors to the observed differences.

      Accordingly, we will revise the Discussion section to more explicitly acknowledge these limitations and clarify that our findings support potential species-dependent differences under the present experimental conditions, rather than definitive intrinsic species-specific mechanisms. Nevertheless, despite the inherent challenges associated with NHP electrophysiological studies, we believe that the present findings provide an important initial framework for understanding the translational relevance of synaptic tagging and capture mechanisms across species.

      (4) Statistical reporting is incomplete

      Many comparisons report exactly Wilcoxon p = 0.0313 and U-test p = 0.0022, across numerous experiments. This suggests very small sample sizes and discrete nonparametric distributions. The manuscript should report exact n values for each comparison, effect sizes, and confidence intervals.

      Second, many genes and proteins are tested. No correction for multiple testing is described. The authors should state whether corrections were applied, and if not, justify this choice.

      We thank the reviewer for this important comment regarding statistical reporting and interpretation. We agree that the repeated occurrence of identical exact p-values in several nonparametric analyses reflects the relatively small sample sizes and the discrete nature of the statistical distributions. This issue is particularly relevant for the NHP experiments, where biological replication is inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with obtaining and processing primate tissue.

      In the revised manuscript, we will provide exact n values for all comparisons, including the number of biological replicates (animals) and slices where applicable. We will also include additional statistical details, including effect sizes and confidence intervals where appropriate, to improve transparency and facilitate interpretation of the reported findings. Furthermore, we are currently in the process of obtaining additional NHP samples and will attempt to include more biological replicates in the revised version to further strengthen the robustness of the analyses.

      We also agree that the issue of multiple testing should be addressed more explicitly, particularly because multiple genes and proteins were examined. In the revised manuscript, we will clearly state the statistical correction methods applied for multiple comparisons where appropriate. For analyses in which corrections were not applied, we will provide justification, noting that several experiments were based on hypothesis-driven candidate targets rather than exploratory large-scale screening analyses. These statistical considerations will be clarified in the Methods and Results sections.

      (5) Interpretation and significance

      The study addresses an important and understudied question: whether associative synaptic plasticity mechanisms differ between rodents and primates. The finding that TBS can support STC in the primate hippocampus is potentially novel and impactful. However, the mechanistic evidence remains incomplete, the molecular analyses are underpowered, and several key controls are missing. At present, the data support the conclusion that under the specific experimental conditions tested, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than in rat slices.

      The stronger claims regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and redundant BDNF/PKMζ architecture require additional experimental support.

      We thank the reviewer for this thoughtful and balanced assessment of our work. We agree that the present data primarily support the conclusion that, under the specific experimental conditions examined, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than that observed in rat slices. We also agree that broader interpretations regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and potentially redundant BDNF/PKMζ-related mechanisms require additional mechanistic investigation and experimental validation.

      Accordingly, we will moderate these interpretations throughout the revised manuscript and clearly state that these conclusions remain preliminary. We will further emphasize that additional experiments, including increased biological replication, expanded temporal analyses, and further mechanistic investigations, will be necessary to more conclusively define the basis of the observed species-dependent differences. Within our current experimental capacity, we are actively working to obtain additional non-human primate samples and plan to incorporate additional biological replicates and key follow-up experiments in the revised version to further strengthen the robustness of the findings.

      At the same time, we believe the present study provides an important initial contribution to an understudied area by directly examining synaptic tagging and capture mechanisms in the primate hippocampus. Given the limited availability of non-human primate electrophysiological data in the field, these findings may offer a valuable framework for future studies investigating the translational and evolutionary relevance of associative synaptic plasticity mechanisms across species.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors have undertaken an investigation of differences between two mammalian species, the brown rat and the crab-eating macaque, in the mechanisms supporting a well-established model of long-term Hebbian synaptic plasticity, Schaffer collateral to CA1 Long-term potentiation (LTP) in the hippocampus. LTP has been long-studied and deeply characterised due to its potential importance in modeling a strong candidate process for the central mechanism of learning and memory. LTP was first discovered in lagomorphs (rabbits), but has since been much more widely studied in rodents (mostly rats and mice), and there has been some complementary work revealing LTP in non-human primates and even in humans, revealing largely overlapping canonical mechanisms of induction, expression, and maintenance. More specifically, this study puts a particular focus on the fascinating associative features of this form of lasting synapse-specific modification, in which a synaptic input can be stimulated with a relatively weak induction protocol that will not produce lasting plasticity on its own, but can undergo lasting LTP if paired with stronger stimulation on a separate synaptic input to the same neuron. This associativity mechanism is particularly attractive within the Hebbian synaptic plasticity framework as it provides a candidate mechanism for associative forms of learning in which stimulus-stimulus, stimulus-reward, stimulus-punishment, or action-outcome associations are formed. A particularly attractive feature of this associative LTP is that there can also be a substantial time-lag between the strong stimulation of one pathway and the weaker stimulation of the other synaptic input, which only undergoes lasting LTP by hijacking the proteins synthesized as a result of strong stimulation elsewhere. This observation has led to the famous tagging and capture hypothesis as an explanation of how such synapse-specific change can be achieved on both stimulated inputs but not on other synaptic inputs, given the potential requirement for cell-wide protein synthesis. This theory, for which there is very strong experimental evidence, posits that a protein tag is left at synapses that have been stimulated with sufficient vigor in recent history, serving as a key mechanism to ensure that those weakly stimulated synapses will undergo change when a larger-scale LTP event occurs due to stronger stimulation elsewhere within a relevant time window. Again, this idea is attractive as it can explain how we might form associations between events that occur slightly separated in time. The manuscript goes on to show that an induction protocol that is particularly physiologically relevant, theta burst stimulation, produces this tag and capture associative effect in ex vivo slices of Macaque hippocampus, much more readily than in side-by-side ex vivo slices of rat hippocampus. Moreover, the manuscript delves into the importance of well-characterised LTP maintenance mechanisms, including PKMzeta and BDNF, which are key factors that ensure that altered synaptic change is maintained for long periods of time despite substantial molecular turnover in the neuron. The observation in this manuscript is that a degree of redundancy for these mechanisms exists in the primate species but not the rodent species, as both mechanisms need to be inhibited to return LTP to baseline in the Macaque, but only one needs to be inhibited to have that effect in the rat. A major emphasis of this study is that there may be a step-wise difference in associative learning mechanisms between rodents and primates that may contribute to their differing cognitive capacities, although I believe a lot more evidence would be required to reach that conclusion.

      Strengths:

      The strengths of this study are that it is technically very proficient and is from a laboratory that has a long history of seminal work on synaptic tagging and capture. The cross-species comparison, particularly involving non-human primates, is also very hard to achieve, and a major strength here is the side-by-side comparison of slices from rat and monkeys. Further strengths of the study are the use of a number of experimental strategies, including both observation and intervention, to demonstrate differential involvement of LTP maintenance mechanisms. A final major strength is conceptual, as it is undoubtedly useful not only to identify shared mechanisms of plasticity between commonly used model organisms and either humans or much more closely related species such as old world monkeys, but also to reveal differences that have the potential to contribute to differences in memory/cognition.

      Weaknesses:

      The findings of this study are a very useful building block for understanding how generalisable mechanisms of LTP are. However, arriving at really substantial conclusions from these findings is challenging, as there are a number of variables that are unaccounted for in this study that may explain the differences that have been observed between rats and monkeys. One example of a potential confound to these interpretations is that rats are nocturnal/crepuscular animals, and macaques are diurnal animals. Thus, to undertake a like-for-like comparison, it would be necessary for the rats to be on a reversed light-dark cycle to ensure that the wake cycle of the rat (dark) is being compared with the wake cycle of the monkey (light). It is possible that the authors have done this, but it is not mentioned in the methods section. The reason this is important is that there is a substantial body of work indicating that different mechanisms are at play in hippocampal LTP during wake and sleep. Transcripts and proteins related to synaptic function are dramatically differentially regulated during sleep-wake cycles, and phosphorylation states of key proteins involved in plasticity are also altered. Moreover, synaptic tagging and capture are specifically disrupted by sleep deprivation. Perhaps the authors have already considered this factor and appropriately reversed the light-dark cycle of their rat subjects, in which case a clarification in the manuscript would be useful. Nevertheless, I have used this as an example because there is a variety of potential confounds that may explain the difference between SC-CA1 TBS LTP in rats and monkeys, e.g., circadian rhythms, degree of enrichment, natural light vs indoor lighting, diet, degree of inbreeding, strain, etc. Thus, to make strong conclusions about the potential for differences in plasticity rules/mechanisms and how those may contribute to differences in cognition, I think it would be necessary to compare a wider variety of species, including a good representation of each order (e.g., nocturnal rats and diurnal squirrels, new and old world primates) and not just a single exemplar. I understand, of course, that this is really pushing the boundaries of practicality, but I see no other way to make a strong conclusion or to generalise to mechanisms or properties of plasticity in rodent’s vs primates. Thus, while I believe the manuscript presents really admirable work, I am not sure the findings are at all easy to interpret.

      We thank the reviewer for this thoughtful and insightful comment, as well as for the encouraging appreciation of our long-duration plasticity recordings and associative plasticity experiments, which are both technically demanding and time-intensive. We fully agree that interpretation of cross-species differences in synaptic plasticity requires careful consideration of multiple biological and environmental variables, including circadian state, enrichment conditions, strain differences, diet, lighting conditions, and species-specific behavioral ecology.

      Regarding the specific concern related to circadian phase and sleep-wake state, the reviewer raises an important point. Rats are nocturnal animals, whereas macaques are diurnal, and hippocampal plasticity mechanisms are known to be influenced by circadian rhythms and sleep-dependent regulation of synaptic proteins and signaling pathways. Previous studies have demonstrated modulation of LTP, synaptic tagging and capture and protein synthesis in rats across normal sleep-wake cycles. We therefore agree that these factors may influence plasticity outcomes and should be carefully considered in comparative studies.

      Studies have further shown that theta frequency is highly sensitive to sleep-related manipulations. Specifically, theta frequency decreases immediately after sleep, remains elevated during sleep deprivation, and rapidly declines following recovery sleep. In aged animals, these effects appear comparatively attenuated, suggesting reduced sleep-dependent modulation of theta dynamics with aging. Therefore, disruption of normal circadian or sleep-wake patterns may significantly alter theta activity and associated plasticity mechanisms within a species and may not accurately reflect physiological baseline states (Utku Kaya et al., 2026).

      In our experiments, recordings from rats and macaques were performed during their respective active phases under standardized laboratory housing conditions, and we will further clarify these details in the revised Methods section. Nevertheless, we acknowledge that circadian state and related physiological variables cannot be completely excluded as contributing factors to the observed differences between species.

      More broadly, we agree with the reviewer that the present study does not permit definitive conclusions regarding universal “rodent versus primate” rules of synaptic plasticity. Our intention was not to propose a generalized dichotomy between rodents and primates, but rather to report that, under the experimental conditions used here, SC-CA1 TBS-LTP and associated synaptic tagging mechanisms differed between rats and macaques. We agree that broader evolutionary or cognitive interpretations would require systematic comparative analyses across multiple species, including both nocturnal and diurnal rodents as well as diverse primate species. Such studies would provide a stronger framework for distinguishing conserved versus species-specific mechanisms of plasticity.

      At the same time, we believe the present findings remain important because they provide one of the first direct experimental comparisons of SC-CA1 TBS-LTP-associated plasticity mechanisms between rodents and non-human primates under controlled ex vivo conditions. Although the interpretation should be done cautiously, the observed differences raise the possibility that certain metaplastic or protein synthesis-dependent mechanisms may not be fully conserved across species. Accordingly, we will revise the Discussion section to better emphasize the exploratory and comparative nature of the study, while explicitly acknowledging the limitations and potential confounding factors highlighted by the reviewer.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      This article describes a very ambitious metascience project aimed at testing the reproducibility of a corpus of publications conducted in Brazil. The strength of the approach lies in its systematic, multicenter replication design. The authors focus on three commonly used experimental paradigms in biology: the MTT assay, RT-PCR, and the elevated plus maze.

      The effort is commendable and reveals a rather low rate of reproducibility, in line with findings from fields considered less reproducible in the life sciences, such as cancer biology.

      Strengths:

      The study is supported by a substantial dataset, incorporating multiple independent replication attempts and the use of stringent, well-defined protocols, which strengthens confidence in the overall conclusions.

      We thank the reviewer for the comments.

      Weaknesses:

      (1) Being neither an expert in metascience nor in statistics, I cannot fully judge the methodological aspects of the article or its extensive supplementary material. I will therefore focus my comments on readability. I found the manuscript difficult to digest. The authors should improve readability if they wish to reach a broad audience of experimental biologists. In particular, they should simplify the description of protocols and highlight the key findings more clearly, using accessible language. See specific points below

      We can try to simplify the description of protocols at specific points for example, by providing an overarching description of the study design in the beginning of the Methods, rather than citing our previous eLife paper (Amaral et al., 2019), as suggested below. The methods are indeed quite extensive, but the this may be inevitable in a large-scale project such as this and we note that Reviewer #2 thought that part of the supplementary material should be incorporated back in the main text, which is a suggestion in the opposite direction. It may thus be hard to strike a balance between readability and comprehensibility that can address both reviewers’ opinions.

      (2) The article appears to oscillate between:

      (i) a description of the approach and the inherent challenges of such a multicenter replication program

      (ii) an estimation of reproducibility.

      These could potentially form two separate articles: one aimed at a broad audience emphasizing key results, and another focused on methodological aspects for a more specific metascience audience. The Results section currently contains redundancies and is difficult to follow for non-experts in statistics. I also find it challenging to extract the main findings.

      There is a bit of redundancy between tables and text, but this was intentional to make both of them self-explanatory. We also think stating the results in the text can allow us to make each of the replication criteria clearer, a concern that was also mentioned by the reviewer.

      As for requiring particular expertise in statistics for understanding, we mostly disagree. The main results (Tables 1 and 2, Figure 2) are expressed as percentages, and the only statistical concepts needed for interpreting these results are understanding prediction and confidence intervals. For this, we could provide a bit more guidance on their interpretation in the Methods section. Beyond that, most of the secondary results (e.g. Figure 3 and Figure 4) involve linear correlations, which is about as simple as statistical analysis gets.

      Of the results presented in the main manuscript, only Table 3 contains anything beyond percentages and correlations. We do agree that the meaning of each ratio in this table could be more clearly described, but there are essentially no expert-level statistics involved in their calculations.

      Other than that, the main statistical issues are the ideal way to aggregate the results from different replications for which we use different strategies for robustness purposes. However, all of these results are already in the supplementary material, so we don’t feel they interfere to much with the readability of the main manuscript.

      A possible improvement would be to include an initial section clearly describing the protocol (replication of a single experiment, across several labs, for three types of assays), followed by a concise presentation of the main results regarding reproducibility in Brazilian science with subsections.

      This is indeed a good idea, and we plan to include an initial overarching description of the project in the Methods section of the revised manuscript.

      Methodological details could be moved either to a Supplementary Information or to a more specific article, while being summarized in the Discussion.

      Again, this is the opposite of what was suggested by Reviewer #2, so we would rather keep the Methods section more or less at its current level of detail.

      (3) This study evaluates the reproducibility of a single experiment from each article, taken out of its broader context. While this provides an estimate of reproducibility, it does not directly contribute to resolving uncertainties within a specific field. This may represent a limitation compared to other reproducibility projects that attempt to replicate multiple key claims within a given study (e.g., in cancer biology or Drosophila immunity). I found that a weakness is that it does play a role in cleaning a field of wrong statements.

      The reviewer is correct in his interpretation. Evaluating the main findings of articles or cleaning a field of wrong statements was never a goal of our study (and we were clear about this from the start). Our aim with the project was metascientific (i.e. evaluate the reproducibility of biomedical experiments with a set of common methods) rather than driven by a particular interest in the findings themselves. This is reflected by our choice of selecting experiments from a random sample of articles from multiple fields, rather than filtering by area of interest or importance. It also underlies our choice to evaluate experiments rather than claims, as this was more statistically tractable and potentially more objective as a meta-research goal.

      To be clear, we don’t feel this approach is inherently better or worse than evaluating claims in the literature, as in the Drosophila immunity article case (i.e. Westlake et al., 2026), which is also an important goal. They are merely approaches that answer different questions. Ultimately, we probably made our choice based on (a) our expertise/interest in meta-research rather than in the fields the replications stemmed from and (b) an attempt to engage Brazilian researchers in the project in a way that was non-confrontational and minimized backlash from their peers. We feel this was valuable for many of the lessons learned, although it also meant learning less about the research findings in question.

      Even though this was not a goal of the study, there is some knowledge obtained about the findings that is indeed largely absent from the current manuscript. We do not feel the current format allows for much discussion of 45 different findings, but we do have plans to address these in future articles (as outlined in our response to point 5). In the meantime, qualitative descriptions of each experiment can be found at https://osf.io/w5z9a. This is already mentioned in the Methods but could be reiterated in the results as well.

      (4) The observation that external observers can predict which experiments are likely to be reproducible is interesting and should be more clearly emphasized.

      We did not go too deep into that finding because we are publishing a separate article focused on the prediction project, which should look into factors that correlate with prediction accuracy, both at the level of predictors (e.g. research field, career level) and of individual predictions (e.g. information taken into account for each answer). We also feel that, given the multiplicity of predictors in the prediction analyses, these findings are a bit tentative, as the strongest predictors may be subject to effect size inflation from the “winner’s curse” effect (as outlined by Reviewer #2). We can try to emphasize it a little more in the discussion (although it already merits a whole paragraph on pages 23-24), but we feel we would be able to discuss it more critically in a follow-up article.

      (5) The manuscript frequently refers to future publications. It would be helpful to clarify what is included in the present article versus what is deferred to subsequent papers.

      Indeed, some of our results did not fit this overarching analysis and were left for future publications. One of them is already available as a preprint, while the others are currently in preparation. Specifically, other results from the project should be spread about across five different articles.

      (a) A narrative article focused on challenges and lessons learned with the project, already published as a preprint at https://osf.io/preprints/metaarxiv/8y3tg_v1 (Amaral et al., 2026).

      (b) An article analyzing the prediction survey and markets results in detail (following the pre-analysis plan detailed in https://osf.io/6av7k/files/pjhgd and adding some exploratory analyses on prediction rationales).

      (c) Three articles describing the results of specific experiments with each experimental method (MTT, PCR, elevated plus maze) along with a discussion of aspects inherent to the method that seem to influence reproducibility.

      We can add this information more explicitly to the Methods section, including the links to the papers that have already been published at the time the manuscript is revised.

      Reviewer #2 (Public review):

      Summary:

      This is an important contribution to science, not only because large-scale replication studies remain rare despite their value, but also because this one focuses on research that was underrepresented in previous large-scale efforts. The findings reveal concerningly low replicability in this field, pointing to a problem that warrants immediate attention. Particularly noteworthy is the study's sampling strategy: by randomly selecting experiments from a wide range of publications based on methods, rather than filtering by research area, importance, or citation counts, the authors have produced results that are potentially more representative of the broader literature than those of previous large-scale replication projects in this and other fields. Overall, this is a fantastic contribution that I will be recommending and using in all my open science talks, and from which I have learned a great deal. Congratulations to the team!

      Thanks!

      Strengths:

      A study of this scale inevitably requires an enormous amount of work and methodological care, and this one is clearly both robust and thoughtfully designed. I want to particularly acknowledge the considerable efforts the authors have made to ensure the robustness of their findings. The use of multiple approaches to estimate replicability, combined with a substantial battery of sensitivity analyses, including a multiverse approach on top of everything else, clearly reflects the authors' genuine commitment to understanding their results and the limits of their conclusions. The transparency and sharing of all protocols, materials, and challenges and limitations encountered is also outstanding.

      We once more thank the reviewer for the compliments.

      Weaknesses:

      There were several instances during my reading of the methodology where I felt the authors relied too heavily on the external supplementary materials, at the expense of basic detail in the main manuscript. I appreciate how overwhelming it can feel to integrate more into an already substantial paper, but without some minimum integration, the reading experience and overall comprehension are too often compromised, at times posing more questions than answers. And it is unrealistic to expect most readers to engage with the extensive supplementary materials provided. Please see the comments below for specific suggestions.

      We do acknowledge that the article currently includes a lot of supplementary material. This includes both supplementary figures/tables relating to the paper and many supplementary methods files (mostly hosted at the Open Science Framework). However, we also note that this is already a rather long paper as it stands and that Reviewer #1 has made the opposite suggestion of simplifying it. Thus, it may be hard to strike a balance that will suit all preferences, and we feel that maybe our attempt has landed somewhere in the middle of both reviewers’ ideal versions of the paper.

      Additionally, I found the discussion rather underdeveloped. There is relatively little engagement with the broader literature, not only with replicability studies from other fields, but more generally with relevant meta-research work on publication bias, blinding, risk of bias, citation practices, etc. Some of the most novel and interesting findings in the paper also receive less attention than they deserve, and the discussion at times reads as a repetition of the results section rather than a critical engagement with them. I would encourage the authors to engage more deeply here, as the study clearly has much more to say. Doing so would further highlight why this study is important for the answers it provides and the questions it can spur. Again, please see the comments below for specific suggestions.

      We can try to engage with some of the above-mentioned literature in more depth in particular replication studies from other fields (some of which have appeared after our preprint (e.g. Tyner et al., 2026) and with the risk of bias and transparency literature (e.g. Serghiou et al., 2021). That said, we note once more that the article (and the Discussion section) are already quite long, and that analyzing each of these articles in depth is likely to be unfeasible.

      Specific suggestions:

      Page 1, abstract: "while t values for replications were positively correlated with researcher predictions about replicability, and negatively correlated with the rate of publications by the original article's last author" - I need to address the question: why t values and not effect sizes, p values, or something else? Update after reading the study: although the authors used others, they seem to place more emphasis on t values, which is not well explained. Without a clear explanation, it just left me wonder why, given that effect sizes would, in principle, be more information.

      Our original plan was to use p values as a predictor (see protocol at https://osf.io/9rnuj), but we later realized this was inadequate as it did not account for effect direction (i.e. significant effects in the opposite direction as the original may yield low p values, but this should not count as replication success). We thus switched to t values to be able to assign positive and negative signs depending on effect size direction. We note that, as we are using non-parametric Spearman coefficients (in which the module of t correlates negatively with the p value), the two approaches are effectively equivalent when original and replication effects have the same direction. This change was accounted for and justified in our list of protocol deviations at https://osf.io/9hj7t.

      Effect size (in relative terms) is already being used in the second predictor in the analysis (i.e. effect size decrease), as our idea was to use one significance-based predictor and one effect size-based predictor, to match what was done for the replication rates). We feel that using relative effects (e.g. response ratios) by themselves may not be as adequate, as for experimental methods with large coefficients of variation and/or low sample sizes (especially PCR ones), one can find large relative effects that are nevertheless far from statistical significance. This also makes relative effects not very commensurable between methods.

      We do believe there is a fair argument, however, to use standardized effect sizes as an alternative to t values (i.e. difference measured in standard errors of the mean) to measure significance/evidence strength. As some replications ended up underpowered, low t values may sometimes be due to insufficient statistical power/low sample size rather than replication failures. Using standardized effect sizes is not devoid of pitfalls (e.g. they can be quite variable when sample size is low), but it is worth doing as a robustness analysis.

      That said, there are a few statistical issues to be decided on how to calculate this (e.g. whether studies should be meta-analyzed using standardized mean differences rather than relative ones for this purpose, or whether an analog of the standardized effect size should be calculated for the log ratio of means). We would have to look more carefully into the multiple possibilities to decide on the best approach (and we do accept suggestions!).

      In the meantime, we note that running the prediction analysis using only experiments with ≥80% power yields a slightly higher correlation of t scores with researcher predictions (ρ = 0.49, p = 0.005), so we do not think that these underpowered experiments affect the trend too much. If anything, they could be masking a higher correlation between researcher predictions and replicability.

      Page 2, paragraph 2: "reproducibility (defined here as reaching the same results when analyzing a set of data)" - In my opinion, this definition is vague enough that it encompasses not only reproducibility (same data, same methods) but also robustness (same data, different methods), and I would therefore recommend providing a more precise definition. The same applies to replicability (different data, same methods), since the definition used does not highlight the importance of using the same methods, and thus also encompasses generalisability (different data, different methods). Explicitly clarifying these distinctions is particularly important as the field grows and the terms become increasingly mixed up and confusing.

      We agree that we should make the description more precise (e.g. “reaching the same results when analyzing a set of data in the same way” for reproducibility and “finding similar results with new data collected under similar conditions” for replicability). We will update these definitions in the revised manuscript.

      Page 2, paragraph 3: "All of these issues raise concerns about the replicability of published results - something that has not been evaluated systematically in the country" - I would suggest providing more information about why those factors may lead to expected lower replicability, ideally with a couple of sentences supported by references. As it stands, less experienced readers may not follow the argumentation and may consider it speculative.

      We would argue that the reader would be correct in this case: the argument is a bit speculative. It does go in the direction of what is generally accepted within the field (i.e. that publication pressure can lead to lower reproducibility for a range of factors), but we’re not sure this connection has been demonstrated empirically, except for indirect evidence (such as the lower reproducibility in papers stemming from top institutions and “trophy journals” in, the higher frequency of positive results in US states with more researchers in Fanelli, 2010, or the higher number of problematic images for highly productive researchers in some countries in Fanelli et al., 2022. We could cite this evidence in the introduction and make the speculated connection more explicit, perhaps adding modeling work as well (e.g. Ioannidis, 2005; Smaldino & McElreath, 2016) to explain why this could be the case. But essentially, our opinion is that the connection remains a speculation.

      Page 3, paragraph 2: "We then opened a public call for Brazilian labs that could replicate experiments using these methods and models, advertised by email, social media and lectures in conferences and institutions, to which 73 labs initially responded" - Since recruiting is an important component of this study, I would recommend providing additional details so the reader can better assess how comprehensive and unbiased the recruitment process was. AND Page 5, paragraph 2: Please provide more information about this open call: how was it advertised, where, and when? This is needed so that the reader can assess its comprehensiveness and potential biases. Even the link provided is not specific enough to understand the process, as it only states: "Calls were open to participants > 18 years old with current or previous experience in experimental research in any field and were advertised via e-mails, lectures and social media."

      We can offer a more detailed description of the recruitment process (e.g. number and distribution of lectures, social media strategy used, etc.), although we would rather do this in a supplementary document so as not to make the Methods section even lengthier. We note, however, that we never aimed to recruit a “representative sample” of labs from the country: we were busy enough trying to get enough labs for the project to happen, and aware that the call would be inevitably biased by our own communication capabilities and personal networks.

      That said, the response rates for different regions of Brazil do generally match the distribution of research labs and graduate programs within the country (with some distortions likely caused by our personal networks, such as the large number of labs in Rio de Janeiro state), and seem to indicate a rather wide dissemination of the call. One way to visualize this would be to present the distribution of corresponding articles from the original studies selected for the replication (or even from the whole sample of articles obtained for experimental selection) along with the distribution of labs at different stages of the project in Figure S3, which generally show similar patterns. This would actually lend support to our statement that “the population of labs that performed replications was largely similar to the one that produced the original results” in the discussion.

      Page 3, paragraph 2: "Based on the expertise of respondents and a feasibility analysis by the coordinating team, we selected 3 outcome assessment methods for replication" - Since this choice determined what was ultimately studied and who could participate, I would like to see more information to understand it: was it based on the most common expertise among respondents? How was feasibility defined and estimated?

      We tried to find the combination of methods that would maximize the number of labs that would be included in the project. This is explicitly stated in our Methods Selection document at https://osf.io/qxdjt, but could be stated more explicitly in the paper as well.

      Page 3, paragraph 3: How was the manual screening performed? Was it done by one or more people? Was there double-screening to ensure reliability of the screening protocol? Did the authors use a specific decision tree or tool? How were conflicts between observers resolved? Were any other validation steps taken to ensure reliability? The same comments apply to the data extraction (who, how many, validation, protocol, etc.).

      We initially used single screening by three different reviewers (see https://osf.io/6av7k/files/u5zdq for criteria), as we were merely looking for a sample of experiments; thus, comprehensive inclusion of all eligible studies was not a priority. After this initial screening step, inclusions were confirmed in a consensus meeting with the three reviewers involved.

      Data extraction was also done by a single individual, but the resulting data led to a protocol that was later checked by two reviewers who had access to the paper and were explicitly oriented to judge whether the protocol consisted in a valid replication. Thus, discrepancies between what was in the paper and what was included in the protocol could potentially be flagged at these stages (as they were in many cases). We do note, however, that this is likely not as effective to prevent errors as having data extracted independently, as reviewers may overlook mistakes more easily when comparing two documents rather than extracting data anew. We did find that some errors in extraction slipped by, such as an MTT experiment where treatment concentration was inadvertently changed from mM to μM in a particular protocol step; this was picked up and corrected by 2 out of the 3 labs, but not by the third one, leading the latter replication to be invalidated.

      Page 3, paragraph 3: As a non-expert, I would need more context about the expected average cost of experiments in this field; otherwise, I cannot assess how representative this sample is or whether potential biases may exist (e.g., cheaper experiments perhaps being expected to be less replicable than more expensive ones). Could expected costs also have affected the reduction in geographical coverage eventually observed in this study (Figure S3)?

      As stated in the manuscript, we initially capped experiments at a predicted cost of R$ 5.000 (around USD 1336 at that time), considering reagent cost alone (as equipment and labor was provided by labs), as mentioned in the manuscript. Exclusion rates for that reason were 12/74 (16%) for MTT experiments, 36/132 (27%) for PCR ones and 4/40 (10%) for EPM ones. This is stated at

      This turned out to be an underestimation in many cases, especially as it did not account for pilot experiments, need for repetition, etc; thus, many experiments ended up costing considerably more than that ceiling. As we had included a contingency fund for those cases which we expected would occur , we avoided removing experiments from the sample for this reason as much as possible. Nevertheless, one elevated plus maze experiment ended up not being replicated for cost reasons, as the necessary rat strain was provided by a single facility in the country, meaning that a large number of rats would have to be acquired and transported to all labs at a cost that we were not able to cover.

      As these costs were covered by the coordinating team, we do not feel that this is likely to underlie the reduction in geographical coverage. Other reasons related to lab structure could have led to labs in less well-resourced regions to leave the project, but they probably has nothing to do with the experiments selected.

      That said, the cost cap does mean that the selection of experiments is not completely representative of the literature, but is enriched in relatively cheap and simple experiments which were able to perform (which was our next step for selecting the final sample of experiments. Exclusion rates due to lack of lab expertise and/or infrastructure to perform the experiment were 21/56 (37%) for MTT experiments, 67/89 (75%) for PCR ones and 7/34 (21%) for EPM experiments.

      We will try adding some of this information to the flowchart in Figure 1, as we agree it provides more context on the representativeness of the selected experiments.

      Page 6, paragraph 2: "(on a scale of 1 to 5)" - Could you clarify whether 1 means no deviations and 5 means everything deviated? Is that how it was phrased to participants? Was there a threshold used by the coordinating team to decide how many deviations were acceptable? (I would briefly clarify all scales mentioned below to allow easier interpretation throughout.)

      The scale ranged from 1 (No relevant differences) to 5 (Very relevant differences that prevent considering the study as a direct replication). This scale was used for both the lab and the validation committee scores, and is described at https://osf.io/xgth2 (debriefing protocol) and https://osf.io/e3fjg (validation protocol).

      For the validation committee, we did use a threshold (any score of 4 or a sum of scores of 10 or more among 3 evaluators) to decide what had to be discussed to decide on inclusion, as mentioned on Page 7 of the Methods. For the labs, we used no threshold labs answered the protocol deviation question as a scale, but the decision of whether to consider the study a valid replication or not was not tied to this score.

      We can make both of these points (meaning of the scale and connection to lab’s decision to consider the replication valid) clearer in the Methods section.

      Page 6, paragraph 4: How were long-text answers (e.g., justifications) reviewed? Was this done manually by one or more members of the coordinating team, or using any text interpretation tool? What steps were taken to ensure the interpretation of these answers was as objective as possible?

      For the initial analysis of justifications, one reviewer read all answers and flagged those that seemed to concern reproducibility of the methods (e.g. “we replicated the protocol exactly as planned”) rather than results reproducibility (e.g. “effects went in the opposite direction”). We then revised these answers among the whole coordinating team to decide whether we should contact the lab asking them to revise them. We can add this information to the Methods section.

      For classifications of the justification into categories (i.e. Table S7), justifications were classified by two independent reviewers based on categories created after an initial inspection of the data, and discrepancies were resolved by consensus. We can add this information to the table legend.

      Page 8, paragraph 1: "If issues were found, the lab and coordinating team reviewed them via email until the sources of errors were identified and corrected (see https://osf.io/58vsx for details)." - Could you please provide information about how often these disagreements arose and briefly explain their causes? I am struggling to understand why these discrepancies occurred and how frequently. Without more detail, the error rate presented in the next paragraph is a little concerning.

      After we extracted data from the lab spreadsheets and summarized the results by code, labs received the results by e-mail and were asked to fill in a form on whether the results were in agreement with what they had found (see details at https://osf.io/nfr6y). Discrepancies in results at least 1 experiment were noted by 36% of the 53 (out of 56) labs that responded. Many of these stemmed from the coordinating team misunderstanding issues such as group identity or experimental unit identification in the spreadsheet. Others had to do with different ways to perform calculations (e.g. relative gene expression or % time spent in open arms). In some cases, simple errors in data transcription or typos caused the discrepancy.

      We were also surprised (and concerned) by the number of experiments in which we later found data errors that were not detected by this process (e.g. 18% of total). Our best understanding of this is that not every lab checked the results with the necessary care, as some errors were quite obvious, as in experiments in which sample size was different, or in which group labels were reversed. Ultimately, agreeing with a form that says “did you find any discrepancies?” may have been performed as a box-ticking exercise with little attention, and was probably not the ideal way to check data which led us to start reviewing results in live meetings afterwards. This is discussed in more detail in our challenges article (Amaral et al., 2026)

      Page 8, paragraph 4: Please provide the version of any package or software used throughout, and make sure to cite R appropriately (R Core Team XXX).

      R 4.5.1 was used for the analysis. We can add this information (which was present in the data repository in the R session info.txt file) and provide the R reference in the manuscript as well.

      In addition, did the authors calculate the log ratio of means (ROM/lnRR) using escalc()? If so, please report this.

      If not, I would recommend doing so, as escalc() implements recommended small-sample adjustments that produce slightly different values compared to a simple manual calculation of log(mean1/mean2).

      Yes, we did use the escalc() function for this calculation (for both the replications and the original effect sizes). We can mention this in the manuscript.

      Page 10, paragraph 1: "Coefficients of variation from the original study were compared to the mean coefficient of variation of its replications using Wilcoxon's signed rank test" - I wonder how these CVs were calculated - whether simply as SD/mean or using escalc() from the R package metafor, which includes a correction for small-sample size. This may affect the fairness of the comparison, particularly since CVs from original studies are expected to be slightly overestimated given their smaller sample sizes relative to the replications.

      We calculated the coefficients of variation as the pooled SD divided by the mean of both group means. The reviewer is correct about the possibility of small-sample effects in this case (which we were not aware of). We will thus look into the possibility of implementing this via the escalc () function in the analysis of the revised manuscript.

      We also acknowledge that this could be a source of bias in the comparisons between original and replication CVs (albeit likely a minor one). That said, we note that sample sizes are not always larger in the replication for some experiments with large original effects, power calculations sometimes yielded lower sample sizes in the individual replication, albeit infrequently. On average, though, replication sample sizes were indeed larger.

      I also have concerns about using the mean CV of all replications and comparing it to a single CV value, as this ignores the uncertainty around that mean.

      This is indeed the case; that said, the CV of the original effect also has random error relative to the true population CV and in that case, there is no way to estimate the uncertainty, as we have a single measure of that parameter. So there is probably no way around ignoring uncertainty in this case.

      We also note that we are looking for evidence of systematic CV inflation across all experiments (rather than for a statistically robust comparison between the CVs of any individual replication). For the sake of measuring this systematic inflation, the use of multiple experiments does allow us to estimate variability at the experiment level which should incorporate the lower-level variability between individual replications if this is not included in the model. Thus, we do not feel that our procedure introduced a systematic bias in the analysis at the experiment-level (although one could argue that it may lead to less precision).

      An additional check could involve calculating the log coefficient of variation ratio (lnCVR; Nakagawa et al. 2015, Methods in Ecology and Evolution; implemented in escalc()) between the original CV and each replication CV, and running a random-effects (or multilevel) meta-analysis that accounts for shared-control non-independence. I believe this would provide a more robust approach, as it does not ignore the uncertainty around the mean CV of the replications - uncertainty that, if neglected, is expected to increase the likelihood of false positive findings. This concern would also apply to the subsequent analysis on absolute means.

      We thank the reviewer for this suggestion, which indeed seems like an option in this case. We will look into this possibility, although we cannot guarantee at the moment that we will implement it, as we were not previously familiar with the method and will have to study it in more detail.

      Page 10, paragraph 2: The change in geographical distribution shown in Figure S3 appears rather striking, with western states disappearing step by step. Should the reader be concerned about the eventual geographical representability of the sample?

      Yes, but there are likely different reasons for that. Labs leaving after being included may have been due to those in less privileged regions of Brazil (e.g. the northern and western regions of Brazil, generally speaking) having more difficulty in persisting in the project. That said, most of the “disappearance” happens between registration and inclusion which usually has to do with the labs not working with the methods that were ultimately included in the project. We also note that most of the states that lose representation were those that had a single lab to begin with, which may make the visual pattern more striking than the actual trend (as states in the South/Southeast also lose labs, but don’t disappear from the map).

      We note again that we never planned to achieve geographical representativeness when recruiting the labs on the contrary, we were aiming to maximize the number of available labs to run the project. That said, we do agree that for the sake of examining whether the population of labs is similar to the one that generated the original experiments (a claim that we do make in the discussion), this representativeness is important to assess. Once more, to allow the reader to evaluate this, we plan to add an additional map to Figure S3 to describe the Brazilian states where the original experiments came from (based on corresponding author affiliations) in which a similar bias towards the South and Southeast Region can be observed.

      Page 15, Figure 3A: I wonder whether adding 95% CIs calculated from the sampling variance of each ratio would improve interpretation and help readers appreciate the real differences between the dots (i.e., means) - along the lines of a forest plot.

      We agree that this would be useful information, and can experiment with the possibility, but our feeling is that the figure will likely become too noisy in cases where the 95% CIs overlap (which are quite frequent). If this is indeed the case, an option to allow the reader to examine this would be better to add an explicit link to the forest plots for each individual experiment (https://osf.io/sx9gv) in the figure legend.

      Page 17, section "Predictors of replication success": It is unclear to me how the decision was made about which results from Figure 4 to present in the text. Intuitively, given that correlations were calculated for both t values and lnRR (and other metrics), I would have expected that whenever a result is highlighted in the text, the authors also report how it changes depending on the metric used - for example, the interesting result regarding the 5-year number of publications, whose correlation is notably lower when using lnRR (−0.31 vs. −0.18). Presenting this nuance in the text would reduce the risk of inadvertently giving the impression of cherry-picking.

      We selected the highest correlation values for each continuous outcome (t score and lnRR) and presented these separately in the text. This is a systematic way to perform the selection, but is obviously subject to the “winner’s curse” effect. We agree that adding both metrics for each predictor would be a fair way to keep this in perspective for the reader, but we would have to think about how to do this without sounding too confusing (as results for the two main outcomes are quite different).

      We do note, however, that the outcomes are indeed different and are expected to vary independently in some cases. For the correlation with replication probability predictions, for example, the effects in opposite directions would likely be expected, as larger original effect sizes will likely lead to larger probabilities to be assigned, but also to a higher possibility of effect size decrease. This low correlation between outcomes is probably something that should be pointed out and discussed in the revised manuscript.

      Page 23, paragraph 1: (this comment should have come during the first % reported, but only in the discussion I realized how important this would be for comparing estimates) I wonder whether the authors should calculate 95% confidence intervals for all their percentages (and those of Errington et al.) using the Wilson method via the function binom.confint() in R, which handles extreme proportions (0% or 100%) more gracefully. This would ensure that uncertainty around these percentages is not neglected and would aid interpretation when comparisons are made.

      We had given this some thought when writing the manuscript – but ultimately opted not to include confidence intervals for our replication percentages and to use the replication rates as descriptive measures only (as done in other replication studies such as (Errington et al., 2021).

      Even though we aimed for our sample of original experiments to be as systematic as possible, it is ultimately constrained by many factors (the choice of methods, the particular expertise of the labs, etc.) thus, adding confidence intervals represents the uncertainty around the replication rate of a very specific population of experiments, which is not directly comparable to those included in other replication efforts in any case.

      We will reconsider whether we should include confidence intervals for replication rates: although doing this for every replication rate in Table 1 and Table 2 may end up being too much information, it could probably be done at least for the replication rates of the main analysis in the text. We note that calculating confidence intervals for percentages is straightforward, requiring only the numbers that are in the table thus, any reader that wants to estimate uncertainty for those rates should be able to do it easily.

      We will also point out the uncertainty around the percentages mentioned in the discussion when comparing our replication rates with those of other studies, which we agree is an important issue to touch on.

      In addition, in the next sentence, the authors are comparing correlation coefficients, at least verbally, these could in principle be transformed into Pearson's r and assigned 95% confidence intervals following meta-analytic workflows, which would better allow us to assess whether these correlations are meaningfully larger or smaller, and help avoid potentially misleading arguments.

      Both correlations in that case are non-parametric (e.g. Spearman’s ρ), so they cannot be directly transformed into Pearson’s r without making assumptions about the distribution (which we would probably avoid doing given the very marked outlier in our own). We can calculate a non-parametric confidence interval for our own correlation coefficient by resampling, but we will have to investigate whether this can be done using the available data from (Errington et al., 2021) (which is probably the case if effect sizes for all experiments have been shared).

      Page 24, paragraph 2: The following result is really interesting and I would love for the authors to expand on it a little. There must be other meta-research studies that, despite not studying replicability directly, have explored a similar predictor: "Other features of the original article were generally uncorrelated with replication outcome, although large rates of publications by the last author were associated with lower replicability, suggesting that incentivizing publication volume may be counterproductive for the reliability of results."

      It is indeed interesting, and seems to confirm an intuition that has long been present in the reproducibility field, but actually has little evidence to support it: if anything, there is evidence in the opposite direction in psychology (Youyou et al., 2023), although they looked at cumulative publication number, while we used number of publications in a fixed interval.

      We can expand a bit further on that finding: that said, we do note that the correlation is relatively weak and has a p value of 0.04. Thus, given the multiplicity of predictors would not be that unlikely to occur by chance, even though it seems intuitive. Thus, even though the relationship seems intuitive, we think it should be considered tentative at best and would refrain from discussing it in too much detail.

      Page 25, paragraph 1: I believe the authors could explore if there is evidence for "incorrect labeling of error bars (Cumming et al., 2007; Vaux, 2004)" by plotting log(SD) vs log(mean) across all original studies, and exploring if large outliers (i.e., points largely deviating from the positive regression) exist. That should provide some insights into whether some values reported as SD in the original studies were indeed SE, which I am assuming is what the authors of the study are referring to when they say "incorrect labelling of error bars" here.

      Yes, that is what we mean by “incorrect labeling of error bars” (as can be grasped from the cited references).

      We can perform this regression, which seems relatively straightforward to do. That said, we note that another likely cause for outliers at least for cell line studies would be the use of different (and eventually inadequate) experimental units (e.g. having error bars that represent technical replicates of the same measurement rather than truly independent experiments). We suspect that this may have an even greater effect in terms of causing error bars not to express the same thing and the regression will not help in differentiating the two causes.

      We should also note that different types of experiments may be expected to have very different SDs, so the regression is likely to have a lot of error associated with it. In particular, it’s probably worth doing separate regressions for each method, to account for the likely difference in CVs between animal and cell line experiments, for example. This could also help tease apart the two causes above, as the experimental unit problem mentioned above will likely only be observed for cell experiments.

      Code: I could not engage with the data and code, but I would like to highlight that the organisation and clarity of the GitHub repository is of high quality.

      Thanks!

      Reviewer #3 (Public review):

      Summary:

      The authors conducted a large-scale replication effort of lab-based biomedical experiments with an emphasis on the country of origin and who conducted the replication experiments. The authors aimed to understand this context in both the outcomes produced, but also in the approach. Finally, the authors aimed to conduct multi-lab replications to provide richer data from the replications. Overall, the authors find replication rates that are like other large-scale replication efforts in the biomedical space. The authors provide rich detail into the three experimental techniques that were the focus of this effort, potential moderators of replication success, and challenges in conducting replications and coordinating a large-scale crowd-sourced effort.

      Strengths:

      The paper is outstanding in being transparent and calibrated in how the results are presented. While the authors were challenged by mundane aspects (e.g., difficulty with logistics), unexpected aspects (e.g., COVID pandemic), and very insightful aspects unique to conducting replications (e.g., experimental issues). The authors also provide variation in how they present the results, including confirmatory, multiverse, and exploratory analysis. A unique strength for this study is the rich in-depth insights about the process and interpretation of conducting replications, including predicting replication success in the lab-based biomedical space.

      We thank the reviewer for the compliments. Again, a more extensive list of insights can be found in our challenges article (Amaral et al., 2026), which we will cite in the revised version.

      Weaknesses:

      The study has weaknesses that the authors acknowledge in their discussion, such as lower number of replications than originally planned that limited the intended effort to compare multiple experiments with multiple attempts against a single original experiment. Another weakness is the limited discussion connecting these findings to the Brazilian research ecosystem.

      We acknowledge the missing replications as a weakness, and we hope we have made that point clear in the discussion.

      Concerning the Brazilian research ecosystem, we could try to explore this in more detail in the introduction. In particular, we believe that a better understanding of the Brazilian academic system, including its regional disparities and the general composition of its workforce (which is largely composed of undergraduate and graduate students), can be useful in interpreting some of the findings.

      We can try to provide a bit more context at the end of the introduction (perhaps between the last 2 paragraphs, which would also address a point made by Reviewer #1), and also in different points of the discussion including those comparing replication rates with other studies or discussing infrastructural difficulties, some of which may be specific to the Brazilian context (such as difficulties in acquiring specific reagents or licenses). Still, we reiterate that, due to the lack of studies with comparable samples in other regions, we cannot tease apart the factors that are specific to Brazil from those affecting lab biology as a whole from the data alone.

      References:

      Amaral OB, Neves K, Wasilewska-Sampaio AP, Carneiro CF. 2019. The Brazilian Reproducibility Initiative. eLife 8:e41602. DOI: https://doi.org/10.7554/eLife.41602

      Amaral OB, Valério B, Carneiro CFD, Mota GPS, Neves K, Abreu M, Tan PB. 2026. Challenges for building up confirmatory science in lab biology: lessons learned from the Brazilian Reproducibility Initiative. MetaArXiv, DOI: https://doi.org/10.31222/osf.io/8y3tg_v1

      Errington TM, Mathur M, Soderberg CK, Denis A, Perfito N, Iorns E, Nosek BA. 2021. Investigating the replicability of preclinical cancer biology. eLife 10:e71601. DOI: https://doi.org/10.7554/eLife.71601

      Fanelli D. 2010. Do pressures to publish increase scientists’ bias? An empirical support from US states data. PLoS One 5:e10271. DOI: https://doi.org/10.1371/journal.pone.0010271

      Fanelli D, Schleicher M, Fang FC, Casadevall A, Bik EM. 2022. Do individual and institutional predictors of misconduct vary by country? Results of a matched-control analysis of problematic image duplications. PLoS One 17:e0255334. DOI: https://doi.org/10.1371/journal.pone.0255334

      Ioannidis jpa. 2005. why Most Published Research Findings Are False. PLoS Medicine 2. DOI: https://doi.org/10.1371/journal.pmed.0020124

      Serghiou S, Contopoulos-Ioannidis DG, Boyack KW, Riedel N, Wallach JD, Ioannidis JPA. 2021. Assessment of transparency indicators across the biomedical literature: How open is open? PLOS Biology 19:e3001107. DOI: https://doi.org/10.1371/journal.pbio.3001107

      Smaldino PE, McElreath R. 2016. The natural selection of bad science. R Soc Open Sci 3:160384. DOI: https://doi.org/10.1098/rsos.160384, PMID: 27703703

      Tyner AH, Abatayo AL, Daley M, Field S, Fox N, Haber NA, Hahn KM, Struhl MK, Mawhinney B, Miske O, Silverstein P, Soderberg CK, Stankov T, Abbasi A, Aberson CL, Aczel B, Adamkovič M, Albayrak N, Allen PJ, Andreychik M, Awtrey E, Axxe E, Azevedo F, Bader MD, Bago B, Bailey J, Bakker M, Banik G, Banks GC, Baskin E, Batruch A, Beatteay A, Behr SM, Berente N, Berry Z, Białkowski J, Bodroža B, Boeschoten L, Bognar M, Bokhove C, Bonfiglio D, Bouwman R, Brady TF, Braithwaite SR, Briceño Jiménez G, Brick C, Bricka T, Briker R, Brown AN, Brown GDA, van Aert RCM, Caldwell K, Capitan S, Capitán T, Chandler J, Charles T, Chartier CR, Chawdhary R, Cheng KJ, Chopik WJ, Clark B, Colvin VE, Comer CC, Costantini G, Coupé T, Cummins J, Czernatowicz-Kukuczka A, de Leeuw J, Dobolyi D, Druckman JN, Duan J, Dujmović M, Dunleavy DJ, Durkee PK, Emery C, Esterling KM, Evans TR, Fedor A, Fernández-Castilla B, Fiala N, Field JG, Fong N, Fonseca MA, Freeman ALJ, Freese J, Geiger SJ, Geng J, Getz LM, Geven LM, Gleibs IH, Gonzales DP, Gooty J, Gourdon-Kanhukamwe A, Greculescu C, Griffin SM, Grigoryan L, Grunow M, Gunby N, Hall B, Hanel PHP, Hannon EE, Harper S, Held MJ, Hickman L, Higgins NC, Hippel S, Hoeppner S, Hong S, Hostler TJ, Inzlicht M, Izydorczak K, Jaeger B, Jankowsky K, Jarke-Neuert J, Jensen M, Jokić B, Jolles D, Jolly P, Jones AM, Juanchich M, Kačmár P, Kapoor H, Keljanovic A, Koirala S, Kołczyńska M, Kouroupaki D, Kühnen U, Landgrave M, Larson MJ, Laulié L, Lawrence ACE, Le Forestier JM, Leahy KE, Lee S, Leslie J, Lewis SC, Limnios C, Lin H, Liu A-C, Lloyd JW, Ludvig EA, Lynott D, MacDonald J, Mallik P, Mallinson DJ, Marinazzo D, Martarelli CS, Matacotta J, McBride A, McHugh C, McMillan G, Méndez E, Metzger M, Michaelides MP, Michalak J, Micheli L, Miller JK, Milyavskaya M, Molden DC, Monjaras AG, Moreau D, Morrow A, Moya C, Mudrik L, Mulder LB, Munt KA, Nandi A, Nason K, Nast C, Nave G, Nax HH, Neubauer F, Nguyen PLL, Nichols AL, Nilsonne G, O’Boyle E, Oettinghaus J, Oh J, Oshana A, Ostermann T, Ostrowski RP, Oyebanjo A, Panczak R, Patrianakos J, Pavez I, Pavlov YG, Persson S, Perugini M, Peters K, Pieters C, Ponizovskiy V, Porter ND, Prenoveau JM, Purić D, Purol MF, Puthillam A, Quinn KA, Ramljak M, Reed WR, Ritchie M, Ritzau M, Roche SP, Rodela R, Röer JP, Ropovik I, Rothschild J, Saal J, Safadi H, Samaha J, Sanchez M, Sankaran S, Santos D, Sargent AC, Sauter M, Schmidt K, Schnabel L, Schroeder AN, Schuetz SW, Schuetze BA, Schulte-Mecklenbeck M, Schütz A, Sevigny EL, Shackleton E, Shafranek RM, Shaki S, Shakya S, Sirota M, Sisco MR, Sitnikov MM, Slevc LR, Smalarz L, Smith CT, Snyder JS, Sommet N, Sonmez F, Spellman BA, Stanulewicz-Buckley N, Stock G, Street CNH, Strømland E, Sundelin T, Syed M, Szabelska A, Szaszi B, Szumowska E, Tagat A, Täuber S, Tay L, Thapa S, Thatcher J, Tsaklakidou D, Tummers L, Turkovich E, Tutor MV, Urbanska K, van ’t Veer AE, van Assen M, van de Ven N, van den Goorbergh R, Vargo EJ, Vaughn LA, Vazire S, Vermeulen JM, Vo DTH, Volkman V, Wagenmakers E-J, Wagner D, Walasek L, Walter F, Warmelink L, Wei L, Weißflog MI, Weller N, Wichman AL, Wilbiks J, Williams JR, Wolfe K, Wort F, Wright R, Wulff JN, Xue X, Yan VX, Yang Y, Yoon S, Žeželj I, Zhang Y, Ziano I, Zogmaister C, Zupan Z, Zwaan RA, Nosek BA, Errington TM. 2026. Investigating the replicability of the social and behavioural sciences. Nature 652:143–150. DOI: https://doi.org/10.1038/s41586-025-10078-y

      Westlake H, David F, Tian Y, Krakovic K, Dolgikh A, Juravlev L, Bournonville TE de, Carboni A, Melcarne C, Shan T, Wang Y, Mu Y, Kotwal A, Pirko N, Boquete JP, Schüpfer F, Rommelaere S, Poidevin M, Liu Z, Kondo S, Ratnaparkhi GS, Chakrabarti S, Liu G, Masson F, Xiaoxue L, Hanson MA, Jiang H, Cara FD, Kurant E, Lemaitre B. 2026. Reproducibility of scientific claims in Drosophila immunity: A retrospective analysis of 400 publications. eLife 15. DOI: https://doi.org/10.7554/eLife.108404.1

      Youyou W, Yang Y, Uzzi B. 2023. A discipline-wide investigation of the replicability of Psychology papers over the past two decades. Proceedings of the National Academy of Sciences 120:e2208863120. DOI: https://doi.org/10.1073/pnas.2208863120

    1. Author Response:

      We thank you for this assessment of our work and the positive assessment of the overall theoretical framework. We can fully answer the concerns, in particular regarding data quality, and will provide detailed answers in the following directions:

      “insufficient description of the data”: We will describe the data as much as possible and will share the data and the analysis code.

      “lack of included equations and code”:  We will share the mathematical equations in the supplementary material, and the full code on an online repository. The reason why our data repository (10.5281/zenodo.18480481) is not yet public is that it cannot be changed after publication. For the review process, we provide a github link to data and code here https://github.com/oliviercotto/eLife_epidR. We will ultimately share the link to the final version of the files on Zenodo.

      “definitions of antibiotic use that are not complete”: We will complete the definition of antibiotic use, which is the use of any antibiotic between 7 days and 3 months before sampling. Children who used any antibiotic 7 days before sampling were not included in the study. The type of antibiotic used is given in supplementary material S1: 93% of the antibiotics prescribed are beta-lactams (amoxicillin, amoxicillin/clavunalate, oral 3rd generation cephalosporins).

      “low sensitivity of assays for carriage”: the carriage study conducted in Sweden is used to get plausible estimates of carriage duration parameters in infants.

      • Strain definition is based mainly on randomly amplified polymorphic DNA (RAPD), not colony morphology. Strains with distinct morphology but the same RAPD profile are considered one strain. Conversely, it was checked that strains of the same timepoint with the same morphology most often had the same RAPD profile.

      • We did check that these data are not much affected by imperfect sampling: observations of a strain ‘disappearing’ from sampling then ‘reappearing’  at later timepoints are rare (14 out of 273 strains). This is why we did not correct these occurrences in the previous version of the analysis. In the revised version, we will add a description of these occurrences and correct them. This correction did not significantly alter the inferred parameters in our preliminary analyses.

      • At a broad level, the fact that E. coli clades vary in their carriage duration is very well established across multiple independent datasets; the precise value of carriage duration difference for “persistent” vs. “transient” that we inferred here (a two-fold difference, supplementary material S3) is actually relatively conservative, in the sense that other studies have detected more important differences. We will create a table summarising available evidence on colonization parameters of E. coli to show that the insights from the Swedish data are qualitatively robust.

      • Yet, we will conduct a range of sensitivity analyses to see how the inferred costs of ESBL resistance vary when varying differences in carriage durations, competition and niche differentiation.

      “technical issues with statistical prior selection and parameter identification”: We disagree there is a “technical issue” with prior selection: The fact that resistance is costly, hence that our priors are left-bounded at 0 for the cost parameters, is a prior expectation based on the observation that resistances do not go to fixation. If resistances only conferred an advantage in treatment, but zero cost, then they would quickly evolve to 100% frequency–contrary to what is observed in virtually all epidemiological studies of resistance. That said, we will relax the definition of these priors to test that the data is also compatible with a strong cost on some traits, and no cost or a “negative cost” on other traits.

      Regarding parameter identification and the specific comments on the inference of colonisation parameters: we re-inferred all colonisation parameters with direct inference assuming specific functional forms. This does not alter much the final colonisation parameters that we then use for our main inference. We will also conduct sensitivity analyses to examine how changing some of the colonisation parameters (carriage duration, competition and niche differentiation)  would alter the main inference.

      “application of non-regional ECDC surveillance data to France”: We will clarify our text, as there is a misunderstanding here: we do not use non-regional ECDC surveillance data for inference. We use ECDC data (i) for illustrative purposes, to show that trends in ESBL in France in this surveillance system are very similar to those observed in France in our focal dataset, thus showing the consistency and representativeness of our data. (ii) to give an overview of the weak and inconsistent association of ESBL with age across Europe, thus supporting the relevance of our approach even if our data concerns infants and children. We will make sure this is clarified in the updated version of the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We have carefully addressed the insightful comments provided by the reviewers which thoroughly increased our comprehension of the dynamics of centriole amplification. The manuscript has been revised accordingly and put in the context of the two papers we published since our last submission, showing that MCC differentiation is a genuine cell cycle variant. A point by point answer to all reviewer comments is provided below.

      Briefly:

      We have streamlined terminology and nomenclature in text and figures / better define experimental conditions with nocodazole

      We have tested the role of dyneins in the dynamics of centriole amplification

      We have done correlative light and electron microscopy on the early stages of centriole amplification

      We have analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors

      Collectively, this allowed us to make a clearer parallel with what occurs during centriole duplication and to demonstrate that centriole biogenesis in the MCC cell cycle is marked by the superimposition of 2 canonical centriole cycles.

      We believe the manuscript will interest a broader readership since it now provides more fundamental insights on the mechanism of centriole biogenesis.

      Public Reviews:

      Reviewer #1 (Public Review):

      The manuscript by Boudjema et al. describes the cellular events underlying centriole amplification and apical migration to allow the assembly of hundreds of motile cilia in multi-ciliated cells. For this, they use cell culture models in combination with fixed and live cell imaging using antibody staining and fluorescence from endogenously tagged centriole and deuterostome markers, respectively. The work is largely descriptive and functional analyses are restricted to treatment with the microtubule depolymerizing drug nocodazole. The imaging is state-of-the-art including confocal microscopy, live imaging with optical sectioning and high optical and temporal resolution, as well as super-resolution imaging by ultra-expansion microscopy.

      The study does a good job of providing a very detailed description of the dynamics of centrioles and deuterostomes that lead to centriole amplification and apical migration in multiciliated cells. This detailed view was missing in previous work. It also reveals the involvement of microtubules at multiple steps: the formation of a cloud of deuterostome precursors, the nuclear envelope tethering of newly formed centrioles, their separation, and their migration to the apical surface.

      It would have been useful to expand the analysis of the role of microtubules by including analyses of the requirement for specific microtubule motors, for a better understanding and additional evidence that microtubule-based transport is involved. A weak point is that there is no visualization of microtubules together with deuterosomes and centrioles at the different steps of centriole amplification and migration, to directly address how these structures may interact with and move along microtubules.

      Overall, apart from experimental aspects and since this is largely a descriptive study, the manuscript would benefit from more precise language and a better description of the complex events underlying centriole amplification and movements.

      We have streamlined terminology and nomenclature, clarified the description of the complex events, and test the role of dyneins in centriole amplification. Microtubules density in MCC does not allow to extract information from imaging. In addition, we have done correlative light and electron microscopy on the early stages of centriole amplification and analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors. We also replied points by points to the reviewer specific comments.

      Altogether, our new data allowed to demonstrate that centriole biogenesis in the MCC cell cycle is marked by the superimposition of 2 canonical centriole cycles. We believe the manuscript will interest a broader readership since it now provides more fundamental insights on the mechanism of centriole biogenesis.

      Reviewer #2 (Public Review):

      This important work will be of interest to centriole and cilia cell biologists. It describes in detail how microtubules control multiple aspects of centriole amplification in brain multiciliated cells. This study provides a greater time-resolved and molecular proteomic mapping of the different steps involved, with or without microtubule disruption. Boudjema et al. show that microtubules are important throughout the centriole amplification process, from the early stages, where the procentrioles emerge from a pericentriolar "nest", through the growth stage where microtubules maintain the perinuclear localisation, to the detachment stage, where microtubules assist in perinuclear disengagement and apical migration. The results are generally well supported by the evidence, but the manuscript would benefit significantly from some heavy editing to introduce more niche terms, standardize abbreviations in text, and labels on figures to help bring the readers, especially non-specialists, along with them - increasing the accessibility of their work.

      We thank the reviewer for his/her enthusiasm. We have streamlined terminology and nomenclature and clarified the description of the complex events to increase the accessibility of our work. We also replied points by points to his/her specific comments.

      Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Boudjerna and Balagé et al. aim to elucidate the spatial origin of centriole amplification and the mechanisms behind the formation of an apical-basal body patch in multiciliated cells (MCCs). To this end, they focused on the role of microtubules and developed new tools for spatiotemporal and high-resolution analysis of different stages of centriole amplification, including the centrosome stages, A-stage, G-stage, and MCC-stage. Among these tools, the MEF-MCC cells grown on micropatterns stands out for its versatility as it is not tissue-specific and does not require epithelial cell-to-cell contact for differentiation. Additionally, the CEN2-GFP; mRuby-DEUP1 knock-in mouse model was used to study different stages of centriole amplification in physiological brain MCCs. This model offers an advantage over the previously described CEN2-GFP model by enabling the resolution of early events in centriole amplification through the visualization of DEUP1-positive structures and their dynamics. Finally, the authors leveraged powerful imaging techniques, including super-resolution microscopy, the U-ExM, and high-resolution live cell imaging in order to detect and track centriole amplification, elongation, disengagement, and migration.

      By combining the MEF-MCC and knock-in mouse model with spatiotemporal imaging in control and nocodazole-treated cells (treated acutely or chronically), the authors define the sequence of events during centriole amplification, revealing the critical roles of microtubules for the first time. Initially, the centrosome-mediated microtubule network forms, organizing a pericentrosomal nest from which procentrioles and deuterosomes emerge. Their findings indicate the importance of microtubules in recruiting and maintaining pericentriolar material clouds that contain DEUP1, PCNT, SAS6, PLK1, PLK4, and tubulins. Following the amplification stage, the procentrioles mature, leading to cells displaying numerous MTOCs, as demonstrated by regrowth experiments. Mature centrioles then disengage from deuterosomes, attach to the nuclear envelope, and migrate to the apical surface facilitated by microtubules.

      Strengths:

      The manuscript provides new insights into the regulatory function of microtubules in centriole amplification. Addressing the role of microtubules during different stages of centriole amplification required the development of new tools to study brain MCCs, which will be useful in future studies of MCCs. A notable strength of this manuscript is the authors' thorough and quantitative analysis of highly dynamic processes in MCCs. The precision and detail in describing these dynamic events are impressive. This comprehensive analysis advances our understanding of MCC biology.

      Weaknesses:

      The role of microtubules and other molecular players during different stages of centriole amplification in brain MCCs can be further studied and strengthened using the tools developed in the manuscript. A more quantitative description of some of the analysis performed in the manuscript is required to strengthen the conclusions.

      We thank the reviewer for his/her enthusiasm. We have tested the role of dyneins in the dynamics of centriole amplification, done correlative light and electron microscopy on the early stages of centriole amplification and analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors. We also replied points by points to the reviewer specific comments.

      Recommendations for the authors:

      As you will see, all reviewers felt that the analyses of the involvement of microtubules should be strengthened by including controls and additional experiments. Also, they agree that significant text editing would help to improve the manuscript's accessibility and readability.

      Specifically, they would suggest (1) streamline terminology and nomenclature in text and figures; (2) better define experimental conditions with nocodazole (concentrations used, effect on microtubules, effect on canonical centriole duplication); and (3), in the absence of other complementary genetic perturbation experiments, add a limitations paragraph in the discussion about conclusions drawn from nocodazole treatment alone.

      Reviewer #1 (Recommendations For The Authors):

      Main issues:

      (1) The authors use variable terminology to describe the same or similar events/structures. For example, in Figure 1 they refer to "centrosome stage" where they observe a pericentrin "cloud", which they later refer to as a "nest". In all other figures the first stage is not referred to as the "centrosome stage" but as the "cloud stage". Again, they also describe the "cloud" as a "nest" occasionally, but not always. In the cartoon, the nest is termed "centrosome cradle". The variable and inconsistent use of terms is confusing and the authors do not provide any explanation for the use of one vs. another.

      The text is now corrected. The centrosome stage corresponds to the stage preceding the beginning of centriole amplification in MCC progenitor. The pericentrosomal cloud of centriole and deuterosome elements forms later on, during the amplification A-stage. The formation of this cloud marks the beginning of A-stage, and persists up to G-stage where it dissolves. When we show that the cloud hosts the first stages of centriole biogenesis, we defined it as a “nest”. We do not use anymore the term craddle.

      (2) What prompted the authors to use the term "nest"? It gives the impression that they describe aspecific physical entity/structure (also depicted in this way in Figure 3P, with microtubules outside of this structure), but what is the evidence for this?

      The cloud is the spatial entity and the term “nest” is used to define a function of this transient compartment. We decided to keep the term “nest” as we now identified it with correlative light and electron microscopy, in addition to U-ExM, and show that the accumulation of centriole and deuterosome elements is accompanied by the formation of immature procentrioles, deprived of MT walls, as well as immature and empty deuterosomes. The scheme with MT outside the cloud/nest is misleading as we see MT organized by the mother centriole. We have now changed this.

      (3) The "nest" may simply be a dynamic accumulation of precursor particles around the centrosome, similar to what has been described for centriolar satellites. Rather than proposing a new entity, I suggest testing whether the "nest" particles may colocalize with PCM1 and thus may be related to centriolar satellites. Based on the data, the nest would simply be the centrosomal MTOC that organizes a radial microtubule array on which particles move around its center. In the absence of other evidence, I am not convinced that a new term is needed.

      We totally agree with the reviewer: the centrosome, as MTOC, concentrates centriolar and deuterosome components. This cloud is consistently dissolved when MT are depolymerized or dyneins inhibited. So, the physical entity is a “cloud”. We used the term “nest” to propose one function for this cloud which is to form deuterosomes and centrioles, before they move away for maturation. In fact, deuterosome and centriole formation are hindered when the cloud is dissolved. We have tried to edit the text all over the manuscript to make it clearer.

      (4) Role of MTs: are microtubules required or do they just facilitate some of the investigated events?

      The reason why the role of MT has not been tested yet during centriole amplification is probably because MT not only constitute the cell cytoskeleton on which molecular motors ride to transport cargos or distribute forces, they are also the core component of the structures we are studying. This is why we have tested a range of nocodazole concentrations and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (Fig. 4 Supplementary 1A-B). This may lead to an underestimation of the role of MT but we cannot study the role of MT on centriole amplification if centrioles cannot be formed.

      Does multi-ciliation in these models eventually occur normally under the concentrations and treatment conditions used here? This should be tested and discussed in the context of whether microtubules are indeed required and at what step of the entire process (amplification, migration, ciliogenesis) they may be critical.

      We did both chronic and acute treatments.

      Chronic treatments were done to test the overall efficiency of centriole amplification when MT (or dyneins) are perturbed. Chronic treatments were used to assess the role of MT (or dyneins) on the global efficiency of centriole and deuterosome formation (number of cells able to amplify, number/size/loading of deuterosomes, final number of centrioles (Fig. 4H-I, Fig. 4 Supplementary 2 B-D). In these chronic treatment, we focused on centriole amplification and not ciliation since it was the scope of this study. Also, we did not take ciliation as a readout of amplification because ciliation is relying on MT polymerization.

      Then, we also did acute treatments to test the role of MT (or dyneins) at each stage of amplification (A-amplification, G-growth, D-disengagement, M-migration; Fig. 4, 5, 7, 8 and associated supplementary figures). Since one stage is dependent on the precedent one, this enabled us to decipher the direct role of MT (or dyneins) on each single stage. We have now edited text, methods, legends and pictograms to be clear on whether acute or chronic treatment was done.

      (5) Can the authors include control (non-amplifying) progenitors in their analyses? It would be useful to know what the signal and distribution of each specific marker are before differentiation begins (before the cloud stage).

      Non amplifying progenitors are analyzed and constitute the so-called “centrosome stage”. We have now precised it and called it the “progenitor stage”.

      (6) Figure 2: Again, the terminology is confusing, since the authors describe that DEUP1 forms a "cloud" with centrin during the A stage.

      Corrections have been done as explained in point 1.

      (7) Description Figure 3: the authors introduce yet another term: "halo" A-stage. Is this the early A stage? Again, this is not explained and confusing. More systematic and consistent description is needed.

      Corrections have been done as explained in point 1. The term halos is used un the lab as it was the first term we used in our Nature paper in 2014 in reference to the halo described by Erich Nigg when they overexpressed Plk4. It was an error to use it in the manuscript.

      (8) Nocodazole treatments: the used concentrations are quite high.

      MCC develop a very dense and stable MT network that is not comparable to cycling cells. MT are very difficult to depolymerize entirely (Fig. 4 Supplementary 1A-B).

      (a) To avoid non-specific effects the authors should test what the minimal concentration is that completely depolymerizes microtubules in their cell model and perform analyses at this concentration.

      We have of course tested a range of nocodazole concentrations at the beginning of the study (Fig. 4 supplementary 1A-B), and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4). In case it was not clear, we refer to this now several time and more clearly in the text and methods.

      (b) They should demonstrate depolymerization of microtubules by microtubule staining in the acute and chronic noc treatments and at the different noc concentrations used.

      This is, and was, in supplementary material (same, Fig. 4 supplementary 1A).

      (c) The authors should demonstrate that the used nocodazole concentrations do not impair normal centriole biogenesis during the cell cycle in these cells; if so, impaired assembly of centriole wall MTs may contribute to the observed effects in Figure 4.

      As mentioned in point 8b, we have of course tested a range of nocodazole concentrations at the beginning of the study (Fig. 4 supplementary 1A), and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4). The ability of the cells to form centrioles during chronic treatments were always assessed using immunostainings of SAS6 and/or CEN2-GFP signals (now exemplified in Fig. 4 Supplementary 1B). We also did EM analysis on cells treated with the highest doses of nocodazole (Nocodazole 10 uM for 24h) and this showed that centrioles can form with, what seems to be MT walls, in cells totally deprived of cytoplasmic MT fibers (Fig. 4 Supplementary 3-4). However, this does not show that all the cells can, because the number of cells that can be analyzed by EM are not sufficient to conclude. Also, one cannot assess whether MT walls are properly polymerized. However, the absence of MT walls should not change the results of the Figure 4, which are based on DEUP1, SAS6 or CEN2-GFP signals for deuterosomes and centrioles. Also MT depolymerization affects the formation of deuterosomes, which should not be altered by MT wall defects as it is not affected, even when centriole formation is blocked (LoMastro et al., 2024). Last but not least, we now show that blocking dyneins, as a comparable and even greater effect, on the formation of the cloud, deuterosomes and centrioles (Fig. 4C-I and Supplementary Fig. 4), which confirms that MTOC function, rather that MT wall formation, explain the centriole biogenesis alteration shown in Figure 4.

      (9) The authors repeatedly refer to the centriole-to-centrosome conversion of amplified centrioles and how this resembles centriole-to-centrosome conversion during the cell cycle. However, they incorrectly claim that this occurs at the G2/M transition. PLK1-dependent modification occurs at this stage, but conversion and PCM recruitment only occur after mitosis (see original work by the Tsou lab, which needs to be cited here).

      We agree with the reviewer. We have now added additional data to show clearly that centriole biogenesis, which requires two cell cycles to proceed in cycling cells, is accelerated during the MCC cell cycle variant where the elongation and maturation cycles are superimposed. This is now clearly shown in Fig. 3, 5, 9 and discussed.

      (10) Figure 6H-J: the authors claim that at low noc concentration, more D-stage cells showed incomplete disengagement than in controls, but the effect is shown only for the highest 10 µM concentration. Do any eof the phenotypes in Figure 6 also occur at the lowest noc concentration (assuming it depolymerizes MTs)? Again, it is crucial to demonstrate this, to exclude unspecific effects not linked to MT depolymerization.

      An error was made on the figure (but not in the legend). In Figure 6, chronic treatments are at 1 or 5 µM. Only acute treatments were done using 10 µM. In both cases, MT are not entirely depolymerized in these experiments (Fig. 4 supplementary 1A).

      (11) Disengagement, Figure 7: The authors describe that DEUP1 signal spreads all over the cytoplasm and becomes diffuse during this process, but one cannot see a diffusive signal throughout cells in the figures.

      We pushed the contrast to make it clearer but the deuterosomes are still bright at this stage and it is difficult to have both signal clear (now in Fig. 6B). We have also changed the example in video (now video 19) to show it more clearly with DEUP1 channel alone.

      (12) Figure 7: localization of disengaged centrioles at microtubule "nodes" is not clear from the images. There are many centrioles and random colocalization may be expected simply based on the high number. Higher resolution and/or magnification and quantification would be needed.

      We have edited and now say that centrioles “colocalize” with MT which, since centrioles nucleate MT, seems normal. We agree that it could be random, but given the density of MT, and the number of centrioles, it does not seem opportune to us to quantify. We can just say that we never see centrioles is regions that are deprived of MT.

      (13) The term "diffusive" to describe slow centriole movements in Figure 8 suggests that it is not motor or force-dependent, but there is no evidence for that. Movement based on opposing forces could produce a similar result, but would not be considered diffusive.

      We agree. We have changed “diffusive” by “diffusive-like”.

      (14) The manuscript would greatly benefit from the analysis of some candidate motor activities that may drive the movement and migrations of centrioles in this system. This would support the importance of the microtubule network for the specific steps in these processes, and better define its role beyond "being required". Dynein may be a candidate or minus end-directed kinesins. Since chemical inhibitors are available, these types of experiments would be straightforward.

      We formerly tested ciliobrevin but had hard time because of the small stability of the drug. Since our submission to eLife, we tested dynapyrazol and dynarestin and found dynapyrazol very efficient in dissolving the Golgi, a good readout of dynein inhibition. We sought to test the role of dyneins, using dynapyrazol, on (i) the formation of the pericentrosomal cloud in A-stage, (ii) the oscillation of DEUP1+ structures during A-stage, (iii) the number, size, loading of deuterosome, (iv) the final number of centrioles, (v) the migration to the nuclear membrane and (vi) the final apical migration of centrioles. The results are now inserted in main and associated Fig. 4, 5, 7, 8, 9.

      (15) Discussion:

      "the role microtubules" lacks "of"

      This is now edited.

      "This lack is..." Lack of what?

      This is now edited.

      "reflexive link" - meaning of "reflexive" is not clear in this context

      We have removed it.

      In my opinion, the study does not identify a nest composed of DEUP1, PCNT, and Centrin2; it only shows that these components accumulate as particles around the centrosome, which functions as MTOC. Consequently, it seems that the "nest" does not exist when MT is depolymerized. One could consider the center of the centrosomal MT array as a nest in this context, but there is no evidence of a specific new structure as suggested by the way the term is used in the manuscript.

      This is what we want to say: the center of the MT array become a nest in this context. We do not state that there is a specific new structure. We just say that MT and dynein dependent concentration of centriole and deuterosome components exists and that this region nests the birth of centrioles and deuterosomes. Also, this compartment is restricted in time and space, which justifies to use a specific term. The MTOC exists in the progenitor cell, while this compartment, marked by DEUP1, Centrin, PCNT accumulation, appears at the beginning of amplification and grows during A-stage to be dissolved at G-stage when all the deuterosomes and centrioles have moved away.

      What is the evidence that "DEUP1 is a centrosomal protein before building deuterosome structures"? It would be good to refer to the specific experiment. Does DEUP1 localize at centrioles also in the absence of microtubules? If not, I would not consider it a centrosomal protein.

      We have removed this statement to avoid misinterpretation.

      "This reminds the centriole-to-centrosome conversion..." the sentence is missing an "of"; also, again the authors confuse the order of events during the cell cycle, where centrosome conversion occurs after completion of mitosis, not at G2/M transition.

      We have removed this statement to avoid misinterpretation. Also, see Point 9.

      "microtubule dependent nuclear migration" should be rephrased; it sounds as if the nucleus migrates.

      This has been changed

      The following discussion of disengagement being linked to association with the nuclear envelope and resembling the process in cycling cells is misleading. In cycling cells movement of centrioles along the nuclear envelope occurs at G2/M and drives centrosome separation (separation of centriole pairs) in preparation for mitosis, not centriole disengagement.

      We are now clearer. We compare centriole-loaded deuterosome organization around the nuclear membrane to the migration of new centrosomes during early prophase (Fig. 5F-H, Fig. 5 Supplementary 2G-K).

      Regarding the possibility that forces by microtubules generated by the daughter centriole drive disengagement also in cycling cells, I would argue that this is unlikely since the daughter centriole can only nucleate microtubules after disengagement has occurred (and conversion to centrosome/PCM recruitment). Once this happens, it may physically separate the disengaged centrioles, which is a different type of activity. Indeed, originally the term "disengagement" was coined to specifically describe the loss of the perpendicular engagement of daughter centrioles with their mothers (Tsou and Stearns, Nature, 2006).

      We have removed this statement to avoid misinterpretation. The perpendicular engagement is difficult to assess on deuterosomes but we do see by live imaging, that attachment changes during D-stage, before centrioles detach clearly from deuterosomes.

      "high resolutive" should be "high resolution"

      Edit done.

      "splitted" should be "split"

      Edit done.

      "Consistently, when the mitotic oscillator is dis-inhibited and cells enter pseudo-mitotic events, centrioles show clear and rapid cell-cycle like clustering" This sentence is not understandable without further explanation; what does mitotic oscillator refer to? What are pseudo-mitotic events? What is cell cycle-like clustering?

      We have removed this statement.

      Minor:

      (1) Abstract: "Centriole number must be restricted to two..." Since cells are born with two centrioles and have 4 centrioles (2 pairs) when they enter mitosis, this sentence is inaccurate.

      The sentence has changed.

      (2) Abstract: "reflexive link"; I am not sure what the term "reflexive" refers to?

      We have removed this statement to avoid misinterpretation.

      (3) Figure 1C, D: it should be described better that the larger magnification panels represent overlays of many cells and what marker they show. This is not obvious since the smaller single-cell panels always show two different markers. Also, it would be more useful to show also single cells in the magnified view. The overlay does not allow us to see if a marker forms a cloud or a single dot, which is as important as the cell-to-cell variation in distribution.

      We have clarified this in the text and the legend. The cell-to-cell variation cannot be estimated with the overlay, but the projection from several cells (number precised) allows to see that the signal is confined in a restricted region. Or not. Which is what we wanted to analyze.

      Related to the above, the authors say that pericentrin forms a cloud at the top left in panel D, but there is only one confined centrosomal dot in the single-cell panel.

      The sentence has changed.

      (4) Results, Figure 2F; video 4: The authors claim connection and disconnection of DEUP1 aggregates with centrosomal centrioles; can the authors comment on the spatial resolution including in z in this movie to support this claim? Can they exclude that the structures are in proximity of each other rather than "connected"?

      This is a single z-section of 500nm. The resolution in xy is 128nm/pixel. Given the sizes of deuterosomes and a mature centriole, and given the fact that we observed this dynamics in several cells in live, we can state that the structures are connected. This is consistent with deuterosomes frequently observed “kissing” the daughter centriole by EM in the present manuscript (Fig. 2D, Fig. 2 supplementary 3 and 4 and Fig. 4 Supplementary 3-4). One has to look carefully at the daughter centriole (marked “dc”) and span in on the serial sections to see the connected deuterosome (marked by a star): this is at very early stage and therefore it is small. We have not zoomed in since previous manuscript have already described this at later stages with bigger deuterosomes. You can refer to main or supplementary figures in previous manuscripts (Al Jord 2014, Khoury Damaa 2024) where serial sections span the entire deuterosomes and daughter centrioles and show, with nanometric resolution, that both structures are frequently sticked to each others on tens of nanometers.

      (5) The term "dynamics" as used in the manuscript should be plural.

      It has been used plural, except when for “dynamic microtubules” and “dynamic attachment to the nucleus”, which we think is ok? We have not found any other singular uses in our manuscript.

      (6) Figure 5: what does "YL1/2 procentriole intensity" refer to in panel F? This should be the intensity of microtubule asters.

      This has been modified.

      (7) Figure 6 - supplement 1B: contrary to the claim in the text, one cannot see tight colocalization with the nuclear pore marker. This seems to be a very small subset of particles and even in those cases colocalization is not tight. Also, what is the relevance of nuclear pore colocalization?

      We edit and change the phrasing as ‘colocalization with NPC’ is not the good term. What we want to say is that there is a tight connection with the nuclear envelope as shown by the localization of NPC on the same z-section as centrioles. This is why we present a single z, to show that centrioles and NPC are on the same z-plane of 500nm. NPC are stained to outline the nuclear membrane. This is also clearly visible for G-stage centrioles in the XY plane. We have now added an entire z-stack on video 18.

      Reviewer #2 (Recommendations For The Authors):

      To improve accessibility of their manuscript, we would suggest making the following edits:

      (1) Define 'specialist' or 'niche' terms each time you introduce them, such as 'pericentrosomal nest', or 'flower-like structures'.

      This has been clarified.

      (2) Have a think about abbreviations, again ones that work for people outside the project- this paper uses 'PC' for 'procentriole' but for many 'PC' is 'Parental centriole' or Figure 6J talks about 'D total' or 'D partial', leaves readers confused.

      This has been clarified.

      (3) Standardize your abbreviations throughout particularly for your treatments- sometimes Noco sometimes, NOCO, or your imaging experiments sometimes Cen-GFP, sometime CEN2-GFP (Figure 7A, D vs. Figure 6) or DEUP1- mRuby, DEUP1-mRuby3 or mRuby3-DEUP1?

      We now use Nocodazole or Noco in the text and the figure respectively, CEN2-GFP and mRubyDEUP1.

      (4) About 10% of the population, including several key figures in this field, are red-green color blind. Although 4 colour fluorescence is difficult to get right for everyone, choosing palettes (especially for two colour panels) is inclusive. More so, greyscale or inverted monochrome images make it easier for everyone to visualize changes in localization, size, and intensity. Red on black small foci is particularly difficult to discern. For example, Figure 3 - more individual channels in grayscale with arrows to mc, dc, and cilia would be helpful - difficult to distinguish stainings.

      We thank the reviewer for this comment and for this recommendation of being more inclusive. We have done the changes.

      To improve the conclusions drawn, we suggest some revisions below:

      (1) Since the paper really hangs on it, a clearer description of the rationale for when, how long and how much nocodazole treatment was done is needed. The logic currently is difficult to follow seemingly random jumps 10x concentration are used. Microtubules control many aspects of cell biology and could be impacted. For example, I particularly found Figures 6D and H difficult to follow i.e. the timing for 6H seems off.

      MCC develop a very dense and stable MT network that is not comparable to cycling cells. MT are very difficult to depolymerize entirely. We have of course tested a range of nocodazole concentrations at the beginning of the study and shown the extent of MT depolymerization under each treatment. We used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4 reviewer 1). The level of perturbation of MT and consequences on centriole formation at the different timings and doses were done for each experiment and are exemplified in Fig. 4 supplementary 1A-B. This figure was already present in the first version of the manuscript but we have now edited text, methods and pictograms to clarify this.

      (2) Perhaps an extension of this point- in general how interdependent are the processes? If there is a defect at the nest stage, how much are the later defects secondary to this, or do MTs genuinely play direct roles at all stages or are these knock-on effects? How do the authors rule this out? Defects in the nest, lead to smaller and more DEUP1+ foci, with defects in concentrating procentriole factors and centrin, which lead to... For example, Figure 4B looks like centrin is reduced upon noco treatment? Does noco treatment affect Cetn2GFP levels globally? Individual channels grayscale would help visualise this better.

      See also our answer to reviewer 1 point 8c.

      The stages are indeed interdependent. This is why we did both chronic and acute treatments. Chronic treatments were done to test the overall efficiency of centriole amplification when MT are perturbed. We typically used low dose of 1µM because nocodazole remains 48h in the culture medium. Acute treatments were done to test the role of MT at each stage of amplification (A-amplification, G-growth, D-disengagement, M-migration). Most of the acute treatments were done live and nocodazole was applied after the first time point of live monitoring. We used 10µM to have a rapid effect, and because nocodazole remains only several hours in the culture medium. This allowed to monitor the stage “n”, in cells where the stage “n-1” was completed without any drug which allowed to analyze a stage without having perturbed the precedent one.

      We now also test the consequences of dynein inhibition using both acute and chronic dynapyrazole treatments. We show that except for centriole migration, dynein inhibition phenocopies MT depolymerization (centriole number, perinuclear organization and disengagement as well as deuterosome number/loading/size).

      Nocodazole chronic treatments do affect intensity of CEN2-GFP at G-stage centrioles suggesting an altered A-to-G transition. In D-stage, CEN2-GFP signal seems normal. We now mention this in the text and in the Fig. 4 Supplementary 1B.

      (3) The authors nicely show the importance of MTs in the structure of the nest from which procentrioles and DEUP1 positive structures emerge. They suggest this nest may be what supports procentriole generation in the absence of DEUP1 and parental centrioles. Firstly how does this nest look in the absence of DEUP1 and/or parental centrioles (centrinone treatment)? This may be what they are trying to show in Figure 5 Supplement 1 but it currently is very difficult to digest what it is showing relative to controls and whether this is significant in the way it is plotted.

      The nest is conserved in the DEUP1KO with or without centrosomal centrioles, as shown by accumulation of Centrin and PCNT at the center of the self-organised MT network (Mercey et al., 2019). This is in fact what motivated our study on the role of MT in centriole amplification. We have edited the legend to precise the quantification done, which is not related to this question. In this quantification, we show that the increased propensity to accumulate PCNT by centriole-loaded deuterosomes between A and G-stage is maintained in the absence of deuterosomes, indicating that centrioles themselves accumulate/recruit PCNT.

      (4) Can you do CLEM on DEUP1-Ruby and these early foci at the cloud stage to see if they are visible at the ultrastructural level, relative to procentrioles, microtubules, and other electron-dense structures?

      We thank the reviewer for this question. We have done CLEM on the pericentrosomal cloud during very early steps of centriole amplification. This showed that DEUP1 early accumulation at the centrosome corresponds to a region rich in fibro granular aggregates, suggesting that DEUP1 may be translated here, through locally concentrated centriolar sattelites, known to be involved in local translation. Then, small deuterosomes and immature centrioles are formed, within this cloud of sattelites, confirming that the pericentrosomal cloud is a nest for centriole biogenesis (Fig. 2C-D + Fig. 2 Supplementary 2-6 for control and Fig. 4 Supplementary 3-4 for nocodazole treated cells). This also shows that immature deuterosomes are not necessarily round shaped, and can be deprived of centriole loading.

      (5) Check the scale bars- see Fig 4E. Check throughout.

      Done.

      (6) Figure 3 Supplement 1 and 2 don't match the legend and are likely reversed - which one is right?

      Done.

      (7) Technical issue - I couldn't play videos 6 or 16? Check these work.

      Done.

      (8) Nomenclature mammalian proteins- mouse or human- should be all caps DEUP1, PLK4, SAS6,etc. Watch your units- space between number and unit.

      This has been done.

      (9) Many of the graphs involve three biological replicates but why not plot the mean of each of the three experiments and do stats? The number of events measured may conflate the significance. Try using Superplots.

      Here is how we proceed: we count the number of occurrence of the phenotype we monitor, and the total number of cells. We apply a X<sup>2</sup> to test whether there is a significative difference between our replicates in each condition. If not, we pool the number of occurrence of the phenotype we monitor and the total number of cells for the 3 replicates, and for each condition. Finally we apply a X<sup>2</sup> between the different conditions. This is how we usually proceed to avoid comparing a mean of percentages. This is now explained in the methods.

      Minor points:

      (1) "DEUP1 is a centrosomal protein and assembles deuterosomes in the pericentrosomal region in brain MCC". I am not sure you have evidence that DEUP1 is a centrosomal protein. You don't seem to study the relationship between centrosomes and DEUP1? Rewrite this title and tone down this claim.

      This has been modified.

      (2) Why the crossbow micropattern (versus some other shape) - seems very specific but not discussed?

      We wanted a shape where centrosome is not localized at the center of mass of the nucleus. Among the corresponding patterns, the crossbow was the one where differentiating cells had less propensity to detach.

      (3) Figure 2 - are the foci of DEUP1 at the cloud stage smaller than at A stage? How do they grow? Measure the diameter at cloud stage, just after they leave the cloud and then once they move away from centrosomal cloud and each other. If so, and they do indeed grow in size from the cloud stage to the growth stage which I think your images suggest - do you envision this happening with the gradual addition of DEUP1 rather than fusion?

      Early deuterosomes are not easy to detect by light microscopy, because of accumulation of DEUP1 in the cloud. We did CLEM on the cloud of early A-stage cells to resolve the earliest deuterosomes which are often very small (see Fig. 2D, Fig. 2 Supplementary 2-6) suggesting that they grow, either by fusion, which we never observe in our movies at later A-stage, or by accretion of DEUP1. However, by light microscopy, we can detect very early but big deuterosomes, which we see splitting later on into smaller ones. So, we cannot conclude on the mechanism that regulate deuterosome size. This is now discussed in the discussion of the manuscript.

      You say in the discussion:

      "Consistently, we never observed fusion events of DEUP1 condensates in our time-lapse experiments. More importantly, we did FRAP experiments on endogenously tagged mRuby-DEUP1 in cells at the different stages of centriole amplification, and did not find significant recovery, supporting that centrosomal DEUP1+ foci and deuterosomes are not liquid-like structures (Figure 8 Supplementary 2)." How do you prove there is no fusion of deuterosomes?

      It is always difficult to prove the absence of something, we agree! But we did tens of movies with high temporal resolution and never observed fusion events. But, as we say in the previous question, the very early deuterosomes can be very small and we do not distinguish them from the DEUP1+ cloud by live imaging. So at this stage, we cannot say. But later on, during A- or G-stage and when deuterosomes are outside the cloud to be easily observed, we very often observe deuterosomes bumping into each others and stay in close contact for minutes, but then moving away. This, for us, supports the lack of fusion properties. But the question remains open. We now explain this in the manuscript and have added an example in video 28.

      If they are getting bigger as I think your imaging suggests from cloud to growth stage, then how is this happening?

      MT depolymerisation and dynein inhibition leads to the formation of very small deuterosomes. Dynein inhibition can even lead to a block in the formation of new deuterosomes suggesting that DEUP1 concentration is a crucial parameter for condensation into deuterosomes. Deuterosome growth may happen through oligomerization of DEUP1 molecules allowed by their dyne-independent concentration. Sorokin in 1968 proposed that a supersaturation of deuterosome components may lead to their solid crystallization into deuterosomes. Deuterosome size can also be regulated by a more complex molecular cascade, involving post-translational modifications of DEUP1 or PCM, such as phosphorylations driven by the cell cycle machinery. This would be consistent with the fact that deuterosomes are very big in the absence of CCNO, a cyclin required for entering the MCC cell cycle variant. This will need further investigations.

      I'm not sure FRAP actually proves fusion doesn't happen.

      Agreed, this is not what we wanted to say, we clarified. The FRAP experiment just suggests that it is not liquid-like.

      It is technically difficult to laser ablate individual or only subsets of deuterosomes...

      This is what was done but anyway, FRAP does not firmly show that deuterosome compartments are not liquid-like as we now precise.

      (4) How do you fix your cells for expansion as you have no preservation of cytoplasmic microtubules? You are saying that there is a "nest" of MTs but beta tubulin ONLY stains the cilia and centriole - why is this? Tyrosinated tubulin on regular confocal shows strong cytoplasmic staining. See Figure 3.

      Cytoplasmic microtubules do not preserve well through the expansion process. We did try a few different fixations and pre-extraction methods but they come at a trade-off to preserving centrioles. i.e. we could either preserve cytoplasmic tubes or centrioles but not both with the same processing method.

      (5) "PCNT puncta partially overlap with centrin (Figure 3 Supplementary 2C). At this stage, PLK4, the master regulatory kinase, and SAS6, one of the first centriolar components are either absent or present as small foci within the cloud, often on the wall of the parent centrioles (Figure 3B-C)." some arrows to highlight this would be useful - difficult to see?

      We have tried to make arrows on what is now Fig. 3 Supplementary 1 G, but there is to many CENTRIN colocalizing with PCNT. We have enhanced the contrast of the merge to make it more visible.

      (6) Figure 3I legend - what are the arrows pointing at? Yellow and white on inserts? ". Around the same time as tubulin, centrin is also recruited to procentrioles (Figure 3I). This stage is probably the stage that we previously documented as A"

      However you see centrin at DEUP1 foci in D, and you don't show any eg. SAS6 or PLK4 positive DEUP1+ structures lacking centrin specifically, centrin seems to be present on all the procentrioles in Figure 3I. Did I miss it where you show centrin negative procentrioles in the cloud?

      Fig. 3I (now Fig. Supplementary 1J), yellow arrows are pointing at centrioles with non-acetylated MT while white arrows point at acetylated MT. This is now indicated in the legend.

      Regarding CENTRIN, it is present as a diffuse staining around the centrosome since the very beginning of amplification (now in Fig. 3 Supplementary 1A with different contrasts), in addition to compose the parental centrioles. This staining can therefore overlap with DEUP1 staining when DEUP1 appears (Fig. 3 Supplementary 1B, E) but not necessarily. In live we observe that CENTRIN and DEUP1 foci can move independently at early stages (Fig. 2 Supplementary 1B, video 2). This is later on, as shown now in Fig. 3 Supplementary 1J (previously Fig. 3I), that procentrioles are all strongly positive for CENTRIN.

      A new paper (Laporte et al., Cell 2024) recently showed that the recruitment of CENTRIN on duplicating procentrioles first occurs at the distal end, visible by a small dot, and then appears gradually at the level of the inner scaffold when procentriole reach 160nm, the stage where POC5 appears, which corresponds to the A-to-G transition in our MCC progenitors (Al Jord et al., 2014). One can therefore consider that the same is happening in our cells, and that, with the CENTRIN cloud, we have difficulties to detect the distal CENTRIN dot. We have changed the text to add this reference and discuss CENTRIN apparition in MCC procentrioles.

      (7) " The DEUP1 asymmetry previously described at the centrosomal daughter centriole (Al Jord etal., 2014) becomes visible in some cells during the cloud stage (Figure 3B, N; Figure 3 Supplementary 2B) and in a majority of cells" difficult to see - maybe enlarge and single channel from Figure 3F-H in the supplemental Figure 3 to emphasise this?

      We have either changed the pictures or the contrast to be more representative with the quantifications. This is visible in Fig. 3A, D, E, G; Fig3. Supplementary 1E and now using correlative light and EM in Fig. 2 Supplementary 2, 3, 4 and Fig. 4 Supplementary 3-4. One has to look carefully at the daughter centriole (marked “dc”). We have not zoomed in since previous manuscript have already described this at later stages with bigger deuterosomes. You can refer to main or supplementary figures in previous manuscripts (Al Jord 2014, Khoury Damaa 2024) where serial sections span the entire deuterosomes and daughter centrioles and show, with nanometric resolution, that both strutures are frequently sticked to each others on tens of nanometers.

      (8) Do you have videos of DEUP1 oscillations with nocodazole to show a lack of oscillations?

      We have now added videos of DEUP1 oscillations under nocodazole and dynapyrazole treatments.

      (9) "In addition, co-staining of centrioles and nuclear pore proteins show a tight colocalization(Figure 6 Supplementary 1B)." I see the colocalisation in panel 1 but less obvious with panel 2 maybe have some more zoomed in panels and some quantification of the colocalization? Is it more striking at the G stage than the D stage?

      We edit and change the phrasing as ‘colocalization with NPC’ is not the good term. There is too many centrioles and NPC, they cannot do otherwise than colocalize… What we want to say is that there is a tight connexion with the nuclear envelope. This is why we present a single z, to show that centrioles and NPC are on the same z-plane. This is also clearly visible for centrioles that are loaded on deuterosomes that are around the nuclear membrane in the XY plane. We also added a video to show an entire z-stack of this kind of staining.

      (10) "Indeed, SAS6 normally disappears from procentrioles when centrioles are docked, just beforeciliation (Al Jord et al., 2014). This suggests that centrioles were able to degrade SAS6, a process also dependent on APC/C (Strnad et al., 2007), but failed to disengage from deuterosomes." Figure 6 Supplement 1E-F - are you sure it wasn't that Sas6 wasn't loaded correctly at the earlier stage and so is reduced recruitment rather than premature disengagement of Sas6? If it is indeed premature disengagement of Sas-6 - what about CP110 - does the CP110 get loaded and is it still present in noco treated cells arrested in the D phase?

      We do not observe SAS6-negative procentrioles on deuterosomes at G-stage but only on deuterosomes in D-stage cells (cells with partly disengaged procentrioles). This is why we hypothesize that, because of the long duration of D-stage and knowing that SAS6 is finally degraded at the end of amplification (Al Jord et al., 2014), we are in the presence of cells where SAS6 has been degraded but where centrioles did not manage to disengage. This is now clarified in the text.

      (11) Can you track deuterostome splitting live? Maybe not enough spatial or time resolution?

      One has to monitor in 3D (multiple z because deuterosomes move a lot), 2 colors, high temporal resolution (dt=2-5’; to be able to track a single deuterosome), and long duration (deuterosomes are sometimes touching each other and then moving away, giving the impression that they split). This eventually leads to the bleaching of the mRuby fusion protein… We have put an example of what we think is a deuterosome splitting in Fig. 6E (former Fig. 7D). But we decided to finally monitor with low temporal resolution (dt=40’) to avoid photobleaching, and analyze numerous deuterosomes and cells to quantify the number and size of deuterosomes over time in single cells.

      (12) The MT nodes - can you segment the tyrosinated MTs and define nodes and then quantify theDEUP1 presence on them?

      Please see answer to reviewer 1 regarding this point.

      (13) Figure 8 supp 1 (E): Representative XY distribution of CEN2-GFP+ centrioles at the end of migration (Sas6 negative) in brain MCCs treated with DMSO, Nocodazole 1µM and 5µM (48h). Scale bar, 5µm Bit more detail on how you define fully migrated vs still migrating centrioles in z. You say you are using Sas-6 negativity to define fully migrated cells in the legend, yet you say noco treatment leads to premature sas-6 negativity, and yet the apical migration takes longer upon noco treatment?

      Nocodazole does not lead to premature SAS6 negativity but to a partial disengagement which lead to SAS6 negative “mature” centrioles being still connected to deuterosomes. We define complete migration when all the centrioles are on the apical side of the nucleus. We now clearly define what “apical” migration stands for in the main text and changed the pictograms in Fig. 8G to clarify this.

      (14) Figure 8H and video 18 - it isn't obviously clear to me that the noco-treated cells are "more erratic" or how you decide what counts as apically migrated successfully. How do you control for drift in z? Can you track individual centrioles as you did in untreated and define what is "erratic about their movement?

      Erratic means that the centrioles are moving away from each others, and back, in a non-predictable way, instead of migrating up and gathering. The drift in z of the whole cell is visible because there is always some centrioles, that are apically located at the beginning, that remains on the apical membrane, probably because they are already docked.

      We have indeed followed the centrioles individually in the nocodazole condition. However, in the control, the XYZ coordinates of one of the centrioles of the centrosome, which normally don’t move, are substracted to the coordinates of all the other centrioles as explained in the method section. This allows to have a subcellular reference, and to circumvent the movements of the cell, which are non-negligible at all at this timescale. In the nocodazole treated cells, the centrosomal centrioles share the erratic movements of the other centrioles and can migrate up and down, which exclude them as a reference. Since the nucleus is also moving a lot, we were left with no reference point.

      (15) Figure 8 supplement 1E can you quantify the final area of centriole patch in XY upon noco treatment?

      It was in main Fig. 8J and is now in Fig. 8 Supplementary 1F.

      (16) Figure 8J legend- MBB is never defined as an acronym.

      Thank you for pointing this.

      (17) Define what is the frequency and how is it calculated - Figure 8J.

      This is the MBB patch area in µm<sup>2</sup>

      Text edits:

      (1) "Altogether, these results suggest that, in this non-tissue-specific proxy of MCC progenitors, microtubules organize the onset of centriole amplification in the pericentrosomal region."

      Sentences have changed.

      (2) "Increasing the temporal resolution to 5-15s reveals that DEUP1+ foci observe an exhibit oscillatory dynamics to at the centrosome (Figure 2E, colored arrows, Video 3, 5/10 cells observed for 1-4min)."

      Sentences have changed.

      (3) "stage procentrioles were involved in this perinuclear migration and distribution. In fact, this dynamic is reminiscent of the centrosome migration that occurs during the G2-to-M progression in cycling cells in preparation for mitotic spindle organization. In cycling cells, this" Grammar - maybe change to "stage procentrioles were involved in this perinuclear migration and distribution. This is reminiscent of the centrosome migration that occurs during the G2-to-M".

      Sentences have changed.

      (4) "We then wondered whether these microtubule-dependent dynamics was were required for an efficient subsequent centriole disengagement during the following D-stage."

      Sentences have changed.

      (5) "Then, monitoring tens of disengagement movies, we identified a transient stage during which disengaging procentrioles redistribute isotropically in the 3 dimensions, along the nuclear membrane (Figure 6A, 4:30, Video 7) before losing its contact to migrate to the apical surface (Figure 6A, 6:30 to 14:00)."

      Sentences have changed.

      (6) Discussion: "Since pioneer electron microscopy studies on basal body production in quail oviduct MCC 35 years ago (Boisvieux-Ulrich et al., 1987, 1990; Boisvieux-Ulrich et al., 1989), this work is the first to assess the role of microtubules in the now finely described centriole amplification process. This"

      Sentences have changed.

      (7) "Using live imaging on brain MCC, we highlight the existence of a nest composed of DEUP1, PCNT and Centrin2, pre-assembled before the onset of centriole amplification onset."

      Sentences have changed.

      (8) "Recently, formation of DEUP1 pure condensates in solution as well as FRAP experiments after overexpression of DEUP1 in MCC progenitors suggested that deuterosomes where are not liquidlike structures (Yamamoto & Kitagawa, 2019). Consistently, we never observed fusion events of DEUP1."

      Sentences have changed.

      (9) "This reminds is reminiscent of the centriole-to-centrosome conversion occurring at the G2-M transition followed by the associated microtubule dependent nuclear migration of new centrosomes at mitosis onset (Agircan et al., 2014)."

      Sentences have changed.

      (10) "Following individual trajectories requires high resolutive resolution spatio-temporal live imaging while avoiding excessive light exposure which disturbs centriole migration (Boudjema et al., 2024)."

      Sentences have changed.

      (11) "Using high temporal resolution microscopy, we further identify that individual dynamics is are complex and can be splitted between divided into the baso-apical migration, where centrioles move in a processive and more..."

      Sentences have changed.

      Reviewer #3 (Recommendations For The Authors):

      (1) Growing MEF-MCCs on micropatterns has successfully mimicked the dynamics of centriole amplification in brain MCCs, allowing the authors to study the spatial origin of procentrioles. Since this is a powerful system, a more quantitative description of the system will be informative and beneficial for future studies. For example: What is the efficiency of this system? Do the cilia that form in MEF-MCCs motile?

      The system of MEF-MCCs has been described in a previous paper from the Kintner lab. It seems that growing the MEF-MCCs on micropatterns did not ameliorate the ciliation which is partial, probably due to the absence of an apico-basal polarity.

      (2) Figure 2: The analogy drawn by the authors between DEUP1 oscillatory dynamics and centriolar satellites is intriguing. In early amplifying cells within the cloud, do these DEUP1 structures co-localize with the satellite marker PCM1?

      We have added immuno stainings of PCM1 in mRuby-DEUP1 / CEN2-GFP cells in Fig. Supplementary 2E. Within the centrosomal cloud, DEUP1 colocalizes with PCM1. Interestingly, this PCM1 concentration at the centrosome is dependent, at least in part, on dyneins. Then, PCM1 can localize around the deuterosomes, but it is never colocalized with deuterosomes (not shown). This is also showed by immuno-EM in Zhao et al., 2019. Although it was shown that PCM1 is a proximity interactor of DEUP1 (called ccdc67 at that time) by Firat-Karalar et al., 2014., absence of PCM1 staining on deuterosomes does not favor the hypothesis of PCM1 and DEUP1 being part of the same entities. One could hypothesizes that DEUP1 is transcribed locally within the satellites, explaining the colocalization of the 2 proteins and the + BioID results, and then form PCM1negative deuterosomes.

      (3) The authors propose a physical link between deuterosomes and centrosomes based on their oscillatory behavior. How are the oscillatory dynamics of DEUP1 affected by nocodazole treatment or inhibition of microtubule motors (i.e ciliobrevin treatment)?

      These oscillations are inhibited by nocodazole (Fig. 4D). They are also inhibited by dynapyrazole (Fig. 4D). We never succeeded in having a nice disruption of the Golgi apparatus with ciliobrevin and therefore we did not used it.

      (4) In addition to nocodazole treatment, it would be important to determine the consequences of microtubule stabilization by taxol and inhibition of microtubule motors during critical stages of centriole amplification where microtubules are reported to play a role for the first time in this manuscript. Another interesting area of investigation will be to study the extent to which microtubule PTMs contribute to these processes.

      We now blocks dyneins during the different stages of amplification. The results are in main and associated Fig. 4, 5, 7, 8. The role of microtubule PTM, is not in the scope of this manuscript.

      (5) Describing microtubule dynamics along with Centrin/DEUP1 dynamics will be informative in assessing whether these structures associate and/or move along microtubules? Have the authors performed their imaging experiments with SIR tubulin?

      Yes, we have tried hard! But we have encountered different obstacles:

      3-color video microscopy is phototoxic,

      siRTubulin is bleaching very rapidly

      The density of microtubules in MCC makes the observation hardly informative

      (6) Figure 5: The role of PLK1 in centriole-centrosome conversion and generation of multiple MTOCs can be tested with a PLK1 inhibitor for further confirmation.

      We have also tried but inhibiting Plk1 blocks the A-to-G and G-to-D transitions so it was not possible to uncouple the role of Plk1 in stage transitions versus centriole maturation.

      (7) Figure 6: The tight co-localization of nuclear pore proteins with centrioles poses questions about the role of nuclear pore proteins or other nuclear proteins that are associated with centrioles during centriole disengagement and migration. Considering the existing literature on centrosome-nucleus attachments, can there be a way to test this question within the scope of this manuscript?

      We have tried to deplete Nup133 but it’s killing the cells. Our additional experiments now show that the nuclear migration of centrioles during G-stage is dynein dependent, reinforcing the parallel with centrosome migration in prophase. We also added results from our scRNA sequencing (Fig. 5 Supplementary 1) showing that some key players of centriole migration to the nuclear membrane are conserved in the MCC cell cycle variant, and expressed with a comparable dynamics as to the canonical cell cycle.

      (8) Figure 8: Manually tracking a subset of migrating centrioles to define their dynamics during centriole migration and docking provides valuable analysis for determining the molecular mechanism of these processes. In addition to microtubules, does actin contribute to this process? Since centrioles eventually migrate to the apical side in nocodazole-treated cells, there should be other molecular players involved in this process.

      We did block actin polymerization but we found that the different stages were affected and that it would be better to dedicate a whole manuscript on the role of actin during each stage of amplification. We discuss the migration mechanism, and the putative role of actin, in the discussion.

      (9) The legends for Supplementary Figures 1 and 2 in Figure 3 are mixed and need correction.

      Figures have been remodelled.

      (10) In Figure 3P, the term "PLK4+" is labeled in bright green, which is not clearly visible. It maybe beneficial to change the color of this label for better visibility.

      We have tried to correct this.

      (11) Figure 6F quantifies "% tethered flowers" on the nuclear membrane. When quantifying, is the3D localization of DEUP1 flowers in both DMSO- and Noc-treated cells considered? A flower may appear to be on the nucleus in 2D, but it could be detached from the membrane in a 3D view.

      The quantifications are done in 3D. However, flowers that are below or above the nucleus are not quantified since the space is confined and the resolution in z to small to see whether they are connected or not. This is now precised in the legend.

      Before the editors proceed with an updated assessment, they've requested that we pass on some of the comments that have arisen as part of the evaluation of your revised manuscript. They feel that these concerns should be addressed before we proceed with issuing a formal assessment and publishing the revised Reviewed Preprint:

      We thank the reviewers and the editors for the corrections and insighfull comments. We apologize for our delayed answer and hope our corrections in the main text and some of the figures will give them satisfaction.

      The revised manuscript is greatly improved with nice new data regarding the role of microtubules. It also has changed quite a bit including the title. The new focus is on the cell and centriole cycle variants in MCC. While this helped to focus the study, there remains an important issue related to the interpretation of the data and the proposed 2-in-1 cycle model. Before providing the final updated assessment, we ask you to address the following points (which were raised already in the first round of review): The manuscript still contains statements that are not aligned with published work and the current view in the field regarding the timing of events during canonical centriole biogenesis. These timings are in conflict with your model that 2 centriole cycles are "superposed" in the MCC cell cycle variant, as currently presented. An alternative straightforward interpretation would be that multiciliogenesis uses an accelerated centriole duplication cycle where key steps occur concomitantly or in short succession instead of being separated by mitotic divisions as in the canonical cycle.

      We do agree with the acceleration of all steps into only one cycle, this is actually what we think we have proposed. When correcting our confusions as regard to centriole-to-centrosome conversion (as explained below) and putting the events in a scheme, this reveals that the events of the two canonical cycles nicely superpose, both in term of molecular composition and dynamics (corrected Fig. 9). We therefore maintain that the null hypothesis is that the acceleration is done through a superposition of events that; although driven by the same molecular machinery, are normally occuring in two consecutive cell cycle. We explain ourself briefly in two paragraphs, before answering point by point to the questions of the reviewers.

      As regard to centriole-to-centrosome conversion:

      We thank the reviewer for pointing out that we used “MTOC conversion” for what is normally called “centrosome maturation”. We have removed the term “centriole-to-centrosome conversion” during the first round of revision but we now realize that “MTOC conversion” leads to the same misinterpretation as regard to the literature on centriole duplication.

      The reviewer asks us to refer to the work of the Tsou lab (Wang 2011, reference now added in the manuscript) showing that daughter centrioles are “modified” (e.g. recruit PCM, become competent for MT nucleation and duplication) during late M/early G1. This “centriole-to-centrosome conversion” can’t occur for our procentrioles at this stage since they are not even born during the mitosis that precedes MCC differentiation. Also, in our cells, such modification does not include the capacity to become competent for duplication since we know that procentrioles become basal bodies without making any round of duplication (Al Jord et al., 2014).

      Also, we have not done the experiments to tackle the question on when our centriole become “modified-like”. What we can say is that during A-stage, they become progressively positive for PCM (Fig. 5 Supplementary 2) and a weak signal shows that some MT are seen emerging from them (Fig. 5 and Fig. 5 Supplementary 2, and see point by point answer).

      What we do see is that, at the A-to-G transition, they increase their PCM recruitment, show clear and strong MTOC ability (sometimes as strong as the centrosomal centrioles), and that this is associated with migration and separation of centrosome/deuterosomes around the nuclear membrane (Fig. 5). We therefore connect this to what occurs at the G2/M transition which is an increased recruitment of PCM protein, an increased ability to nucleate MT, associated with centrosome migration and separation at the nuclear membrane. Since this process in the canonical cell cycle is called “centrosome maturation”, we therefore should refer to this term in our study. However, centrioles in the MCC variants are not organized in centrosomes, so we now compare what we see to the “centrosome maturation” of the canonical cell cycle with an associated reference (Joukov et al., 2018), but name it “centriole maturation”.

      We have modified the text (track changes visibles) and the schemes (Fig. 5, Fig. 5 Supplementary 1 and 2, Fig. 9, Fig. 9 Supplementary S1; new versions uploaded) accordingly.

      As regard to 1.5 or 2 cell cycles

      Except for the “MTOC conversion” that we have now changed, as explained above, we think our work does suggest (depicted on Fig. 9) what the reviewer states for centriole duplication: “In the current view, centriole biogenesis starts in early S, elongation proceeds through G2/M and by early G1 it is complete. During M/early G1 centrioles disengage and newly formed daughters recruit PCM (centrosome conversion). Then these centrioles go through another complete cell cycle and when they reach early G1 again they have acquired DAs and SDAs. Key here is that biogenesis and disengagement/centrosome conversion are separated by the first mitosis (ensuring duplication occurs only once), and acquisition of DAs and SDAs is separated by another mitosis (ensuring that cells only form a single cilium)”.

      We feel that going from early S to a G1 phase, after 2 mitosis, is what one can call “2 cell cycles”. One of the paper that inspired us a lot when studying how the cell cycle machinery can drive centriole amplification in MCC is a paper from Jadranka Loncarek team (Kong et al., 2014) where they also state that “nascent centrioles gradually mature through 2 cell cycles”. Very interestingly, in this study they show that when they enhance Plk1 activation, they could erase centriole age and new procentrioles are able to recruit PCM and appendages within only 1 cell cycle, without mitotic progression, like what we see in MCC. We have added the reference in our discussion.

      Point by point answer

      (1) Original work on canonical centriole disengagement and centriole-to-centrosome conversion should be cited (e.g. PMID: 16862117, PMID: 21576395)

      As explained earlier, we used the wrong term since the begining. We do not speak about the centriole-to-centrosome (nor MTOC) conversion since we do not test when centriole modification (Wang et al., 2011) occurs in the MCC cell cycle variant. We know that PCNT is present on the procentrioles during A-stage (as shown in Fig. 5 Supplementary 2B), but we do not know when it is recruited (UExM did not work properly with this antibody). We quantify a weak MT staining in regrowth experiment during A-stage and see that procentrioles can be connected to MT in both brain MCC and MEFs (as shown in Fig. 5D, E for brain MCC and Fig. 5 Supplementary 2F for MEFs) , but we do not know when during A-stage they become competent for nucleation. We therefore did not speak about this process that we do not document. What we clearly document/quantify is the enhanced MT nucleation capacities at the A-to-G transition, concomitent with the nuclear migration (easily defined with Cen2-GFP or GT335 stainings) and that we compare to centrosome maturation occuring at the canonical G2/M transition.

      (2) The authors state in several places that canonical centriole formation and maturation takes two iterations of the canonical cell cycle. This is imprecise. Based on the above work and work by others, the broadly accepted view is that it takes 1.5 cell cycles. This difference matters for the final proposed model (see below). Reviewed e.g. here: PMID: 20869612; PMID: 30601682

      Our answer is in the preamble.

      (3) "Centriole maturation cycle superposes with centriole elongation cycle in the MCC cell cycle variant": Your description of the canonical cycle differs from the current view in the field. In the current view, centriole biogenesis starts in early S, elongation proceeds through G2/M and by early G1 it is complete. During M/early G1 centrioles disengage and newly formed daughters recruit PCM (centrosome conversion). All this occurs in 0.5 cycles. Then these centrioles go through another complete cell cycle and when they reach early G1 again they have acquired DAs and SDAs (total of 1.5 cell cycles). Key here is that biogenesis and disengagement/centrosome conversion are separated by the first mitosis (ensuring duplication occurs only once), and acquisition of DAs and SDAs is separated by another mitosis (ensuring that cells only form a single cilium).

      (4) Fig 5A, B and Fig. 9

      (a) Are 2 separate figures needed for the model? They seem redundant.

      We find it easier not to wait Fig. 9 to have the first part depicted.

      (b) The model shows loss of SAS6 throughout G1, but this already occurs during M/early G1

      Thanks. It was already ok in Fig. 9, we have modified for Fig. 5.

      The model shows "MTOC capacity/conversion" during S phase, but this occurs during early G1

      Thanks a lot, as explained earlier, we used the term MTOC conversion occurring in G1 for what is normally called centrosome maturation occurring in G2/M, as explained earlier. We do not speak anymore of MTOC conversion since we have not tackled this question (explained above). We have therefore removed MTOC conversion in the texts and the schemes and replaced it by “centrosome maturation” for the duplication cycle, and by “enhanced MT nucleation capacity” for the MCC cycle. To be clearer and schematize that procentrioles are competent for MT nucleation before G2/M or A/G transitions, we have added some MT nucleated from G1 procentrioles during the canonical cycle, and from late A-stage procentrioles during the MCC cycle.

      The model shows disengagement only in the second M phase, but this occurs already at the first M phase, directly following centriole biogenesis, right before centosome conversion.

      This is a big edition error in both Fig. 5 and 9. Of course the daughter centriole disengage during the first M-phase. This has been changed. Thanks a lot for spotting it. This, however does not contradict the hypothesis of superposition.

      We also added the acquisition of distal appendage which was written in Fig. 5 but not in Fig.

      9 for duplication during the second M-phase.

      When the correct timings are incorporated in the figure, the proposed superposition of two cycles is not an accurate description of the events. Instead, your data seem consistent with a model where MCC incorporates all steps in one cell cycle variant that lacks mitoses, so that disengagement and MTOC conversion occur together with centriole elongation, followed immediately by acquisition of DAs and SDAs.

      We do agree with the acceleration of all steps into only one cycle, this is actually what we tried to propose. When putting the events in a scheme, this reveals that the events of the two canonical cycles nicely superpose, both in term of molecular composition and dynamics (Fig. 9). We therefore maintain that the null hypothesis is that the acceleration is done through a super opposition of events that; although driven by the same molecular machinery, are normally occurring in two consecutive cell cycle. This is notably consistent with the findings of Kong et al., 2014 cited previously.

      (5) While all reviewers felt that there was no need to introduce the new term "nest", they leave it to the authors to keep it. However, the authors may want to consider that the term is still not introduced and explained properly, which may confuse readers. For example, while this section reads like an introduction to the term: "Correlative DEUP1 live-imaging and EM highlights the existence of a pericentrosomal "nest" in brain MCC", the term is already used two times before without explanation. The first mentioning is at the beginning of the results section and is followed by citations, which gives the impression that these studies describe the nest, which is not the case.

      The first mention of “nest” is in the end of introduction resuming the findings of the paper where the term is in the following context: “we found that centriole amplification emerges in a pericentrosomal “nest” concentrating core centriole/deuterosome elements”. We looked at nest definition in the Collins Dictionnary : “a structure or other place where creatures, esp. birds, give birth or leave their eggs to develop”, we felt this was clear. We added quotation marks around the term nest.

      Then, the result section opens with this sentence: “The origin of amplified centrioles in MCC remains controversial. Some live imaging experiments and electron microscopy suggest that the centrosome could constitute a nest for centriole and deuterosome biogenesis (Al Jord et al., 2014; Kalnins et al., 1972; Mori et al., 2017), but others have proposed that procentriole-loaded deuterosomes emerge independently from the centrosome location, all over the cytoplasm (Nanjundappa et al., 2019; Sorokin, 1968; Zhao et al., 2013, 2019).”. Here, the term nest is again used as a place of birth for centrioles and deuterosomes which is what is actually proposed in these papers. First, Kalnins el al., in 1969 (we made an error on the reference date, this has been changed), resume in their abstract “This observation suggests that all of the clusters may form initially in close association with the diplosomal centrioles”. Then, not to mention Al Jord 2014 which comes from our lab, the title of Mori et al. is “Cytoplasmic E2f4 forms organizing centres for initiation of centriole amplification during multiciliogenesis”, and in the paper, they show that E2F4 accumulates at the centrosome. This is now also proposed by collaborators for MCIDAS (Lu et al., 2025). We feel that these references, which are often omitted, are appropriated at this location.

      Then we continue with: “To test whether microtubules drive the organization of a centrosomal nest from which procentrioles emerge”, which keeps the notion of the place of birth.

      Then the title "Correlative DEUP1 live-imaging and EM highlights the existence of a pericentrosomal "nest" in brain MCC" arrives. In this section we first speak about a pericentriosomal cloud on which we zoom in using CLEM, to then conclude at the end of the section “Altogether live imaging mRuby-DEUP1/CEN2-GFP during early A-stage suggests that core deuterosome and centriole components are concentrated in a primordial cloud around the centrosome, which constitutes a nest where centrioles and deuterosomes concomitantly form before they move away from the centrosomal region (Fig. 2F)”.

      Finally, we begin the discussion section regarding the nest by: “We named this transitory compartment a “nest” since deuterosomes and procentrioles emerge specifically in this region and grow while moving away from it.”

      During the first revision, we tried to make it clearer. If this is still not the case after and the reviewer has another proposition of definitions/phrasing, we will be glad to consider it.

      As replied to the other reviewer, the term “nest” does not need to be retained as a new terminology. It is just a way for us to identify the transitory region and to best define one of its function/characteristic which is to host the birth of new deuterosomes and centrioles.

      The following comments from Reviewer #3 may also provide further context regarding the editors' remaining concerns:

      The authors have done an excellent job addressing the points I raised overall, and the revision is substantially improved in focus and clarity. That said, some concerns raised by other reviewers, particularly regarding terminology and statistical analysis, could have been addressed more fully. One issue remains insufficiently resolved. Several quantitative analyses (for example Fig. 5C and 5E) still appear to rely on pooled single-event measurements collected across three independent experiments. This approach can overstate statistical significance. The authors indicate in their rebuttal that they use chi-square tests to compare proportions and to justify pooling across replicates. However, I am not convinced this addresses the issue for the intensity-based and single event distributions shown in the panels specified above. I recommend that these key analyses be represented with biological replicates shown explicitly (superplot-style, with replicates distinguished).

      Our reply was for the comparison of proportions and not the intensity-based and single event distributions shown in the panels Fig. 5C and Fig. 5E. We have now changed our plots to represent biological replicates explicitly (superplot-style, with replicates distinguished). As for the statistical analysis: we evaluated differences in marker intensity between A-stage and G-stage samples using a linear regression model, with stages as the main effect and replicate as a fixed covariate, to account for batch variation. Statistical significance was assessed using Type II ANOVA.

      Separately, I continue to feel that some newly introduced terminology (for example, the "nest") may not be necessary at this stage. It may be sufficient to describe these structures and focus on their spatiotemporal behavior, composition, and measurable features, rather than assigning new names. Having read the authors' response, I understand that they would like to retain this terminology, which is acceptable; however, it may not be readily adopted by the field.

      The term “nest” does not need to be retained as a new terminology. It is just a way for us to identify the region and to best define one of its function/characteristic which is to host the birth of new deuterosomes and centrioles.

      Minor correction (remove "in MCCs" part from the following sentence):

      In MCC, PCM1 depletion alters deuterosome formation and centriole production in brain and airway MCC (Hall et al., 2023; Zhao et al., 2021).

      Done

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work convincingly shows that, rather than gradually "evolving" throughout interphase, global chromatin architecture undergoes unexpectedly sharp remodeling at G1-S (and to a lesser extent, S-G2) transitions. By applying "standard" Hi-C analyses on carefully sorted cells, the authors provide an excellent temporal view of how global chromatin architecture is changed throughout the cell cycle. They show a surprisingly abrupt increase in compartmentation strength (particularly interactions between the "active" A compartments) at G1-S transition, which is slightly weakened at S-G2 transition. Follow-up experiments show convincingly that the compartment "maturation" does not require the DNA synthesis accompanying S phase per se, but the authors have not identified the responsible factors (work for future publications). The possible biological ramifications of these architectural changes (setting up potential replication "factories", and/or facilitating transcription-replication conflict resolution, both more pertinent for the active A compartments, which are most affected) have been well discussed in the article, but still remain speculative at this stage.

      We thank Reviewer #1 for their positive and constructive assessment of our work, and we agree that the questions of responsible factors and biological ramifications are important directions for future studies.

      My major criticism of this article is aimed more at the state of the field in general, rather than this specific article, but it should be discussed to give a more balanced view: what actually is a chromatin compartment? Chromosomal tracing and live tracking experiments have shown that the majority of "structures" identified from Hi-C experiments are statistical phenomena, with even "strong" interactions only being infrequent and transient. A-B compartments are "built up" from multiple very low-frequency "interactions", so ascribing causal effects for genome functions is even tougher. As a result, I have very little confidence in the results of the authors' polymer simulations and their inferred "peninsula" A compartment structures without any other supporting experimental data.

      We thank the reviewer for raising this important conceptual point. This issue extends beyond the scope of the present study but reflects an important ongoing discussion in the 3D genome field regarding the biological interpretation of chromatin compartments.

      We agree that Hi-C interactions should not be interpreted as stable pairwise contacts present in every cell. A growing body of evidence from chromatin tracing and live-cell imaging studies has demonstrated that many chromatin interactions identified by Hi-C are probabilistic and dynamic, with substantial cell-to-cell variability. Relatively speaking, however, A/B compartment organization represents a robust population-level property of genome organization that is highly reproducible across biological replicates and closely correlates with multiple independent genomic features. In particular, replication timing (RT) correlates very well with A/B compartment organization, with early and late RT domains corresponding to A and B compartment domains, respectively.

      Furthermore, single-cell DNA replication sequencing (scRepli-seq) analyses have revealed remarkably low cell-to-cell variability in RT, suggesting that RT profiles and A/B compartment organization reflect biologically meaningful and relatively stable features of nuclear architecture rather than purely statistical artifacts. Thus, while individual chromatin contacts may be transient and probabilistic, the megabase-scale compartment organization inferred from them appears sufficiently reproducible to support reproducible RT programs and other genome functions. Additional support comes from decades of work on DNA replication demonstrating that spatiotemporal replication patterns, visualized as replication foci following short EdU pulses, are remarkably reproducible between individual cells throughout S-phase progression. These patterns reveal clear spatial segregation between early-replicating A-compartment regions and late-replicating B-compartment regions even at the single-cell level.

      To directly address the reviewer’s concern that A/B compartment organization might represent only an ensemble-level statistical phenomenon without biological relevance at the single-cell level, we performed L1/B1-EdU DNA FISH on asynchronous mESCs and MC12 embryonic carcinoma cells. L1 elements are enriched in B compartment domains, while B1 elements are enriched in A compartment domains, allowing visualization of compartment segregation in individual nuclei across the cell cycle. This single-cell analysis confirmed our Hi-C findings: compartment segregation increased from G1 to early S, remained elevated throughout S phase with reduced cell-to-cell variability, and then weakened in G2. Thus, compartment segregation is detectable in single cells, and the temporal dynamics of compartment maturation identified by population Hi-C were independently recapitulated at single-cell resolution. We have added a new Results section describing these findings titled “Stepwise A/B compartment reorganization during interphase is conserved at single-cell resolution”, including new Figure panels 2D–H and Figure S5.

      Regarding the polymer simulations, we agree that these models should be interpreted with caution. We do not view them as direct representations of individual nuclei, but rather as heuristic models that help visualize structural trends present in the Hi-C data. To make this point explicit, we have added the following statement to the revised manuscript: “We note that these models are derived from population-averaged Hi-C data and should therefore be interpreted as a heuristic framework for understanding A/B compartment dynamics, rather than as definitive representations of individual nuclei.”

      That said, we did try to provide orthogonal experimental support for the "A peninsula" model by performing DNA FISH. In brief, we measured distances between probe pairs spanning two A domains on chromosomes 2 and 15 across different cell-cycle stages. We observed significant increases in inter-probe distances from G1 to early/mid S, with the most pronounced changes involving the central probes (i.e., probes located near the domain center), consistent with physical extension of the A domain during S phase. While these data do not prove the exact geometry depicted by the model, these findings provide independent experimental support for the peninsula model as a simplified but biologically grounded interpretation of the Hi-C data. These results are described in the Results section titled “A-compartment consolidation during S-phase involves enhanced long-range contacts and structural reorganization” and are presented in new Figure panels 5D–F and Figure S12.

      We thank the reviewer again for raising this important conceptual issue, which prompted us to better clarify both the biological interpretation and the limitations of our analyses.

      Specific minor points:

      (1) A better explanation for how Figure 1E was generated is required, because this figure could be very misleading. Figure 1F and all other cis-decay plots (and the Hi-C maps themselves) show that the strongest interactions are always at smaller genomic separations, so why should there be more "heat" at the megabase ranges in Figure 1E?

      We appreciate the reviewer's observation. The apparent discrepancy is simply due to the fact that the decay plot (Fig. 1E in the original submission, now Fig. S2C) does not include the shortest-range interactions. The lowest distance plotted is 25 kb, following the method originally described in Nagano et al. (Nature, 2017), which we used as a reference. The shortest-range interactions (below 25 kb) are indeed the most enriched, as seen on the diagonal of the Hi-C maps (Fig. 2A) and in the standard cis-decay plot (Fig. 1F in the original submission, now Fig. S2F). With the 25 kb cutoff in place, the "heat" observed at megabase distances (specifically 12–50 Mb) in early/mid G1 corresponds to the dark, non‑specific band around the diagonal visible in the Hi-C maps at the same time points. This is also reflected in the cis-decay plot (Fig. S2F), where distances in that range appear above the expected curve (a "bump" rather than a linear decay).

      To avoid confusion, we have updated the figure legend accordingly (Fig. S2C): “(C) Contact decay profiles for all cell cycle phases, plotted from 25 kb to 50 Mb, illustrating a continuum of cis-interactions and a progressive shift from long-range (> 12 Mb) to short-range (< 1 Mb) interactions during the G1-to-S phase transition.”

      We hope this explanation clarifies the figure.

      (2) An ultra-high-resolution Hi-C study (Harris et al., Nat Commun, 2023) identified very small A and B compartments, including distinctions between gene promoters and gene bodies, raising further questions as to what the nature of a compartment really is beyond a statistical phenomenon. It is unreasonable to expect the authors to generate maps as deep as this prior study, but how much do their conclusions change according to the resolution of their compartment calling? The authors should include a balanced discussion on the "meaning" of A/B compartments.

      We thank the reviewer for highlighting recent ultra-high-resolution work, such as Harris et al. (Nat Commun, 2023), which reveals compartment-like features at much finer genomic scales. We agree that these findings raise important questions regarding the scale-dependence and interpretation of A/B compartmentalization.

      In our study, we specifically focus on coarse-grained compartment organization, analyzed across multiple resolutions (from ~1 Mb to sub‑megabase scales). Importantly, the key conclusions, including the abrupt strengthening of compartmentalization at the G1/S transition, are robust across these resolutions.

      We also note that fine-scale compartment-like features likely operate under different rules than larger-scale compartments. Recent evidence suggests that these "micro‑compartments" are more dynamic and transient (Harris et al., Nat Commun, 2023; Goel et al., Nat Struct Mol Biol, 2025), whereas the large-scale compartments analyzed here capture more stable, global segregation patterns. Understanding how these two regimes relate to one another remains an important open question.

      We have added the following statement in the Discussion acknowledging the scale-dependent nature of compartmentalization: “At the same time, recent ultra-high-resolution Hi-C studies [36,37] have revealed compartment-like features at much finer genomic scales, emphasizing that A/B compartmentalization is, to some extent, inherently scale-dependent. Understanding how these fine-scale, often transient micro-compartments relate to the more stable, large-scale segregation patterns described here will be an important direction for future studies.”

      Reviewer #2 (Public review):

      Summary:

      This manuscript by Choubani et al presents a technically strong analysis of A/B compartment dynamics across interphase using cell-cycle-resolved Hi-C. By combining the elegant Fucci-based staging system with in situ Hi-C, the authors achieve unusually fine temporal resolution across G1, S, and G2, particularly within the short G1 phase of mESCs. The central finding that A/B compartment strength increases abruptly at the G1/S transition, stabilizes during S phase, and subsequently weakens toward G2 challenges the prevailing view that compartmentalization strengthens monotonically throughout interphase. The authors further propose that this "compartment maturation" is triggered by S-phase entry but occurs independently of active DNA synthesis, and that it involves a consolidation and large-scale reorganization of A-compartment domains.

      Strengths:

      Overall, this is a thoughtfully executed study that will be of broad interest to the 3D genome community. The data are of high quality, and the analyses are extensive, albeit not completely novel. In particular, previous work (Nagano et al 2017 and Zhang et al 2019) has shown that compartments are re-established after mitosis and strengthened during early interphase, and single-cell Hi-C studies have reported changes in compartment association across S phase. In particular, Nagano et al show that DNA replication correlates with a build-up of compartments, similar to what is presented here, with the authors' conclusion that compartment strength peaks in early S. The idea that it weakens toward G2, rather than continuing to strengthen, appears to be novel and differs from the prevailing framing in the literature.

      We thank Reviewer #2 for their thoughtful assessment and critique. We address their specific concerns below.

      Weaknesses:

      That said, several aspects of the conceptual framing and interpretation would also benefit from further clarification, and the mechanistic interpretation of the reported compartment dynamics requires more careful positioning relative to established models of genome organization. Specific concerns are outlined below:

      (1) One of the major conclusions of the study is that compartment maturation does not require ongoing DNA replication. However, the interpretation would benefit from more precise wording. Thymidine arrest still permits licensing, replisome assembly, and other S-phase-associated chromatin changes upstream of bulk DNA synthesis. Therefore, their data, as presented, demonstrate independence from DNA synthesis per se, but not necessarily from the broader replication program. Please clarify this distinction in the text and interpretations throughout the manuscript.

      We thank the reviewer for this important distinction. We agree with their point and have never claimed that compartment maturation is independent of the broader replication program. That is why we carefully used the term "active DNA synthesis" rather than "replication" throughout the manuscript.

      However, we acknowledge that one sentence in the text was ambiguous. The original sentence read: “These results confirm that the cell population was successfully synchronized at the G1/S boundary, representing a pre-replicative state where replication had not yet initiated, although cell-cycle markers indicated entry into S-phase.”

      We have now revised it to: “These results confirm that the cell population was successfully synchronized at the G1/S boundary, representing a state where the replication program (including origin licensing, replisome assembly, and helicase activation) has been initiated, as indicated by cell-cycle markers, but ongoing DNA synthesis (elongation) is blocked. ”

      This clarifies that compartment maturation is independent of active DNA synthesis (elongation) but not necessarily independent of upstream replication-associated processes. The change has been made in the manuscript.

      (2) A major conceptual issue that is not addressed at all is the well-established anti-correlation between cohesin-mediated loop extrusion and A/B compartmentalization. Numerous studies have shown that loss of cohesin or reduced loop extrusion leads to stronger compartment signals, whereas increased cohesin residence or enhanced extrusion weakens compartmentalization. Given this framework, an obvious alternative explanation for the authors' observations is that the abrupt increase in compartment strength at G1/S, and its decline toward G2, could reflect cell-cycle-dependent modulation of cohesin activity rather than a compartment-intrinsic "maturation" program.

      The manuscript does not explicitly consider this possibility, nor does it examine loop extrusion-related features (such as loop strength, insulation, or stripe patterns) across the same cell-cycle stages. Without discussing or analyzing this widely accepted model, it is difficult to distinguish whether the reported compartment dynamics represent a novel architectural mechanism or an indirect consequence of known changes in extrusion behavior during the cell cycle. I strongly encourage the authors to analyze their data to determine if they observe anti-correlated loop changes at the same time they observe compartment changes. Ideally, the authors would remove loop extrusion during interphase using well-established cohesin degrons available in mESCs and determine if the relative differences in compartment dynamics persist.

      We thank the reviewer for raising this interesting point. We agree that there is a well-established anti-correlation between cohesin-mediated loop extrusion and A/B compartment strength in the literature.

      To test whether cell cycle compartment dynamics, particularly compartment maturation at the G1/S transition, could be explained by changes in loop extrusion, we analyzed insulation at RAD21/CTCF sites (mESC data from Hansen et al., eLife, 2017) across the cell cycle. During normal cycling, we indeed observed an anti-correlation: insulation dropped as compartment strength increased at the G1/S transition. However, in G1/S-arrested cells, insulation did not drop compared to late G1 (it even slightly increased) even though compartment maturation still occurred, indicating that the two processes can be uncoupled. This is consistent with other studies showing that loop extrusion and compartment dynamics are driven by independent mechanisms (Nora et al., Cell, 2017; Zhang et al., Nat Commun, 2021), although we cannot fully rule out some contribution from loop extrusion dynamics without direct cohesin degron experiments.

      We have added a new Results section describing these findings titled “Compartment maturation is independent of cohesin-mediated loop extrusion”, including new Figure panels 3H, I, and Figure S7.

      (3) The proposed "peninsula-like" A-domain structures are inferred from ensemble Hi-C data and polymer modeling, rather than directly observed physical conformations. That is, single-cell imaging data clearly have shown that Hi-C (especially ensemble Hi-C) cannot uniquely specify physical conformations and that different underlying structures can produce similar contact patterns. The "peninsula" language, as written, risks being interpreted as a literal structural model rather than a conceptual visualization. Instead of risking this as just another nuanced Hi-C feature in the field, the authors could strengthen the manuscript by either (i) explicitly framing the peninsula model as a heuristic description of contact redistribution rather than a definitive physical architecture, or (ii) discussing alternative structural scenarios that could give rise to similar Hi-C patterns. Clarifying this distinction would improve the rigor and help readers better understand what aspects of A-compartment consolidation are directly supported by the data versus model-based extrapolations. For example, it would be useful to clarify whether the observed increase in long-range A-A contacts reflects spatial extension of internal A regions, changes in loop extrusion dynamics, increased compartment mixing within the A state, or population-averaged heterogeneity across alleles.

      We thank the reviewer for this important clarification. We agree that the "peninsula" model should be framed as a heuristic description. As detailed in our response to Reviewer #1 (see above), we have added a disclaimer to the manuscript and provided orthogonal DNA FISH support for physical extension of A-domains during S phase. We have also ensured that the language emphasizes the conceptual nature of the model.

      (4) The extension of the analysis to additional cell types using HiRES single-cell data is a valuable addition and supports the idea that compartment maturation is not unique to mESCs. However, the limitations of these data, in particular, the limited phase resolution, in addition to the pseudo-bulk aggregation and variable coverage, should be emphasized more clearly in the main text. Framing these results as evidence for conservation in principle, rather than definitive proof of identical dynamics across tissues, would be a more appropriate framing.

      We agree with the reviewer. We have already explicitly acknowledged the limited temporal resolution and variable coverage of the HiRES dataset in the main text. To better reflect its supporting role, we have moved the HiRES figure (previously Fig. 4) to Fig. S10 and merged the corresponding results section with the previous one titled: “Formation of a consolidated A compartment in S-phase”.

      We have also revised the language to avoid overstatement. The original conclusion read: “Together, these findings strongly indicate that compartment maturation and the accompanying A compartment consolidation represent a robust and universally observed feature across different developmental contexts.”

      This has been changed to: “Together, these findings support the notion that compartment maturation and the accompanying A-compartment consolidation are not unique to mESCs and may represent a broadly conserved feature of mammalian chromatin organization.”

      Similarly, the abstract has been adjusted from: “Moreover, compartment maturation was not limited to mESCs but was also observed across different developmental contexts in mice.” to: “Moreover, compartment maturation was not limited to mESCs but was also evident across different developmental contexts in mice.”

      These changes frame the results as evidence for conservation in principle rather than definitive proof of identical dynamics across tissues.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please address the minor points in the public review.

      In addition, on page 7, line 285: "In contrast, interactions showed minimal change across all distances though interphase". Do the authors mean "In contrast, B-B interactions..."?

      We thank the reviewer for catching this. The sentence has been corrected.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Lu and colleagues demonstrates convincingly that PRRT2 interacts with brain voltage-gated sodium channels to enhance slow inactivation in vitro and in vivo. The work is interesting and rigorously conducted. The relevance to normal physiology and disease pathophysiology (e.g., PRRT2-related genetic neurodevelopmental disorders) seems high. Some simple additional experiments could elevate the impact and make the study more complete.

      Strengths:

      Experiments are conducted rigorously, including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

      We thank the reviewer for these positive comments and for the thoughtful evaluation of our work.

      Weaknesses:

      There are a few missing experiments and one place where data are over-interpreted.

      (1) An in vitro study of Nav1.6 is conspicuously absent. In addition to being a major brain Na channel, Nav1.6 is predominant in cerebellar Purkinje neurons, which the authors note lack PRRT2 expression. They speculate that the absence of PRRT2 in these neurons facilitates the high firing rate. This hypothesis would be strengthened if PRRT2 also enhanced slow inactivation of Nav1.6. If a stable Nav1.6 cell were not available, then simple transient co-transfection experiments would suffice.

      We thank the reviewer for raising this point. In our previous work, PRRT2 produced broadly similar effects on Nav1.2 and Nav1.6. Therefore, in the initial version of this study, we focused primarily on Nav1.2 as a representative neuronal Nav channel isoform and placed greater emphasis on testing whether PRRT2-dependent regulation of slow inactivation extends across additional Nav isoforms.

      We have now performed new heterologous expression experiments to test whether PRRT2 modulates Nav1.6 slow inactivation. Consistent with our findings for other Nav isoforms, PRRT2 significantly enhances the slow inactivation of Nav1.6. We have incorporated these data into the revised Results and Figures, please refer to Page 8, Lines 211-215; Figures 4E and J.

      (2) To further demonstrate the physiological impact of enhanced slow inactivation, the authors should consider a simple experiment in the stable cell line experiments (Figure 1) to test pulse frequency dependence of peak Na current. One would predict that PRRT2 expression will potentiate 'run down' of the channels, and this finding would be complementary to the biophysical data.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we performed a pulse-train protocol in the stable Nav1.2 cell line and quantified the use-dependent attenuation (“run-down”) of peak sodium current across successive depolarizations (Figure 1-figure supplement 1C). Compared with control cells, PRRT2-expressing cells exhibited a larger decline in peak current during trains, indicating greater reduction in channel availability during repetitive depolarizations (Figure 1-figure supplement 1C). This pattern is consistent with our observations above showing that PRRT2 enhances Nav channel slow inactivation. These new data have been incorporated into the revised manuscript. Please refer to Page 5, Lines 133-140; Figure 1-figure supplement 1C.

      (3) The study of one K channel is limited, and the conclusion from these experiments represents an over-interpretation. I suggest removing these data unless many more K channels (ideally with measurable proxies for slow inactivation) were tested. These data do not contribute much to the story.

      We agree with the reviewer’s assessment. To avoid over-interpretation and to maintain focus on PRRT2-dependent regulation of Nav channel slow inactivation, we have removed the potassium channel dataset and the associated conclusions from the revised manuscript.

      (4) In Figure 2, the authors should confirm that protein is indeed expressed in cells expressing each truncated PRRT2 construct. Absent expression should be ruled out as an explanation for the enhancement of slow inactivation.

      We thank the reviewer’s concern regarding expression of the truncated PRRT2 constructs in the Nav1.2 stable cell line, particularly PRRT2(1-266), which shows little effect on slow inactivation of Nav1.2 channels. In the revised manuscript, we conducted western blot to verify expression of the PRRT2(1-266)-HA construct in the Nav1.2 stable cell line. We have added these results to the revised manuscript, please refer to Page 6, Lines 171-173; Figure 2-figure supplement 1A and B.

      Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is primarily expressed in the nervous system and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 directly interacts with Nav1.2 and Nav1.6, modulating channel properties and neuronal excitability.

      In this study, Lu et al. reported that PRRT2 is a physiological regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect can be replicated by the C-terminal region (256-346) of PRRT2, and is highly conserved across species from zebrafish, mouse, to human PRRT2. TRARG1 and TMEM233, the other two DspB family members, showed similar effects on Nav1.2 slow inactivation. Co-IP data confirms the interaction between Nav channels and PRRT2. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared to WT mice.

      Strengths:

      (1) This study is well designed, and data support the conclusion that PRRT2 is a potent regulator of slow inactivation of Nav channels.

      (2) This study reveals similar effects on Nav1.2 slow inactivation by PRRT2, TMEM233, and TRARG1, indicating a common regulation of Nav channels by DspB family members (Supplemental Figure 2). A recent study has shown that TMEM233 is essential for ExTxA (a plant toxin)-mediated inhibition on fast inactivation of Nav channels; and PRRT2 and TRARG1 could replicate this effect (Jami S, et al. Nat Commun 2023). It is possible that all three DspB members regulate Nav channel properties through the same mechanism, and exploring molecules that target PRRT2/TRARG1/TMEM233 might be a novel strategy for developing new treatments of DspB-related neurological diseases.

      We thank the reviewer for careful evaluation and insightful suggestions.

      Weaknesses:

      (1) Previously, the authors have reported that PRRT2 reduces Nav1.2 current density and alters biophysical properties of both Nav1.2 and Nav1.6 channels, including enhanced steady-state inactivation, slower recovery, and stronger use-dependent inhibition (Lu B, et al. Cell Rep 2021, Fig 3 & S5). All those changes are expected to alter neuronal excitability and should be discussed.

      We thank the reviewer for this suggestion. Although the present study focuses on PRRT2-dependent regulation of slow inactivation, we agree that PRRT2 may influence excitability through additional Nav-dependent mechanisms, including reduced current density and shifts in the voltage dependence of channel inactivation (Fruscione et al., 2018; Lu et al., 2021; Valente et al., 2023). Notably, because PRRT2 facilitates entry of Nav channels into slow-inactivated states both from closed states and from open states during prolonged depolarization, some of these previously reported effects may partly reflect enhanced slow inactivation and the resulting reduction in Nav channel availability. We have expanded the Discussion to integrate these prior findings and to clarify that these additional PRRT2-dependent effects may converge to shape neuronal excitability. Please refer to Page 16, Lines 445-452.

      (2) In this study, the fast inactivation kinetics was examined by a single stimulus at 0 mV, which may not be sufficient for the conclusion. Inactivation kinetics at more voltage potentials should be added.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we expanded our analysis of Nav1.2 fast-inactivation kinetics to include a range of test potentials (-20, -10, 0, +10, +20 and +30 mV) in the presence and absence of PRRT2. These experiments showed that PRRT2 expression did not significantly affect Nav1.2 fast-inactivation kinetics under these conditions. We have incorporated these new results into the revised manuscript. Please refer to Page 4, Lines 100-103; Figure 1C.

      (3) It is a little surprising that there is no difference in Nav1.2 current density in axon-blebs between WT and Prrt2-mutant mice (Figure 7B). PRRT2 significantly shifts steady-state slow inactivation curve to hyperpolarizing direction, at -70 mV, nearly 70% of Nav1.2 channels are inactivated by slow inactivation in cells expressing PRRT2 when compared to less than 10% in cells expressing GFP (Figure supplement 1B); with a holding potential of -70 mV, I would expect that most of Nav channels are inactivated in axon-blebs from WT mice but not in axon-blebs from Prrt2-mutant mice, and therefore sodium current density should be different in Figure 7B, which was not. Any explanation?

      We thank the reviewer for raising this point. In our axonal bleb recordings, although the holding potential was -70 mV, sodium current density was measured after a hyperpolarizing pre-pulse to -110 mV, which was applied before the test depolarization to relieve inactivation as much as possible (as described in the Methods). Therefore, the current density measurement in Figure 7B reflects the available current after this recovery step, rather than the steady-state availability at -70 mV. The lack of a difference in Figure 7B does not contradict the PRRT2-dependent shift in steady-state slow inactivation. In the revised manuscript, we have clarified this point explicitly in the Results and figure legend to avoid confusion. Please refer to Page 10, Lines 294-295.

      (4) Besides Nav channels, PRRT2 has been shown to act on Cav2.1 channels as well as molecules involved in neurotransmitter release, which may also contribute to abnormal neuronal activity in Prrt2-mutant mice. These should be mentioned when discussing PRRT2's role in neuronal resilience.

      We thank the reviewer for this suggestion. In addition to the Nav-dependent mechanisms, previous studies have shown that PRRT2 also regulates synaptic vesicle cycling (Valente et al., 2016; Coleman et al., 2018; Tan et al., 2018) and presynaptic surface expression of Cav2.1 channels (Ferrante et al., 2021). These effects are also expected to influence neurotransmitter release and, consequently, neuronal and network excitability. In the revised manuscript, we have expanded the Discussion to acknowledge that these additional PRRT2-dependent mechanisms may also contribute to cortical resilience. Please refer to Page 16, Lines 452-457.

      Reviewer #3 (Public review):

      This paper reveals that the neuronal protein PRRT2, previously known for its association with paroxysmal dyskinesia and infantile seizures, modulates the slow inactivation of voltage-gated sodium ion (Nav) channels, a gating process that limits excitability during prolonged activity. Using electrophysiology, molecular biology, and mouse models, the authors show that PRRT2 accelerates entry of Nav channels into the slow-inactivated state and slows their recovery, effectively dampening excessive excitability. The effect seems evolutionarily conserved, requires the C-terminal region of PRRT2, and is recapitulated in cortical neurons, where PRRT2 deficiency leads to hyper-responsiveness and reduced cortical resilience in vivo. These findings extend the functional repertoire of PRRT2, identifying it as a physiological brake on neuronal excitability. The work provides a mechanistic link between PRRT2 mutations and episodic neurological phenotypes.

      We thank the reviewer for this positive evaluation of our work and for the constructive comments.

      Comments:

      (1) The precise structural interface and the molecular basis of gating modulation remain inferred rather than demonstrated.

      We thank the reviewer for this comment. To avoid over-interpretation, we have removed the AlphaFold-based interaction prediction from the revised manuscript. We have also expanded the Limitations section to emphasize that direct structural and biochemical mapping of the PRRT2-Nav channel interface—through approaches such as targeted mutagenesis, crosslinking, and structural determination—will be required to define the binding interface and establish the molecular basis of gating modulation. Please refer to Page 16, Lines 465-468.

      (2) The in vivo phenotype reflects a complex circuit outcome and does not isolate slow-inactivation defects per se.

      We agree with the reviewer. Impaired slow inactivation in Prrt2-mutant mice is one plausible contributor to reduced cortical resilience. PRRT2 has also been reported to regulate surface exposure of Nav and Cav2.1 channels (Ferrante et al., 2021), as well as neuronal synaptic vesicle cycling (Valente et al., 2016; Coleman et al., 2018; Tan et al., 2018). Each of these PRRT2-associated processes could influence cortical excitability in vivo. We have therefore expanded the Discussion to clarify that the cortical phenotype likely reflects the combined contribution of multiple PRRT2-dependent mechanisms, rather than an isolated defect in slow inactivation alone. Please refer to Page 16, Lines 446-458.

      (3) Expression of PRRT2 in muscle or heart is low, so the cross-isoform claims are likely of limited physiological significance.

      We thank the review for this comment regarding physiological relevance. In the revised manuscript, we clarify that the cross-isoform analysis was intended to assess mechanistic generality at the channel level, rather than to imply equivalent physiological relevance across tissues. The functional consequence of PRRT2 depend on the Nav isoform composition and cellular context of each tissue. We also note that the broad isoform activity of the PRRT2 should be considered in any future attempt to manipulate PRRT2 function therapeutically. Please refer to Page 14 and 15, Lines 414-416; Lines 429-430.

      (4) The mechanistic separation between the trafficking effect of PRRT2 and its gating effects is not clearly resolved.

      We thank the reviewer’s concern regarding the possible contribution of trafficking effects to PRRT2-dependent regulation of Nav channel slow inactivation. Previous studies in heterologous overexpression systems have shown that PRRT2 can influence Nav channel trafficking and surface expression, raising the possibility that the observed effects on slow inactivation regulation might be secondary to altered channel abundance or localization. However, slow inactivation develops on a timescale of tens of milliseconds to seconds, whereas detectable changes in Nav channel trafficking and surface abundance generally occur over much longer intervals (minutes to hours) (Freal et al., 2023; Higerd-Rusli et al., 2023). These distinct temporal profiles argue against trafficking as the primary basis for the effects of PRRT2 on Nav channel slow inactivation described here, although direct quantification of dynamic changes in Nav channel surface expression will be required to fully exclude such a contribution (Liu et al., 2022; Tyagi et al., 2025). We have incorporated this point into the Discussion section. Please refer to Pages 13, Lines 378-388.

      (5) Additional studies with Nav1.6 should be carried out.

      We thank the reviewer for this suggestion. We have performed experiments to directly examine the effects of PRRT2 on Nav1.6 slow inactivation and incorporated these new data into the revised Results and figures, please refer to Page 8, Lines 211-215; Figures 4E and J.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for future experiments (not for this paper)

      (1) Exploit the lower protein expression in V5-PRRT2 mice to examine the effects of a hypomorphic allele.

      We thank the reviewer for this insightful suggestion. We note that the V5 epitope knock-in reduced PRRT2 protein expression, which may functionally resemble a hypomorphic allele. Accordingly, in addition to its utility for biochemical experiments (e.g., co-immunoprecipitation), this line could serve as a genetic tool to interrogate PRRT2 dose-dependent effects in vivo. We have added this point to the revised manuscript, please refer to Page 9, Lines 265-267.

      (2) Examine disease-causing PRRT2 mutations.

      We thank the reviewer for this constructive suggestion. Testing disease-associated PRRT2 variants for their ability to regulate Nav channel slow inactivation would be an important next step to strengthen the disease relevance of the mechanism proposed here. Moreover, identifying missense variants that selectively disrupt slow-inactivation regulation could help pinpoint residues that are critical for PRRT2-Nav functional coupling and thereby inform future structure-function studies. We plan to pursue this direction in follow-up work.

      (3) Investigate spreading depolarization in PRRT2-deficient mice.

      We thank the reviewer for this suggestion. Although we have shown that PRRT2 deficiency facilitates spreading depolarization in the cerebellum, whether PRRT2 exerts similar control over spreading depolarization susceptibility in the cerebral cortex remains to be determined. We plan to address this in an independent study and to test how cortical spreading depolarization relates to other PRRT2-associated neurological disorders.

      Reviewer #2 (Recommendations for the authors):

      This study is, in general, well executed, and the manuscript is well written. However, I do have some questions.

      (1) The authors' previous works have shown that PRRT2 regulates both Nav1.2 and Nav1.6, considering the wide expression Nav1.6 in CNS and its role in neuronal activity, what makes the authors not include Nav1.6 in this study?

      We thank the reviewer for raising this question. In our previous work, PRRT2 produced broadly similar effects on Nav1.2 and Nav1.6. Therefore, in the initial version of this study, we focused primarily on Nav1.2 as a representative neuronal Nav channel isoform and placed greater emphasis on testing whether PRRT2-dependent regulation of slow inactivation extends across additional Nav isoforms. In response to reviewers’ concern, we have now performed new experiments to directly examine the effect of PRRT2 on Nav1.6 slow inactivation. These results have been incorporated into the revised manuscript. Please refer to Page 8, Lines 211-215; Figures 4E and J.

      (2) Please explain why you chose 0 mV rather than -70 mV (closer to membrane potential) in the slow inactivation protocol.

      We thank the reviewer for raising this question. Nav channels can enter into slow inactivation from both resting/closed states and activated/open states. In our steady-state slow-inactivation assays, we found that PRRT2 enhances Nav1.2 slow inactivation under both conditions (Figure 1-figure supplement 1A and B). In whole-cell recordings, Nav1.2 channels typically begin to activate at command voltages more depolarized than approximately -60 mV. Accordingly, a conditioning voltage of -70 mV predominantly probes entry into slow inactivation from closed states, whereas 0 mV drives channel activation and more effectively induces slow inactivation. We therefore chose 0 mV as the primary conditioning potential because it is widely used in conventional slow inactivation protocols and induces slow inactivation more robustly than conditioning voltages at -70 mV. We have added this explanation in Methods section of revised manuscript, please refer to Page 20, Lines 569-571.

      (3) The authors mentioned that the insertion of V5 markedly reduced the PRRT2 protein level; thus, Prrt2-V5 knock-in mice could be considered as PRRT2 knock-down mice. Is there any noticeable difference in phenotype between Prrt2-V5 knock-in mouse and Prrt2-mutant mouse? In other words, is PRRT2 knockdown sufficient to affect neuronal excitability, or is a complete PRRT2 ablation required?

      We thank the reviewer for raising this concern regarding the functional consequences of reduced PRRT2 expression in the Prrt2-V5 knock-in mice. Given that PRRT2 protein levels are markedly reduced in this line, and that cerebellar stimulation-induced dystonia is a characteristic phenotype of PRRT2 deficiency, we tested whether Prrt2-V5 knock-in mice also exhibit this phenotype. We found that electrical stimulation of the cerebellar cortex induced dystonia-like attacks in a subset of Prrt2-V5 knock-in mice. These dystonic behaviors resembled those previously observed in Prrt2-mutant mice, whereas no such behaviors were induced in wild-type mice (Figure 6-figure supplement 1). These findings indicate that a substantial reduction of PRRT2 expression (approximately 80%) is sufficient to impair neuronal function and elicit a disease-relevant phenotype in a subset of animals, supporting the interpretation that the V5 knock-in allele is hypomorphic. We have incorporated these results into the revised manuscript, please refer to Page 9, Lines 265-267; Figure 6-figure supplement 1.

      (4) In Discussion (Page 13, lines 358-361), the authors mentioned a putative interaction between PRRT2 and the Nav channel by modeling, while there is no related data. Please either add modeling data or remove those sentences.

      We thank the reviewer for this suggestion. To avoid over-interpretation, we have removed the statements regarding the AlphaFold-based interaction model from the revised manuscript. We agree that the interaction interface remains to be demonstrated experimentally, and we now discuss this point in the Limitations section. Please refer to Page 16, Lines 465-468.

      (5) Typo: Page 14, line 399, "TMEM232" should be "TMEM233".

      We thank the reviewer for pointing out this typo. We have corrected it in the revised manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Mechanistic depth: While the functional data show altered slow-inactivation kinetics, the mechanistic explanation remains superficial. The AlphaFold-based prediction of PRRT2 interaction with DIV-S3 is speculative. The authors should clarify their illustrative rather than evidential intent and avoid over-interpretation.

      We thank the reviewer for this comment. To avoid over-interpretation, we have removed the AlphaFold-based interaction prediction from the revised manuscript. We have also expanded the Limitations section to emphasize that direct structural and biochemical mapping of the PRRT2-Nav interface, including targeted mutagenesis, crosslinking, and structural determination, will be necessary to elucidate the molecular basis of this interaction and its effect on channel gating. Please refer to Page 16, Lines 465-468.

      (2) Separation of trafficking vs. gating effects: Previous studies showed PRRT2 influences Nav trafficking and surface expression. Here, surface expression changes are not systematically quantified. Such an analysis would strengthen the argument that gating effects are not secondary to altered channel abundance or localization.

      We thank the reviewer’s concern regarding the possible contribution of trafficking effects to PRRT2-dependent regulation of Nav channel slow inactivation. We agree that direct analysis of Nav channel surface localization during prolonged depolarization and hyperpolarization would provide stronger evidence to distinguish gating effects from trafficking-dependent mechanisms. However, such experiments are technically challenging in this context: conventional surface biotinylation assays do not provide the temporal resolution required for these rapid protocols, and live-cell imaging approaches to monitor dynamic changes in Nav channel surface expression during slow-inactivation paradigms have not yet been established in our laboratory.

      Although PRRT2 has been reported to regulate Nav channel surface expression in heterologous systems, we consider it unlikely that trafficking is the major determinant of the slow-inactivation effects described here. Slow-inactivation develops on a timescale ranging from tens of milliseconds to seconds, whereas detectable changes in Nav channel trafficking and surface abundance generally occur over much longer timescales (minutes to hours) (Freal et al., 2023; Higerd-Rusli et al., 2023). We have expanded the Discussion in a revised manuscript. Please refer to Pages 13, Lines 378-388.

      (3) Isoform generalization: Data on other Nav channel subtypes are presented as evidence of a conserved mechanism. However, given tissue-specific expression of PRRT2, these findings may be of limited in vivo relevance. At the very least, additional studies with Nav1.6 should be carried out.

      We thank the review for this suggestion. In response, we conducted new experiments to examine the effect of PRRT2 on Nav1.6 slow inactivation. These results show that PRRT2 promotes entry of Nav1.6 channels into slow-inactivated states and delays their recovery, consistent with its effects on the other Nav isoforms examined in this study. We have incorporated these new data into the revised manuscript. Please refer to Page 8, Lines 211-215; Figures 4E and J.

      Furthermore, we clarify that the cross-isoform analysis was intended to assess mechanistic generality at the channel level, rather than to imply equivalent physiological relevance across tissues. The functional consequence of PRRT2 depend on the Nav isoform composition and cellular context of each tissue. We also note that the broad isoform activity of the PRRT2 should be considered in any future attempt to manipulate PRRT2 function therapeutically. Pages 14 and 15, Lines 414-416 and 429-430.

      (4) In vivo functional link: The EEG after-discharge threshold assay suggests decreased cortical resilience, but causality between slow-inactivation impairment and hyperexcitability remains indirect. Complementary in vivo recordings would strengthen the physiological link.

      We thank the reviewer for this helpful suggestion. To further link impaired slow-inactivation to the hyperexcitability, we applied a repetitive stimulation protocol in corpus callosum slices, a white-matter region of brain enriched in both PRRT2 and Nav channels. During high-frequency stimulation (e.g., 20 Hz), the amplitude of the compound action potential progressively decreased over the course of the stimulus train. This phenomenon, often referred to as adaptation, reflects activity-dependent reduction in Nav channel availability (Fleidervish et al., 1996; Mickus et al., 1999; Kim et al., 2012). Compared with wild-type mice, Prrt2-mutant mice exhibited less adaptation during high-frequency stimulation, consistent with impaired slow inactivation during repetitive activity, which may contribute to hyperexcitability (Figure 7-figure supplement 2). We have added these results to the revised manuscript. Please refer to Pages 11, Lines 311-322; Figure 7-figure supplement 2.

      (5) Structural interaction: It remains unclear whether PRRT2 binds the α-subunit directly or through accessory proteins. Crosslinking or detergent-solubilization controls of different stringencies could clarify this.

      We thank the reviewer for raising this important issue. We agree that our co-immunoprecipitation data do not distinguish whether PRRT2 associates with the Nav channel α-subunit directly or through other components of the protein complex. To avoid over-interpretation, we have revised the relevant text in the manuscript to remove any implication of direct binding and now describe the result as an association between PRRT2 and Nav channels.

      We have also expanded the Limitations section to note that additional experiments, such as crosslinking and structural studies, will be required to define the interaction interface between PRRT2 and Nav channels. Please refer to Page 16, Lines 465-468.

      (6) Comparisons to other regulators: The paper positions PRRT2 as distinct from FHFs and β-subunits. The data support this, but the discussion could more critically assess whether PRRT2 acts by stabilizing a pore-based inactivated conformation, as suggested for other slow-inactivation modulators.

      We thank the reviewer for this insightful suggestion. At present, relatively few modulators have been characterized in detail with respect to their effects on Nav channel slow-inactivation kinetics. Moreover, even for compounds such as lacosamide, which has been proposed to act as a slow-inactivation modulator, the underlying mechanism remains under debate (Errington et al., 2008; Jo and Bean, 2017). Therefore, in the revised manuscript, we discussed the possible mechanism of PRRT2 in the context of current models of Nav channel slow inactivation.

      Previous studies suggest that entry into the slow-inactivated state involves at least two coupled processes: conformational changes in the voltage-sensing domains and structural rearrangements in the pore region, including the selectivity filter and intracellular activation gate (Catterall et al., 2024; Silva, 2014). During prolonged depolarization, voltage sensors become stabilized in the up-state, while the pore undergoes progressive rearrangements associated with slow inactivation (Balser et al., 1996; Vilin et al., 1999). Thus, mechanisms that further stabilize voltage sensors in the up-state and/or facilitate pore-based inactivated conformations could enhance slow inactivation.

      Within this framework, PRRT2 may enhance slow inactivation by facilitating one or both of these processes, although direct evidence is still lacking. We have incorporated this discussion in relative section of revised manuscript. Please refer to Page 14, Lines 389-404.

      Response references:

      Jo S, Bean BP. Lacosamide Inhibition of Nav1.7 Voltage-Gated Sodium Channels: Slow Binding to Fast-Inactivated States. Mol Pharmacol. 2017 Apr;91(4):277-286.

      Errington AC, Stöhr T, Heers C, Lees G. The investigational anticonvulsant lacosamide selectively enhances slow inactivation of voltage-gated sodium channels. Mol Pharmacol. 2008 Jan;73(1):157-69.

      (7) Behavioral/clinical link: Given the strong human genetics background of PRRT2 disorders, a brief analysis or reference to electrophysiological phenotypes in patient neurons would contextualize the cortical findings.

      We thank the reviewer for this suggestion. Previous studies showed that iPSC-derived excitatory neurons from a patient carrying a homozygous PRRT2 mutation exhibited increased sodium currents and neuronal hyperexcitability (Fruscione et al., 2018). Given that slow inactivation regulates Nav channel availability and thereby influences neuronal excitability, these electrophysiological abnormalities in patient-derived neurons may, at least in part, reflect impaired PRRT2-dependent regulation of Nav channel slow inactivation. We have added this point to the relative section of the revised manuscript. Please refer to Pages 15, Lines 432-437.

      Minor comments

      (1) Figures should include statistical sample sizes (n) and ideally overlay data points rather than only means {plus minus} SEM.

      We thank the reviewer for this suggestion. In the revised manuscript, we present both individual data points and mean ± SEM in the column graphs. For the line graphs, individual data points were not overlaid because of space and readability constraints, and these panels therefore display mean ± SEM only. Sample sizes for each group are provided in the corresponding figure legends.

      (2) The AlphaFold model should be provided as a supplementary figure with confidence scores indicated.

      We thank the reviewer for this suggestion. However, because the predicted Nav1.2-PRRT2 interaction interface has not yet been experimentally validated in our study, we chose to remove the AlphaFold-based model from the revised manuscript to avoid over-interpretation.

      (3) Clarify whether TTX sensitivity was verified in the axonal bleb preparation.

      We thank the reviewer for raising this point. We verified the identity of the sodium currents in the axonal bleb preparation by their sensitivity to TTX, and this information has now been added to Figure 7A in the revised manuscript. Please refer to Page 10, Line 290; Figure 7A.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In this valuable study, the authors developed long-term imaging tools to simultaneously monitor the temporal and spatial dynamics of excitatory and inhibitory synapses and reported that excitatory and inhibitory synapses need to develop synergistically during synaptogenesis to maintain balance. While the analysis and quantification of the imaging data are incomplete, there is convincing evidence that the developed tools are feasible. If these tools can function stably in vivo, their applications will be much broader.

      We have completely overhauled our analysis and quantification methods and generated custom-made drift correction and tracking pipelines. Also, we have tested these tools ex vivo.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      By imaging the dynamics of synaptic proteins in cultured neurons, this study presents significant findings regarding the dynamics of excitatory and inhibitory synaptic proteins during development. The evidence shows that the ratios of excitatory and inhibitory synaptic proteins are stable during synapse development. This discovery advances our understanding of the complex mechanisms governing synapse formation. The strength of the evidence is robust, as it is supported by a combination of biological assays and endogenous labeling.

      Strengths:

      This research sheds light on the dynamics of the excitatory and inhibitory synapses during development. It is crucial to understand that while excitatory synapses and inhibitory synapses are developed independently, the ratio of their number is relatively stable during development, maintaining a stable excitatory/inhibitory ratio.

      Important findings and implications in the research include:

      (1) Persistent Synapse Dynamics: Excitatory and inhibitory synapses remain highly dynamic even in mature neurons (DIV12-14), challenging the dogma that synaptic structures are stable after the synaptogenesis stage.

      (2) Maintained E/I Balance: Despite ongoing synapse turnover (formation/elimination) and presynaptic terminal reduction, the overall density and ratio of excitatory-to-inhibitory synapses remain relatively stable during circuit maturation (Figure 7).

      (3) Developmental Shifts: While presynaptic compartments decrease over time, postsynaptic sites increase, suggesting independent regulation of pre- and postsynaptic elements within a stable E/I framework.

      We thank the Reviewer for their positive feedback and careful review of our study.

      Weaknesses:

      This study focuses on specific synaptic proteins within synapses, which may not fully represent the dynamics of other synaptic machinery; also, whether similar observations exist in vivo is still unknown. Further research is needed to explore the implications of these findings in more complex neuronal environments.

      We also thank the Reviewer for their insights and suggestions. We have added discussion of this important point to the Discussion section. Furthermore, we have tested the applicability of our tools ex vivo (new Figures 1, 4, and 6). While using these tools in vivo for live imaging is the eventual goal, we started in a reduced culture system given the relative simplicity. Our current study now provides a framework for future experiments applying these approaches in more complex in vivo systems.

      Reviewer #2 (Public review):

      Summary:

      The Garbett et al. identified a critical need to begin to understand the interplay between the assembly, maturation, and elimination of excitatory and inhibitory synapses. They also detail the lack of reliable tools to address this gap in knowledge. Here, the authors developed synaptic reporters expressed by lentiviruses (mClover3-Homer1c, HaloTag-Syb2, and tdTomatoGephyrin). They combined these reporters with resonance scanning confocal imaging to measure synapses over a 15-hour period during neuron development and in mature neurons in primary hippocampal cultures. Using these reporters in the same neuron, the authors compared the ratios of postsynaptic excitatory and inhibitory specializations that co-localize with presynaptic terminals during development and in mature neurons and found that they are stable across time points. Finally, the authors developed CRISPR/Cas9 tools (TKIT) to knock-in endogenous fluorescent tags (GFP/tdTomato-Gephyrin) or epitope tags (HA-Bassoon and HAHomer1) to begin to study synapse dynamics using endogenous proteins. I believe this paper highlights an important gap in knowledge and begins to offer methodologies to determine the dynamic coordination between excitatory and inhibitory synapses.

      Strengths:

      (1) The experiments are well-designed and carefully controlled.

      (2) The authors carefully validated the reporter and TKIT constructs.

      (3) The authors provide strong proof-of-principle for the use of the reporter constructs to track synapse formation, maintenance, and elimination over a 15-hour period.

      (4) Ingenious use of technologies (reporters, TKIT, and resonance scanning confocal microscopy) to develop a platform for future studies of synapse dynamics.

      (5) Strong evidence supporting that the ratio of excitatory and inhibitory synapses (those that oppose syb2) stays constant through development.

      We thank the Reviewer for their positive assessment of our study.

      Weaknesses:

      Overall, this is a well-executed study that develops tools to simultaneously image excitatory and inhibitory synapse dynamics and represents an important first step to address the fundamental question regarding the coordination between these two types of synapses.

      Minor weaknesses of the manuscript include:

      (1) The lack of a characterization of endogenous Homer1-positive excitatory synapses using TKIT.

      We attempted to perform live imaging of endogenous Homer1-positive synapses using the TKIT approach by tagging endogenous Homer1 with mClover3 but encountered low signal/noise while live imaging. This prompted us to focus our current study on live imaging endogenous Gephyrin. Future studies using more robust tags (e.g. StayGold, HaloTag) for TKIT tagging of endogenous Homer1 will likely help circumvent this issue.

      (2) Discussion about other approaches to study excitatory and inhibitory synapses using endogenous proteins (e.g., intrabodies - FingR or nanobodies) should be included.

      This important point was also raised by other Reviewers. We have now significantly expanded the Discussion section, including discussion of this point.

      (3) The activity state of a neuron and/or a synapse might alter the dynamic properties (formation, maintenance, and/or elimination). A discussion on whether the overexpression of Homer1 and/or gephyrin might alter synapse/neuron activity would provide greater interpretability of the results. A discussion of the potential limitations and benefits of the reporter and TKIT approaches would be beneficial.

      We agree and have added discussion of these points to the Discussion section.

      (4) A description and interpretation of the computational approach to calculate particle tracking would be helpful. I found that particle tracking figures, while elegant, are difficult to interpret.

      As discussed in more detail below, we have generated drift correction and particle tracking approaches for the revised manuscript. We now elaborate on these new approaches in the paper.

      We thank the Reviewer again for their very helpful input and suggestions.

      Reviewer #3 (Public review):

      In the present study, the authors describe the development of new tools and imaging strategies to assess the concomitant development of excitatory and inhibitory synapses in dissociated neuron cultures. To this end, they generate fluorescently tagged constructs of excitatory and inhibitory synapse marker proteins using either conventional overexpression or CRISPR-based strategies. They then image these marker proteins over a timespan of 15 hours to assess synaptic dynamics at different developmental timepoints. Based on their data, they conclude that excitatory and inhibitory synapse development occur in concert to maintain a functional balance despite individual synapse turnover.

      Overall, this study addresses an interesting question, i.e., the interplay between the development of excitatory and inhibitory synapses, which has important implications, particularly for neurodevelopmental disorders in which the balance of excitation and inhibition is disrupted. The experiments are technically solid and well-executed, and the individual images are highly compelling.

      We thank the Reviewer for their positive assessment of our study.

      However, a number of aspects remain to be addressed in order for the study to support the claims made by the authors. First, the novelty aspect of the development of the fluorescently tagged synaptic proteins is unclear, since reporters of this nature are in routine use in many labs. Second, the analysis of the acquired images often seems incomplete, with only example images but no quantification shown, or the distinction between spatial and temporal dynamics appearing unclear. Third, given this incomplete analysis, the interpretations of the authors are not always convincingly supported by the data presented. In conclusion, substantial improvements are required to render the main messages of the study clear and compelling.

      We agree and have incorporated all of the Reviewer’s suggestions in the revised manuscript (please see below).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This is an interesting study. This reviewer has the following questions/comments for the authors:

      (1) Please provide evidence that the gRNAs targeting each gene of synaptic protein have no offtarget effects.

      We now include analysis of off-target effects for the TKIT tools (new Figure S6).

      (2) While structural E/I balance is shown, functional electrophysiological validation (e.g., mEPSC/mIPSC ratios) is absent. It is interesting to know whether the balanced functional structural changes translate to functional?

      We thank the Reviewer for this insightful suggestion and now include these recordings in the revised paper (new Figure 8).

      (3) In lines 217-218, please define thresholds for "stable" vs. "dynamic" puncta (e.g., temporal and spatial criteria).

      We more clearly define our categorization parameters (e.g. new Figure 2).

      (4) In Figure 5B: The low co-localization between endogenously tagged Bassoon and antibodystained Bassoon is likely due to the low TKIT efficiency. Quite a few HA-tagged Basson signals are insensitive to Basson-antibody. The authors are suggested to explain those.

      We thank the Reviewer for identifying this and add discussion to the Results section.

      (5) For the data analysis. If each n represents an independent neuronal culture, should the authors are suggested to provide the number of neurons/dendrites analyzed for each independent culture?

      We have added these important details to the manuscript.

      (6) Regarding the title, the author used the term "coordinated dynamics". This reviewer finds it is a bit over-claim because the stable ratios of the number of excitatory synapses and inhibitory synapses are likely an association, not actively "coordinated". I suggest that the authors rephrase this.

      We agree that we cannot argue that excitatory and inhibitory synapses are causally coordinated in our current study. Their levels are likely associated by either association or direct coupling, which we now discuss further in the first paragraph of the Discussion. We have rephrased the title accordingly.

      Reviewer #2 (Recommendations for the authors):

      I have only minor suggestions that I think will improve the manuscript:

      (1) Please define Syn1/2 on line 129.

      We have defined this in the revised paper.

      (2) For Figures 2B, C, and 4B, C: are the puncta in panel C from the dendrites in panels B? If so, it would be helpful to identify the ROIs selected in panels C.

      We now include this in new Figure 2.

      (3) For the particle tracking figures, while the ability to track all synaptic puncta is very impressive, it is sometimes difficult to clearly track the lifespan of a synaptic puncta from the current figures. I believe that it would be helpful if the authors selected specific examples of synapses formed, maintained, and eliminated.

      We agree and now include more examples.

      (4) I believe that more detail about the computational approach and analysis for the particle tracking (Figs 2E and 4E) would help the interpretability of the figure.

      This important point was also raised by the other Reviewers. We generated custom tools during the revision that significantly expand the capabilities of our tracking approaches and more clearly describe them in the revised manuscript.

      (5) Similar to the rigorous gephyrin TKIT analysis (Fig. 6), did the authors perform a similar analysis for Homer1c TKIT? This might be valuable to confirm that overexpression of the Homer1 reporter does not indirectly alter synapse dynamics.

      We attempted to perform live imaging of mClover3 TKIT-tagged endogenous Homer1 but encountered low signal/noise with live imaging. We now add discussion that optimization of more robust tags (e.g. StayGold, HaloTag) will likely be necessary for live imaging of different target proteins.

      (6) The tools developed by Garbett et al. have the potential to be broadly utilized in the field to provide new insight into the coordination of excitatory and inhibitory synapses. It would thus be helpful for the authors to include a discussion about the strengths and limitations of the reporter and TKIT methods relative to other approaches used to live image synapses (e.g., intrabodies (FingR and nanobodies)).

      We have now significantly expanded the Discussion to include these important points.

      (7) In the discussion, can the authors elaborate on whether it is experimentally feasible to apply their TKIT labeling of gephyrin and Homer1c in the same neuron to assess the endogenous excitatory and inhibitory synapse dynamics from the same neuron?

      We have added discussion of this point and also proof-of-concept data supporting tagging of two postsynaptic targets within the same neuron (new Figure S5D).

      Reviewer #3 (Recommendations for the authors):

      (1) While the new tools described in the current manuscript can undoubtedly be used for the described purposes, the novelty of these tools is unclear to me. Viral vectors expressing fluorescently tagged versions of Homer1, synaptobrevin, and gephyrin are commercially available, e.g., via Addgene, and they are in routine use in many labs. CRISPR-mediated strategies for this purpose have also been previously reported (e.g., Willems et al. 2020, PLOS Biology; Fang et al. 2021, eLife). It is not clear to me how the tools reported here present a significant improvement over existing resources, other than that they use different fluorescent tags. If this aspect is a central part of the current manuscript, it should be expanded on in the discussion, including a direct comparison with available tools to highlight the novel aspects.

      We agree and have significantly expanded the Discussion to include these important points. Also, rather than argue that our tools are superior to pre-existing approaches, we adjust the text to argue that our tools and analytical approaches have been designed and optimized for the purposes we apply them to.

      (2) In addition to generating new tagged constructs, the authors also state that they have developed new imaging and analysis strategies to facilitate long-term assessment of synaptic dynamics. However, in many figures, they present only sample images, with little quantification to allow assessment of the wider relevance of the imaged synapses. For example, in Figures 2C and 4C, they present one example each of, e.g., a stable, nascent, transient, or eliminated synapse. However, they do not provide any quantification on how frequently any of these events occur, or whether they can be reliably quantified at all. These quantifications (i.e., percentage of each event type across a large population of synapses) would be necessary and should be added to demonstrate that this tool can be used for more than single example images.

      We have generated custom-made drift correction and particle tracking approaches for the revised manuscript. Based on the reviewer’s suggestion, we have quantified the relative frequencies of stable, nascent, transient, and eliminated synapses (Fig 2B-G, Fig3A-F, Fig 5A-F, Fig 7B-C). These metrics greatly enhance the biological interpretation of our results. We have also added a supplemental movie with an example image with corresponding categorized tracks for each puncta type (Movie S3)

      (3) The authors do present an automated visual representation of spatial track length across the neuron, e.g., in Figure 2E and 4E, although this is also not quantified. Moreover, the track lengths appear surprisingly short, despite the authors' claims that their analyses 'highlight the dynamic nature of excitatory synapses over these timescales'. It is not clear to me whether these short tracks are more than just jitter, either in the synapses themselves or in the images due to technical limitations. E.g., in panel 2E, I see very few examples in which the track is not simply centered around one point, but actually expands over a distance. Quantification of the distance between start and end points of the tracks would be important to support the claim that these synapses are dynamic in terms of spatial translocation (if that is what the authors meant). Or if the 'dynamic nature' of the synapses referred to temporal dynamics, it is unclear to me how this information can be gained from the represented tracks.

      We thank the reviewer for these excellent points. To accurately access spatial motion, we drift-corrected our images with a custom correction algorithm to eliminate stage or microscope drift as a source of contaminating motion (See Methods, Movie S2), in addition to collecting time-lapse imaging with Nikon perfect focus. We noticed heterogeneity in our cultures such that some areas contained very mobile neurites, while other remained stationary (Fig. S1). We binned movies into either moving or still neurites and assessed spatial metrics as suggested (Fig. S1A). Consistent with our binning, puncta on moving neurites showed larger net displacement (distance between start and end points), but puncta on still neurites also showed ~1 µm net displacement (Fig. S1D). We also quantified puncta speed and found that puncta on moving neurites generally moved faster (Fig. S1C). We appreciate the reviewer’s insight that track length were surprisingly short, and after employing our drift correction and revised tracking methods, we now see substantially longer track lengths (Fig 2E, Fig 3C & F, Fig S2B & C). We additionally see a large fraction of tracks that persist throughout the imaging session (Fig 2E, Fig S2B & C).

      (4) In Figure 3, the authors now quantify track length, but in this case in the unit 'minutes', from which I would interpret that this is now meant to assess the temporal dynamics rather than the spatial dynamics. The lack of a clear distinction between spatial dynamics and temporal dynamics is very confusing to me, since these are entirely independent measures. 'Track length' to me indicates spatial dynamics, and I would expect the units to be a measure of distance. 'Track duration', which the authors also use in some places, but inconsistently as far as I can tell, makes sense to me for the assessment of temporal dynamics, with the units being a measure of time. I would strongly recommend being very clear about this distinction, since the current representation of the data is very difficult to follow and interpret.

      In addition to new spatial metrics, we have clarified in the text when we are referring to spatial dynamics (distance) versus temporal dynamics (time). As suggested, we use duration when referring to time, and speed or distance when referring to spatial metrics.

      (5) The images from the newly generated CRISPR-based tags in Figures 5-7 are striking and very compelling - these will be very useful tools. However, here too, it seems that the interpretation of the data does not really match the results. All quantification indicates that there is very little change in synapse density or other assessed parameters over the time course of the imaging, and yet the authors emphasize the dynamic nature of visualized synapses. More compelling quantification would be needed to support this claim.

      We have quantified spatial and temporal metrics for live neuron culture imaging for all tools developed including CRISPR-based tags (Figure 7).

      (6) The discussion is extremely short and provides almost no integration of the results of the study into the framework of existing knowledge. Instead, it focuses almost exclusively on unanswered questions and future perspectives, which are also important, but not helpful in interpreting the findings from the current study. The latter aspects should be added to provide essential context for the current findings.

      We agree and have added additional discussion of our current findings to help contextualize their significance.

      We thank the Reviewers again for their positive feedback and insightful input, which has undoubtedly strengthened our study.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates how human temporal voice areas (TVA) respond to vocalizations from nonhuman primates. Using functional MRI during a species-categorization task, the authors compare neural responses to calls from humans, chimpanzees, bonobos, and macaques while modeling both acoustic and phylogenetic factors. They find that bilateral anterior TVA regions respond more strongly to chimpanzee than to other nonhuman primate vocalizations, suggesting that these regions are sensitive not only to human voices but also to acoustically and evolutionarily related sounds.

      The work provides important comparative evidence for continuity in primate vocal communication and offers a strong empirical foundation for modeling how specific acoustic features drive TVA activity.

      Strengths:

      (1) Comparative scope: The inclusion of four primate species, including both great apes and monkeys, provides a rare and valuable cross-species perspective on voice processing.

      (2) Methodological rigor: Acoustic and phylogenetic distances are carefully quantified and incorporated into the analyses.

      (4) Neuroscientific significance: The finding of TVA sensitivity to chimpanzee calls supports the view that human voice-selective regions are evolutionarily tuned to certain acoustic features shared across primates.

      (4) Clear presentation: The study is well organized, the stimuli well controlled, and the imaging analyses transparent and replicable.

      (5) Theoretical contribution: The results advance understanding of the neural bases of voice perception and the evolutionary roots of voice sensitivity in the human brain.

      Weaknesses:

      (1) Acoustic-phylogenetic confound: The design does not fully disentangle acoustic similarity from phylogenetic proximity, as species co-vary along both dimensions. A promising way to address this would be to include an additional model focusing on the acoustic features that specifically differentiate bonobo from chimpanzee calls, which share equal phylogenetic distance to humans.

      (2) Selectivity vs. sensitivity: Without non-vocal control sounds, the study cannot determine whether TVA responses reflect true selectivity for primate vocalizations or general auditory sensitivity.

      (3) Task demands: The use of an active categorization task may engage additional cognitive processes beyond auditory perception; a passive listening condition would help clarify the contribution of attention and task performance.

      (4) Figures and presentation: Some results are partially redundant; keeping only the most representative model figure in the main text and moving others to the Supplementary Material would improve clarity.

      We thank the reviewer for contributing to the improvement of the present study and for the extremely constructive criticism. Concerning the identified weaknesses of our work, we provide here some general answers while the detailed review (below) addresses point-by-point the reviews in high detail.

      (1) We totally agree that acoustics and phylogeny cannot be disentangled in our study, which is a limitation. We now provide the suggested analysis on the acoustic specificities of chimpanzee and bonobo calls.

      (2) This point on selectivity vs. specificity is indeed crucial, and we now provide a more careful viewpoint and phrasing on this aspect, since our study can only provide partial arguments for this important distinction.

      (3) Task demand following species categorization might rightfully yield to the engagement of distinct brain network compared to merely listening to the stimuli. We discuss this aspect and put forward the argument that, while we cannot control for this aspect, our attentional control study performed by an independent sample, N=28 provides clear evidence that no species triggered an attention bias. In other words, task demand might play a role, but at least in the study we know that attentional resources were not biased towards one species in particular since no effects were observed.

      (4) We agree that results were not articulated in a clear fashion and that figures were redundant. We addressed this aspect and regrouped the figures where appropriate while we include the rest in the supplementary material now.

      Reviewer #2 (Public review):

      Summary:

      This study investigated how the human brain responds to vocalizations from multiple primate species, including humans, chimpanzees, bonobos, and rhesus macaques. The central finding - that subregions of the temporal voice areas (TVA), particularly in the bilateral anterior superior temporal gyrus, show enhanced responses to chimpanzee vocalizations - suggests a potential neural sensitivity to calls from phylogenetically close nonhuman primates.

      Strengths:

      The authors employed three analytical models to consistently demonstrate activation in the anterior superior temporal gyrus that is specific to chimpanzee calls. The methodology was logical and robust, and the results supporting these findings appear solid.

      Weaknesses:

      The interpretation of the findings in this paper regarding the evolutionary continuity of voice processing lacks sufficient evidence. A simple explanation is that the observed effects can be attributed to the similarity in low-level acoustic features, rather than effects specific to phylogenetically close species. The authors only tested vocalizations from three non-human primate species, other than humans. In this case, the species specificity of the effect does not fully represent the specificity of evolutionary relatedness.

      We want to thank the reviewer for the constructive criticism and for evaluating the manuscript.

      Concerning the principal weakness highlighted, we provide new analyses behavioral, acoustics, model-based fMRI that improve our understanding of the influence of both phylogeny and bioacoustics in our data. We argue that the explanation proposed by the reviewer cannot explain our results, as also observed in several other research from us and others. We discuss this aspect and emphasize that including stimuli from more species would greatly improve the understanding of phylogeny and bioacoustics in this context.

      Reviewer #3 (Public review):

      Summary:

      Ceravolo et al. employed functional magnetic resonance imaging (fMRI) to examine how the temporal voice areas (TVA) in the human brain respond to vocalizations from different nonhuman primate species. Their findings reveal that the human TVA is not only responsible for human vocalizations but also exhibits sensitivity to the vocalizations of other primates, particularly chimpanzee vocalizations sharing acoustic similarities with human voices, which offers compelling evidence for cross-species vocal processing in the human auditory system. Overall, the study presents intellectually stimulating hypotheses and demonstrates methodological originality. However, the current findings are not yet solid enough to fully support the proposed claims, and the presentation could be enhanced for clarity and impact.

      Strengths:

      The study presents intellectually stimulating hypotheses and demonstrates methodological originality.

      Weaknesses:

      (1) The analysis of the fMRI data does not account for the participants' behavioral performance, specifically their reaction times (RTs) during the species categorization task.

      (2) The figure organization/presentation requires significant revision to avoid confusion and redundancy.

      We thank the reviewer for evaluating our manuscript and for the constructive criticism as well as the many suggestions. Concerning the weaknesses of the study, we provide here some quick answers while more detailed responses can be found below.

      (1) We now include behavioral data analysis (accuracy data controlled for reaction times and acoustics of existing Model 3, using mixed-effects logistic regression) in addition to a new, 4th model for fMRI data. This 4th model was computed in a model-based fashion by modeling the probability of correct categorization within the TVA (fitted regression coefficients, per Participant, Species, Trial) and revealing the neural correlates of this modulator.

      (2) We totally agree that figure redundancy was a problem and we now reduced confusion by combining congruent aspects while pushing other results to the supplementary material.

      Recommendations for the authors:

      Reviewing Editor Comments:

      With additional analyses and discussions, the work has the potential to offer important insight into the evolutionary continuity of voice processing.

      We thank the Reviewing Editor for this additional motivation and for offering us the possibility to revise our manuscript. We will now provide our point-by-point reviewing, referring to manuscript modifications by section and/or line number(s). All modifications are also highlighted in light grey in the text.

      Reviewer #1 (Recommendations for the authors):

      The manuscript is clearly written and addresses an important comparative question about the specificity of human TVA responses. The acoustic analyses are well designed, and the imaging work is careful and thorough. However, several conceptual and methodological issues need clarification or tempering of claims, particularly regarding (i) the distinction between sensitivity and selectivity, (ii) the confounding of acoustic and phylogenetic factors, and (iii) the interpretation of "chimpanzee-specific" TVA activity.

      (1) Introduction

      Line 48: cite more recent infant EEG evidence for early voice sensitivity (Calce, Curr Biol).

      The reference and explanation were added, lines 46-48.

      Line 53: mention recent data on voice processing in marmosets (Jafari, Cell Rep; Dureux, Curr Biol).

      We added the references and the mention of these interesting studies on common marmosets, lines 53-54.

      Line 59: Fecteau et al. (2004) already explored cross-species selectivity; please integrate and discuss.

      We now mention here the work from Fecteau and colleagues and its relevance, see lines 57-59.

      Line 70: clarify that in [27] (Bodin et al., 2021) human TVA responded similarly to human nonverbal vocalizations and macaque coos, likely due to acoustic similarity.

      We added this important aspect, thank you for this precision. See lines 71-72.

      Clarify why an active species-categorization task was chosen instead of passive listening, which is standard in TVA research. Were participants familiarized with stimuli beforehand?

      We added a sentence on this aspect, but basically to summarize it here: we wanted to be able to test human recognition of nonhuman primate species’ calls. From the start, we wanted to test the frontal mechanisms related to decision-based processes of humans when categorizing non-human primate calls hence the 2023 article we published. See lines 75-77 and we also added information on familiarization to the stimuli in the Methods, lines 679-682.

      The 16 acoustic features mentioned should be briefly defined earlier, as they are central.

      We feel like describing 16 acoustic parameters in the introduction would be heavy on the reader, so we instead added a reference to the supplementary table (Table S1) in which these are named and described. See line 80.

      Explain why only chimpanzees and bonobos were selected among the great apes, and discuss the value of including both, given their equal phylogenetic proximity but largely dissimilar acoustics.

      The stimuli were obtained by Thibaud Gruber and his team and through collaborations with Katie Slocombe and Zanna Clay. Unfortunately, at the time we could only use chimpanzee and bonobo calls for the great apes. Therefore, it was mainly a material constraint rather than a deliberate choice to exclude other great apes. We now discuss this aspect and present the absence of other great apes as a limitation (lines 587-591).

      Rephrase references to "recruitment" of TVA - this term implies general activation, while the key question concerns selectivity (stronger responses to voices vs. non-vocal controls).

      We rephrased throughout the manuscript, thank you for this suggestion.

      The hypothesis section should more clearly separate the acoustic and phylogenetic predictions, and clarify which earlier data motivate each.

      We now explicitly categorize the hypotheses according to either Bioacoustics or Phylogeny to clarify. We also added references motivating each hypothesis. See lines 114-120.

      (2) Methods

      Clarify whether stimuli were RMS-normalized or otherwise balanced for energy (line 128).

      Sound pressure level was kept constant but the stimuli were not normalized, specifically to avoid a negative impact on their naturality. We added a sentence (lines 131-132) including a reference on this aspect.

      The task design could benefit from reporting accuracy in addition to reaction times for the 4AFC species classification task.

      We agree this aspect was missing. We now report accuracy data (controlled for reaction times and acoustics of Model 3) for the species categorization task (lines 147-165; Fig.1B), and in the Methods (lines 769-786). The fitted regression values of this analysis are also used for a new fMRI model (Model 4), to uncover within-TVA correlates of the probability of correct species categorization (lines 309-325; Fig.4).

      Please note that previously, the behavioral data of the species categorization task were completely absent (N=23), and the reaction times data previously part of Fig.1 were for the species attentional bias task (independent sample of N=28). Since this aspect was not clear at all (same remark by all reviewers—apologies for that), we now include a clear separation in Fig.1, with newly added panels D & E part of a distinct figure area named: “Control task: Testing for Species attentional bias (N=28)”. Panel D illustrates the control task paradigm (each species as exogenous cue; “dot-probe” paradigm) while panel E shows the results (target sine wave tone or “bip” detection reaction times), showing that no species triggered more attentional capture than the others (Species effect non-significant).

      The acoustic parameters used in Models 2 and 3 should be explicitly listed in the Methods (even if already published elsewhere).

      In addition to their description in Table S1, we now include the 16 acoustic parameters used to calculate acoustic distance between the species in the Methods, see lines 828-844.

      Consider simplifying the presentation of the three models: a figure summarizing their relationships would help.

      We now include only one figure (Fig.2) for Model 3, and we pushed model 1&2 to the supplementary material. We also simplified Fig.3 for a clearer view of the overlaps between the 3 models within the TVA.

      The description of “systematic and thorough control of phylogeny” (line 119) is overstated, given that only three nonhuman species were included.

      We agree with the reviewer and we suppressed both “systematic” and “thorough” from the sentence.

      Provide rationale for not including a nonvocal control category (e.g., scrambled vocalizations or environmental sounds) to assess TVA selectivity.

      The main objective of the study was to uncover whether human participants could recognize the vocalizations from nonhuman primates—from both great apes and monkeys—as compared to the human voice. We therefore did not include nonvocal or noise stimuli. We added this point as a limitation in the Discussion (lines 593-596 and 609-611).

      Even though we did not include such stimuli for the reason mentioned above, the delineation of subtypes of nonvocal material within the TVA of our participants (Fig.2) are, in our opinion, clarifying the message: chimpanzee-selective activations are fully within ‘voice vs. animal’ and ‘voice vs. nature’ TVA subareas, while it is not the case in ‘voice vs. music’ and ‘voice vs. noise’ TVA subareas.

      Clarify if participants were trained or had a practice session to recognize the four species before scanning.

      The participants were indeed trained on 3 stimuli per species before entering the MRI scanner. These stimuli were discarded from the species categorization task. We added a sentence about this aspect, see lines 131-132.

      Specify what is meant by "no good or bad response" in the attentional control task (line 724).

      We suppressed this wording as it was highly confusing.

      (3) Results

      Behavioral accuracy should be reported to complement reaction times.

      We now added behavioral data for the species categorization task as well as the neural correlates of accurate species categorization. See our previous response above (‘‘‘).

      Figures 2-4 largely overlap; consider merging or simplifying to reduce redundancy.

      We agree and this point was raised by the other reviewers as well. Task-based results are now presented only for Model 3 as Fig.2, while Fig.3 (previously Fig.5) summarizes the overlap between the three models. Figures for Models 1 & 2, previously labelled Fig.3 and Fig.4, were moved to the supplementary material.

      Figure 2: Please indicate more clearly where "chimp-selective" areas are located (perhaps with zooms).

      We agree, we now modified Fig.2 with zoomed-in panels and a clearer outline of chimp-selective areas (solid blue outline). This outline is also referenced in the text (lines 236-237).

      Correction for multiple contrasts: With many pairwise tests, adjustments (Bonferroni or FDR) should be mentioned explicitly.

      We now specify ‘FDR correction at the voxel level’ at the beginning of the Results section (lines 195-198) as well as in each figure.

      Replace "specific to chimpanzee" with "selective for chimpanzee" to avoid implying exclusivity.

      We made the suggested replacement throughout the manuscript.

      Discuss whether the small macaque-related clusters might simply reflect acoustic overlap rather than true category selectivity.

      We added a section on this important aspect, including results that support the role of mid-STG/STS regions for more noise-like stimuli, including the use of macaque coos. See lines 450-461.

      (4) Discussion

      The discussion overstates claims of "chimpanzee-selectivity" in TVA. The evidence shows relative preference, not absolute selectivity.

      We now specify from the start of the Discussion that we are not interpreting the results as absolute selectivity but rather as more relative preference, see lines 371-373.

      The authors repeatedly conflate acoustic and phylogenetic factors; this should be explicitly acknowledged as a limitation.

      We agree, and we completed the limitations section already dedicated to this aspect by a more explicit account of the confound, see lines 609-611.

      Clarify what is meant by "recruitment" and "selectivity" (lines 411-419, 577). TVA activity often reflects enhanced responses to voices compared to non-vocal sounds, not exclusive activation.

      We clarified this wording in the Discussion (lines 377-378) and replaced another instance by “activated the […]” to make it clearer what we imply, namely enhanced activity triggered by chimpanzee calls within human TVA.

      The lack of non-vocal control conditions should be discussed as a major interpretive limitation.

      We added this point as a limitation in the Discussion (lines 593-596).

      The statement that "chimpanzee-selective activity" arose in humans who have never been exposed to chimp calls (line 450) invites evolutionary speculation but should be more cautiously phrased.

      We agree, and we rephrased by: “[…] with chimpanzee calls triggering responses in the anterior STG/TVA of our human participants […]”. See lines 432-433.

      The comparison to recent macaque data (Giamundo et al., 2024 PNAS) is crucial: these findings of human-voice-selective neurons in macaques directly parallel the present human-chimp result.

      We agree with the reviewer, and we are hopeful to read similar results for other apes/great apes in the future.

      Reviewer #2 (Recommendations for the authors):

      (1) The primate vocalizations used in this study were recorded in diverse social and emotional contexts, which may have contributed to the observed differences in TVA activation. Since the temporal voice areas are known to be sensitive to affective and socially relevant cues, these contextual differences could confound the interpretation of species-specific neural responses. Therefore, I suggest that the authors conduct a post-hoc analysis to quantify and compare the affective valence, arousal levels, and social contexts associated with each stimulus set.

      We agree that the TVA are sensitive to social—or socially relevant—cues, motivating the very thorough work of the expert reserve personnel on-site to accurately categorize the calls according to the very specific context they were produced in. If the reviewer meant presenting these stimuli to non-expert participants and asking them to categorize the context or valence, we think it would make no sense since the ratings would be completely below chance level and therefore uninformative. The newly added behavior—and model-based fmri—data include this crucial point, a factor that we named ‘Context’ in our analyses. In fact, for each species’ 18 stimuli, we control for agonistic and affiliative production context—split evenly, per species. Also, computing an additional posthoc analysis by splitting the stimuli according to Context would result in too few trials to get sensible and reliable fMRI results.

      That being said, our study targets this specific aspect by extracting the acoustic features that characterize our stimulus set the best, across context-species-valence-arousal, which is exactly what we want. Through the three types of modeling we used—from more simplistic to more elaborate the results converge only for one species: chimpanzee calls.

      We think the addition of behavioral data, model-based fMRI data, and the specific analysis on acoustic differences between chimpanzee and bonobo calls strengthens the message and the validity of our findings.

      (2) Although the author mentioned that the behavioral effects triggered by these vocalizations have been reported previously, the behavioral responses of the participants in the current study are also crucial for our understanding of the results. If the MRI data can be combined with the participants' behavioral responses for comprehensive analysis, the conclusions of this study will be more compelling.

      We agree with the reviewer, and we added the behavioral data—controlling for reaction times, production context and acoustics of interest—and we also included a model-based fMRI modeling of the probability of correct species categorization as Model 4, Fig.4. See, respectively: lines 147-165, Fig.1B; Methods, lines 769-786; Neuroimaging results, lines 309-325.

      (3) I am still not convinced that phylogenetic proximity drives the observed neural selectivity. While chimpanzee vocalizations do elicit stronger responses in anterior STG, the claim that this reflects evolutionary relatedness lacks evidence. If the acoustic features of a certain call from a particular species are similar to those of human voices, it may also lead to similar effects.

      We agree with the reviewer that generalizing our results in terms of phylogenetic proximity alone is not a viable option. Including many more primate species including other great apes would be necessary, and we mention this crucial aspect in the limitations section. We also insist in the Discussion on the interdependence between phylogeny and acoustics in our data, since: 1) we cannot fully disentangle these factors here, 2) we cannot attribute our results to either one or the other. See lines 387-390, 410-411, 473-477, 587-591.

      If the acoustic features of a certain call from a particular species are similar to those of human voices, it may also lead to similar effects.

      We agree, and nobody could disagree: if an auditory object is extremely similar to the human voice in terms of acoustics, it would therefore potentially activate the TVA. This is exactly our message: in the natural ‘auditory world’, the calls from chimpanzees seem to be among the very few animal auditory signals that are sufficiently close, acoustically, to the human voice and therefore trigger TVA activity. They also happen to be the calls from a species which is phylogenetically the closest to humans with minimal differences with other great apes. Our results are in that sense very aligned with work from the laboratory of Pascal Belin, namely on ‘voice patches’ in the primate brain located in the (anterior) TVA, cited in our manuscript.

      We therefore think our interpretation does not exclude that in the near future, similar results within the TVA could be observed for other auditory objects, and if animal, from a species potentially much more distant phylogenetically or from vocal signals of other great apes.

      We added a key limitation point in the Discussion on the absence of auditory control stimuli in our design, such as scrambled or spectrum shifted per-species stimuli, which would have made the interpretation clearer identical acoustics but alteration/destruction of the species auditory object. See lines 593-596 and 609-611.

      Reviewer #3 (Recommendations for the authors):

      While the manuscript presents intriguing results, several concerns are raised for further consideration, detailed below.

      We thank the reviewer for evaluating the manuscript and for the constructive criticism and suggestions.

      Major concerns:

      (1) This study claims that bilateral anterior superior temporal gyrus (aSTG) in humans can be specifically activated by chimpanzee vocalizations rather than all other primate species after regressing out relevant acoustic parameters using three distinct analyses. I am wondering if a control stimulus (e.g., scrambled chimpanzee vocalizations) were presented, would the activation patterns in these same temporal voice areas (TVA) exhibit significant differences compared to the natural chimpanzee vocalizations?

      We completely agree with the reviewer, and this point was also raised by the other reviewers. We therefore added a key limitation point in the Discussion on the absence of auditory control stimuli in our design, such as per-species scrambled or spectrum shifted stimuli, which would have made the interpretation clearer—identical acoustics but alteration/destruction of the species auditory object. See lines 609-611.

      (2) The figure organization/presentation requires significant revision to avoid confusion and redundancy. E.g:

      Figure 1C is the same as Figure S1. In addition, Figure 1C lacks a figure legend and descriptive label.

      The scatter plots in Figures 2D, 2H, 3D, 3H, and 4D, 4H are same as those in Figures S2, S3, and S4. However, some of these duplicate plots even have inconsistent axis labels.

      In several panels, the main figures appear to be summaries derived from the supplementary figures. The authors should organize these figures well to eliminate redundancy.

      Please double-check all the figures to make sure of accuracy.

      We agree that the figures were badly organized and were too crowded and redundant. We now suppressed the redundancy between Fig.1 and Fig.S1, and we reduced fMRI results to one figure for statistical Model 3 while the other models are in the supplementary data—we also justify this decision in the text by highlighting that model 3 is the most elaborate and sensitive one. Fig.3 (previously ‘Fig.5’) shows the overlaps between models and was simplified and clarified as well.

      (3) The analysis of the fMRI data does not account for the participants' behavioral performance, specifically their reaction times (RTs) during the species categorization task. It is possible that processing vocalizations from certain species requires more cognitive effort or induces higher decision uncertainty. Could the observed neural effects be confounded by the decision-making process itself?

      We now include behavioral data analysis (accuracy data controlled for reaction times and acoustics of existing Model 3, using mixed-effects logistic regression) in addition to a new, 4th model for fMRI data. This 4th model was computed in a model-based fashion by modeling the probability of correct categorization within the TVA (fitted regression coefficients, per Participant, Species, Trial) and revealing the neural correlates of this modulator. We now display these results in Fig.4 and we introduce the motivation factor for including a categorization task rather than more traditional passive listening (lines 75-77), as well as limitations, lines 595-596.

      (4) One interesting attempt of this study is to dissociate biologically salient information in animal vocalizations from their low-level acoustic properties. This presents a fundamental conceptual challenge: how to rigorously disentangle a vocalization's species-specific attributes from its inherent acoustic correlates. More precisely, what essential biological information persists in a species' vocal signal after statistically accounting for all quantifiable acoustic features? I recommend that the authors address it in the discussion.

      We thank the reviewer for this very important comment, and for suggesting we discuss it in the manuscript. We completely agree: we cannot fully orthogonalize species and acoustics, and this aspect relates also more broadly to cognitive and affective neuroscience studies involving vocal material. Namely: “What is an auditory object without acoustics?”

      We included a full paragraph on this aspect, see Discussion, lines 570-584.

      (5) If a brain region, such as TVA, is responsive to both acoustic parameters and biological meanings of animal vocalizations, the method used in this study might be inadequate by setting covariates to zero. It is possible that species information is embedded within a specific acoustic pattern. The current modeling approach may not capture such complex information and could potentially introduce bias when estimating the species effect. I recommend that the authors address this issue in the discussion.

      We thank the reviewer for this point once again, we addressed it in the Discussion, lines 581-584, and also in the section dedicated to study limitations, lines 609-613.

      (6) In the discussion, non-human primate vocalizations are "unreadable" to humans. If this is the case, what is the fundamental perceptual difference between these vocalizations and those from the other animal species? An alternative and highly plausible explanation for the findings is the differential familiarity of the participants with the various species, driven by media exposure (e.g., documentaries) or zoo visits and interactions. The authors need to provide a stronger justification for their control stimuli and directly address, either through discussion or additional analysis, how the factor of familiarity might explain their results better than the proposed "evolutionary distance" hypothesis.

      We now discuss this important aspect, see lines 560-569.

      We thought about doing additional analyses on this aspect but we concluded that we did not have any reliable indicators of familiarity for our participants, and additionally they were all recruited for being ‘unfamiliar’ with great apes or old-world monkeys’ vocalized communication.

      Also, frequent mismatches in the media between images of apes and the associated vocal signals (for instance, the depiction of a chimpanzee but with background audio of macaque coos) are not helping this cause.

      Minor:

      (1) No figure legend and result description for Figure 1.

      Figure 1 has a legend, maybe it was cut out during the uploading process, but it is present and verified now.

      (2) In the main text, three statistical models were referenced. Was the data used in each subsequent statistical model derived from the processed data of the preceding model? Please clearly explain this in the main text.

      We now specify this aspect in the Methods and the Results section to clarify that each model is independent from the others (lines 964-966 and 189-191, respectively).

      (3) In Figure 5, the two dashed lines representing Model 1 and Model 2 are confusing for readers.

      We modified the figure (now Fig.3) and simplified it by removing some outlines and clarifying the colors, therefore improving readability.

      (4) Lack of reaction times in the species categorization task.

      We clarified behavioral data, including the results for the species categorization task and for the control, exogenous cueing task, see modified Fig.1 and behavioral results section of the Results.

      (5) Figures 2, 3, 4, 5, Please keep the font size of the figure title consistent.

      Figure title font size were uniformized.

      (6) Line 201, Line 224, and so on, (EFG) → (E, F, G).

      We modified this aspect in every figure legend, including the supplementary material.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Argunşah et al. describe and investigate the mechanisms underlying the differential response dynamics of barrel vs septa domains in the whisker-related primary somatosensory cortex (S1). Upon repeated stimulation, the authors report that the response ratio between multi- and single-whisker stimulation increases in layer (L) 4 neurons of the septal domain, while remaining constant in barrel L4 neurons. The authors attribute this divergence to differences in short-term synaptic plasticity, particularly within somatostatin-expressing (SST<sup>+</sup>) interneurons. This interpretation is supported by

      (1) The increased density of SST+ neurons in L4 of the septa compared to barrel domain,

      (2) The stronger response of (L2/3) SST+ neurons to repeated multi- vs single-whisker stimulation and

      (3) the reduced functional difference in single- versus multi-whisker response ratios across barrel and septal domains in Elfn1 KO mice, which lack a synaptic protein that confers characteristic short-term plasticity, notably in SST+ neurons.

      Consistently, a decoder trained on WT data fails to generalize to Elfn1 KO responses. Finally, the authors report a relative enrichment of S2- and M1-projecting cell densities in L4 of the septal domain compared to the barrel domain, suggesting that septal and barrel circuits may differentially route information about single vs multi-whisker stimulation downstream of S1.

      Strengths:

      This paper describes and aims to study a circuit underlying differential response between barrel columns and septal domains of the primary somatosensory cortex. This work supports the view these two domains contribute distinctly to the processing single versus multi-whisker inputs and highlight the role of SST+ neuron and their short-term plasticity. Together, this study suggests that the barrel cortex multiplexes whisker-derived sensory information across its domains, enabling parallel processing within S1.

      Weaknesses:

      Although the divergence in responses to repeated single- versus multi-whisker stimulation between barrel and septal domains is consistent with a role for SST<sup>+</sup> neuron short-term plasticity, the evidence presented does not conclusively demonstrate that this mechanism is the critical driver of the difference. The lack of targeted recordings and manipulations limits the strength of this conclusion: SST<sup>+</sup> neuron activity is not measured in L4, nor is it assessed in a domain-specific manner. The Elfn1 knockout manipulation does not appear to selectively affect either stimulus condition, domain or interneuron subtype. Finally, all experiments were performed under anesthesia, which raises concerns about how well the reported dynamics generalize to awake cortical processing.

      We thank the reviewer for their careful reading of the manuscript and their balanced assessment of both its strengths and limitations. We acknowledge the reviewer’s concerns regarding the lack of direct, layer- and cell-type–specific recordings and manipulations of SST<sup>+</sup> interneurons, as well as the use of anesthesia. As noted in the Discussion, these factors limit the extent to which causal mechanisms can be established and the degree to which the reported dynamics can be generalized to awake cortical processing. For this reason, we intentionally frame the Elfn1–SST mechanism as a working model supported by converging anatomical, developmental, physiological, and genetic evidence, rather than as definitive proof. We believe this conceptual framing appropriately reflects the scope of the current data while highlighting clear directions for future work.

      Reviewer #2 (Public review):

      Summary:

      Argunsah and colleagues demonstrate that SST expressing interneurons are concentrated in the mouse septa and differentially respond to repetitive multi-whisker inputs. Identifying how a specific neuronal phenotype impacts responses is an advance.

      Strengths:

      (1) Careful physiological and imaging studies.

      (2) Novel result showing the role of SST+ neurons in shaping responses.

      (3) Good use of a knockout animal to further the main hypothesis.

      (4) Clear analytical techniques.

      Comments on revisions:

      The authors have effectively responded to my initial critiques - I have no further concerns.

      We thank the reviewer for their positive evaluation of our work and for recognizing the novelty of the findings, the careful physiological and imaging approaches, the use of the Elfn1 knockout model, and the clarity of the analytical framework. We are pleased that the reviewer has no further concerns and appreciates the contribution of this study to understanding the role of SST<sup>+</sup> interneurons in shaping sensory processing in the barrel cortex.

      Reviewer #3 (Public review):

      Summary:

      This study investigates the functional differences between barrel and septal columns in the mouse somatosensory cortex, focusing on how local inhibitory dynamics (particularly involving SST<sup>+</sup> interneurons) may mediate temporal integration of multi- whisker (MW) stimuli in septa. Using a combination of in vivo multi-unit recordings, calcium imaging, and anatomical tracing, the authors propose a model in which Elfn1-dependent synaptic facilitation onto SST<sup>+</sup> interneurons contributes to the distinct sensory responses to MW input in barrels and septa, enabling functional segregation between these domains.

      Strengths:

      The study presents a thought-provoking and useful conceptual model for understanding sensory processing in the somatosensory cortex. While barrel columns have been widely studied, septal regions remain relatively understudied in mice. If septa indeed act as selective integrators of distributed sensory input, this would suggest a novel computational role for cortical microcircuits beyond the classical view focused on barrels. Although still hypothetical, the proposed model in which SST<sup>+</sup> interneurons contribute to domain-specific sensory responses between barrel and septal domains is intriguing and opens new avenues for investigating inhibitory circuit mechanisms.

      Weaknesses:

      The primary limitation of this study lies in the spatial and cellular specificity of the recording techniques. The physiological data rely predominantly on unsorted multi-unit activity (MUA) recorded with lowchannel-count silicon probes. Because MUA aggregates signals from multiple neurons over a radius of approximately 50-100 µm (often wider than the typical septal width in mice), this approach makes it difficult to confidently isolate activity originating strictly from within septal domains. The manuscript would benefit from additional analyses to validate the spatial specificity of these recordings, such as systematically varying spike detection thresholds to test the robustness of domain attribution, as suggested by the reviewer. Furthermore, although the authors now appropriately frame their findings in the Elfn1 knockout mice as indirect evidence, it is worth emphasizing that the study lacks direct in vivo, cell-type-specific recordings and manipulations to more definitively test the proposed mechanism.

      We thank the reviewer for their thorough and constructive evaluation of the manuscript and for highlighting both the conceptual strengths of the study and its technical limitations. We agree that the spatial and cellular specificity of unsorted multi-unit recordings imposes inherent constraints on the interpretation of domain-specific activity, particularly given the narrow width of septal compartments in mice. As now clarified in the manuscript, we do not claim absolute cellular specificity of “septal” recordings but rather interpret them as septal-enriched populations. To directly address this concern, we performed additional threshold-based analysis demonstrating that the key domain-specific effects persist selectively in Layer 4 under stricter spike-detection criteria, supporting a local circuit origin of the critical findings. Further, the more stringent detection criteria (Suppl Fig 3A) collapse the divergence seen in Layer2/3 (Suppl Fig 4C), suggesting that this divergence arises in Layer 4, where SST+ interneuron distributions diverge between barrel and septa.

      We further agree that the Elfn1 knockout results provide indirect, rather than definitive, evidence for causal involvement of SST<sup>+</sup> interneurons and therefore intentionally frame the Elfn1–SST mechanism as a working model supported by converging anatomical, physiological, developmental, and genetic observations. We believe this explicitly moderated interpretation appropriately reflects the scope of the current data while establishing a clear conceptual framework and motivation for future studies employing cell-type-specific recordings and manipulations to directly test the proposed mechanism.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Major comments

      (1) Interpretation of "septal" recordings: The authors claim that the activity recorded from electrodes placed in the septa can be confidently attributed to septal neurons. In my previous review, I raised a major concern that such "septal" recordings likely include spikes from adjacent barrels, given the broad spatial resolution of MUA and the narrowness of the septa in the mouse S1. In fact, the intermediate properties observed in septal recordings from wild-type mice could be explained by a mixture of activity from principal and neighboring barrels-an interpretation that contrasts with the authors' conclusion. Upon reviewing the probe model used (A8x8-Edge-5mm-100-200-177), I noticed a discrepancy between the manufacturer's design and the schematic provided in the manuscript. The electrodes are located near the right edge of the probe rather than the center, suggesting that neurons in adjacent barrels could easily be sampled. In my previous review, I therefore suggested alternative approaches, such as calcium imaging, to more convincingly support the authors' claims. However, the revised manuscript does not include new experiments or additional analyses addressing this issue. Instead, the authors argue that using a high spike detection threshold (SD > 7.5) ensures that recorded activity originates from septal neurons, even though this value does not appear particularly conservative, as it was merely adopted from a previous study without justification in the present context. While I agree that a higher threshold may reduce contamination from distant sources, it does not guarantee that only septal neurons contribute to the signal. By nature, MUA reflects activity from multiple neurons within a radius of at least 50-100 µm. To more rigorously support the claim of spatial specificity, I strongly encourage the authors to reanalyze their existing dataset by systematically varying the spike detection threshold and quantifying how the properties and selectivity of detected units change. If neurons closer to the electrode indeed exhibit distinct domain-specific properties, they should become more prominent as the threshold increases. Such an analysis would strengthen the authors' interpretation and improve the manuscript's impact, even in the absence of new experimental data. Alternatively, the authors could revise their claims to acknowledge that the "septal" electrodes likely record from a population that includes septal neurons as well as neurons located at the periphery of principal and adjacent barrels.

      We agree with the reviewer that, by nature, MUA reflects the activity of multiple neurons within a spatial radius and that recordings obtained from electrodes positioned in the septa may include contributions from neurons located at the periphery of adjacent barrels. This concern is further compounded in superficial layers by probe geometry and orientation: given the narrow width of septa and the lateral spread of processes in upper cortical layers, recordings in L2/3 are inherently more susceptible to spatial mixing than those in layer 4, where columns are more compact and cytoarchitecturally distinct. To directly address these issues, we reanalyzed the same dataset using a more stringent spike detection threshold (SD > 9.5), compared to the originally reported SD > 7.5. Importantly, increasing the threshold selectively reduced or eliminated effects in L2/3, while the key domain-specific differences in L4 responses both the differential MW/SW dynamics in wild-type animals and their attenuation in Elfn1 knockout mice remained robust (the new Supp. Fig. 3. In the manuscript). This threshold-dependent dissociation is consistent with the interpretation that the critical effects reported in L4 arise from neurons spatially closer to the electrode and are less influenced by probe orientation or distant sources, rather than reflecting simple mixing of barrel signals. While this analysis does not claim absolute cellular exclusivity of septal neurons, it provides empirical support that the principal conclusions of the study are robust to stricter spatial sampling criteria and are particularly anchored in L4 circuitry. Accordingly, we now explicitly acknowledge in the manuscript that “septal” recordings likely represent septal-enriched populations rather than purely septal neurons, while emphasizing that the persistence of L4 effects under higher spike-detection thresholds strengthens the conclusion that local L4 inhibitory dynamics underlie the reported functional differences between barrel and septal domains.

      The greater sensitivity of L2/3 results to spike-detection threshold is also expected based on both anatomical considerations and probe geometry. Neurons in L2/3 possess broader horizontal dendritic and axonal arbors and participate in more laterally distributed integration across columns, making population signals in these layers intrinsically less spatially focal. As a result, conservative spike-detection criteria preferentially suppress L2/3 effects, particularly when recordings are obtained with probes optimized for deeper layers. Importantly, our two-photon calcium imaging data while similarly limited to L2/3 demonstrate that SST<sup>+</sup> interneurons show locally measurable and stimulus-specific responses at the single-cell level, providing independent support that L2/3 SST<sup>+</sup> activity is stimulus-modulated rather than artifactual. Taken together, these observations suggest that L2/3 results reflect more distributed and integrative network activity, whereas the L4 effects that persist across thresholds are more directly attributable to local circuit mechanisms. This layer-specific dissociation further supports our interpretation that the central findings of the study are driven by local inhibitory dynamics in L4, with L2/3 activity reflecting downstream integration rather than primary domain-specific computation.

      (2) Interpretation of the Elfn1 KO data: The authors' interpretation that Elfn1-dependent facilitation of SST<sup>+</sup> interneurons underlies the differential sensory responses between barrel and septal domains is conceptually appealing and supported by several converging, albeit indirect, lines of evidence. Specifically, the consistent correspondence among the differential activation of SST<sup>+</sup> neurons upon SWS and MWS, the late development of the barrel-septa differences in the responses to SWS and MWS, and the attenuation of this difference in Elfn1 knockout mice lends plausibility to the proposed model. However, it should be emphasized that the data remain indirect: the study does not include direct recordings of SST<sup>+</sup> neuronal activity from the knockout mice, nor cell-type- specific manipulations to demonstrate causal involvement. The mechanistic explanation therefore represents a hypothesis rather than definitive proof. That said, the authors clearly acknowledge these limitations in the Discussion and appropriately moderate their claims by presenting the SST-Elfn1 mechanism as a working model. Given this careful framing, the current manuscript can be regarded as a valuable conceptual contribution that advances our understanding of how inhibitory dynamics may shape temporal processing in the barrel cortex. Further experiments, as mentioned above, will be essential to test the causal role of this mechanism directly.

      We thank the reviewer for this thoughtful and balanced assessment. We fully agree that the Elfn1 knockout experiments provide indirect rather than definitive evidence for a causal role of SST<sup>+</sup> interneurons in mediating the domain-specific MW/SW response dynamics between barrels and septa For this reason, throughout the revised manuscript we explicitly frame the Elfn1–SST mechanism as a working model rather than a proven mechanism.

      Minor comments:

      The authors have adequately addressed my previous minor comments. In this round, I carefully reviewed the revised manuscript and identified several issues related to references. I would also like to add a brief comment regarding the Discussion section:

      (1) Stachniak et al., 2021 is included in the reference list but is not cited anywhere in the main text. Please either remove this entry or cite it appropriately in the manuscript.

      Removed.

      (2) Yamashita et al., 2018 is cited in the main text (Line 767), but it is not included in the reference list.

      Fixed.

      (3) Sylwestrak and Ghosh, 2012 is cited at Line 261 and Line 270, but likewise absent from the reference list.

      Fixed.

      (4) At Line 497, Chen et al., 2015 is cited, but, the appropriate and original reference would be Chen et al., 2013 (PMID: 23792559), which should either replace or precede the 2015 citation.

      Added.

      (5) At Line 221, El-Boustani et al., 2018 is cited. However, this study is based on the visual cortex, whereas the manuscript concerns the barrel cortex. A more relevant citation (e.g., Lefort et al., 2009 [PMID: 19186171]) would better support the discussion of cellular organization in the barrel cortex. Please consider updating the citation.

      Thank you for this suggestion. We agree with the reviewer and now we have changed El-Boustani with Lefort et al. 2009 as suggested by the reviewer.

      (6) Furthermore, Chakrabarti & Alloway (2006) performed tracer-based mapping of projections from barrel and septal columns in rat S1 and similarly suggested differential organization of M1- and S2projection neurons in the barrel and septal regions.

      Although the current study thoroughly analyzed the layer-specificity of the location of these projection neurons, the lack of explicit discussion of this relevant prior work is a notable omission.

      The authors should incorporate a comparison with these results to better contextualize their findings.

      The following text is added to the discussion: “Our retrograde labeling data supports and expands on previous work proposing similar models (Alloway, 2008; Chakrabarti and Alloway, 2006).”

    1. Author response:

      (1) Introduction & Roadmap

      We are grateful to the Reviewers for engaging with outstanding questions relating to our findings’ connections to multiple subdisciplines of cognitive neuroscience. Noting that Reviewers 1 and 2 interpreted our findings differently, we welcome the opportunity to engage in what Reviewer 1 characterised as “an interesting debate”. To promote a shared understanding and discussion of our findings, we have organised our response to address more technical comments first.

      Our provisional response is organised as follows: Section 2 addresses selected technical comments relating to our Results. Section 3 addresses comments related to the design of our behavioural paradigm. Section 4 focuses on the broader interpretation of our findings. Section 5 concludes our provisional response with potential future directions and a summary of the significance of our findings.

      (2) Selected technical comments related to our Results

      We apologise to Reviewer 2 for the confusion in relation to the meaning of “attended” and “unattended” trials. What we said was “Positive Pref values indicate a higher response rate to the contralateral side than the ipsilateral side (relative to the electrode)” (Figure 3c caption), “we indexed all contralateral whisker vibrations according to their associated Perf and Pref” (Results text), and “we divided trials into (contralaterally) attended (Pref<sub>C/L</sub>: Pref>0) and unattended (Pref<sub>I/L</sub>: Pref<0) groups” (Results text). We can confirm that we defined an “unattended trial” (Pref<0) as a contralateral stimulus trial in the centre of an epoch (10-15 trials) within which the mouse responded (licked) more frequently to ipsilateral stimuli. Critically, we did not define an unattended trial as an ipsilateral stimulus trial. Furthermore, attention thus defined (i.e. Pref>0) can vary independently of the whisker stimulus associated with rewards. Indeed, while we initially did not include this result in our paper for the sake of brevity, even unrewarded “attended” trials (Pref>0) evoked significantly greater neuronal responses than unrewarded “unattended” (Pref<0) trials. We note that this is an analysis suggested by Reviewer 1, and we will include and discuss this result in our revised manuscript (e.g. in relation to literature suggested by Reviewer 2). For additional clarity, we use “Performance” (Perf) in relation to overall stimulus detection, consistent with the analysis of Lee et al. (2020), which found this measure was correlated with pupil diameter in a vibrissal target detection task.

      We thank Reviewer 1 for noticing that the axes on Figure 3e should be labelled “Pref>0” (Y axis) and “Pref<0” (X axis), as suggested by the figure caption. We will correct this in our revised submission. The yellow point on Fig 3e shows the unit from Fig 3d, while the yellow line in Fig 3e shows the magnitude of that unit’s (non-normalised) gain modulation. While this is alluded to in the Results text (“The example unit in Figure 3d is in the 93rd percentile of units for raw modulation depth (ΔHits(attended – unattended) = 3.3 spikes/second; yellow line in Fig.3e)”, this should be explained in the Figure caption, and it will be in our revised manuscript. We would also like to clarify that Figures 3g–3h display results for all units, not just the top 25%. We agree this is not sufficiently clear and we will rectify this in our revised manuscript. Addressing Reviewer 2, while we acknowledge that mice responded less to both stimuli in the second block, they also meaningfully adjusted their behaviour to the reversal in reward contingencies: their responses to the previously rewarded stimulus reduced significantly more than those to the previously unrewarded stimulus.

      (3) Design of the behavioural paradigm

      We made a deliberate design choice to maximise the ecological validity of our behavioural paradigm, and note that there are advantages to doing so. For example, our paradigm can be used to show that even unrewarded “attended” trials (Pref>0) evoke significantly greater neuronal responses than unrewarded “unattended” (Pref<0) trials (see Section 2, above). Indeed, it is precisely this finding that makes our paradigm uniquely suited to the investigation of value-driven attentional capture (Anderson et al., 2011): in this instance attention directed to stimuli that are no longer rewarded despite equal availability of rewarded stimuli. This finding also demonstrates that our paradigm dissociates attention from stimulus-reward contingency at least as well as other paradigms which have been successfully used to study spatial attention in mice. As noted in Section 2, we will discuss this result in relation to other relevant research (e.g. Ramamurthy et al., 2025) in our revised manuscript.

      Briefly, the direct manipulation of reward contingencies is one of two noteworthy methodological distinctions between our own paradigm and that of Ramamurthy and colleagues (2025). The task of Ramamurthy et al. (2025) associated all whisker stimuli with rewards and delivered stimuli to different whiskers on a single whisker pad. These methodological distinctions may have reduced the relevance of the spatial differences between stimuli to the mice undertaking the task. Indeed, it is not certain that a mouse would treat the unilateral variation in whisker stimulation Ramamurthy and colleagues delivered as primarily spatial or featural. The psychophysical and neural differences between spatial and featural attention in humans suggest dissociable underlying mechanisms, and the same may be true in mice. Thus, our own paradigm may more effectively isolate spatial attention from featural attention. Conversely, to the extent that the findings of Ramamurthy and colleagues do reflect spatial attention, our combined findings and paradigms help elucidate the associated mechanisms across spatial scales in mice.

      We acknowledge that spatial cueing is well-suited to isolating the effects of covert attention from other forms of attention. However, it should be noted that spatial cueing in rodents is subject to its own challenges, including limitations in trial numbers due to the required manipulation of stimulus intensity (Reynolds et al., 2000; Herrmann et al., 2010), cue validity and associated trial probabilities (Peterson & Gibson, 2011; Girardi et al., 2013). Such experiments are further complicated by the duration and efficacy of training (i.e. the number of mice that learn the task; Wang & Krauzlis, 2018; Hu & Dan, 2022). It is also worth noting that trial probability manipulations introduce the same limitation in trial numbers with block-type attention tasks (You & Mysore, 2020; Kanamori & Mrsic-Flogel, 2022).

      While there are clear differences between our own paradigm and those mentioned above, there are also important similarities. First, these tasks are all goal-directed, stimulus-driven, and reliant on learned task contingencies (e.g. Peterson & Gibson, 2011; Girardi et al., 2013). Furthermore, these paradigms are all operant conditioning protocols which leverage learned stimulus-reward contingencies to train attention-related behaviours in mice. A noteworthy similarity between our findings and those of authors using block-type attention tasks in particular (e.g. You & Mysore, 2020; Kanamori & Mrsic-Flogel, 2022) is the observation of apparent attentional biases in behavioural responses independent of the experimental manipulations (i.e. stimulus probability / reward contingency).

      (4) Comments relating to the broader interpretation and discussion of our findings

      Fundamentally, attention involves dedicating limited processing resources to some stimulus events at the expense of others. The design of our behavioural paradigm was informed by existing literature on spatial attention in humans, non-human primates, and mice. Our choice of behavioural and neuronal measures as proxies for attention in mice is consistent with this literature. It is technically possible “an animal could pay ‘more’ (rather than less) attention to the stimulus delivered on the unrewarded side, to make sure it suppresses the incorrect response”, but this seems unlikely given what is known about how attention is typically allocated in such tasks, based on the previously mentioned literature.

      With respect to the interpretation and discussion of our findings, Reviewer 1 describes them as “a behavioral phenomenon that can reasonably be interpreted as spatial attentional capture” but suggests they do not clearly distinguish whether this attentional capture is covert or overt. We respectfully disagree for three reasons. First, as discussed in our paper, whisker motion during detection tasks has consistently been associated with reduced detection performance (Ollerenshaw et al., 2012; Kyriakatos et al., 2017; Vandevelde et al., 2023), suggesting that a “receptive” strategy (Diamond & Arabzadeh, 2013) of whisker immobilisation is more applicable to the current data than a “generative” strategy of asymmetric whisker movement (O'Connor et al., 2010; Dominiak et al., 2019). Second, if our behavioural and neuronal findings were due to the mice moving their whiskers to maximise contact with the meshes, we would expect increased evoked neuronal responses to be associated with greater Perf, not just with greater Pref. This pattern was not observed. Of course, the mice might have employed different whisker movement strategies during epochs of high Pref and Perf, but this seems unlikely and is not a parsimonious explanation for our findings. Third, as noted in the Methods section of the paper, we deliberately positioned the meshes close to the base of the whiskers, limiting the impact of whisker movements on stimulus detectability and the incentive to make them.

      In contrast, Reviewer 2 questions the interpretation of our findings as evidence of spatial attention and suggests they might reflect working memory instead. Current research suggests attention and working memory are intimately related integrative brain functions. Indeed, some researchers have even proposed that working memory might be a form of internally directed attention (Awh & Jonides, 2001; Chun, 2011; Gazzaley & Nobre, 2012; Kiyonaga & Egner, 2013; or vice versa: Libedinsky & Fernandez, 2019). Consistent with the comments of Reviewer 2, more recent work seems to emphasise the coordination of attention and working memory (e.g. Joe & Kim, 2023; Zhu et al., 2026; for reviews see Huynh Cong & Kerzel, 2021; van Ede & Nobre, 2023), along with shared mechanisms (Kiyonaga et al., 2021; Panichello & Buschman, 2021), and nuanced dissociations (Liu et al., 2025). Attention is difficult to dissociate from working memory partly because there are multiple definitions (and/or types) of attention. We did not discuss the various definitions and/or forms of attention at length in our paper, but we will briefly discuss this in the revised manuscript.

      The “interesting debate” to which Reviewer 1 refers could also be described as vigorous, despite approximately three decades of research. This debate broadly relates to the degree to which attentional control is driven by exogenous (e.g. colour contrast) versus endogenous factors (e.g. the focus of spatial attention, see Fig.2 in Belopolsky et al., 2007; see also: Liesefeld & Mueller, 2020; Manini et al., 2021; Beffara et al., 2022), and the degree to which this is a function of experimental context. The review article by Luck et al. (2021) entitled “Progress toward resolving the attentional capture debate” provides a striking illustration of this debate, as do the twenty-two commentaries (and three commentary responses) associated with it. Admittedly, this debate largely revolves around human attention experiments, and human cognition may be more complex than mouse cognition. However, the complexity of human cognition may also be easier to study and appreciate because complex behavioural experiments can be explained to, understood, and performed by human participants with relative ease.

      (5) Comments relating to future directions and the significance of our findings

      The complexity of the attentional capture debate underscores the importance of developing accessible and scalable animal experiments which can be used to provide mechanistic insights. If the human attention literature is any indication, a diversity of rodent experimental paradigms will be necessary to thoroughly map the neuronal implementation of spatial attention. Returning to our paradigm, Reviewer 1 noted that valuable insights into the mechanisms of vibrissal spatial attention might be obtained from comparing the magnitude of attentional modulation we observed between putative regular and fast-spiking categories of units, and between units located in different cortical layers. We agree it is important to understand spatial attention with cell-type and circuit (including laminar) specificity. However, because we could not persuasively cluster our units based on waveform width, and because of the lack of histological data, segregating units on the basis of such variables is not feasible. Despite our assertion that our findings reflect the effects of covert attention (contra Reviewer 1), we agree that future experiments will be required to conclusively rule out overt attention. Noting the proximity of the meshes to the base of the whiskers in our paradigm, and the difficulty of tracking whiskers in this context, Botulinum toxin injections (as in Ramamurthy et al., 2025) might be a means of achieving this.

      The above notwithstanding, our findings provide multiple contributions to the literature on spatial attention (and perhaps working memory). We detected significant attentional gain modulation across a population of 1461 responsive units. While the gain modulation exhibited by the median unit was modest (albeit statistically significant), the top 25% of responsive units showed a ~12% response modulation (relative to firing rate range for each unit), and ~21% of responsive units were suppressed by the average vibrissal stimulus in the unattended state. Our experimental framework offers an accessible platform for future studies leveraging genetic and circuit-level interventions to dissect the cell-type specific mechanisms of spatial attention. Our work is timely, noting the recent focus of human research on the nexus of attention, selection history, and valence (e.g. Serences, 2008; Della Libera & Chelazzi, 2009; Della Libera et al., 2011; van den Berg et al., 2014; Kim & Anderson, 2019, 2023). Our work is also uniquely poised to stimulate new interdisciplinary research into the circuit mechanisms of value-driven attentional capture, with translational relevance to psychopathologies such as ADHD, addiction, and depression; where value-driven attentional capture is altered (for a review see Anderson, 2021).

      References

      Anderson, B. A. (2021). Relating value-driven attention to psychopathology. Curr Opin Psychol, 39, 48-54. https://doi.org/10.1016/j.copsyc.2020.07.010

      Anderson, B. A., Laurent, P. A., & Yantis, S. (2011). Value-driven attentional capture. Proceedings of the National Academy of Sciences of the United States of America, 108(25), 10367-10371. https://doi.org/10.1073/pnas.1104047108

      Awh, E., & Jonides, J. (2001). Overlapping mechanisms of attention and spatial working memory. Trends Cogn Sci, 5(3), 119-126. https://doi.org/10.1016/s1364-6613(00)01593-x

      Beffara, B., Hadj-Bouziane, F., Ben Hamed, S., Boehler, C. N., Chelazzi, L., Santandrea, E., & Macaluso, E. (2022). Dynamic causal interactions between occipital and parietal cortex explain how endogenous spatial attention and stimulus-driven salience jointly shape the distribution of processing priorities in 2D visual space. Neuroimage, 255. https://doi.org/10.1016/j.neuroimage.2022.119206

      Belopolsky, A. V., Zwaan, L., Theeuwes, J., & Kramer, A. F. (2007). The size of an attentional window modulates attentional capture by color singletons. Psychonomic Bulletin & Review, 14(5), 934-938. https://doi.org/10.3758/Bf03194124

      Chun, M. M. (2011). Visual working memory as visual attention sustained internally over time. Neuropsychologia, 49(6), 1407-1409. https://doi.org/10.1016/j.neuropsychologia.2011.01.029

      Della Libera, C., & Chelazzi, L. (2009). Learning to Attend and to Ignore Is a Matter of Gains and Losses. Psychological Science, 20(6), 778-784. https://doi.org/10.1111/j.1467-9280.2009.02360.x

      Della Libera, C., Perlato, A., & Chelazzi, L. (2011). Dissociable Effects of Reward on Attentional Learning: From Passive Associations to Active Monitoring. PLoS One, 6(4). https://doi.org/10.1371/journal.pone.0019460

      Diamond, M. E., & Arabzadeh, E. (2013). Whisker sensory system - from receptor to decision. Prog Neurobiol, 103, 28-40. https://doi.org/10.1016/j.pneurobio.2012.05.013

      Dominiak, S. E., Nashaat, M. A., Sehara, K., Oraby, H., Larkum, M. E., & Sachdev, R. N. S. (2019). Whisking Asymmetry Signals Motor Preparation and the Behavioral State of Mice. J Neurosci, 39(49), 9818-9830. https://doi.org/10.1523/JNEUROSCI.1809-19.2019

      Gazzaley, A., & Nobre, A. C. (2012). Top-down modulation: bridging selective attention and working memory. Trends Cogn Sci, 16(2), 129-135. https://doi.org/10.1016/j.tics.2011.11.014

      Girardi, G., Antonucci, G., & Nico, D. (2013). Cueing spatial attention through timing and probability. Cortex, 49(1), 211-221. https://doi.org/10.1016/j.cortex.2011.08.010

      Herrmann, K., Montaser-Kouhsari, L., Carrasco, M., & Heeger, D. J. (2010). When size matters: attention affects performance by contrast or response gain. Nat Neurosci, 13(12), 1554-1559. https://doi.org/10.1038/nn.2669

      Hu, F., & Dan, Y. (2022). An inferior-superior colliculus circuit controls auditory cue-directed visual spatial attention. Neuron, 110(1), 109-119 e103. https://doi.org/10.1016/j.neuron.2021.10.004

      Huynh Cong, S., & Kerzel, D. (2021). Allocation of resources in working memory: Theoretical and empirical implications for visual search. Psychon Bull Rev, 28(4), 1093-1111. https://doi.org/10.3758/s13423-021-01881-5

      Joe, J., & Kim, M. S. (2023). Spatial Attention in Visual Working Memory Strengthens Feature-Location Binding. Vision (Basel), 7(4). https://doi.org/10.3390/vision7040079

      Kanamori, T., & Mrsic-Flogel, T. D. (2022). Independent response modulation of visual cortical neurons by attentional and behavioral states. Neuron, 110(23), 3907-3918 e3906. https://doi.org/10.1016/j.neuron.2022.08.028

      Kim, H., & Anderson, B. A. (2019). Dissociable neural mechanisms underlie value-driven and selection-driven attentional capture. Brain Research, 1708, 109-115. https://doi.org/10.1016/j.brainres.2018.11.026

      Kim, H., & Anderson, B. A. (2023). Primary Rewards and Aversive Outcomes Have Comparable Effects on Attentional Bias. Behavioral Neuroscience, 137(2), 89-94. https://doi.org/10.1037/bne0000543

      Kiyonaga, A., & Egner, T. (2013). Working memory as internal attention: toward an integrative account of internal and external selection processes. Psychon Bull Rev, 20(2), 228-242. https://doi.org/10.3758/s13423-012-0359-y

      Kiyonaga, A., Powers, J. P., Chiu, Y. C., & Egner, T. (2021). Hemisphere-specific Parietal Contributions to the Interplay between Working Memory and Attention. J Cogn Neurosci, 33(8), 1428-1441. https://doi.org/10.1162/jocn_a_01740

      Kyriakatos, A., Sadashivaiah, V., Zhang, Y., Motta, A., Auffret, M., & Petersen, C. C. (2017). Voltage-sensitive dye imaging of mouse neocortex during a whisker detection task. Neurophotonics, 4(3), 031204. https://doi.org/10.1117/1.NPh.4.3.031204

      Lee, C. C. Y., Kheradpezhouh, E., Diamond, M. E., & Arabzadeh, E. (2020). State-Dependent Changes in Perception and Coding in the Mouse Somatosensory Cortex. Cell Rep, 32(13), 108197. https://doi.org/10.1016/j.celrep.2020.108197

      Libedinsky, C. D., & Fernandez, P. F. (2019). Graded Memory: A Cognitive Category to Replace Spatial Sustained Attention and Working Memory
 Yale J Biol Med, 92(1), 121-125. https://www.ncbi.nlm.nih.gov/pubmed/30923479

      Liesefeld, H. R., & Mueller, H. J. (2020). A theoretical attempt to revive the serial/parallel-search dichotomy. Attention Perception & Psychophysics, 82(1), 228-245. https://doi.org/10.3758/s13414-019-01819-z

      Liu, Y., Fu, Y., Tang, E., Wu, H., Han, J., Xie, M., Zhang, Y., Peng, B., Huang, J., Liu, H., Chen, H., & Qin, P. (2025). Neural dissociation of attention and working memory through inhibitory control. Nat Commun, 17(1), 22. https://doi.org/10.1038/s41467-025-66553-7

      Luck, S. J., Gaspelin, N., Folk, C. L., Remington, R. W., & Theeuwes, J. (2021). Progress toward resolving the attentional capture debate. Visual Cognition, 29(1), 1-21. https://doi.org/10.1080/13506285.2020.1848949

      Manini, G., Botta, F., Martin-Arevalo, E., Ferrari, V., & Lupianez, J. (2021). Attentional Capture From Inside vs. Outside the Attentional Focus. Frontiers in Psychology, 12. https://doi.org/10.3389/fpsyg.2021.758747

      O'Connor, D. H., Clack, N. G., Huber, D., Komiyama, T., Myers, E. W., & Svoboda, K. (2010). Vibrissa-based object localization in head-fixed mice. J Neurosci, 30(5), 1947-1967. https://doi.org/10.1523/JNEUROSCI.3762-09.2010

      Ollerenshaw, D. R., Bari, B. A., Millard, D. C., Orr, L. E., Wang, Q., & Stanley, G. B. (2012). Detection of tactile inputs in the rat vibrissa pathway. J Neurophysiol, 108(2), 479-490. https://doi.org/10.1152/jn.00004.2012

      Panichello, M. F., & Buschman, T. J. (2021). Shared mechanisms underlie the control of working memory and attention. Nature, 592(7855), 601-605. https://doi.org/10.1038/s41586-021-03390-w

      Peterson, S. A., & Gibson, T. N. (2011). Implicit attentional orienting in a target detection task with central cues. Conscious Cogn, 20(4), 1532-1547. https://doi.org/10.1016/j.concog.2011.07.004

      Ramamurthy, D. L., Rodriguez, L., Cen, C., Li, S., Chen, A., & Feldman, D. E. (2025). Reward history guides focal attention in whisker somatosensory cortex. Nat Commun, 16(1), 5580. https://doi.org/10.1038/s41467-025-60592-w

      Reynolds, J. H., Pasternak, T., & Desimone, R. (2000). Attention increases sensitivity of V4 neurons. Neuron, 26(3), 703-714. https://doi.org/10.1016/s0896-6273(00)81206-4

      Serences, J. T. (2008). Value-Based Modulations in Human Visual Cortex. Neuron, 60(6), 1169-1181. https://doi.org/10.1016/j.neuron.2008.10.051

      van den Berg, B., Krebs, R. M., Lorist, M. M., & Woldorff, M. G. (2014). Utilization of reward-prospect enhances preparatory attention and reduces stimulus conflict. Cognitive Affective & Behavioral Neuroscience, 14(2), 561-577. https://doi.org/10.3758/s13415-014-0281-z

      van Ede, F., & Nobre, A. C. (2023). Turning Attention Inside Out: How Working Memory Serves Behavior. Annu Rev Psychol, 74, 137-165. https://doi.org/10.1146/annurev-psych-021422-041757

      Vandevelde, J. R., Yang, J. W., Albrecht, S., Lam, H., Kaufmann, P., Luhmann, H. J., & Stuttgen, M. C. (2023). Layer- and cell-type-specific differences in neural activity in mouse barrel cortex during a whisker detection task. Cereb Cortex, 33(4), 1361-1382. https://doi.org/10.1093/cercor/bhac141

      Wang, L., & Krauzlis, R. J. (2018). Visual Selective Attention in Mice. Curr Biol, 28(5), 676-685 e674. https://doi.org/10.1016/j.cub.2018.01.038

      You, W. K., & Mysore, S. P. (2020). Endogenous and exogenous control of visuospatial selective attention in freely behaving mice. Nat Commun, 11(1), 1986. https://doi.org/10.1038/s41467-020-15909-2

      Zhu, P., Guan, C., Fu, Y., Shen, M., & Chen, H. (2026). Working memory encoding of attended information is adaptive to future relevance. J Exp Psychol Learn Mem Cogn. https://doi.org/10.1037/xlm0001582

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors present comprehensive experimental observations and a theoretical framework to explain the heterogeneous behaviour of sarcomeres in cardiomyocytes. They show that a stochastic component exists in their contractile activity, which may act as a feedback mechanism regulating physiological function.

      Strengths:

      Experiments and data analysis are robust and valid. The rigorous statistical analysis and unbiased methods enable the authors to draw well-supported conclusions that go beyond the existing literature. Their outcomes inform about cellular activity at the individual level and the authors explain how the transient dynamics of single sarcomeres are governed by a force-velocity relationship and lead to the complex contractile patterns. The similarity of the results to the study cited in [24] demonstrates the validity of the in vitro setup for answering these questions and the feasibility of such in-vitro systems to extend our knowledge of out-of-equilibrium dynamics in cardiac cells.

      Very interesting the suggestion that the interplay between intrinsic fluctuations and the dynamic instability are part of a feedback mechanism for maintaining structural and functional homeostasis.

      The addition of the theoretical model and the new text of the manuscript improves the clarity of the study.

      Reviewer #2 (Public review):

      Summary:

      Sarcomeres, the contractile units of skeletal and cardiac muscle, contract in a concerted fashion to power myofibril and thus muscle fiber contraction.

      Muscle fiber contraction depends on the stiffness of the elastic substrate of the cell, yet it is not known how this dependence emerges from the collective dynamics of sarcomeres. Here, the authors analyze contraction time series of individual sarcomeres using live imaging of fluorescently labeled cardiomyocytes cultured on elastic substrates of different stiffness. They find that a reduced collective contractility of muscle fibers on unphysiologically stiff substrates is partially explained by a lack of synchronization in the contraction of individual sarcomeres.

      This lack of synchronization is at least partially stochastic, consistent with the notion of a tug-of-war between sarcomeres on stiff sarcomeres. A particular irregularity of sarcomere contraction cycles is 'popping', the extension of sarcomers beyond their rest length. The statistics of 'popping' suggest that this is a purely random process.

      Strengths:

      This study thus marks an important shift of perspective from whole-cell analysis towards an understanding the collective dynamics of coupled, stochastic sarcomeres.

      Reviewer #3 (Public review):

      The manuscript of Haertter and coworkers studied the variation of the length of a single sarcomere and the response of microfibrils made by sarcomeres of cardiomyocytes on soft gel substrates of varying stiffness.

      The measurements at the level of a single sarcomere are an important new result of this manuscript. They are done by combining the labeling of the sarcomeres z line using genetic manipulation and a sophisticated tracking program using machine learning. This single sarcomere analysis shows strong heterogeneities of the sarcomeres that can show fast oscillations not synchronized with the average behavior of the cell and what the authors call popping eveents which are large amplitude oscillations. Another important result is the fact that cardiomyocyte contractility decreases with the substrate stiffness, although the properties of single sarcomeres do not seem to depend on substrate stiffness.

      The authors suggest that the cardiomyocyte cell behavior is dominated by sarcomere heterogeneity. They show that the heterogeneity between sarcomere is stochastic and that the contribution of static heterogeneity (such as composition differences between sarcomeres) is small.

      Strengths:

      All the results are, to my knowledge, new and original. The authors also made a theoretical model where each sarcomere is described by a Langevin equation based on a non-linear coupling between force and velocity of the sarcomeres. This model accounts well for the experimental results including the observation of what the authors call popping events.

      We thank you and the reviewers for the positive evaluation of our revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Origin of the 3-Hz oscillation and required model extension. These oscillations are reproduced by our model, and their origin is already discussed in the manuscript (see lines 403–406).

      (2) Inclusion of all 5085 LOIs vs. the selected 2321. We have expanded the explanation of the LOI selection criteria in the manuscript and clarified that the main conclusions are not sensitive to this choice (lines 161-166)

      (3) Fig. 3G caption — popping rate. The caption has been updated to clarify the units and normalization. 

      (4) Fig. 4G — "Length x" vs. ΔL. Notation corrected for consistency.

      (5) Fig. 4G — gray data points. Confirmed: these represent the mean, and the caption has been updated accordingly.

      (6) Relation of k_l to the true substrate stiffness. We have added the following clarification: "The model evaluation compared the distributions of sarcomere length changes and velocities from simulations with representative experimental LOIs from substrates (5, 15, and 85 kPa, mapped to k_l = 0.5, 1.5 and 8.5 in our 1-D model; k_l is unitless, so only the ratios between values are meaningful — rescaling k_l leaves model output unchanged under correspondingly rescaled parameters) covering the full range of mechanical loads." (lines 365-369)

      (7) Could a simpler model fit the data? The cubic polynomial in Eq. (3) was deliberately chosen as a generalist ansatz rather than imposed: its coefficients were obtained by data-driven inference via Differential Evolution, and if lower-order terms within this family had sufficed, the higher-order coefficients would have been driven toward zero. The inferred nonmonotonic force–velocity relation has two extrema separated by an unstable negative-slope branch, which sets a lower bound on the polynomial order — a linear F–v is monotonic and a quadratic admits only a single extremum, so cubic is the minimum polynomial order capable of producing the observed shape. Furthermore, the qualitative phenomena we report — popping events, dynamic instability, and stochastic heterogeneity — cannot arise from any monotonic force–velocity relation, as discussed in the section on the non-monotonic instability. With 10 parameters covering complex contractile dynamics at the individual sarcomere and myofibril level across different substrate stiffnesses, the present model is parsimonious within the family of polynomial force–velocity ansätze; we have not exhaustively searched alternative non-polynomial functional families, but any such alternative would still need to reproduce the same non-monotonic shape that the data require.

      (8) Lines 497–507 in the Discussion. On reflection, we feel these lines provide useful context for the broader interpretation and would prefer to retain them.

      (9) Line 331 — motivation of Eq. (3). We have added citations to prior work motivating this form of the equation for the broader readership.

      (10) Line 427 — "scaled". Corrected.

      Reviewer #3 (Recommendations for the authors):

      We thank the reviewer for the recommendation of a theoretical appendix. The full model code, with the formulation and implementation documented in detail, is publicly available in our GitHub repository accompanying the paper, which we believe provides a complete reference for readers wishing to explore the model further. We therefore feel an additional appendix is not necessary within the scope of this revision.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the Editors for the positive assessment on our manuscript. We also thank the Reviewers for their positive remarks and constructive comments. Based on the Reviewers’ feedback, we have conducted additional experiments and provided supporting data to address Reviewers’ comments. Particularly, we provided quantitative measurement for rotational polarity of ependymal cells in Agbl5<sup>M1/M1</sup> mutants and assessed the microtubule polarization. We quantified the intensity of apical actin network in ependymal cells to strength the role of CCP5 in organizing actin network. Using scanning electron microscopy, we demonstrated the affected polarity of trachea multicilia in Agbl5<sup>M1/M1</sup>. We co-immunostained ependymal cilia with GT335 and acetylated tubulin to address the effects on their length in cilia in the mutant. We assessed the presence and length of primary cilia in ependymal cell progenitors to identify their potential contribution to the defective polarity in Agbl5<sup>M1/M1</sup> ependymal cells. We feel that these revisions have much strengthened this MS.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Dad et al. explored the roles of cytosolic carboxypeptidase 5(CCP5)in the development of ependymal multicilia in the brain. CCP family are erasers of polyglutamylation of ciliary-axoneme microtubules. The authors generated a new mutant mouse of Agbl5 gene, which encodes CCP5, with deletion of its N-terminus and partial carboxypeptidase (CP) domain (named AGBL5M1/M1).

      Strengths:

      The mutant mice revealed lethal hydrocephalus due to degeneration of ependymal multicilia. Interestingly, this is in contrast with the phenotype of Agbl5 mutants with disruption solely in the CP domain of CCP5 (named AGBL5M2/M2) that did not develop hydrocephalus despite increased glutamylation levels in ependymal cilia as observed for AGBL5M1/M1 mutants. The study has been well-performed and the findings suggest a unique function of the N-domain of CCP5 in ependymal multicilia stability.

      Weaknesses:

      The content of this article is relatively descriptive and lacks molecular insights.

      We thank the Reviewer’s positive comments. To address the molecular insights of the dysregulated planar cell polarity (PCP) in Agbl5<sup>M1/M1</sup> ependyma, we have conducted additional experiments to assess the microtubule polarization in ependymal cells (Figure 7O-P). We quantified the intensity of actin networks around BB patches to better understand how it is affected in the ependyma of the mutants and contributes to the dispersion of BBs (Figure 4M-N), (Please see Recommendations for the authors).

      We also assessed trachea multicilia in Agbl5<sup>M1/M1</sup> mutants using SEM and found that the polarity of trachea multicilia was affected as well (Figure S2).

      Reviewer #2 (Public review):

      Summary:

      This study analyzed the consequences of Agbl5 mutation on ependymal cell development and function. The authors first characterize their mutant mouse line reporting a reduced lifespand and severe hydrocephalus. Next, they report a defect in ependymal cell cilia number and motility. They provide evidence for impaired basal body organisation and cilia glutamylation.

      Strengths:

      Description of a mutant mouse which implicates Cytosolic Carboxypeptidase 5 (the product of Agbl5 gene) for proper ependymal cells.

      Weaknesses:

      Description of phenotype is incomplete:

      We thank the Reviewer’s constructive comments. We have performed additional quantitative analysis of the phenotypes in Agbl5<sup>M1/M1</sup> that we feel strengthen this study.

      Figure 3G - the sequence from the movie is not really informative. Providing beating frequencies as quantification of the data would be more informative.

      We have provided the beating frequency as well as the mean vector length of cilia beating directions (that reflects the coordination of cilia) in Figure 3H and 3I respectively in the revised manuscript.

      Figure 3 - the quantification of actin network would strengthen the message.

      We agree with the Reviewers. We have quantified the total intensity of actin around BBs and the actin intensity normalized to signals of the BB marker (CEP164). The data have been provided in Figure 4M and 4N respectively. The quantitative analysis showed that both the total intensity of apical actin network and the intensity of F-actin per BB are reduced in Agbl5<sup>M1/M1</sup> ependymal cells compared to that in wild-type mice, suggesting that CCP5 is involved in organizing actin network around BB. This analysis certainly improves the clarity of this message.

      Lines 219 -220 - the authors conclude «Taken together, in Agbl5M1/M1 ependymal cells, the expression of genes promoting multiciliogenesis were not impaired but certain proteins associated with differentiated ependymal cells are not properly expressed». However, they do not assess gene but protein expression (IF). In addition, their quantification shows differences in the number of FoxJ1 positive cells which indeed is an impaired expression.

      We will clarify this statement and emphasize the number of FoxJ1-positive cells.

      Microtubules are involved in the local organization of ciliary basal bodies (see Werner et al., Vladar et al.,2011; Boutin et al., 2014). It would be interesting for the authors to check whether the subapical network of microtubules is glutamylated or not during ependymal cell differentiation and how this network is affected in their mutants.

      We thank the Reviewer’s constructive comments. We conducted an immunostaining on whole-mount lateral walls of lateral ventricles for GT335 and Centrin1, the position of the latter being used to localize the subapical layer. While the GT335 signal in multicilia is increased in Agbl5<sup>M1/M1</sup> ependyma (Figure S8E), its signals underneath BBs are not much different between the mutant and wild-type (Please see Figure S8C, D, G, H).

      Showing the data mentioned in the discussion on Cep110 would be a nice addition to the paper.

      These data have been provided in Supplementary Figure S9.

      Line 354: "The latter serves as a component of tissue polarity that is required for asymmetric PCP protein localization in each cell (Boutin et al., 2014; Vladar et al., 2012)." The cited reference did not demonstrate that this microtubule network is required for asymmetric PCP localization.

      We thank the Reviewer for critical reading. The cited reference (Bountin et al., 2014) has been removed.

      Reviewer #3 (Public review):

      Summary:

      The authors developed a new Agbl5 KO allele, extending the deletion to the N-terminus of CCP5 to explore its function in mouse ependymal cells.

      Strengths:

      They show that the KO mice exhibit severe hydrocephalus due to disorganized and mislocated basal bodies. Additionally, they present evidence of both impaired beating coordination and a reduction in ciliary beating.

      Weaknesses:

      The manuscript is well-written but lacks specific interpretations of the results presented. Further experiments are needed to be fully convincing.

      We thank the Reviewer’s comments. We have performed further analysis and conducted additional experiments to strengthen this study.

      (1) We have quantified the intensity of actin staining around BB patches and its intensity relative to the number of BBs to assess to which extent the actin networks in Agbl5<sup>M1/M1</sup> ependymal cells are affected (please refer to the above response to the comments of Reviewer 2#). The results were shown in Figure 4M-N.

      (2) We Co-stained tdTomato with an ependymal cell-specific markers to strengthen the expression of Agbl5 in ependymal cells (please see Figure 6C-E).

      (3) We have conducted co-immunostaining of GT335 and Ac-Tub and compared the length of their signals in ependymal multicilia between WT and Agbl5<sup>M1/M1</sup> mice (please see Figure 6O, P, R, S).

      (4) We quantified the area of ependymal cells in the wild-type and Agbl5<sup>M1/M1</sup> mice. Indeed, the area of ependymal cells is increased in the mutants. However, the primary cilia are present in the ependymal cell progenitors of Agbl5<sup>M1/M1</sup> mice and have similar length with that in the wild-type (Please see Figure 7M, N and our response to this point below).

      (5) We performed additional analysis to address the affected rotational polarity in the Agbl5<sup>M1/M1</sup> mutant mice (please see Figure 3I, Figure 7E).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors showed that the actin networks were severely affected, leading to impaired stability of basal bodies and that the intensity and length of acetylated tubulin signal in the multicilia were dramatically reduced in AGBL5M1/M1mutant mice (Figures 3 and 5). Data also suggested the dysregulation of planar cell polarity. Are expression and localization of other planar cell polarity proteins such as tyrosinated tubulin and Fzd6 affected in mutant mice?

      We thank the Reviewer’s recommendations. We have assessed the expression of tyrosinated tubulins and found they are similarly polarized in ependymal cells from wild-type and Agbl5<sup>M1/M1</sup> mice. The results are presented in Figure 7O, P in the revised MS. We also tried to assess the expression of Fzd6. However, with the antibody we tested, Fzd6 signals were not convincing. Therefore, we prefer to not showing the results and drawing a conclusion on it.

      (2) The phenotype of multiciliated cells in tracheas should also be examined in mutant mice. It is important to elucidate whether AGBL5 commonly functions in multiciliated cells of other organs.

      We thank the Reviewer’s suggestion. We have assessed the multicilia in the tracheas of P30 mice using scanning electron microscopy. Indeed, unlike the multicilia in wild-type mice that orientate to the same direction, those in the tracheas of Agbl5<sup>M1/M1</sup> mice often radiate to different directions in individual cells (Figure S2). Therefore, Agbl5 appears commonly involved in the alignment of multicilia.

      (3) According to Figure 1B, AGBL5 is highly expressed in the brain. Which cells in the brain express it besides ependymal cells?

      Based on the localization of tdTomato tracer engineered in Agbl5 mutant alleles (Figure 5B), Agbl5 is broadly expressed in the brain, including most if not all neurons, but its expression is much weaker in the subventricular zone (Please see Figure 5B). We clarified this in the revised MS.

      (4) From a mechanistic point of view, it is necessary to identify binding proteins with the N-domain of AGBL5 and perform functional analyses.

      We agree with the Reviewer. We feel that identification of the binding partners of CCP5 N-domain and functional analysis may be more suitable to go along with other mechanistic analysis on the function of CCP5 in ependymal cell polarities in our future study.

      Reviewer #2 (Recommendations for the authors):

      (1) Movie 3: The authors could comment on beating direction that seems impaired at the cell scale here, analysis of rotational polarity would be a plus.

      We thank the reviewer’s recommendation. We have analyzed the beating directions of cilia in individual cells and presented their consistency in each cell using mean vector length. These results indeed demonstrated defective rotational polarity in the cell level in Agbl5<sup>M1/M1</sup> mice (please refer to Figure 3I). We also analyzed the beating directions of ependymal multicilia in earlier stage in tissue level (Figure 7E). The mean vector length of cilia beating direction in Agbl5<sup>M1/M1</sup> mice is significantly reduced compared to that in wild-type, suggesting an aberrant rotational polarity in the tissue level in the mutant (Figure 7E).

      (2) Line 166 : ref to Werner et al., 2011 is not correct (no ependymal cells in that paper).

      We thank the reviewer’s critical reading. This reference has been removed.

      (3) Figure S4: B and D look similar picture to me same for C and F.

      We apologize for using the wrong images in this Figure. It has been corrected (Revised Figure S5).

      (4) Line 328: "Therefore, CCP5 apparently contributes to the establishment of both translational and tissue polarities in ependymal cells." Should be rephrased since translational polarity is also a tissue-level parameter which is the coordinated positioning of the ciliary patch. Cf Mirzadeh et al., 2010; Boutin et al., 2014.

      We thank the Reviewer’s comments. The sentence has been rephrased. This concept has been clarified where else needed in the revised manuscript. 

      (5) Line 348: "Planar cell polarity (PCP) pathway is essential for the establishment of rotational and tissue polarities in ependymal cells" Rotational polarity also has a tissular component (ie coordination of beating direction across tissue which is reflected by coordination of basal body polarities across tissue).

      We thank the Reviewer’s comments. We have clarified this point in the revised MS.

      (6) Incomplete bibliography citation (ie Walentek et al. without date).

      We thank the Reviewer’s critical reading. This bibliography citation has been fixed.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 3: The authors assert that the mutant's apical actin networks are significantly disrupted. However, the cell shown in Figure 3Q-R exhibits less compact centrioles than the controls, which could account for the reduction in phalloidin staining. Because centriole dispersion is variable in the mutant, quantifying actin staining in representative cells would be necessary to support such a statement.

      We thank the Reviewer’s comments. To address this concern, we have quantified the total intensity of actin network around BBs as well as the intensity of F-actin signals normalized to the level of immunosignals of BBs ((revised Figure 4M, N) please also refer to our response to Reviewer 1#). The results indicated the intensity of actin signal per BB is reduced in the mutant compared to that of wild-type mice. We feel that this analysis strengthened our statement.

      (2) Figures S3 and 4A-B show that the authors examine tdT expression to show that Agbl5 is expressed in ependymal cells but not in the SVZ. However, the tdT signal intensity is very low, and cells are very dense in this brain region. Double staining with specific markers of ependymal and/or SVZ cells would help convince readers that tdT is not expressed in SVZ cells.

      We agree with the Reviewer that the intensity of tdT signal is low, but broadly detectable in brain. Compared with its expression in ependymal cells, that in SVZ is much lower if any (Figure 4B’). To further confirm the identity of tdT-positive cells along the surface of ventricles, we have co-stained the brain sections of Agbl5<sup>WT/M1</sup> mice for tdT and S100b, a marker of mature ependymal cells (Figure 5C-E). The signal of tdt is colocalized with that of S100b and is much lower in cell layers next to S100b-positive cells.

      (3) Figure 4C-D and S4: The authors demonstrate that the number of FoxJ1+ cells per section increases at P7 (4C-E), while the number of S100β+ cells per mm decreases. Quantifications should be carried out in a similar manner to ensure comparability (number of positive cells per mm). Additionally, it remains unclear how to interpret these results, as S100β and FoxJ1 are two markers of differentiated cells, yet they exhibit opposite trends compared to controls. Is this a direct or indirect effect of Agbl5 mutation? The increase in the number of FoxJ1+ cells is particularly surprising given that the number of GT335 multicilia per mm remains unchanged (Figure 5).

      We agree with the Reviewer that quantifications should be carried out in a similar manner. In the revised MS, the quantification of Foxj1-positive cells is presented in number per mm (Figure 5I). To be noted, the expression of Foxj1 was assessed at P7 when ependymal cells are differentiating. while the expression of S100β was assessed at P17 when ependymal cells are supposed to be fully mature. Although S100b is used as a marker of mature ependymal cells, given its unclear function, we removed the results of S100b-positiving cell counting to avoid confusion in the revised manuscript.

      (4) Figure 5: In this figure, the authors analyze the labeling obtained with GT335, Acetylated Tubulin, and Arl13b antibodies. They show that the area of the cilium labeled by GT335 has increased, while the area labeled by the Acetylated Tubulin antibody has decreased in the knockout (KO) compared to the control. However, the length of the cilia observed through labeling with the Arl13b antibody remains unchanged. These observations are intriguing, but the low-magnification images in Figure 4 do not allow for the differences in ciliary axoneme labeling to be seen. Double GT335/AcTub labeling and higher magnifications are necessary for improved visualization of the differences in labeling along the axonemes.

      We thank the Reviewer comments. We have co-stained the cilia with GT335 and Ac-Tub antibodies, re-quantified cilia length labeled with respective antibodies and provided high magnification images. Please see the revised Figure 6O,P,R,S.

      (5) Figure 6: An analysis of ciliary beats using a high-speed camera shows no difference in ciliary beat frequency between the control and KO groups. At least, 3 animals should be analyzed. According to Figure 5, these findings indicate that the decrease in ciliary acetylation and the increase in ciliary glutamylation do not affect the beat frequency; instead, they disrupt the orientation of the beats. While these results are intriguing, they require further confirmation. Analyzing ciliary beats with a high-speed camera is informative, but at least three animals per genotype should be examined to ensure rigor. Furthermore, if the coordination of ciliary beats is impaired within the cells, this should be validated by double-labeling centrioles and basal feet to demonstrate that the orientation of cilia within the cells is abnormal.

      We thank the Reviewer’s comments. Sections shown in Figure 5 (currently Figure 6) are from P7 mice, while the ciliary beating analysis shown in Figure 6 (currently Figure 7) is from P15 mice. As the PTM changes in cilia were also observed in Agbl5<sup>M2/M2</sup>, we don’t think this is the cause that disrupts the orientation of the beats. The rotational polarity of Agbl5<sup>M1/M1</sup> ependymal cells is affected. Please refer to the analysis in Figure 3I and Figure 7E in the revised manuscript.

      (6) Figure 6F-G: β-Catenin labeling reveals cells of varying sizes in the KO. This phenotype is typical of ciliary mutants that lack primary cilia (Mirzadeh et al., 2010). Hence, it is essential to examine the mutation's impact on the presence, length, and positioning of the primary cilium in ependymal cell progenitors.

      We thank the Reviewer’s constructive comments. We assessed the area of ependymal cells labeled with β-Catenin. Indeed, the ependymal cells in the mutant showed larger area than that of wild-type. The ratio of the area of BB patch over that of cell surface is reduced (please see Figure 7O, P in the revised manuscript). However, primary cilia are present in ependymal cell progenitors in the mutant and exhibit comparable length with those in the wild-type (Figure S8). Due to some technique problems, we were unable to get convincing results from whole-mount ventricle walls for the primary cilium positioning at this time. We speculate that the localization of certain sensory proteins in primary cilia or the positioning of primary cilia might be affected in Agbl5<sup>M1/M1</sup> mice. We discussed this possibility and will certainly systemically assess this intriguing aspect in our future investigation.

      (7) Given the regular beating frequency in the KO at P15, how do the authors explain the complete absence of ciliary beating in the adult? How many animals were analyzed? One would expect ciliary beating to remain unaffected as it was at P15 unless the cilia structure was specifically altered at the adult stage. Is that the case?

      We thank the Reviewer’s critical questions. We do think that the ciliary structure of Agbl5<sup>M1/M1</sup> ependymal cells is likely altered during aging. Given that only Agbl5<sup>M1/M1</sup> but not Agbl5<sup>M2/M2</sup> mice develop hydrocephalus, we speculate the N-domain of CCP5 may contribute to the integrity of ependymal multicilia. We have added this in the Discussion section. For each genotype, 2 mice were analyzed.

      (8) Line 264 of the manuscript: replace intercellular with intracellular.

      It has been revised.

      (9) Indicate the number of animals analyzed in each experiment

      It has been included in figure legends.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Gruskin and colleagues use twin data from a movie-watching fMRI paradigm to show how genetic control of cortical function intersects with the processing of naturalistic audiovisual stimuli. They use hyperalignment to dissect heritability into the components that can be explained by local differences in cortical-functional topography and those that cannot. They show that heritability is strongest at slower-evolving neural time scales and is more evident in functional connectivity estimates than in response time series.

      Strengths:

      This is a very thorough paper that tackles this question from several different angles. I very much appreciate the use of hyperalignment to factor out topographic differences, and I found the relationship between heritability and neural time scales very interesting. The writing is clear, and the results are compelling.

      We thank Reviewer 1 for their kind words and enthusiastic support of our manuscript.

      Weaknesses:

      The only "weaknesses" I identified were some points where I think the methods, interpretation, or visualization could be clarified.

      (1) On page 16, the authors compare heritability in functional connectivity (FC) and response time series, and find that the heritability effect is larger in FC. In general, I agree with your diagnosis that this is in large part due to the fact that FC captures the covariance structure across parcels, whereas response time series only diverge in terms of univariate time-point-by-time-point differences. Another important factor here is that (within-subject) FC can be driven by intrinsic fluctuations that occur with idiosyncratic timing across subjects and are unrelated to the stimulus (whereas time-locked metrics like ISC and timeseries differences cannot, by definition). This makes me wonder how this connectivity result would change if the authors used inter-subject functional connectivity (ISFC) analysis to specifically isolate the stimulus-driven components of functional connectivity (Simony et al., 2016). This, to me, would provide a closer comparison to the ISC and response time series results, and could allow the authors to quantify how much of the heritability in FC is intrinsic versus stimulus-driven. I'm not asking that the authors actually perform this analysis, as I don't think it's critical for the message of the manuscript, but it could be an interesting future direction. As the authors discuss on page 17, I also suspect there's something fundamentally shared between response time series and connectivity as they relate to functional topography (Busch et al., 2021) that drives part of the heritability effect.

      We agree that investigating the heritability of ISFC (or stimulus-driven functional connectivity) would make for a very interesting future direction. Ultimately, we chose to analyze FC (vs. ISFC) profiles to allow for direct comparison with the sizable existing literature on the heritability of FC (such as in our Movie vs. Rest FC analysis) and decided to refrain from analyzing ISFC data in order to keep the present manuscript focused. ISFC analysis of this dataset will be a focus of future work.

      (2) The observation that regions with intermediate ISC have the largest differences between MZ, DZ, and UR is very interesting, but it's kind of hard to see in Figure 1B. Is there any other way to plot this that might make the effect more obvious? For example, I could imagine three scatter plots where the x- and y-axes are, e.g., MZ ISC and UR ISC, and each data point is a parcel. In this kind of plot, I would expect to see the middle values lifted visibly off the diagonal/unity line toward MZ. The authors could even color the data points according to networks, like in Figure 3C. (They also might not need to scale the ISC axis all the way to r = 1, which would make the differences more visible.)

      We thank R1 for this helpful suggestion- we originally set the y-axis limits to r = 1 in order to facilitate comparison between ISC (Fig. 1B) and FC profile (Fig. 6B) similarity, but we agree that this renders the group differences harder to discern and have updated the plot accordingly (along with thicker lines to enhance readability). We prefer to keep the line plots in the main body as they allow for direct comparison of all three groups on the same plot, but we have included the scatter plot version in Fig. S2 for those who are interested.

      (3) On page 9, if I understand correctly, the authors regress the vector of ISC values across parcels out of the vector of heritability values across parcels, and then plot the residual heritability values. Do they center the heritability values (or include some kind of intercept) in the process? I'm trying to understand why the heritability values go from all positive (Figure 2A) to roughly balanced between positive and negative (Figure 2B). Important question for me: How should we interpret negative values in this plot? Can the authors explain this explicitly in the text? (I also wonder if there's a more intuitive way to control for ISC. For example, instead of regressing out ISC at the parcel/map level, could they go into a single parcel and then regress the subject-level pairwise ISC values out when computing the heritability score?).

      We indeed included an intercept in this model using MATLAB’s fitlm function. This means that the model estimates the best-fitting line of the following form: heritability<sub>i</sub>=β0+β1ISC<sub>i</sub> +ε<sub>i</sub>. We agree that the interpretation of these ε<sub>i</sub> values and alternative approaches to controlling for ISC should be clarified. As such, we have added the following passages to the text:

      Methods: “Because the heritability of ISC is constrained by the degree of synchronization in a given area, we also sought to identify areas in which BOLD time courses were more/less heritable than would be expected based on ISC alone by fitting a linear model of the form heritability<sub>i</sub>=β0+β1ISC<sub>i</sub>+ε<sub>i</sub> and plotting the residuals. Regarding alternative approaches to controlling for ISC, although the heritability model introduced by Ge et al. allows for the inclusion of covariates defined at the subject level (e.g., age), it does not allow for covariates that are defined at the dyad level (e.g., pairwise ISC).”

      Results: “Here, negative values in the residual map indicate parcels where heritability is lower than expected based on ISC, while positive values indicate higher-than expected heritability.”

      (4) On page 4 (line 155), the authors say "we shuffled dyad labels"- is this equivalent to shuffling rows and columns of the pairwise subject-by-subject matrix combined across groups? I'm trying to make sure their approach here is consistent with recommendations by Chen et al., 2016. Is this the same kind of shuffling used for the kinship matrix mentioned in line 189?

      Briefly, shuffling the kinship matrix involved permuting the rows and columns of the matrix in the same manner (also known as the quadratic assignment procedure), whereas shuffling the dyad labels involved random permutations of the three group labels (MZ, DZ, unrelated), which could not be done through matrix operations as the age- and gender matching precluded the use of a complete similarity matrix. However, given concerns raised by Reviewer 2, we have removed our significance claims from this (and similar) sections, which we discuss in more detail in response to Reviewer 2’s weakness A.

      (5) I found panel A in Figure 4 to be a little bit misleading because their parcel-wise approach to hyperalignment won't actually resolve topographic idiosyncrasies across a large cortical distance like what's depicted in the illustration (at the scale of the parcels they are performing hyperalignment within). Maybe just move the green and purple brain areas a bit closer to each other so they could feasibly be "aligned" within a large parcel. Worth keeping in mind when writing that hyperalignment is also not actually going to yield a one-to-one mapping of functionally homologous voxels across individuals: it's effectively going to model any given voxel time series as a linear combination of time series across other voxels in the parcel.

      We agree that our efforts to present a simplified depiction of hyperalignment may mislead less familiar readers and have amended Fig. 4A according to this suggestion. We have also added text to the methods section (below) to clarify that the outputs of hyperalignment are time series that reflect linear combinations of other voxels’ time series from that parcel.

      “This approach independently transforms each subject's data within discrete anatomical parcels into the common space, yielding functionally aligned vertex time series that are calculated as weighted linear combinations of the original time series from all other vertices within that same parcel for that subject.”

      (6) I believe the subjects watched all different movies across the two days, however, for a moment I was wondering "are Day 1 and Day 2 repetitions of the same movies?" Given that Day 1 and Day 2 are an organizational feature of several figures, it might be worth making this very explicit in the Methods and reminding the reader in the Results section.

      We agree that this would be helpful and have added the following text to the relevant sections:

      “All clips were only viewed once by each subject, with the exception of the brief montage which was included at the end of each of the four runs for test-retest purposes.”

      “To characterize the heritability of brain responses to complex stimuli, we used 7T fMRI data from 178 HCP Young Adult subjects acquired across two days (using two largely non-overlapping sets of movie stimuli, see Methods)…”

      References:

      Busch, E. L., Slipski, L., Feilong, M., Guntupalli, J. S., di Oleggio Castello, M. V., Huckins, J. F., Nastase, S. A., Gobbini, M. I., Wager, T. D., & Haxby, J. V. (2021). Hybrid hyperalignment: a single high-dimensional model of shared information embedded in cortical patterns of response and functional connectivity. NeuroImage, 233, 117975. https://doi.org/10.1016/j.neuroimage.2021.117975

      Chen, G., Shin, Y. W., Taylor, P. A., Glen, D. R., Reynolds, R. C., Israel, R. B., & Cox, R. W. (2016). Untangling the relatedness among correlations, part I: nonparametric approaches to inter-subject correlation analysis at the group level. NeuroImage, 142, 248259. https://doi.org/10.1016/j.neuroimage.2016.05.023

      Simony, E., Honey, C. J., Chen, J., Lositsky, O., Yeshurun, Y., Wiesel, A., & Hasson, U. (2016). Dynamic reconfiguration of the default mode network during narrative comprehension. Nature Communications, 7, 12141. https://doi.org/10.1038/ncomms12141

      Reviewer #2 (Public review):

      Summary:

      The authors attempt to estimate the heritability of brain activity evoked from a naturalistic fMRI paradigm. No new data were collected; the authors analyzed the publicly available and well-known data from the Human Connectome Project. The paper has 3 main pieces, as described in the Abstract:

      (1) Heritability of movie-evoked brain activity and connectivity patterns across the cortex.

      (2) Decomposition of this heritability into genetic similarity in "where" vs. "how" sensory information is processed.

      (3) Heritability of brain activity patterns, as partially explained by the heritability of neural timescales.

      Strengths:

      The authors investigate a very relevant topic that concerns how heritable patterns of brain activity among individuals subjected to the same kind of naturalistic stimulation are. Notably, the authors complement their analysis of movie-watching data with resting-state data.

      Weaknesses:

      The paper has numerous problems, most of which stem from the statistical analyses. I also note the lack of mapping between the subsections within the Methods section and the subsections within the Results section. We can only assess results after understanding and confirming the methods are valid; here, however, Methods and Results, as written, are not aligned, so we can't always be sure which results are coming from which analysis.

      (A) Intersubject correlation (ISC) (section that starts from line 143): "We used nonparametric permutation testing to quantify average differences in ISC for each parcel in the Schaefer 400 atlas for each day of data collection across three groups: MZ dyads, DZ dyads, and unrelated (UR) dyads, where all UR dyads were matched for gender and age in years." ... "some participants contributed to ISC values for multiple dyads (thus violating independence assumptions)"

      This is an indirect attempt to demonstrate heritability. And it's also incorrect since, as the authors themselves point out, some subjects contribute to more than one dyad.

      Permutation tests don't quantify "average differences", they provide a measure of evidence about whether differences observed are sufficient to reject a hypothesis of no difference.

      Matching subjects is also incorrect as it artificially alters the sample; covarying for age and sex, as done in standard analyses of heritability, would have been appropriate.

      It isn't clear why the authors went through the trouble of implementing their own nonparametric test if HCP recommends using PALM, which already contains the validated and documented methods for permutation tests developed precisely for HCP data.

      The results from this analysis, in their current form, are likely incorrect.

      We appreciate that permutation tests do not quantify average differences and intended to write “We used non-parametric permutation testing to quantify [the significance of] average differences…”. Our intention with this analysis was not to demonstrate heritability, but rather to quantify group differences in ISC in a manner that is interpretable for readers who are unfamiliar with h<sup>2</sup> (e.g., “identical twins’ BOLD time courses were 59% more similar than those from pairs of unrelated individuals”) and motivate the formal heritability analysis used later in the paper. Indeed, all of the heritability analyses in this paper leveraged a validated multidimensional heritability method first introduced by Ge et al. (2016) and used by many other investigators since then. Furthermore, we covaried for age and sex at the subject level in all our heritability analyses, and always tested the significance of these heritability values using a validated permutation procedure (the quadratic assignment procedure; Hubert & Schultz, 1976) that respects the non-independence of dyadic data.

      Regarding the shuffling procedure used for Figure 1, while PALM is the standard for univariate, subject-level GLMs in the HCP pipeline and can accommodate nested designs (i.e., subjects within families), it is not designed to handle the unique relational dependencies of dyadic ISC analysis (i.e., the same subject contributing to multiple dyads). Although the element-wise resampling approach was the most appropriate approach available, it is known to inflate the false positive rate (Chen et al., 2016; doi:10.1016/j.neuroimage.2016.05.023); given that this analysis was simply meant to motivate our later hypothesis testing heritability analyses, we have removed significance claims from this section of the manuscript. Still, we emphasize that this has no bearing on the validity of our conclusions which were supported by our formal heritability analyses; throughout our paper we have correctly used the appropriate methods to back the stated claims.

      (B) Functional connectivity (FC) (section that starts from line 159): Here the authors compute two 400x400 FC matrix for each subject, one for rest, one for movie-watching, then correlate the correlations within each dyad, then compared the average correlation of correlations for MZ, DZ, and UR. In addition to the same problems as the previous analysis, here it is not clear what is meant by "averaging correlations [...] within a network combination". What is a "network combination"? Further, to average correlations, they need to be r-to-z transformed first. As with the above, the results from this analysis in its current form are likely incorrect.

      We regret that R2 had difficulty understanding our analysis and have added the following text to the relevant Methods section to clarify our approach:

      “For example, there are 16 parcels in the Kong et al. Auditory network and 17 parcels in the Language network, so the FC profile for a given subject’s Auditory-Language network combination consists of the (16 * 17 =) 272 correlation coefficients between all unique pairs of one parcel from each network.”

      As we stated in the previous Methods paragraph, “All Pearson r values in this and all other analyses were Fisher z-transformed before averaging (and converted back to Pearson r for visualization)”. Thus, contrary to the reviewer’s assertion, these analyses were performed correctly. Once again, we emphasize that this analysis was not intended to demonstrate heritability, but rather to describe group differences in FC in familiar units.

      (C) ISC and FC profile heritability analyses (section that starts from line 175): Here, the authors use first a valid method remarkably similar to the old Haseman-Elston approach to compute heritability, complemented by a permutation test. That is fine. But then they proceed with two novel, ill-described, and likely invalid methods to (1) "compare the heritability of movie and rest FC profiles" and (2) to "determine the sample size necessary for stable multidimensional heritability results". For (1), they permute, seemingly under the alternative, rest and movie-watching timeseries, and (2), by dropping subjects and estimating changes in the distribution.

      The (1) might be correct, but there are items that are not clearly described, so the reader cannot be sure of what was done. What are the "153 unique network combinations"? Why do the authors separate by day here, whereas the previous analyses concatenated both days? Were the correlations r-to-z transformed before averaging?

      The (2) is also not well described, and in any case, power can be computed analytically; it isn't clear why the authors needed to resort to this ad hoc approach, the validity of which is unknown. If the issue is the possibility that the multidimensional phenotypic correlation matrix is rank-deficient, it suffices that there are more independent measurements per subject than the number of subjects.

      Regarding (1), we have clarified in section 2.6 that the 153 unique network combinations reflect each unique pair of 17 Kong networks. All of our analyses, including this one, were performed separately for each day of data collection, as we state throughout the paper and visualize in our figures (although we acknowledge that, on some occasions, we [conservatively] performed FDR-correction on a combined set of p-values, as discussed in our response to K). Given that the null hypothesis for this analysis is that rest FC and movie FC are equally heritable, we are not sure why permuting rest and movie FC matrices would be invalid. All Pearson r values were z-transformed before averaging, as we stated in our paper.

      Regarding (2), we included this analysis in response to editorial concerns that our heritability analyses were not sufficiently powered, and we chose this approach because it serves as a simple way to demonstrate the stability of our results at various sample sizes whose validity is self-evident. Furthermore, this sort of subsampling approach has been used many times before in our field (e.g., Marek et al., 2022) and others (e.g., Manyara et al., 2024) to demonstrate the sample-size dependence and stability of statistical effects. We have added text explaining this to the relevant Methods section (2.6).

      (D) Frequency-dependent ISC heritability analysis (from line 216): Here, the authors decompose the timeseries into frequency bands, then repeat earlier analyses, thus bringing here the same earlier problems and questions of non-exchangability in the permutations given the dyads pattern, r-z transforms, and sex/age covariates.

      We did not use dyadic permutation testing for any of the frequency-dependent ISC analyses; rather, we used the jackknife SEMs to compare heritability across frequency bands and have added an explicit description of this to section 2.7. We have addressed the r-z transform and covariate concerns in previous comments.

      (E) FC strength heritability analysis (from line 236): Here, the authors use the univariate FC to compute heritability using valid and well-established methods as implemented in SOLAR. There is no "linkage" being done here (thus, the statement in line 238 is incorrect in this application. SOLAR already produces SEs, so it's unclear why the authors went out of their way to obtain jackknife estimates. If the issue is non-normality, I note that the assumption of normality is present already at the stage in which parameters themselves are estimated, not just the standard errors; for non-normal data, a rank-based inversenormal transformation could have been used. Moreover, typically, r-to-z transformed values tend to be fairly normally distributed. So, while the heritabilities might be correct, the standard errors may not be (the authors don't demonstrate that their jackknife SE estimator is valid). The comparison of h2 between dyads raises the same questions about permutations, age/sex covariates, and r-z transforms as above.

      We used jackknife SEs for these analyses to maintain consistency with the multidimensional heritability package used here, which only outputs jackknife SEs. We note that this jackknife approach (and the corresponding multidimensional heritability analysis) was detailed in prior work (Anderson et al., 2021), and that the leave-one-family-out jackknife has a long history of being used to estimate SEs in heritability studies, especially when working with smaller samples (Knapp et al., 1989). We are also not sure what “the comparison of h2 between dyads” means- heritability cannot be compared “between” dyads; rather, it is defined across dyads.

      (F) Hyperalignment (from line 245): It isn't clear at this point in the manuscript in what way hyperalignment would help to decompose heritability in "where vs. how" (from the Abstract). That information and references are only described much later, from around line 459. The description itself provides no references, and one cannot even try to reproduce what is described here in the Methods section. Regardless, it isn't entirely clear why this analysis was done: by matching functional areas, all heritabilities are going to be reduced because there will be less variance between subjects. Perhaps studying the parameters that drive the alignment (akin to what is done in tensor-based and deformation-based morphometry) could have been more informative. Plus, the alignment process itself may introduce errors, which could also reduce heritability. This could be an alternative explanation for the reduced heritability after hyperalignment and should be discussed. An investigation of hyperaligment parameters, their heritability, and their co-heritability with the BOLD-phenotypes can inform on this.

      To help set up our hyperalignment analyses, we have added text to the introduction explaining how hyperalignment would help to decompose heritability. The description in the Methods section included a reference to Bazeille et al., 2021, in which the hyperalignment method used here is discussed in detail. Still, we have added citations to additional papers (also cited in the Bazeille et al. paper, and elsewhere in our paper) in case that might be helpful. We note that it is not the case that all heritabilities were reduced by hyperalignment- as can be seen in Figs. 4D, 8A, and S15, hyperalignment did increase heritability in some voxels and network combinations. This would be expected under the alternative (albeit unlikely) hypothesis that functional topographies are not heritable, such that topographic variation between related individuals would obscure similarities in their (heritable) topography-independent brain responses. Recognizing that this alternative is unlikely, we believe the main novelty of this analysis comes from the magnitude of the hyperalignment effect (up to 40% of brain-wide heritability) and its spatial pattern (e.g., larger heritability decreases in visual vs. auditory cortex, the opposite of our NT result).

      We agree that we would see lower post-hyperalignment heritability if the alignment process itself introduced errors/noise, but this would be deeply surprising as hyperalignment increases ISC by design (and errors/noise could only decrease ISC). To demonstrate this, we have added Figure S7 which shows that (as expected) ISC across all voxels and subject pairs increases after hyperalignment (and that this increase is larger when hyperalignment is performed in larger parcels). Given that hyperalignment increased ISC, and that it is blind to twin status, we are unsure how it could have introduced errors that would have confounded this result.

      (G) Relationships between parcel area and heritability (from line 270): As under F), how much the results are distorted likely depends on the accuracy of the alignment, and the error variance (vs heritable variance) introduced by this.

      We agree that alignment accuracy could potentially impact parcel-level differences in how much heritability changes following hyperalignment, and we included the frequency dependent h<sup>2</sup><sub>residuals</sub> (controlling for differences in ISC) in Fig. 3 for this reason, as more accurate hyperalignment should result in greater increases in ISC, raising the heritability ceiling. We note that we observe similar relationships between parcel rank and frequency dependent changes in these residualized maps, suggesting that our parcel-level differences are not simply the result of better alignment in more sensory parcels.

      (H) Neural timescale analyses (from line 280): Here, a valid phenotype (NT) is assessed with statistical methods with the same limitations as those previously (exchangability of dyads, age/sex covariates, and r-z transforms). NT values are combined across space and used as covariates in "some multivariate analyses". As a reader, I really wanted to see the results related to NT, something as simple as its heritability, but these aren't clearly shown, only differences between types of dyads.

      We have addressed the exchangeability, covariates, and r-z transform comments above (in A). As we explained for our FC strength analyses, we are underpowered to evaluate the heritability of unidimensional traits (like the heritability of NT magnitude), and the heritability of a closely-related measure (BOLD turnover magnitude) has already been established in a larger sample of HCP subjects (https://doi.org/10.1152/jn.00402.2022). Still, we agree that more results related to the heritability of NTs would be of interest to our readers. As such, we have added an analysis in section 3.4 quantifying the heritability of multivariate NT topographies and used SOLAR to quantify the heritability of NT magnitudes, with the disclaimer that this and similar analyses are underpowered (hence the large difference in day 1 and day 2 heritability effect sizes). We also removed significance claims for the dyadic NT similarity analysis.

      (I) Significance testing for autocorrelated brain maps and FC matrices (from line 310): Here, the authors suddenly bring up something entirely different: reliability of heritability maps, and then never return to the topic of reliability again. As a reader, I find this confusing. In any case, analyses with BrainSMASH with well-behaved, normally distributed data are ok. Whether their data is well behaved or whether they ensured that the data would be well behaved so that BrainSMASH is valid is not described. As to why Spearman correlations are needed here, Mantel tests, or whether the 1000 "surrogate" maps are valid realizations of the data under the null, remains undemonstrated.

      We brought up reliability in this section because we show the reliability of our results across the two days of data collection several times in the paper. R2 is correct to point out that BrainSMASH was validated using normally distributed brain maps, and although some of our brain maps contain normally distributed values, others are right skewed (due largely to the fact that many voxels/parcels exhibit low ISC while visual/auditory areas have very high ISC). In preparing our original manuscript, we visualized BrainSMASH’s variogram outputs for one of the most skewed inputs (vertex-wise BOLD time course heritability) and found that the autocorrelation structures of the empirical and null maps were well-matched. We did not include this in the original manuscript as it is not commonplace in the field to report the variograms, see Author response image 1. Furthermore, our use of Spearman (vs. Pearson) correlations renders these distributional differences less relevant, as the Spearman correlation transforms all inputs to a uniform distribution. To empirically check that these distributional differences do not bias our results, we retested the significance of all brain map associations using the spin test (10.1016/j.neuroimage.2018.05.070), an alternative method that does not assume normally distributed inputs, and obtained identical p-values for all analyses (P<.001 in all cases).

      Author response image 1.

      (J) Global signal was removed, and the authors do not acknowledge that this could be a limitation in their analyses, nor offer a side analysis in which the global signal is preserved.

      Although we agree that GSR is a contentious preprocessing step for certain analyses, it has explicitly been shown to increase ISC signal-to-noise without compromising FC fingerprints (Graff et al., 10.1016/j.dcn.2022.101087), and it is uncommon to perform ISC analyses with and without GSR. Still, we have added additional text to our Methods section explaining our rationale for using GSR and that this could affect our results. We also re-ran our main analysis (BOLD time course heritability) with and without GSR and found that GSR had little impact on our results; we have included this in our manuscript as Fig. S4.

      Specifically, we see that GSR resulted in a slight increase in heritability (average Day 1 h<sup>2</sup> with/without GSR = .064/.060; Day 2: .068/.061) and almost no effect on the spatial pattern of our results (With GSR/without GSR Spearman ρ = .99, P<sub>brainSMASH</sub> < .001 on both Day 1 and Day 2).

      (K) FDR is used to control the error rate, but in many cases, as it's applied to multiple sets of p-values, the amount of false discoveries is only controlled across all tests, but not within each set. The number of errors within any set remains unknown.

      We agree that the FDR usage in our original manuscript was inconsistent, in that for two analyses we FDR-corrected p-values from the two days of data collection together (instead of correcting p-values from each day separately and reporting voxels/parcels/etc. that were significant at q<.05 on both days, as in the rest of our analyses). We note that both approaches are more conservative than reporting significant results at q<.05 separately; regardless, to maintain consistency we have updated all analyses such that FDR correction is always performed separately for each day of data collection.

      (L) Generally, when studying the heritability of a trait, the trait must be defined first. Here, multiple traits are investigated, but are never rigorously defined. Worse, the trait being analyzed changes at every turn.

      Here, we analyze the heritability of movie-evoked BOLD time courses (Figures 1-5) as well as FC profiles (Figures 6-8). We defined FC profiles in our Introduction as an individual’s pattern of pairwise FC strengths (and further detailed how we quantified FC profiles in the relevant Methods section), and believe that “BOLD time course” is a well understood phrase in the field and does not need to be further defined. We also used hyperalignment to decompose the heritability of these traits into topography-dependent and independent portions, and (new to this version) also explicitly quantify the heritability of neural timescales, which we defined as the AUC of the ACF until the first negative ACF value in both the relevant Results and Methods sections.

      To make this clearer, we have modified the last paragraph of our Introduction to begin with:

      In the present work, we address these questions by analyzing 7T fMRI recordings of a twin sample acquired by the Human Connectome Project (Van Essen et al., 2013) to quantify the heritability of two distinct high-dimensional traits—stimulus-evoked BOLD time courses and functional connectivity profiles—across the cortex.

      Reviewer #3 (Public review):

      Strengths:

      It's sort of novel to study the heritability of movie-watching fMRI data. The methodology the authors used in the paper is also supportive of their findings. Figures are nicely organized and plotted. They finally found that sensory processing in the human brain is under genetic control over stable aspects of brain function (here referring to neural timescale and resting state connectivity).

      Weaknesses:

      What I am worried about most is the sample size and interpretation of heritability.

      (1) Figure 1. I assumed that the authors just calculated the ISC within each group (MZ, DZ, and UR). Of course, you can get different variations between each group. Therefore, there is heritability. Why not calculate ISC across the whole sample, then separate MZ, DZ, and UR?

      We believe that this question is getting at the difference between pairwise ISC (i.e., correlating one BOLD time course from one subject with that from another subject) and leave-one-subject-out ISC (i.e., correlating one BOLD time course from one subject with the corresponding average time course across all other subjects). We chose to use the pairwise ISC method because it allows us to capitalize on the information contained in the n<sup>2</sup> pairwise ISC matrix (whereas the other approach averages out meaningful information to yield a n<sup>1</sup> ISC matrix) and leverage a more sophisticated multidimensional heritability approach. Also, the leave-one-subject-out approach introduces additional issues re: handling family-level data (e.g., should we include a subject’s twin in the leave-one-subject-out average? If so, how should we handle subjects who don’t have a twin in the dataset, as averaging data from different numbers of subjects will lead to different ISC magnitudes? etc.).

      (2) Heritability scores in the paper are sort of small. If the sample size is small, please consider p-values, which will tell more about the trustworthiness of your heritability.

      We report p-values for heritability throughout our paper (e.g., stating that BOLD time courses are significantly heritable in 99% of parcels in Figure 2), and we believe that the reliability of our spatial maps across days of data collection (also quantified with p-values) further demonstrates the trustworthiness of our results. Finally, as we demonstrate in Figure S5, our sample size is more than sufficient to reliably detect small effects.

      (3) I don't understand the high-frequency signals in fMRI data. It's always regarded as noise, the band 1 here in particular.

      In addition to driving shared neuronal responses (which are captured in BOLD signal oscillations <.1 Hz or so), movies also elicit shared cardiac, respiratory, and motion responses across participants at higher frequencies. Although we used a relatively conservative denoising approach here, we believe some of these non-neuronal signals are still present in our data; alternatively, it is also possible that these signals reflect “fast” BOLD responses at >.15 Hz (as discussed in 10.1016/j.neuroimage.2021.118658). In any case, the fact that information in this frequency band is considerably less heritable than information in slower frequency bands supports the idea that this band is noisier and suggests that our heritability results are driven by canonical neuronal activity-related BOLD signals.

      (4) The statement "we show that the heritability of brain activity patterns can be partially explained by the heritability of the neural timescale" should come from Figure 5. However, after controlling for NT, the heritability decreased max. 0.025 in temporal areas. I am not sure this change supports the statement. If the visual cortex is outlined, and combining ISC changes in the visual cortex, I think this would somehow be answered. Instead of delta h2, adding a new model h2 would be obvious to the readers.

      Although the decrease of 0.025 is small, we note that this constitutes around ~50% of BOLD time course heritability in some voxels (seen in comparison to Fig. 4C), and the spatial pattern of this result is quite consistent across days of data collection, indicating its reliability. Furthermore, the whole-brain distributions of results shown in Fig. 5B are clearly skewed towards negative values, indicating that controlling for NT partially reduces (or “explains”) BOLD time course heritability. Still, we agree that showing raw h<sup>2</sup> values in addition to the difference maps would be helpful for some readers and have added a corresponding supplementary figure (S12) which shows these.

      (5) Figures 7 and 8, when getting the difference of heritability, please also consider the standard errors of the heritability estimates. Then you can compare across networks/regions.

      We did consider adding standard errors for these heritability estimates, but found that visualizing standard errors for each of the 153 unique network combinations in our heatmaps rendered the visualizations difficult to parse, and given that our hypotheses concerned global (e.g., hyperaligned vs. MSM-aligned) or network-level (e.g., sensory vs. associative) patterns, we focused on calculating standard errors/p-values for these analyses (although we note that dyad-level standard errors can be found in Fig. 6B, where they are clearly marginal compared to the group effects).

      (6) I think movie VS resting state is a really important result in this paper. However, there is almost no discussion. Discussing this part would be more beneficial for understanding the genetic control over the neuron arousal and excitation circuits.

      We agree that this result was relatively under-explored in our Discussion section and have added additional text (lines 851-855) to connect this result to recent work on arousal-dependent uniqueness of FC.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Do the authors have any ideas why we see this hotspot of heritability in pMTG/LOTC? It really jumps out in Figure 1A and Figure 2. The more posterior sensory MT+ area seems to drop when regressing out ISC in Figure 2B, but this pMTG area stays hot. Is there anything special about this kind of multimodal biological motion/action observation / social perception area (Pitcher & Ungerleider, 2021)? I don't think this is necessary to discuss in the manuscript, but I'm curious if the authors have any speculation.

      We are not certain as to why BOLD time courses in this parcel are particularly heritable- although this area is associated with biological motion, that particular function tends to be more right lateralized, and here we see nominally higher heritability in the left hemisphere. Per a Neurosynth review (and consistent with the left lateralization), we believe this may have more to do with speech processing, but a more definitive answer will require further investigation.

      (2) Page 3, line 127: "More information on these clips"-it might be worth saying a little bit more here just to make sure people understand that these are audiovisual clips, they include language, they're long enough to convey meaningful social and narrative information, etc.

      We agree and have added additional details on the clip composition to the relevant methods paragraph.

      (3) Figure 1 caption: can you add a sentence reminding readers what's going on with Day 1 and Day 2?

      We thank R1 for this suggestion and have added a sentence to this effect at this location.

      (4) Page 9, line 379: "although these more associative parcels do not encode a substantial amount of stimulus-specific information"-is this really true? I suspect these association areas still have decent ISCs, even if there are many processing stages downstream of the raw stimulus.

      Although these parcels are not the most synchronized by the stimulus, we agree that it is unfair (and vague) to say that they do not encode a substantial amount of stimulus-specific information. We have edited this sentence to make a more specific claim and highlight the relatively lower ISC in these parcels vs. more unimodal sensory areas.

      (5) Page 9, line 417: Can you unpack a bit more what you mean by "supra-BOLD frequency band"?

      Here, we refer to the fact that BOLD signals resulting from neuronal firing events have frequencies below ~.15 Hz (Josephs and Henson, 1999). We have added additional text and the Josephs and Henson citation to this line to further unpack this point.

      (6) Page 18, line 695: This discussion of how attention and gaze might partly shape response time series reminded me of recent work by Borovska & de Haas (2024)-might be worth citing.

      We are grateful to R1 for alerting us to this very relevant work and have included a reference to it in our discussion.

      (7) Page 19, line 755: I'm not sure I'd describe the hyperalignment results here as a "deleterious effects [on] heritability"-my reading was that hyperalignment allows you to say something more specific about heritability of function by allowing you to effectively factor out heritability effects that reduce to individual differences cortical topography; this seems like a good thing!

      We agree that “deleterious” was a poor word choice given its negative connotation, and have edited this sentence to read:

      “With this in mind, future studies investigating genetic correlations between brain function and behavioral variables may benefit from hyperalignment, as it can factor out individual-specific cortical topography and thus yield more precise estimates of functional heritability.”

      (8) I would love to see a ventral view in some of these plots! Not asking you to recreate the figures, but the ventral temporal cortex is an area of interest for many folks in the movie fMRI space (e.g., Haxby et al., 2011).

      We agree that ventral views would be of interest to some readers and have added the corresponding maps for our main results in supplementary figures S3 and S9.

      References:

      Borovska, P., & de Haas, B. (2024). Individual gaze shapes diverging neural representations. Proceedings of the National Academy of Sciences, 121(36), e2405602121. https://doi.org/10.1073/pnas.2405602121

      Haxby, J. V., Guntupalli, J. S., Connolly, A. C., Halchenko, Y. O., Conroy, B. R., Gobbini, M. I., Hanke, M., & Ramadge, P. J. (2011). A common, high-dimensional model of the representational space in human ventral temporal cortex. Neuron, 72(2), 404416. https://doi.org/10.1016/j.neuron.2011.08.026

      Pitcher, D., & Ungerleider, L. G. (2021). Evidence for a third visual pathway specialized for social perception. Trends in Cognitive Sciences, 25(2), 100-110. https://doi.org/10.1016/j.tics.2020.11.006

      Reviewer #2 (Recommendations for the authors):

      (1) To address the common core analytical problems listed under A), B), C), D), E), and basically throughout the methods:

      (a) Conduct permutations with exchangability restrictions to account for the pattern of dyad-relationships as e.g. implemented in PALM.

      (b) Control for age and sex covariates as covariates (e.g. as in SOLAR), rather than by matching.

      (c) Perform r-to-z transforms when conducting further analyses on correlations that assume normality.

      (d) For all analyses that assume normal distributions, e.g. in SOLAR and BrainSMASH, check that this is the case.

      We have explained how PALM is not suited for the study of effects that are defined at the dyad level (A), that we controlled for age and sex covariates in all our formal heritability analyses in our original submission (B), that we always performed r-to-z transforms when indicated in our original submission (C), and that our spatial permutation results don’t hinge on distributional differences (D).

      (2) Replace SEs derived from kacknife approach with those from SOLAR, or provide a comparison and motivation and/or demonstrate that SEs are correct.

      A more thorough explanation of the block jackknife procedure can be found in prior work introducing the multidimensional heritability method used here (Anderson et al., 2021).

      (3) Given problem (F & G):

      (a) Consider studying the parameters that drive the hyperalignment. They can be included as covariates in heritability analyses, and/or their heritability is of interest to understand the reasons for the heritability reduction post-hyperaligment.

      We agree that this would be interesting but the specific parameters that drive hyperalignment are beyond the scope of this study.

      (b) Include the alternative explanation of hyperalignment-induced noise in the discussion.

      We have added a figure showing that hyperalignment does not increase noise in ISC and explained here why “hyperalignment-induced noise” does not constitute a reasonable alternative explanation for our results.

      (4) Add heritability results for NT phenotypes.

      We have added heritability analyses for NT topography and (global) NT magnitude, as detailed above.

      (5) Motivate global signal removal, and acknowledge this process typically alters results substantially.

      We have added an explanation of our rationale for using GSR and shown in this response that it does not in fact substantially alter the results.

      (6) Rephrase and/or clarify the following:

      (a) "permutations quantify average differences" (under A).

      (b) "network combinations" and related analyses (under B & C).

      (c) why some analyses are separated per visit/day and others not (C).

      (d) methods and reasons for sample size estimation (C).

      We have rephrased or clarified all of the above.

      Reviewer #3 (Recommendations for the authors):

      (1) Participants should be recleared. I know HCP 7T data has 184 subjects. How can the authors have 176 twins and 690 unrelated subjects?

      As we reported in our Methods section, 178 subjects had complete movie-watching datasets, and 176 subjects had complete movie-watching and resting-state datasets. Of the 178 subjects with complete movie-watching data, we identified 690 age- and sex-matched dyads.

      (2) Figure 1. I don't find Figure S1A in Figure S1.

      We thank R3 for catching this error- we have amended this reference to read Fig. S1.

      (3) I could also suggest putting Figure 1 and Figure 2 together.

      We thank R3 for this suggestion- ultimately, we prefer to keep these figures separate to reinforce the difference between our dyadic similarity and formal heritability analyses.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We are most grateful to both reviewers for providing valuable feedback on our manuscript.

      Reviewer 1 had solely favorable comments, with no suggestions for revision.

      Reviewer 2 pointed out that experiment evaluating the effect of CP4 on pVHL half-life (originally included as Figure 3c) was difficult to evaluate because of CP4’s effect on pVHL abundance prior to cycloheximide treatment. We agree with this assessment, and we opted to remove this experiment from the revised manuscript since it was not central to our overarching conclusions.

      Reviewer 2 also pointed out that experiment evaluating the effect of CP4.29 on HIF-2α half-life (originally included as Figure 4g) was not very compelling. We agree with this assessment, and we opted to remove this experiment from the revised manuscript since it was not central to our overarching conclusions.

      We agree with Reviewer 2’s suggestion that additional experiments could further solidify that C4.29 downregulates HIF2 in a purely “on-target” manner, however we prefer to reserve such studies for the future.

      Reviewer 2 also made several valuable suggestions for the text itself (awkward wordings / citations / clearer figure legends). We appreciate this feedback and have updated the text accordingly.

    1. Author response:

      We would like to express our gratitude for the thorough evaluation of our manuscript by the editors and reviewers. We are grateful for the overall positive assessment. The suggestions for improvement are reasonable, and we are certain that addressing these points will improve the clarity, accessibility, and scientific integrity of the study. Thus, we plan to conduct a revision of the manuscript, addressing all the points raised. The most important planned adjustments are outlined below.

      (1) Improving the accessibility of the probabilistic modeling framework

      Reviewer 1 kindly stated that our Bayesian modeling framework for testing for species differences 'sets a new standard for our field.' As a new standard, however, the method should be explained in a more accessible way. Hence, we plan to provide additional explanations for the statistical workflow, e.g., by providing comprehensible visuals, to make the workflow easier to understand and easier to apply.

      (2) Statistical validation of qualitative claims

      We acknowledge that a statistical validation of qualitative claims regarding the relationship between seed and tongue movements and between upper and lower beak movements would considerably strengthen the validity of our findings. We thank Reviewer 2 for bringing permutation tests to our attention for quantifying the correlation between time series. Since permutation tests involving index-shuffling of one of the data sets are generally not valid for time-series data [1, 2], we'll consider a variant of a trial-swapping permutation test, such as a permute-match test [3]. Alternatively, the truncated time shift (TTS) test [2] might be an option, as also this method is valid for auto-correlated time series data. At this point, we can't tell yet which method we'll use for the revised manuscript. We need more time to assess the requirements of each method and evaluate which test is most appropriate to answer our specific research questions and best fits our kind of data.

      (3) Adjustments in the discussion

      Following the suggestion by Reviewer 1, we'll refine our discussion on the effects of skull size differences, putting more emphasis on the implications of potential effects for feeding kinematics in small species.

      Furthermore, as suggested by Reviewer 2, we'll soften our discussion on potential functions of lingual papillae in seed processing, as the current literature lacks experimental evidence for the claimed mechanistic roles.

      References

      (1) Yuan, A. E., & Shou, W. (2022). Data-driven causal analysis of observational biological time series. Elife, 11, e72518.

      (2) Yuan, A. E., & Shou, W. (2024). A rigorous and versatile statistical test for correlations between stationary time series. PLoS biology, 22(8), e3002758.

      (3) Yuan, A. E., & Shou, W. (2025). Permute-match tests: Detecting significant correlations between time series despite nonstationarity and limited replicates. eLife, 14.

    1. Author response:

      We thank the reviewers for their time and attention which will significantly improve the paper. Further, we are grateful for their appreciation of our goals and work. In sum, the reviewers point to our overstated discussion of experimental evidence which we will tone down, some slightly confusing points of argumentation which we will clarify, and some discussion points on the role of normative theories that we will add text to address. We believe this will improve the paper significantly and hope you agree!

      Major Concern: Experimental Support for Path-Integration is not as strong as suggested

      The major point raised by all reviewers (reviewer 1 comment 1, reviewer 2 comment 1, reviewer 3’s only weakness) was that our presentation of the experimental perturbation evidence for path-integration is stronger than the reality. On reflection, we agree with this evaluation. We thank the reviewers for raising it; we will moderate our writing and include the sensible caveats raised. In sum, we still think that the convergence of evidence points to path-integration: first, disruptions to grid cells lead to path-integration problems, though these perturbations admittedly aren’t perfectly precise; second, normative theories of path-integration lead to grid cells and predict grid cell behaviour; third, mechanistic models of path-integration match grid cell behaviour and predict connectivity subsequently measured in entorhinal cortex. However, the evidence is not as all-encompassing as we suggested.

      That said, we’d like to further comment on one point. It is argued (reviewer 1, comment 1) that there are other theories of grid cell function, and that we discuss these theories. We discuss efficient-coding only models of grid cells and emphasise strongly why we reject them. We also briefly discuss oscillatory-interference models of path-integration and our reasons for not pursuing them further. As such, the reviewer is correct that our reading of literature strongly points us towards path-integration rather than other theories. We will slightly change the framing of the paper to make it clear that we are making a case. However, we are not aware of other theories the reviewer might be referring to. If the reviewer can point us to the other suggested theories that we do not address we would be happy to evaluate and include them.

      We now turn to the remaining comments, and how we plan to address them.

      Reviewer 1, Comment 2 – There could be multiple roles for grid cells

      The reviewer is indeed right that grid cells might perform multiple functions. This could just mean that the same computational motif (e.g. path-integration) is reused across different computations though that introduces no changes to the required normative theory. A stronger claim would be that grid cells perform both path-integration and some other function. This, according to a normative perspective, would most likely change how grid cells were optimally structured. We use the fact that large parts of the grid cell code can be captured with only path-integration as an argument against additional roles for grid cells. That said, there exist properties of grid cells not well-captured by path-integration which could well be smoking guns for additional roles of grid cells. The review already discusses both discrepancies between grid cells in three and two dimensions, and inhomogeneities in the grid in complex environments, and we will add two more (heading direction and peak-to-peak/angular variability, discussed below) that we are grateful to the reviewers for raising, and we discuss each of these in detail below.

      That said, whether these are necessarily arguments against purely path-integration or a reflection of interesting mappings of the core path-integration mechanism to the measurements we make remains to be seen. We would argue that both 3D grid cells (as explained below: there appear to be 2D slices in which grid cells behave as you’d expect) and spatial inhomogeneities (as explained in the paper: mappings of torus to world can introduce warping) can be explained without reference to additional computational roles of grid cells, which remain to us the most parsimonious explanation. We discuss next the slight update to path-integration only that the heading direction story suggest. But in sum, our view is that these discrepancies are likely not fatal for our path-integration-centric view of grid cells, but may well suggest some very interesting clarifications.

      Reviewer 1, Comment 4 – The system has two heading signals: true & internal, why?

      The reviewer is right to point to the puzzle over true vs. purely internal heading direction and which drives grid cells. We believe recent work from Abraham Vollan has effectively solved this puzzle: there appear to be two parallel circuits, one theta-modulated and following internal heading direction, another theta-unmodulated and aligning more with true heading direction. We will make sure to include discussion of this exciting work in our revised submission. This serves as a good example of an update we concede to the most austere version of the path-integration only view. Rather, it seems there are two parallel path-integrators working with different heading signals. The reasons for this remain unclear, but seem to be related to attention and planning (Vollan et al. 2026).

      Reviewer 2, Comment 3: Real Grid Cells have peak-to-peak variability & Angular variability

      The reviewer is right to point to the discrepancy in peak-to-peak firing rate and angles within a module that we did not adequately address. First, it is Sorscher’s RNN models, not nonnegative PCA that can generate a distribution of grid angles (Redman et al. 2025), which suggests that path-integration and such variability are compatible. We emphasise this point because the non-path-integration results from nonnegative PCA produce grid cells oriented at 30 degree offsets, something not measured even when you’re careful as in Redman et al. 2025. Thus, this becomes an interesting target for future work: perhaps using theories of path-integration up to an error threshold (rather than perfect) such angular diversity would be recovered. We will include this in our discussion. Further, we will include discussion of peak-to-peak variability that, as yet, has no obvious role.

      Reviewer 2, Comment 1: grid cells are inhomogeneous in 3D or complex environments, doesn’t that break the theory?

      Disrupted grid coding in extended or 3D environments indeed deserve more discussion, which we will add. In particular, we will add recent evidence that grid cells in 3D can be understood via the correct sequence of 2D projections(Qi & Yartsev, 2026). These two phenomena seem, to us, consistent with a path-integration only view of grid cells, as discussed above, and we hope to make this position clearer.

      Reviewer 2, Comment 5: Couldn’t there be other reasons for multiple modules?

      We have suggested a consistent normative framework in which multiple modules are explained through their role in non-linear coding. We think this elegant, and the most parsimonious current theory. We could, of course, be wrong. The discrepancies pointed to above might be good clues to follow to work out what else these modules might be doing, but currently these alternative explanations seem not to exist. We will text to clarify this.

      Reviewer 1, Comment 3: The review confuses computational and parameter parts of normative theory

      We disagree with the reviewer’s dichotomisation of normative theory. We view a normative theory as the complete procedure that produces the predictions. Almost all such theories have parameters and hence fitting a theory to data comprises both elements (a) [computational role] and (b) [specific parameters] identified by the reviewer. Occasionally theories have no parameters in the traditional sense, e.g. Rebecca et al.; instead they have heavy assumptions that play an equivalent role. It is true that, as the reviewer says, Sorscher et al.’s work was criticised for producing grid cells only for specific parameter values. We never found this as damning as Schaeffer et al. argued: simply it says that that theory is only correct within the given parameter range. Rather, arbitrating between models, parameters, or assumptions seems the same basic process: see what they predict and keep working with models while they remain useful ways to understand measured phenomena. If a model with very specific parameter values remains useful, that seems okay. In fact, we argued extensively why we think the nonnegative PCA model is not a useful model, but this was for completely different reasons. To us this story just reinforces the importance of hygiene in normative research: perform parameter sweeps and clarify how they constrain the claims you are making, carefully arbitrate what models can capture. Indeed, that is the whole goal of this review. We might be misunderstanding and, if so, we welcome correction.

      Reviewer 2, Comment 4: Normative Models of Cells Beyond Grid Cells

      The reviewer is right that extending these models to other cell types is an interesting area for further work, and that other cell types do seem to be involved in aspects of navigational computations both in RNNs and the brain. We will include a discussion to this effect in the revised manuscript. That said, we think the modularity of grid cells and their tight-linking to path-integration calculations should also be appreciated as a win!

      Reviewer 2, Comment 2: Multi-modularity is not cleanly explained

      We thank the reviewer for the comments, we agree. We will clarify the story regarding multiple modules, and will explain the equation further.

      Reviewer 1, Comment 5: the early introduction of phase-shifted Grid Cells seem the perfect place to normatively argue for Path-integration!

      We agree with the reviewer that this point can be made both normatively (‘oh look! If I try to do this optimally, I get translations!’) or, as we did early in the paper, mechanistically (‘oh look! With these cells I can do this!’). Indeed, a large part of the point of our paper is that path-integration is what is required to normatively derive phase-shifted grid modules, something discussed by Rebecca et al., our earlier work, and RNN studies, and appreciated for two decades. The earlier part of the paper does not discuss these papers as that section is aimed at giving intuition for the solution (mechanism). Later sections then heavily discuss the normative angle. We hope that division of labour makes sense.

      Finally, we will refine our summary of Rebecca et al. The reviewer is right that neurons don’t have to be discrete, we apologise for that error, but our understanding is that the only meaningful role of a neuron in Rebecca et al.’s work is the region in which is active, effectively making every neuron a binary unit, which seems dubious. We will clarify that by “predict velocity from each current and next encoding” we mean that the normative constraint they enforce is axiom 1: sequential activity of sets of neurons i then j can be uniquely interpreted as a trajectory, i.e. a step or velocity. Their work is elegant, and we will try to do more justice to it in the revision.

      To conclude, we thank the reviewers for their extensive comments, and look forward to releasing a version that addresses their concerns.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Weaknesses:

      The pipeline is very complete, but also complex. Workflows (optimal artifact removal, best curation for data from a particular brain area or species) will vary according to experiment. Therefore, a discussion of the adaptability of the pipeline in the “Limitations” section would be helpful for readers.

      We added a dedicated paragraph in the Discussion section under “Limitations” focusing explicitly on the adaptability and flexibility of the pipeline. Furthermore, we took this feedback as an opportunity to make the pipeline itself significantly more modular and customizable with the most recent release (v1.2.0: https://aind-ephys-pipeline.readthedocs.io/en/latest/releases/1.2.0.html).

      Reviewer #1 (Recommendations for the authors):

      (1) In the description of the Phase-shift correction (Line 166-167): The current text reads “As a result, different groups of channels are sampled asynchronously.” A better description would be: “Sample times for different groups of channels are offset in time by a known amount.”

      We replaced the phrase in the manuscript text with the suggested formulation.

      (2) Figure 5 and description of the benchmarking overview (Line 326-336): How were spike trains (times) selected for the injected ground truth units? What was the range of firing rates?

      All injected spike trains were generated as independent Poisson processes featuring a mean firing rate of 15 Hz. We have now incorporated this explicitly into the main text to clarify the ground-truth injection process.

      (3) Figure 6, panel b: Are the gray points in the raster the original spikes in the test recording? From the pattern, it looks like there are 8 recovered ground truth units. Were the other 2 undetected by either sorter?

      That is correct; the two remaining units were undetected by both sorters. To clear up any confusion, we updated the caption for Figure 6 to state: “Note that spikes undetected by any of the sorter are not shown in the plot.”

      (4) Figure 7, panel c: Are all units returned from KS included in these distributions? (i.e., regardless of the KS refractory metric calculated by the sorter) - it would be useful to add that detail to the caption. It would also be helpful for panel C to include a total unit count from the two sorters... Also, since there are multiple ways to calculate the refractory period contamination, it would be good to state the calculation used here.

      Because we rely directly on the hybrid ground-truth for accurate validation, we included all raw units returned by Kilosort for this specific analysis. We have explicitly added a note detailing this to the caption. Panel C does report the total raw unit count returned by the two sorters (N = 3046 for KS2.5; N = 3652 for KS4).

      Additionally, to clarify the evaluation procedure, we appended the following statement to the main text: “For all results, we perform spike train comparisons and compute performance metrics as defined in (Buccino et al. 2020), using all units returned by the spike sorter (without any sorterspecific curation).”

      (5) Comments about the pipeline:

      The paper clearly demonstrates the immense utility of the pipeline in the authors’ work. I did some testing to try to understand its adaptability to workflows at my institution.

      I tested the pipeline on our local cluster running LSF. I’ve worked on a similar pipeline using Nextflow to automate ephys analysis with the same sorters. Questions that came up for me that would be usefully addressed in the ’Limitations’ section:

      (i) Is the pipeline meant to be run only in total? In particular, is it possible to start with preprocesseddata? (aind-ephys-preprocessing/code/params.json does not appear to include any means to turn off filtering, for example). Is the pipeline meant to be run only in total? In particular, is it possible to start with preprocessed data? (aind-ephys-preprocessing/code/params.json does not appear to include any means to turn off filtering, for example).

      To accommodate users who wish to run only parts of the workflow or use external preprocessing setups, we have refactored the codebase to support a custom preprocessing pipeline option. This makes it possible to turn off standard filtering or inject custom workflows.

      (ii) For debugging purposes, is there a means to go from preprocessing or sorting to result collection,so that interim results can be interpreted even when some steps of the pipeline aren’t working?

      The pipeline is designed to be a spike sorting pipeline, so the spike sorting step cannot be skipped. However, we have rewritten the post-sorting architecture to make it highly lightweight and fault-tolerant. The postprocessing step now only requires the random spikes and templates computation and downstream steps have been update to accomodate this lightweight option. As an example, if no quality metrics are computed, the curation step will be skipped. The visualization and QC steps also required updates to be tolerant to missing extensions. This required coordinate updates across several components:

      Postprocessing: PR #12

      Curation: PR #13

      Visualization: PR #21

      Quality Control: PR #20

      (iii) If these options to skip processes and output data ’partway’ are available, it would be great toadd that to the documentation.

      We have fully updated our online documentation for v1.2.0 (release notes: https://aind-ephys-pipeline.readthedocs.io/en/latest/releases/1.2.0.html), introducing a brandnew “Customization” guide page that comprehensively explains how to construct and provide custom preprocessing and postprocessing strategies, as well as how to integrate a new spike sorter in the pipeline: https://aind-ephys-pipeline.readthedocs.io/en/latest/customization.html

      Reviewer #2 (Public review):

      Summary:

      This work presents a reproducible, scalable workflow for spike sorting that leverages parallelization to handle large neural recording datasets. The authors introduce both a processing pipeline and a benchmarking framework that can run across different computing environments (workstations, HPC clusters, cloud). Key findings include demonstrating that Kilosort4 outperforms Kilosort2.5 and that 7× lossy compression has minimal impact on spike sorting performance while substantially reducing storage costs.

      Strengths:

      (1) Extremely high-quality figures with clear captions that effectively communicate complex workflow information.

      (2) Very detailed, well-written methods section providing thorough documentation.

      (3) Strong focus on reproducibility, scalability, modularity, and portability using established technologies (Nextflow, SpikeInterface, Code Ocean).

      (4) Pipeline publicly available on GitHub with documentation.

      (5) Clear cost analysis showing ~$5/hour for AWS processing with transparent breakdown.

      (6) Good overview of previous spike sorting benchmarking attempts in the introduction.

      (7) Practical value for the community by lowering barriers to processing large datasets.

      Weaknesses:

      No significant weaknesses were identified, although it is noted that the limitations section of the discussion could be expanded.

      We thank the reviewer for their constructive feedback on our manuscript.

      Reviewer #2 (Recommendations for the authors):

      The authors could discuss why 2.25 bps is the “lowest supported” level and whether more aggressive compression could be achieved with custom approaches, potentially exploring where performance breakdown occurs.

      The 2.25 bits-per-sample (bps) limit is an inherent constraint of the WavPack lossy compression library itself. While more aggressive, domain-specific, or custom compression schemes could be explored, we focused on WavPack due to its native support in modern neurophysiology ecosystems and its excellent performance in our prior simulated benchmarks (Buccino et al. 2023). We agree that using this hybrid benchmarking framework to explore alternative compression configurations is a highly valuable avenue for future work. We have added the following text to the Discussion: “The benchmarking pipeline will continue to develop as an open evaluation framework, enabling transparent and reproducible comparisons of spike sorting and preprocessing methods across the community. As one example, the work on lossy compression could be extended with additional codecs and parameter settings, exploiting our ability to read out spike sorting degradation directly from the hybrid ground truth spike times.”

      (2) The limitations section would benefit from expansion to include: (i) discussion of how simulated data limitations may affect generalization of benchmarking results to real neural data, and (ii) clarification of the effort required to add new spike sorters, including configuration complexities for coordinating Nextflow processes beyond simple SpikeInterface integration.

      We have expanded the Discussion section to address both items:

      (i) We added a paragraph detailing the specific limitations of hybrid ground-truth datasets (e.g., how idealized template injection might miss extreme multi-unit overlapping dynamics or nonstationary noise properties found in real tissue).

      (ii) We added a structural overview section clarifying the workflow complexity, detailing exactly what steps are required to map a new spike sorter into a Nextflow execution processes beyond its baseline addition to Spike Interface.

      (3) The authors should clarify the terminology of “hypothetical experiment” in the introduction to improve reader comprehension.

      We have removed the word hypothetical from the introduction to ground the explanation more directly.

      (4) The cost analysis could be improved by making it clearer whether “runtime” refers to wall-clock vs. total parallel compute time.

      We mean wall-clock time. While total parallel compute time aggregated across cloud workers remains roughly identical to the overall sequential execution on a lone cloud instance, cluster parallelization slashes the wall-clock time drastically. We have updated the text to explicitly state that reported runtimes represent wall-clock time.

      (5) The authors could address the Nextflow Java dependency limitation by discussing containerized execution options (Docker/Singularity) as a solution, while noting relevant HPC system restrictions.

      We have updated the text to mention the official pre-built Nextflow container images as an elegant workaround for environments where local Java installations are blocked or restricted: “However, one option to bypass installation issues is to run the main pipeline script in container images packaged with Nextflow (https://hub.docker.com/r/nextflow/nextflow).”

      (6) Figure 8 analysis would be strengthened by explicitly noting that compression effects are more substantial for lower-accuracy units, suggesting better preservation of higher SNR units.

      We appreciate this insight. To evaluate this systematically, we generated a new supplementary figure (Figure S3) which shows sorting performance during lossy compression as a function of the Signal-to-Noise Ratio (SNR) of ground truth units. The plot demonstrates that for Neuropixels 2.0 recordings, the slight drop in sorting accuracy is indeed heavily concentrated among low-SNR units. We have integrated this observation into the Results section.

      Reviewer #3 (Public review):

      (1) Could the authors please expand on the statement on line 274, that processing their test dataset serially “on a single GPU-capable cloud workstation... would take approximately 75 hours and cost over 90 USD.” How were these values calculated? I was a bit surprised that this is a ¿4-fold slowdown from their pipeline, but only increases the cost by 1.35x... More context on why this is, and maybe some context on what a g4dn.4xlarge is compared to the other instances, might help.

      We have expanded the cost analysis section in the manuscript methods to explain these figures explicitly. The serial run relies on a single continuous, higher-tier GPU workstation instance (g4dn.4xlarge) running uninterrupted for 75 hours.

      Our distributed pipeline, by contrast, dynamically provisions CPU-only instances to process chunked preprocessing steps concurrently, then spins up short-lived GPU spot instances only when Kilosort executes. While this parallel execution compresses the overall wall-clock time by over 4-fold, the cost is only moderately reduced because the CPU-only instances with many parallel processing cores are only slightly less expensive than GPU instances.

      (2) One of the most commonly used preprocessing pipelines for Neuropixels data is the CatGT/ecephys pipeline from the developers of SpikeGLX at Janelia. It may be worth commenting very briefly... on how the preprocessing steps available in this pipeline compare to the steps available in CatGT. For example, is “destriping” similar to the “-gfix” option in catGT to remove high-amplitude artifacts?

      We have added a section drawing direct comparisons to CatGT preprocessing workflows. We explicitly clarify that our phase-shift correction performs the exact same function as CatGT’s Tshift. We also point out that while our current version lacks a direct equivalent to CatGT’s saturation removal feature (-gfix), this capability is scheduled for incorporation in our upcoming pipeline release.

      (3) Why are there duplicate units (line 194), and how often is this an issue? I understand that this is likely more of a spike sorter issue than an issue with this pipeline, but 1-2 sentences elaborating why might be helpful for readers.

      Duplicate units are primarily an artifact of template-matching sorting routines (such as Kilosort), which can occasionally split a single biological neuron into multiple overlapping spatial templates or over-extract templates in highly active channel regions. We have added two clarifying sentences explaining this phenomenon in the text: “Next, duplicated units, that can arise when using template-matching methods if different templates are consistently fit to the same spikes, are removed based on the fraction of overlapping spikes.”

      Customizability of cluster curation parameters It seems from the parameter files on GitHub that the cluster curation parameters are customizable - correct? If so, it may be worth explicitly saying so in the curation section of the text... A presence ratio of >0.8 could be particularly problematic for some recordings (e.g. state transitions, behavior specific cells).

      (4) Yes, they are completely customizable. We agree that a rigid presence ratio cutoff of 0.8 would erroneously discard highly valid units that are modulated by specific behavioral states, or are active only during sleep vs. wake cycles. We have explicitly added text in the Curation section clarifying that all quality metric thresholds can be modified by the user: “Units are tagged as passing a default_qc when they satisfy the following criteria based on quality metrics thresholds. Thresholds can be user defined, and these are the default”.

      (5) The axis labels in Figures 3d-e are too small to see, and Figure 3d would benefit from a brief description of what is shown.

      We have updated the figures with enlarged, high-visibility axis labels and expanded the caption of Figure 3d to clearly describe the visualization.

      Figure 4 labels (“neural” vs “passing QC”) (6) What is the difference between “neural” and “passing QC” in Figure 4?

      We have updated the figure caption for Figure 4 to include an explicit cross-reference to the Curation methodology section, which defines the strict quantitative boundary between raw neural classification and formal automated QC passage.

      (7) I understand the current paper is focused on spike data... but I am curious about the NP2.0 probes that save data in wideband. Does the lossy compression negatively affect the LFP data? Is software filtering applied for the spike band before or after compression?

      Compression is applied to the raw streams prior to any secondary downstream software processing. For Neuropixels 1.0, compression is executed strictly on the action potential (AP) stream. For Neuropixels 2.0, compression operates directly on the unified wide-band data stream.

      Software filtering to separate bands is conducted post-decompression, as captured in our baseline workflow definitions (e.g., WavPack compression → decompression → preprocessing → Kilosort4). To clarify this, we added the following text: “In all cases, compression was applied before any preprocessing took place. For Neuropixels 1.0, we compressed the AP stream only. For Neuropixels 2.0, we compressed the full wide-band data.”

      Because LFP signals possess inherently smooth continuous dynamics across both space and time, they are much more amenable to lossless or near-lossless compression. Thus, the minor losses introduced by lossy compression are overwhelmingly localized to high-frequency spike band features, leaving LFP components virtually unaffected.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      The superiority of the optimized system might simply be due to insufficient T7 RNA polymerase in the initial lysate.

      We performed a T7 RNA polymerase titration (0–1600 ng/µL) in the initial system to test this hypothesis. Standard CFPS protocols typically utilize T7 RNA polymerase at ~90–100 ng/µL<sup>1</sup>. To fully characterize the concentration-dependent effect and determine the exact saturation threshold of T7 RNA polymerase in our system, we tested an extended range from 0 to 1600 ng/µL. As shown in the revised Figure S3B, the initial system's output reaches a plateau at ~800 ng/µL—a concentration nearly ten times higher than standard protocols. Increasing the concentration further (up to 1600 ng/µL) led to a decline in yield, likely due to inhibitory effects of excess enzyme or buffer components. Even under these T7-saturated conditions, our optimized system achieved ~45-fold higher NLuc output compared to the maximum possible output of the initial system. Notably, when the lysate concentration is increased to 70%, the productivity gap reaches nearly 80-fold, further demonstrating the extraordinary efficiency of our platform.

      As revised in the Discussion, this improvement confirms that the performance gain is not a result of a mere increase in T7 concentration. Instead, it represents a systemic synergy where our streamlined buffer and the optimized metabolic environment of the fast lysate together alleviate the transcriptional bottlenecks inherent in traditional platforms.

      Reviewer #2 (Public review):

      Performance or efficiency claims... needs to be supported by comparisons with typical cell free expression systems.

      We agree that robust benchmarking is essential for validating our claims of high efficiency. Our comparative evaluation was conducted across three levels:

      (1) Literature-based benchmarking: As detailed in Figures 3C, 4A-D, S3A-B, S4, and S5C, we extensively compared our system against the "initial" (35-component) and "PEPbased" platforms, which are established benchmarks widely utilized in CFPS literature. These diverse comparisons consistently demonstrate the superior performance and robustness of our optimized system across various conditions.

      (2) Commercial benchmarking: To provide independent verification, we performed a head-to-head comparison with a high-end commercial E. coli CFPS kit (PePExpress, Shanghai Epizyme, EC010L). As shown in the comparative data provided in this response (See author response image 1), our system exhibited remarkable rapid-expression capability, significantly outperforming the commercial kit in both speed and absolute yield. Our platform reached near-maximum yield within 2 hours, demonstrating a significant efficiency advantage over the commercial alternative.

      (3) Robustness and translational quality: The comparison was extended to challenging targets beyond standard reporters. As shown in Figures 4E-H, the successful synthesis of active BsaI restriction enzyme (a cytotoxic protein) and the functional assembly of vimentin (an aggregation-prone protein) demonstrate that our optimized system maintains superior translational quality and robustness compared to typical platforms that often struggle with such complex targets. By outperforming established academic benchmarks and a leading commercial platform in both yield and the ability to handle challenging proteins, our results provide compelling evidence that the simplified 7component system is highly efficient. In the revised Conclusion, we have explicitly contextualized "efficiency" as the integration of high protein productivity, reduced reaction complexity, and accelerated preparation speed.

      Author response image 1.

      Comparative evaluation of sfGFP yields between our _e_CFPS system (70% lysate) and a commercial kit (PePExpress) over an 8-hour time course.

      Summary of revisions: T7 titration data have been added to Supplementary Figure S3B in the revised manuscript. To provide the additional benchmarking evidence requested, commercial comparison data (PePExpress kit) are provided in Author response image 1, while the main manuscript remains focused on the mechanistic synergy and streamlined architecture of the system.

      We hope that these substantial new data and the corresponding revisions satisfy the reviewers' queries.

      References:

      (1) Kigawa, T. et al. Cell-free production and stable-isotope labeling of milligram quantities of proteins. FEBS Lett. 442, 15–19 (1999).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single-molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states but also enabled the real-time monitoring of protein conformational changes, precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      We thank the reviewer for this accurate and thoughtful summary of our work and its broader significance. We agree that the combination of single-molecule FRET with orthogonal validation approaches enables mechanistic resolution of conformational states and transitions that are not accessible by ensemble measurements. In particular, this framework allows direct discrimination of ATP-free and ATP-bound conformations, real-time tracking of transport cycle progression, and identification of transient intermediates in the heterodimeric ABC transporter TmrAB. We further agree that these capabilities support a generalizable strategy for dissecting conformation dynamics in related ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results, and conclusions supported by the experimental data. The authors have determined the conformational dynamics of TmrAB across different ATP concentrations, including physiological ones, and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies.

      Weaknesses:

      The scientific study needs a bit of in-depth analysis with respect to consistency in K<sub>d</sub> and its implications on the mechanism.

      The apparent K<sub>d,ATP</sub> values were determined using two complementary approaches that report on different aspects of the system. Ensemble FRET measurements yielded values of 51 ± 38 µM (TmrAB<sup>NBD</sup>), 68 ± 25 µM (TmrAB<sup>PG</sup>), and 95 ± 26 µM (TmrAB<sup>PG_EQ</sup>), which are in good agreement with previously reported biochemical estimates (~100 µM for TmrAB<sup>EQ</sup>) (Stefan et al, 2020). The slightly elevated value observed for the E→Q variant may reflect modest perturbation of nucleotide handling in this slow-turnover background. Notably, the close agreement between labeled and unlabeled variants indicates that fluorophore attachment does not measurably affect ATP binding.

      In contrast, smFRET-derived K<sub>d,ATP</sub> values (13 ± 1 µM for TmrAB<sup>NBD</sup> and 2 ± 1 µM for TmrAB<sup>PG</sup>) are systematically lower. This difference likely arises from the difficulty of deconvoluting overlapping FRET populations at sub-K<sub>d,ATP</sub> concentrations, particularly for TmrAB<sup>PG</sup>, where state assignment is less well separated. Despite this quantitative offset, both approaches consistently indicate ATP saturation well below physiological concentrations and therefore support the same mechanistic conclusion that ATP binding drives conformational switching in TmrAB.

      Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATP-bound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I have three major points and a few minor criticisms.

      We thank the reviewer for the thoughtful and constructive evaluation of our manuscript and for highlighting the strength of combining structural and single-molecule approaches. We have addressed all major and minor points in detail below and revised the manuscript where appropriate to clarify limitations, justify analysis choices, and improve transparency.

      Major points:

      (1) The main weakness is that the authors base their conclusions on a very limited set of FRET pairs. While TmrAB has been extensively studied in terms of its structure, the authors should at least acknowledge this limitation more clearly.

      We agree that our conclusions are based on a limited number of FRET reporter pairs, and we now explicitly state this limitation in the revised manuscript. The chosen labeling positions were selected to probe two functionally critical regions—the nucleotide-binding domains and the periplasmic gate—based on prior structural and spectroscopic evidence. While this represents sparse sampling of the full conformational space, it is consistent with typical smFRET studies of membrane transporters, where experimental constraints generally limit the number of simultaneously accessible labeling positions (Asher et al, 2021; Asher et al, 2022; Levring et al, 2023; Wang et al, 2020).

      Importantly, both independent reporter variants yield consistent ATP-dependent population shifts, supporting the robustness of the observed trends. We further clarify that additional labeling sites could, in principle, resolve finer structural sub-states; however, given the already limited population separation in the current variants, such extensions would likely provide diminishing returns in state resolvability under the present experimental conditions. This trade-off is now explicitly discussed.

      (2) Most smFRET distributions were fitted with one, two, or three Gaussians. However, in several cases, additional populations with noticeable amplitudes appear to be present (e.g., Figure 3c at 0.1 mM and 3 mM ATP; Figure 4a, apo; Figure 4c, 0.3 mM R9L). Could the authors clarify why these populations were not included in the analysis?

      We thank the reviewer for this careful observation. Low-amplitude sub-populations are occasionally detected in individual histograms; however, they were not included in the quantitative model because they do not meet criteria for reproducibility, amplitude robustness, or structural assignability. Specifically, these features vary between replicates, contribute minimally to total population, and cannot be mapped to structurally or biochemically defined states based on available cryo-EM (Hofmann et al, 2019), DEER/PELDOR (Barth et al, 2018; Barth et al, 2020), or accessible-volume simulations.

      Similar minor subpopulations have been reported in smFRET studies and often attributed to photophysical or labeling heterogeneity effects (Asher et al, 2022; Husada et al, 2018). To avoid over-parameterization, we therefore restricted analysis to reproducible, structurally supported states. This rationale is now clarified in the revised manuscript.

      (3) Figure 3c (3 mM ATP): Is it truly possible to distinguish the two states in this distribution?

      We agree that state separation in the TmrAB<sup>PG</sup> variant is limited (ΔE = 0.11), and we now explicitly acknowledge this constraint in the manuscript. To improve robustness under these conditions, we used a constrained fitting strategy in which the apo-state distribution was fixed from nucleotide-free measurement, reducing parameter degeneracy during fitting of ATP-bound datasets.

      While single-molecule trajectory-based approaches such as Hidden Markov Modeling would be ideal for resolving dynamic interconversion, this was not feasible due to the low fraction of dynamic traces at the available temporal resolution. We therefore rely on population-level analysis, which remains consistent across replicates and reporter variants.

      Notably, independent measurements from two reporter positions (TmrAB<sup>NBD</sup> and TmrAB<sup>PG</sup>) yield similar ATP-bound population fractions at saturating ATP concentrations (~77% vs. ~80%), supporting the robustness of the inferred state distribution despite partial overlap.

      We have revised the manuscript to more clearly articulate methodological limitations, strengthen the justification of our analytical approaches, and improve the clarity of data presentation. These revisions enhance the transparency and robustness of the study and address the reviewer’s concerns.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here are a few comments that can help to improve the study.

      (1) Line 115: The authors have checked the purity and monodispersity of the protein sample using SDS-Gel and size exclusion chromatography; however, additional characterization using negative stain electron microscopy, which clearly shows the monodispersity, will be useful.

      We agree that negative stain EM can provide an additional assessment of sample homogeneity. Given the extensive prior structural characterization (Hofmann et al, 2019; Nocker et al, 2026; Nöll et al, 2017) and the SEC profiles presented here, we believe that additional negative stain EM would unlikely provide substantial new information regarding sample homogeneity. We have clarified this point in the manuscript by explicitly referencing the relevant cryo-EM studies.

      (2) Line 116: The authors have mentioned that the enzymatic activity of TmrAB was retained after purification. Although smFRET results showing conformational dynamics of TmrAB confirm its ATPase activity, a comment on the effect of labelling on ATPase activity will be useful.

      We appreciate this important point. Previous studies on spin-labeled TmrAB<sup>NBD</sup> demonstrated transport activity comparable to wild-type TmrAB, indicating that cysteine substitution and label conjugation do not substantially perturb this variant (Barth et al, 2018). In addition, AV simulations showed that fluorophores at the TmrAB<sup>NBD</sup> labeling positions do not interfere with ATP- or substrate-binding sites, supporting the conclusion that FRET labeling does not affect ATP binding, hydrolysis, or transport. For TmrAB<sup>PG</sup>, however, equivalent transport data were not available, and AV simulations suggested interference of fluorophores with periplasmic gate dynamics. We therefore directly compared the transport activity of LD555/LD655-labeled TmrAB<sup>PG</sup> and unlabeled wild-type TmrAB using a single-liposome transport assay with the fluorescein-labeled peptide C4F (RRYC<sup>F</sup>KSTEL) (<sup>F</sup>, fluorescein; Fig. 1– Fig. S3a). Both variants showed indistinguishable transport activity, demonstrating that fluorophore conjugation at the periplasmic gate preserves transport function.

      (3) Line 117 and Figure S1c. Please add the reference for consistency of ATPase activity with previous studies on TmrAB.

      We have added a reference to previous biochemical studies reporting comparable ATPase activity and kinetic parameters for TmrAB to support the consistency of our measurements.

      (4) Line 119: It mentions that "Cysteine-maleimide labeling of detergent-solubilized TmrAB achieved site-specific labeling efficiencies exceeding 90%". The legend of Figure S1d mentions about labeling efficiency in the range of 40-50%. A clarification will be helpful for the reader. Also, calculations can be extended to the ratio of LD555 and LD655 labels on the molecule, which can be considered in analyzing results.

      We apologize for the lack of clarity. The reported >90% labeling efficiency refers to the site-specific cysteine labeling efficiency per accessible site, as determined by dye incorporation. In contrast, the 40–50% values shown in Fig.1–Fig. S1d reflect the per-site efficiency for donor-lonely and acceptor-only populations respectively, which together account for the >90% overall labeling efficiency. We have revised the main text and figure legend to clearly distinguish between per-cysteine labeling efficiency and the fraction of correctly double-labeled molecules. We also clarify that only complexes with appropriate donor– acceptor stoichiometry were included in the smFRET analysis.

      (5) Figure 1: Line 627: This line mentions "For all simulations, TmrA is shown in blue with LD655 (orange) and TmrB in yellow with LD555 (green)." Is it (which label on which subunit) known for the experimental setup?

      We thank the reviewer for pointing out this potential source of confusion. In the experimental system, fluorophore attachment occurs stochastically. Therefore, the assignment of donor and acceptor dyes to specific subunits is random. The representation shown in Figure 1 reflects one possible configuration for visualization purposes only. We have clarified this explicitly in the figure legend to avoid misinterpretation.

      (6) Figure S1-2a. Tau value can be better represented in a graph for visual readers instead of in the form of a table, and a dotted line with the threshold (~1 ns) will give a better representation of no change. Values can be included in the graph as well.

      We appreciate this helpful suggestion. We have revised Figure S1-2a to include a graphical representation of fluorescence life times, including a reference line around ~1 ns to facilitate visual comparison. Numerical values are retained alongside the plot for completeness.

      (7) Figure 2a: Each component of the assembly has been pointed with an arrow, which can mix two components and confuse readers. It would be good to make a legend column on the left or right and depict or indicate each component of the assembly clearly.

      We have changed the labeling in Figure 2a to improve clarity by separating the components and introducing a clearer legend layout, ensuring that each element of the assembly is unambiguously labeled.

      (8) The physiological concentration of ATP can range up to 5-10 mM. A comment on choosing the ATP concentration specifically to be 3 mM would be useful for the readers.

      We appreciate this suggestion. While intracellular ATP concentrations can reach up to 5–10 mM, values around 3 mM are commonly used as physiologically relevant conditions in in vitro biochemical and biophysical studies. We selected 3 mM ATP as a representative near physiological concentration that ensures saturation of ATP-dependent conformational transitions while remaining comparable to previous studies on TmrAB (Hofmann et al, 2019; Nocker et al, 2026; Nöll et al, 2017; Stefan et al, 2020). We have clarified this rationale in the manuscript.

      (9) Figure 2c is not cited in the text.

      We thank the reviewer for noting this oversight. Figure 2c is now explicitly cited in the main text.

      (10) Results in Figure 2 and 3 have been analyzed using 2 and 3 Gaussian distributions, respectively. It would be good to explain the rationale for it.

      We appreciate that this important point was brought to our attention. The number of Gaussian components was determined based on the minimal model required to describe reproducible and structurally supported populations. For ATP titration experiments (Figure 2 and Figure 3), two populations (apo and ATP-bound) were sufficient and consistent across replicates. In contrast, three populations were required under trapping conditions (Figure 4), where an additional state (OFF<sup>open</sup>) becomes kinetically stabilized and clearly resolved. We have clarified this rationale in the manuscript.

      (11) Figure 3b: data points do not seem to be saturated with respect to ATP concentration. It needs more points beyond 3 mM. Different K<sub>d</sub> at different sites in the structure could represent differential local dynamics over the structure.

      Previous structural studies demonstrated that 1 mM ATP is sufficient to saturate both nucleotide-binding sites under trapping conditions (Hofmann et al, 2019), indicating that the concentration range used here is adequate. Consistent with this, both ensemble and smFRET measurements approach saturation by 3 mM ATP, a near-physiological condition commonly used in biochemical studies. While additional data points above 3 mM could further define the plateau, they are unlikely to alter the mechanistic conclusion. We have clarified this point in the manuscript.

      (12) Figure 3 and Figure 1 - S1 have two different Kd values with respect to ATP concentration; both of these graphs measure conformational changes using smFRET. A comment specifying these Kd values based on single molecule verses ensemble measurement from will be helpful for readers.

      We appreciate this important point and have clarified it in the manuscript and the response to Reviewer #1 above. The K<sub>d,ATP</sub> values in Fig. 1–Fig. S1 are derived from ensemble FRET measurements, whereas those in Fig. 3 are obtained from smFRET population analysis. This difference likely arises from the difficulty of deconvoluting overlapping FRET populations at sub-K<sub>d,ATP</sub> concentrations, particularly for TmrAB<sup>PG</sup>, where state assignment is less well separated. Despite this quantitative offset, both approaches consistently indicate ATP saturation well below physiological concentrations and therefore support the same mechanistic conclusion that ATP binding drives conformational switching in TmrAB. We now explicitly distinguish these methods and their interpretation in the manuscript.

      (13) Figure 4: Slow-turnover TmrAB mutant has been employed in cysteine mutant on the PG opening side, but not towards the NBD side. Either experimental data or a comment on not pursuing it would be helpful for the reader. Similarly, experiments in the presence of peptide and in the absence of ATP, which can help to understand the role of substrate in conformational dynamics in the absence of ATP, are not pursued in this study. Along similar lines, experiments with wild type, in the presence of MgADP +/- substrate, are not shown in this study.

      We thank the reviewer for these insightful suggestions. The slow-turnover variant was specifically applied to the periplasmic gate reporter (TmrAB<sup>PG</sup>) because this construct provides direct sensitivity to outward-facing conformations, which are central to resolving the OF<sup>open</sup> state. In contrast, the NBD reporter primarily monitors nucleotide-binding domain (NBD) dimerization and is less suitable for distinguishing periplasmic conformational differences.

      Experiments in the absence of ATP but in the presence of peptide, as well as MgADP ± substrate, would indeed be valuable for further dissecting substrate effects. However, these conditions are beyond the scope of the current study, which focuses on ATP-driven conformational dynamics and the identification of kinetically hidden intermediates. We have added a statement in the Discussion to acknowledge these possibilities as directions for future work.

      (14) Figure 4, peptide concentration has been varied in the right panel. The result can also be presented as the % of OFopen and OFoccluded state with increasing concentration of peptide.

      We thank the reviewer for this suggestion. While such a plot would indeed be informative and could improve our understanding of substrate binding and substrate-induced trans-inhibition, the current dataset does not contain sufficient data points to construct a reliable concentration-dependent curve, particularly given that peptide saturation was not reached in our experiments. The characterization of substrate binding is further complicated by the presence of two distinct substrate-binding sites one in the outward-facing and one in the inward-facing state with likely completely different K<sub>d</sub> values and would require a more complex binding model. We have therefore decided against including this plot in the current manuscript. We do acknowledge, however, that future smFRET studies with improved temporal resolution are particularly well suited to investigating substrate binding to TmrAB and its effects on conformational equilibrium, and we have noted this in the Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) In all figures, can you please label the transporter schematics with the conformational states they represent?

      We thank the reviewer for this suggestion. All transporter schematics in the main and supplementary figures have been updated to include clear labels indicating the corresponding conformational states, thereby improving clarity and consistency.

      (2) As a suggestion, it may improve clarity to include the labelling positions (residue numbers) directly in Figure 1a and b, even though they are provided in the legend.

      We appreciate this suggestion. Residue numbers corresponding to labeling positions have now been added directly to Figure 1a and b to improve readability and facilitate interpretation.

      (3) Lines 183-188: This is a key point. It would be helpful to include a reference line for the expected state (0.63). Interestingly, this value coincides with the shoulder observed in Fig. 3c (0.1 mM ATP). Is there an explanation for this (see also point 2)?

      We thank the reviewer for highlighting this point. We considered adding a reference line at 0.63 to the plot; however, we decided against it. While a subpopulation does appear at ~0.63 —consistent with the expected FRET efficiency of the OF<sup>open</sup> conformation—it is only present in a single condition (0.1 mM ATP) and is not observed across other ATP concentrations for this TmrAB variant. It more likely reflects a minor non-reproducible subpopulation or photophysical artefact, in line with our response to Point 2 of the public review (Reviewer #2).

      (4) The final section of the Results section seems like an afterthought, especially since the heading suggests a broader scope.

      We appreciate this comment. We have revised the final section of the Results to improve its structure and ensure that the scope indicated by the heading is fully reflected in the content. This section now more clearly integrates kinetic and thermodynamic aspects of the transport cycle.

      References

      Asher WB, Geggier P, Holsey MD, Gilmore GT, Pa; AK, Meszaros J, Terry DS, Mathiasen S, Kaliszewski MJ, McCauley MD, Govindaraju A, Zhou Z, Harikumar KG, Jaqaman K, Miller LJ, Smith AW, Blanchard SC, Javitch JA (2021) Single-molecule FRET imaging of GPCR dimers in living cells. Nat Methods 18: 397–405. doi:10.1038/s41592-021-01081-y

      Asher WB, Terry DS, Gregorio GGA, Kahsai AW, Borgia A, Xie B, Modak A, Zhu Y, Jang W, Govindaraju A, Huang LY, Inoue A, Lambert NA, Gurevich VV, Shi L, Lefkowitz RJ, Blanchard SC, Javitch JA (2022) GPCR-mediated beta-arrestin activation deconvoluted with single-molecule precision. Cell 185: 1661– 1675 e1616. doi:10.1016/j.cell.2022.03.042

      Barth K, Hank S, Spindler PE, Prisner TF, Tampé R, Joseph B (2018) Conformational coupling and transinhibition in the human antigen transporter ortholog TmrAB resolved with dipolar EPR spectroscopy. J Am Chem Soc 140: 4527–4533. doi:10.1021/jacs.7b12409

      Barth K, Rudolph M, Diederichs T, Prisner TF, Tampé R, Joseph B (2020) Thermodynamic basis for conformational coupling in an ATP-binding cassette exporter. J Phys Chem LeJ 11: 7946–7953. doi:10.1021/acs.jpclett.0c01876

      Hofmann S, Januliene D, Mehdipour AR, Thomas C, Stefan E, Brüchert S, Kuhn BT, Geertsma ER, Hummer G, Tampé R, Moeller A (2019) Conformation space of a heterodimeric ABC exporter under turnover conditions. Nature 571: 580–583. doi:10.1038/s41586-019-1391-0

      Husada F, Bountra K, Tassis K, de Boer M, Romano M, Rebuffat S, Beis K, Cordes T (2018) Conformational dynamics of the ABC transporter McjD seen by single-molecule FRET. EMBO J 37: e100056. doi:10.15252/embj.2018100056

      Levring J, Terry DS, Kilic Z, Fitzgerald G, Blanchard SC, Chen J (2023) CFTR function, pathology and pharmacology at single-molecule resolution. Nature 616: 606–614. doi:10.1038/s41586-023-05854-7

      Nocker C, Pečak M, Nocker T, Fahim A, Sušac L, Tampé R (2026) Single-molecule dynamics reveal ATP binding alone powers substrate translocation by an ABC transporter. Nat Commun 17 doi:10.1038/s41467-026-70021-1

      Nöll A, Thomas C, Herbring V, Zollmann T, Barth K, Mehdipour AR, Tomasiak TM, Bruchert S, Joseph B, Abele R, Olieric V, Wang M, Diederichs K, Hummer G, Stroud RM, Pos KM, Tampé R (2017) Crystal structure and mechanistic basis of a functional homolog of the antigen transporter TAP. Proc Natl Acad Sci U S A 114: E438–E447. doi:10.1073/pnas.1620009114

      Stefan E, Hofmann S, Tampé R (2020) A single power stroke by ATP binding drives substrate translocation in a heterodimeric ABC transporter. eLife 9: e55943. doi:10.7554/eLife.55943

      Wang L, Johnson ZL, Wasserman MR, Levring J, Chen J, Liu S (2020) Characterization of the kinetic cycle of an ABC transporter by single-molecule and cryo-EM analyses. eLife 9: e56451. doi:10.7554/eLife.56451

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      WIPI1 is a PROPPIN family protein that has been implicated in Retromer-mediated membrane fission events. Although the cargos that it has been tested to be important for are diverse, one of the cargos that is unaffected is Beta1-Integrin. This leads the authors to assess another PROPPIN family protein - WIPI2, which is a homolog of WIPI1. KD using siRNA is effective and had no consequences on LAMP1, EGFR trafficking or GLUT1 trafficking. Integrin-B1, however, had a large and significant defect in its recycling from the endosome, with a clear endosomal colocalisation. Complementation experiments with WT WIPI2 recovered the phenotype, but various mutant WIPI2 complements resulted in elongated tubules, and there was also a dominant negative effect of the mutant. Integrin is a classic retreiver cargo, so the authors rationalise that WIPI2 may be playing a role with retreiver that WIPI1 plays with retromer. To assess this, they perform a set of immunoprecipitations. SNX17, the retreiver-associated sorting nexin, co-IPs with WIPI2 in a VPS26C-dependent manner. VPS26C but not VPS26 co-IPs with WIPI2, and the reciprocal with WIPI1. These interactions were not present for the FSSS mutation of WIPI2. WIPI2 localises to Rab11 endosomes mainly, as does retriever. Mutations of WIPI2 not only affected WIPI2 localisation, but also VPS35L mutations, indicating that there is a functional relationship between the two.

      On the whole, I find the manuscript compelling. The manuscript is very clearly written, the results are convincing and well performed. The flow of experiments is logical, and although not comprehensive in the subsequent mechanistic understanding, the fundamental findings are important and convincing. My comments below are, on the whole, minor and are intended to support the communication of the findings to the field.

      We are happy that the reviewer has received our work quite positively.

      (1) The IP interaction data were convincing; however, for me and some others, an interaction is only convincing when performed in vitro, and understood at a structural level. I do not suggest the authors do that in this case; however, I think, at a minimum, some sensible moderation of claims would be useful here.

      Indeed, quantitative in vitro data on the affinities would be a nice addition. However, we have significant trouble to recombinantly express and purify well-behaved WIPI2 in sufficient quantities for such studies. We keep working in this direction but are not there yet.

      We have now inserted a phrase into the discussion section highlighting this limitation: "Our immunoprecipitation assays cannot distinguish and more detailed structural and interaction studies with pure compounds will be necessary to elucidate the nature of this interaction". We nevertheless think that the the isoform specificity of the IPs, the effect of the point mutations in WIPI2 on these interactions, and the functional effects in vivo lend signficant support to the notion of a complex even if there is no proof of direct binding of WIPI2 to Retriever.

      (2) I found the final localisation data and its interpretation confusing. My interpretation of that data would not be that the retreiver is relocalised, but rather that there is less of both recruited to the membrane and the remaining localisation distribution is shifted. In addition, I am not quite sure of the model here - is the idea that WIPI2 recruits retreiver, if that is the case, I find it hard to resolve with its role as a mediator of fission. Clarity would be appreciated here.

      We are not quite sure what "final" localisation data the reviewer refers to, but we guess it is Fig. 9. This figure primarily provides in vivo evidence supporting the connection between Retriever and WIPI2. It does this by showing that the S67 substitution shifts both proteins. In WIPI2 wildtype cells, WIPI2 and VPS35L strongly colocalize in Rab11 compartments. S67 substitutions in WIPI2 abolish this localisation; WIPI2 shifts mainly to Rab5 compartments, where VPS35L shows only a moderate increase, and to Rab7 compartments, where VPS35L shows no increase at all.

      We do not understand the reviewer's interpretation that less Retriever would be recruited to the membranes in the S67 variants. VPS35L remains completely associated with punctate, presumably membrane-bounded structures also in the mutants, providing no evidence for a detachment from the membrane. The same is observed in a WIPI2 knockdown. Therefore, we did not claim that WIPI2 is the main factor recruiting Retriever to the membrane, for which our experiments yield no hints. This does not exclude that the interaction of WIPI2 could strengthen membrane recruitment, or that two pools of Retriever exist, one interacting with Snx17 and another interacting with WIPI2, and that both link to each other in a coat. We did not dwell on this in the discussion because our experiments cannot distinguish these possibilities and were not conceived to analyse membrane recruitment of Retriever.

      (3) I am concerned that the repeats being compared for statistical analysis are not biological repeats but technical repeats (cells in the same experiment). I should think the idea of the statistical comparison is to show experimental reproducibility and variability across biological repeats. Therefore, I would expect an appropriate number of biological repeats (3 or more minimum), to be the data compared in the statistical analysis and graphs. I think it is appropriate to average the technical repeats from each biological repeat. I find these to be useful resources https://doi.org/10.1083/jcb.202401074, https://doi.org/10.1083/jcb.200611141

      The repeats being compared are biological repeats from independent experiments. This is described in Methods, where the reviewer may not have seen it. In order to make the independent experiments more evident in the figures, we have now colour coded the individual cell measurements from the three independent experiments. This allows to visualize both the individual data points, the average from each experiment and the variability across the independent experiments.

      Reviewer #2 (Public review):

      Summary:

      The manuscript from De Leo and Mayer presents evidence that the PROPPIN protein, WIPI2, associates with the Retriever complex, and is required for the proper transport of the SNX17-Retriever cargo, beta1-integrin. This finding fits with prior papers from the Mayer lab, which showed that a related PROPPIN, WIPI1, is required for the transport of some SNX27-Retromer cargo, including GLUT1. The retromer and retriever complexes are architecturally similar. Importantly, they act at the same endosomes, and each transports cargo from endosomes to the plasma membrane. Thus, the possibility that each also requires a structurally related PROPPIN is of interest. However, the manuscript is incomplete, and the main claims are only partially supported.

      Strengths:

      The topic that PROPPIN proteins are important for the function of the Retromer and Retriever complexes expands our view of the trafficking complex.

      Weaknesses:

      Many important controls are missing. Several points that are made in the manuscript are only supported through a single approach.

      We made a serious effort and implemented many suggestions of this reviewer, but orthogonal approaches are not always available or accessible.

      Reviewer #3 (Public review):

      Summary:

      The manuscript of Mayer and colleagues analyzes the function of WIPI proteins in mammalian cells. The authors previously identified CROP as a complex consisting of WIPI1 and the retromer complex, primarily in yeast cells. In mammalian cells, both WIPI1 and WIPI2 exist, whereas retromer has a homologous complex termed retriever. They now find that WIPI2 can form a complex with retriever subunits. They named this complex CROP2. Their data further indicate that CROP2 and CROP1 have distinct substrate specificities as knockdown of CROP2 subunits affects beta1 integrin sorting, whereas knockdown of CROP1 affects EGFR and GLUT1. They further identify a similar sequence (FSSS) in both WIPI1 and WIPI2, which is required for their specific binding to retromer and retriever.

      Strengths:

      CROP1 and CROP2 seem to use similar features for their formation, and have different substrates, which is convincingly shown.

      Weaknesses:

      The analysis lacks information that this is a complex as claimed. It can be deduced from the interaction analysis, but was not shown.

      It is of course desirable to obtain a detailed structural and in vitro characterisation of this interaction, which we have not provided because we currently do not have sufficient amounts of well-behaved source material for this. We nevertheless think that the interaction we show, which is strictly isoform-specific and dependent on single amino acid substitutions in a motif that in CROP1 is necessary for the interaction its recombinant subunits, supports that CROP2 is a similar a complex. We don't show a direct interaction but also don't claim in the manuscript that the interaction between WIPI2 and Retriever is direct and independent of additional factors.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you will see, the reviewers generally value the contribution to the field, but they feel that some claims require additional experimental support.

      (1) I have summarized the major points below.

      (a) Both reviewers 1 and 2 agree that the quality of localization data presented in Figure 9 and S5-S7, and the interpretation of the data, could be improved. See comment 2 from reviewer 1 and comments 23, 24 and 25 from reviewer 2. They not only suggest ways to improve the presentation of the data, but additionally suggest improving the staining of the Rab11 marker and additionally explain the lack of co-localization between VPS35 and Rab5, which has been reported in the literature.

      This impression was due to the fact that some figures showed projections of image stacks, which was not indicated clearly in the figure legend. We have changed this and now show single image planes throughout all figures.

      (b) Both reviewers 1 and 3 note that the evidence supporting a functional WIPI2-Retriever complex in vivo is currently weak. We agree that additional biochemical data demonstrating the presence of the CROP1 and CROP2 complexes in vivo would strengthen the central message of the paper and elevate it to a more fundamental discovery.

      We understood that the reviewers did not ask for further in vivo evidence but would welcome structural characterisation of the complex and quantitative binding data in vitro with purified proteins. Structural characterisation is out of scope of our study and in vitro binding studies have remained hampered by the fact that WIPI2 is hard to express and purify and not well behaved in vitro.

      (c) All reviewers agree that the authors should carefully repeat their statistical analysis to account for the number of biological replicates. Reviewer 1 suggests publications that the authors could refer to.

      The reviewers have probably overlooked the respective description in the methods section, where it had been stated that we analysed biological replicates from independent experiments. In graphs showing measurements from individual cells we now make this evident through colour coded dots, in which each colour represents data points stemming from an independent experiment. This makes it evident that the variance from experiment to experiment is low. The means (n = 3) were generally compared using a two-tailed unpaired t-test.

      (d) Reviewer 2 additionally has various minor points that would greatly improve the readability and presentation of the work, and we recommend addressing (comments 1, 2, 3, 4, 12, 15, 17, 20, 27, 28, 29). All reviewers, in general, provide great minor suggestions. It would be great if the CROP1 and 2 complexes could be clearly introduced in each figure. We also agree that the WIPI2 CT labelling is confused and should be changed to "control" or similar.

      Many of the points raised by this reviewer were actually quite minor or questions of personal preference, not major problems as stated in the review. Nevertheless, we found a number of useful suggestions in this review and have addressed these points as detailed in the response to reviewer 2.

      (2) In addition to the major shared concerns laid out in the points above, reviewer 2 has some further minor suggestions:

      (a) Comment 6. Could the author explain the discrepancies between the example blot shown in Figure 1D and the quantification (1E).

      The two have actually been quite consistent. The reviewer might have mistaken the marker lane as the 0 min reference value to arrive at this impression. We have now removed the marker lane to avoid this.

      (b) Comment 9 - could the authors clarify how surface labelling experiments were carried out?

      This had been clearly described in the methods section, where this reviewer has probably not seen it.

      (c) Comment 11 - The reviewer suggests normalizing the surface levels of markers to the cell area and not per cell. This is a reasonable suggestion.

      The analysis had already been performed as proposed. This had been clearly described in the methods section, which the reviewer may not have looked at.

      (d) Comment 19 "In Figure S4, the authors observe tubular structures. The authors should perform immunofluorescence with endosomal markers such as EEA1, LAMP1 and Retromer to determine the nature of the tubulovesicular structures." The authors could try a Rab4 or Rab11 overexpression plasmid to show whether these are elongated recycling tubules.

      This has now been added.

      Reviewer #1 (Recommendations for the authors):

      Minor comments:

      (1) The figures are not colourblind friendly, and should be changed to be so. Additionally, single colour images should be grayscale.

      That was a good learning opportunity. We adapted the colour schemes of the images to make them more colourblind friendly, now using magenta, green, and white for the overlaps. In doing so we have relied on published recommendations, but we have not found a colourblind colleague to check the efficacy of this change.

      (2) WIPI2^CT labels are confusing, as people may think they are a mutant. I suggest changing to "control" or similar.

      These have been changed.

      (3) "The effect was comparable to that of a knockdown of SNX17 (Figure 3 A, B)." On page 6. Based on this sentence, I was expecting to see a comparison to SNX17 KD, but it was not there as far as I can tell.

      This statement referred to a publication by P.Cullen and collaborators. We have changed the wording and inserted the (missing) reference to make this clear.

      Reviewer #2 (Recommendations for the authors):

      The manuscript is modest. In addition, many of the claims should be better supported by the addition of orthogonal data. Moreover, the quality of some of the data presented needs to be improved. Overall, the manuscript requires better descriptions of the methods. In many figures, it was not clear how the experiments were performed.

      The experimental descriptions that the reviewer refers to had been provided in the Methods section, where this reviewer may have overlooked them.

      The paper should also be better organized. Some less important findings are in the main figures, whereas some critical results are in the supplemental figures. In addition, there were multiple issues with the readability of the paper, and the authors should consider using a professional editor to make the paper easier to read.

      We had given the paper to colleagues who found it clear, and also Reviewer 1 has underlined its clarity. Nevertheless, we have re-phrased the manuscript in some parts to optimise it.

      One of the main claims in the paper is that the FSSS motif of WIPI2, as well as a conserved amphipathic helix, is critical for WIPI2 function in the CROP2 complex. It is notable that these are the same regions that are also critical for the role of WIPI2 in autophagy (Gubas et al., 2024 PMID: 39152217). The authors should include this information in the manuscript and cite the paper.

      Indeed. We mention this now in the introduction of the revised version.

      Additional Major Issues:

      While some of the issues raised below are actually minor and/or matters of personal preference, several comments led us to improve and correct the figures and we thank this reviewer for the constructive suggestions.

      (1) In Figure 1, it appears from the representative images that WIPI2 KD cells have higher levels of EGFR (Figure 1A and 1B). Is this correct?

      To some degree. This increase is not systematic. A moderate increase has been observed only in 2 experiments out of 4. Therefore, we did not investigate this.

      (2) Also in Figure 1, the colocalization is difficult to see. The authors should add the separate channels in addition to the merged images. Since the point is supposed to be that there is no impact on EGFR, all of this data could go into the supplement.

      We had considered this already for the original version but dismissed the idea. The overlap is quantified in Fig. 1C, which provides the relevant values from four experiments. Fig. 1A/B provide only sample pictures, which also permit to see overlap (yellow) 0 and 5 min after the induction of degradation, which vanishes at later timepoints. Separating the channels would quadruple the space that this figure occupies, which would not be practical and not change the point to be made.

      (3) The scale bars for each panel differ from each other. To better assess the data, the exact same magnification should be shown for each panel.

      Corrected

      (4) Figure 1C is confusing. The authors should explain which lines correspond to EEA1 and LAMP1.

      Corrected

      (5) In Figure 1D, the authors show different blots for control and WIPI2 KD. Could the authors compare WIPI2 and EGFR in the same blot? Without a comparison on the same blot, it is impossible to know whether the starting levels of EGFR are the same. Moreover, the quantitation in Figure 1E sets the value for each cell line to 100%. Instead, the starting levels in each cell line should be compared. The authors should use the amount of EGFR at zero time in the control cells to define 100%, and then indicate the relative initial EGFR levels in the WIPI2KD cells.

      A new blot is shown now and the quantification has been performed as proposed.

      (6) The quantification in Figure 1E does not match the representative blot shown in Figure 1D. According to the graph, the rate of degradation of EGFR is similar in both cell lines. But the representative blot shows that there are large differences.

      We do not understand this comment. The representative blot shows similar kinetics for both. Perhaps the reviewer got confused by the fact that a marker lane was still present on the left blot and not labelled as such. The new version of the figure corrects this.

      (7) The blot showing the WIP2 knockdown in Figure 1D has a lot of background. However, the blot of the WIPI2 knockdown in Figure S1 looks very good. The authors should make sure that they load enough sample and use a good antibody for the experiments in Figure 1.

      The new blot that we added in response to comment 5 corrects this.

      (8) In Figure 2 and Figure 3A, the cells are too confluent. This is an issue because the cells might not be metabolically active. In addition, the signal is saturated. The authors should make sure that all of the data is collected on cells that are not too confluent.

      The confluency of the culture cannot be judged from single frames, which were selected to show several cells. We had controlled confluency and underlined in the Methods section that “For microscopy, the cells were plated on 18-mm-diameter glass coverslips on 24-well plates and grown for 2 or 3 days according to the protocol of DNA or siRNA transfection by reaching a confluency of 70-80%”. The reviewer may not have seen this.

      (9) One main issue with these figures, especially the non-permeablized cells, is that it is impossible to assess how much of the signal is on the cell surface. The authors should provide the methods that they used to prevent inadvertent permeabilization of the cells. Were these experiments performed at 4 degrees? The authors should include a control of an antibody to a protein that is not found on the cell surface.

      There is an internal control in that the non-permeabilised WIPI2KD cells, which have been treated with the same antibody, show no much less staining than the control cells (Fig. 3A). In WIPI2KD cells, integrin becomes accessible for antibody staining only upon detergent permeabilization. This demonstrates that our procedure does not lead to significant inadvertent permeabilization of the cells.

      (10) The authors should perform surface biotinylation assays as an orthogonal approach to determine GLUT1 levels and beta1-integrin levels at the cell surface, respectively.

      There is a strong, qualitative difference in the surface labelling of beta1-integrin that is not observed for GLUT1. Given that, it is not obvious to us what additional argument would be provided by surface biotinylation or subfractionation experiments.

      (11) In quantifying surface levels of GLUT1 or beta1-integrin by microscopy, the authors should normalize to the cell area, rather than per cell.

      The reviewer has probably not seen that the Methods section states that the cell area has been used for normalisation.

      (12) In Figure 3, the nuclear DAPI stain in the KD cells is much less bright than in the control cells. The authors should make sure to choose representative images.

      The nuclear DAPI signal has been visible in all cells. Depending on the position of the nucleus, is shape and dimension in the z-direction, individual nuclei can show different degrees of staining. The images shown are representative. We have adjusted the settings now to make the nuclei in the WIPI2KD cells easier to spot.

      (13) For the immunofluorescence studies, the authors should be using single z planes rather than maximum projection.

      Images have been exchanged by single planes.

      (14) For the experiments in Figure 3, the authors should check the total levels of EEA1 and LAMP1 by western blot to test whether WIPI2 KD affects the levels of these proteins. If these organelle marker proteins are impacted, this could impact the colocalization measurements shown in Figures 3C and D.

      We have measured the total fluorescence intensity of EEA1 and LAMP1 in the images. It shows no significant difference between control and WIPI2 knockdown cells (new Fig. 3F, H).

      (15) In Figure 4A, the helical representation is rotated in the WIPI2-Sloop; the orientation of the residues that are not mutated should stay the same.

      Yes. Done.

      (16) In Figure 4B and 4C, cells that were not transfected with WIPI2 WT or WIPI2 Sloop should be shown.

      Since the transfection efficiency is limited, the fields contain both non-transfected (lacking green fluorescence) and transfected cells (showing green fluorescence). We have now marked transfected cells with an asterisk.

      (17) The cells in the lower panel of 4B have an unusual morphology and are much more round. The authors should choose cells that are representative of each experimental condition.

      We now provide another field.

      (18) In Figure 4C, it looks like the magnification of the top panels is different from the bottom panels. The same magnification for all the panels should be shown (and the size of the scale bars should be the same.

      Corrected

      (19) In Figure S4, the authors observe tubular structures. The authors should perform immunofluorescence with endosomal markers such as EEA1, LAMP1 and Retromer to determine the nature of the tubulovesicular structures.

      We have done this (new Fig. S4). Rab4 is on tubules. Rab5 on the structures from which the tubules emanate.

      (20) In Figure 5A, the top scale bar is missing.

      Corrected.

      (21) In Figure 5B, the confluency is too high.

      See our response above. A single field does not permit to judge this. Confluency was controlled for all cultures. The cultures were not confluent.

      (22) The IP studies shown in Figures 6, 7 and 8, should be accompanied by colocalization studies.

      Colocalization measurments have now been integrated into the manuscript (Figs. S5, S6). They are consistent with the IP data.

      (23) Figure 9 was very confusing and should be broken up into multiple figures. Data showing that localization did not change in any of the cell lines can be put in figures that are distinct from figures that show that localization changed in the various mutants. Figures that show no change can go in the supplement.

      Since every panel of Fig. 9 shows a statistically significant difference we left the figure unchanged.

      (23) Representative figures should be shown in the same figure as the corresponding graph. In addition, the order of the colocalization data shown in the graphs and figures should match the order described in the text.

      We consider the graphs of Fig. 9 as the relevant information. Representative images are just illustration. Integrating them with the graphs would make it necessary to split everything up into multiple figures, making it harder to compare the different combinations. Therefore, we left the figures unchanged.

      (24) In Figure S7, the Rab11 signal looks continuous, which makes the colocalization analysis meaningless. The authors should determine how to take images that can be evaluated. On a more minor note, the zoomed panels should be labeled as well.

      This is a result of having shown a projections of multiple planes. The images have now been replaced by single plane images. Zoomed panels have been labelled and the scale bar added.

      (25) The low colocalization of VPS35L with Rab5 is surprising, as SNX17 has been previously shown to co-localize with early endosomes positive for EEA1. This result may have occurred due to overexpression because the authors chose to utilize plasmids that express a tagged protein. There are antibodies to each of the endogenous proteins, and this is what should be used for this set of experiments.

      This comment made us control the analysis performed for these images, which by mistake had been performed on z-projections rather than on single planes. This distorted the values. The re-analysed data shows a higher colocalisation with Rab5, but it remains inferior to colocalisation with Rab11.

      (26) The authors should determine whether β1-integrin colocalizes with WIPI2 in endosomal compartments.

      This was done. WIPI2 colocalizes with beta-integrin on EEA1-and SNX17-positive strcutures but not positive for LAMP1 (Fig. 3E/F).

      Minor points

      (27) In one of the panels in Figure 1A, "30 min" is duplicated.

      Removed

      (28) In Figures 5C and 5D, the y-axis should indicate that this is surface β1integrin.

      Changed and added “surface”

      (29) In Figure 9 there is a typo in panel A. It is VPS35L and not VPS35.

      Corrected

      Reviewer #3 (Recommendations for the authors):

      This is an overall convincing study, which shows that the two complexes, CROP1 and CROP2 function at different membranes and serve different substrates. While I agree with their localization analysis, I have one key issue. The authors claim that each of the two forms a complex and base this on their specific pull-down and western blot analyses.

      I find it important that they show that both indeed form stable complexes in vivo, using pull-down and mass spectrometry approaches. They have all the necessary tools in hand and could use WIPI1 and WIPI2 to demonstrate the existence of the two complexes. The FSSS mutants of each are good controls for such an analysis.

      The manuscript actually presents the demanded in vivo experiments. Figs. 6 to 8 show pull-downs of WIPI1 and WIPI2 from cells, including also the FSSS mutant. While we haven't analysed this interaction by mass spectrometry, the Western blot analysis confirms the analysis. Cooperation of these proteins is further supported by the in vivo phenotypes, where the S67A substitution in WIPI2 produces a similar phenotype on integrin beta1 localisation as inactivation of Retriever.

      A second aspect is the general presentation. The paper would be a lot more accessible if the subunits of each complex (CROP1 and CROP2) were also introduced in the figures of each part. For readers, a final model is helpful to put the data into context and show where each complex operates in the cell.

      We have introduced a scheme of the respective complexes, including the names of the compunds, in Figs. 6 and 7 to avoid confusion.

      Finally, it is not clear how the statistics compare to repeats in their data. This should be clarified.

      This had been described in methods. Statistics has always been done on biological replicates stemming from independent experiments. We have added a cartoon (Fig. 10) depicting the trafficking pathways affected by CROP1 and CROP2.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Comments from Reviewing Editor:

      I want to share that both reviewers appreciated that this revision has appropriately addressed many of the concerns they raised. However, reviewers concurred that additional wet-lab experiments which validated the findings would have made the work much more impactful; and their concerns about the quality of chromatin accessibility data appear not to be fully resolved. Might I suggest a textual revision that specifically points out these caveats, if you are not able to provide additional data? This would then proceed to VOR without additional need to review. Thanks much for your patience while I assessed the manuscript claims and reviewer opinions.

      The changes were very minor (2 sentences in the Discussion and a small section in the Supplementary Notes). It would be great if we could proceed to the VOR stage.

    1. Author response:

      We appreciate the time and attention to our manuscript and the feedback from the reviewers, who were overall supportive of the work. Both reviewers validated the technical approach we used to differentiate the wild-type (WT) and knockout (KO) neurons noting: “The combination of sparse Cre delivery with channel rhodopsin-mediated optotagging in Npas4 fl/fl:Ai32 mice is technically elegant” and “the rigorous optogenetic tagging strategy used to distinguish KO from WT neurons in vivo makes the single-cell comparisons much more convincing.” Furthermore, they note the consistency of the reported results, stating: “The reported phenotype is internally consistent and converges on a coherent story”.

      Both reviewers also pointed out several concerns or points of improvement for the manuscript. Below, we first offer several scientific and methodological clarifications that we believe resolve a number of the reviewers' concerns. We then outline which remaining points we plan to address through revision, and which fall outside the scope of the current study.

      Scientific Clarifications:

      Request for a standard housing control. Both of the reviewers brought up the long-term enrichment paradigm (EE) we opted to use for this study and expressed interest in seeing data from standard housed (SE) animals. This is an approach the lab has taken in its slice physiology work [1-3], where comparing EE and SE conditions has revealed important differences between cellular phenotype. However, the in vivo experiments described here differ in a key way: obtaining these recordings requires extensive handling, training, and daily transport between the vivarium, home cage, and behavior room. These experimental steps themselves constitute the kind of novel, salient experience known to induce NPAS4, making a true SE comparison unattainable within this paradigm. In our experiment, mice were housed in EE as a supplemental, well-established strategy to induce NPAS4 in CA1 pyramidal neurons but we believe the behavior alone would be sufficient. We will describe this more clearly in the text of the manuscript.

      Consistent with this view, place fields recorded from wild-type mice in other studies using SE but undergoing comparable handling and training procedures, are similar in size, spatial information, and stability to the WT place fields we reported here [4,5]. As part of our revisions, we will consider statistical comparisons between our WT neurons and those reported in other studies to quantitatively assess whether a difference exists.

      More broadly, we note that the existing literature on NPAS4 induction does not, to our knowledge, establish a baseline level of NPAS4 expression in CA1 pyramidal neurons in the complete absence of behavioral experience. Reports of NPAS4 expression in CA1 have generally relied on animals exposed to some form of salient or novel experience [3,6,7], consistent with our framework that NPAS4 induction reflects behaviorally-driven activity rather than a constitutive baseline.

      Expression profile of NPAS4. Reviewer #2 brought up a concern about the extent of the NPAS4 expression, referring to the IHC results in Figure 1A stating: “Even under EE, only a few percent of CA1 pyramidal neurons express detectable NPAS4 at any given moment (Figure 1A), yet the AAV strategy deletes the gene in 30 to 60 percent of pyramidal neurons. In effect, the majority of cells classified as KO in this study would not have been expressing the protein under the relevant conditions.” We wish to clarify two points here. First, in the experimental paradigm used to obtain the IHC results, mice were exposed to enrichment for only 90 minutes while in the in vivo physiology paradigm, mice were housed in an enriched environment (with frequent toy changes to ensure novelty) for weeks. Thus, NPAS4 is almost certainly expressed in a much larger percentage of WT neurons in mice that were kept in chronic enrichment and used for the in vivo studies. Second, while the NPAS4 protein is only expressed in cells for several hours following neuronal activity, it initiates an inhibitory synapse phenotype that persists long-term. Thus, even though a small percentage of neurons are NPAS4+ in the IHC results, it is likely that a much larger percentage of them have expressed NPAS4 in the past and now show the inhibitory synapse phenotype. Evidence for this comes from the slice physiology results in Figure 1C (and see similar results from adolescents [1-3]) in which animals were housed in enrichment long-term and differences between inhibition persisted in nearly every WT/KO comparison.

      We also recognize the related possibility that NPAS4 expression may not be uniform across the pyramidal cell population, but may instead concentrate in particular functional subtypes, such as cells with higher firing rates or stronger spatial tuning. As part of our revisions, we plan to test this directly by stratifying the KO population by firing rate and relating it to the magnitude of the observed phenotype. Taken together, we believe that while only a small fraction of CA1 pyramidal neurons are NPAS4+ at any given moment, a much larger fraction have experienced NPAS4 induction and the accompanying synaptic reorganization over the timescale of chronic enrichment making the WT/KO comparison in this study substantially less diluted than the IHC snapshot alone would suggest.

      Timeline of NPAS4 expression and synaptic reorganization. Reviewer #1 pointed out that this study only examines the effects of NPAS4-deletion on longer timescales (weeks to months after the virus expression and subsequent knockout) stating “[the study] is less definitive about the immediate causal sequence by which NPAS4 induction alters inhibition and reshapes spatial and temporal coding”. The reviewer is correct, the temporal relationship between NPAS4 expression, changes in synaptic inhibition, and changes in neuronal firing are important outstanding questions in the field. Currently, we lack molecular tools that would enable us to clearly test these relationships but with our existing, albeit limited information, we have the following working model.

      When an animal is placed into a new context, a subset of CA1 pyramidal neurons will fire action potentials in a spatially refined manner. This activity will drive NPAS4 expression in those neurons, resulting in protein expression that persists for a couple of hours before the protein is degraded.

      Following expression, NPAS4 will bind to various sites in the genome and initiate a genetic program which results in changes in inhibition recruiting CCK basket cell synapses to the soma and destabilizing CCK dendritic synapses. The exact mechanism behind this reorganization of inhibition is unknown, but the phenotype likely emerges over the course of several hours following NPAS4 expression and persists for days following the stimulus that induced NPAS4.

      While our chronic knockout approach does not allow us to resolve the precise timing of events in this sequence, it does allow us to ask a distinct and complementary question: what is the long-term consequence for a neuron that has never been able to execute this program? Our results demonstrate that NPAS4-deficient neurons which cannot initiate NPAS4-dependent inhibitory reorganization regardless of their activity history show systematic degradation in spatial and temporal coding precision. This establishes that the NPAS4-dependent inhibitory phenotype has lasting and functionally meaningful consequences for in vivo information encoding, a question that shorter-timescale or acute manipulations would not be well-positioned to address. Resolving the immediate causal sequence between NPAS4 induction, synaptic reorganization, and changes in firing will be an important goal for future work as new molecular tools become available.

      Behaviors that drive NPAS4 expression. Reviewer #2 pointed out that “NPAS4 is also induced by contextual fear conditioning and other paradigms which would predict context-specific effects rather than a uniform refinement function.” They are correct NPAS4 is expressed in response to different behavioral paradigms, including fear conditioning and environmental enrichment. However, the subregion in which NPAS4 is induced depends critically on the behavioral paradigm. When mice are exposed to contextual fear conditioning, NPAS4 expression is robust in CA3 and the dentate gyrus but negligible in CA1 [6]. This is consistent with the known activity patterns of these subregions: CA3 neurons are strongly recruited during contextually-dependent associative learning, while CA1 neurons are more reliably driven by exposure to novelty and respond in a spatially-refined manner. Consistent with this, studies using fear conditioning have focused on behavioral discrimination and synaptic changes in CA3 and granule cells [6]. To our knowledge no study has examined the relationship between fear conditioning, NPAS4, and CA1 pyramidal neuron function. Whether behavioral paradigms beyond environmental enrichment and spatial navigation can induce NPAS4 in CA1, and what consequences that might have for pyramidal neuron firing, are interesting questions for future work.

      We also wish to address the conceptual framing underlying this concern. In CA1, we do not believe that “context-specific effects” are separable from a “uniform refinement function.” CA1 pyramidal neurons respond in a context-dependent manner. When a mouse is placed onto a linear track, there is a subset of neurons that will increase their activity over the course of that exposure. But within this subset, individual neurons will also show spatially-refined responses firing action potentials as the animal runs through the corresponding place field. The spatial precision NPAS4 confers is always nested within context-dependent mechanisms NPAS4 refines whatever representation a neuron is already computing, rather than overriding the context-dependency of that representation. We therefore do not view these as competing frameworks.

      The role of NPAS4 in shaping CCK synapses. Reviewer #2 made the point that “the CCK to pyramidal cell connectivity that the authors invoke as the mechanistic anchor is also dense in standard housing, so the absence of detectable NPAS4 in SE conditions raises the further conceptual problem of how NPAS4-negative neurons would normally be innervated by CCK+ basket cells in the first place.” We wish to clarify that NPAS4 is not necessary for the formation of CCK synapses onto CA1 pyramidal neurons there are likely a number of NPAS4-independent mechanisms that regulate this synaptic connectivity (for example, see [8]). Rather, we place NPAS4 in the role of an activity-dependent modulator that acts on top of this baseline connectivity: when NPAS4 is expressed in response to neuronal activity, it shifts the balance of CCK inhibitory input along the somatodendritic axis, increasing somatic and decreasing dendritic CCK synaptic strength [1,2]. The question is therefore not how CCK synapses are established in the absence of NPAS4, but rather how experience-dependent activity uses NPAS4 to fine-tune the distribution of those synapses and it is this fine-tuning that our study links to the precision of in vivo spatial and temporal coding.

      Methodological Clarifications:

      Clarification on how stability analysis was performed. Reviewer #2 requested additional analysis for the stability results: “A control analysis using a fixed reference window around the original peak, rather than re-identifying the peak each epoch, would help distinguish a genuine plasticity-like shift from instability driven by noise.” We wish to clarify that this is precisely the methodology that was used in the manuscript. For the stability analysis shown in Figures 4C-E, the activity was aligned to the peak activity in epoch 1 such that 0 always represents the location of the peak in epoch 1. This approach allows us to identify how that activity differs in subsequent epochs, namely whether it has shifted relative to the activity in epoch 1. We will make this more clear in the results and methods sections.

      Request for Ai32 control. Reviewer #2 made the point that “The comparison throughout the manuscript pits Cre+ ChR2+ neurons (NPAS4 KO) against neighboring non-transduced neurons (WT). This is internally elegant, but leaves open the possibility that part of the phenotype arises from chronic ChR2 expression or constitutive Cre activity rather than from NPAS4 loss, especially given that most of the readouts are subtle.” We agree this would be the ideal control and regret that it is no longer experimentally feasible, as the laboratory in which these experiments were conducted is no longer operating. However, we believe several features of the existing dataset make a ChR2 or Cre artifact unlikely. First, the effects of chronic ChR2 expression are not known to produce the specific pattern of phenotypes we observe in particular the redistribution of somatic versus dendritic inhibition, which is recapitulated independently in acute slice recordings from animals that did not undergo optotagging procedures (Figure 1C). Second, the phenotype we report is internally coherent across multiple independent metrics: place field size, stability, signal-to-noise ratio, theta coupling, and phase precession all shift in the same direction, in a manner consistent with a specific change in inhibitory synaptic balance rather than a nonspecific effect of transgene expression. Third, the sparse nature of the Cre expression means that KO and WT neurons share the same local network, same LFP, and same behavioral context any network-level effect of Cre or ChR2 would be expected to affect both populations similarly. We will add a discussion of these points to the manuscript.

      PSTH clarification (unit of opto-response). To quantify the opto-response, we treated each light-on + light-off period (a total of 2 seconds) as the one trial. We aligned the trials by the light-on period, binned the spikes by 1 msec bins, and then summed the responses across trials to produce a histogram. From this histogram we found the maximum response during light off (e.g. the 1 msec bin with the greatest response which should be reported as number of spikes). We subtracted this from the maximum response during light on. Thus, the unit of opto-response should be spike counts. We will clarify this in the text and figures.

      Use of male mice. Reviewer #1 rightfully pointed out that this study only used male mice. In this study, we only used mice that were larger than 20 grams to ensure the mice could carry the weight of the implanted drives while performing the behavior. As this genetic line of mice is on the smaller size, only male mice were above this weight threshold. Importantly, slice work conducted in the Blood good lab has not identified sex differences in NPAS4 phenotypes [3,9]. Future studies would benefit from the use of both male and female mice. We will state this more explicitly in the text and expand on the potential implications of excluding female mice from our study.

      Future planned changes to manuscript:

      As the reviewers suggested, we intend to add the following analyses and make the following changes to the manuscript:

      Stratify key analyses (stability, theta coupling, phase precession) by FR to determine whether there is a dependency on the firing rate of cells.

      Apply hierarchical bootstrapping and add per-animal color-coding to supplementary figures to assess animal-level variability and protect against pseudoreplication.

      Add a circular-linear phase-position correlation analysis as an additional quantification of phase precession strength, complementing the existing slope-based analysis.

      Improve discussion around the temporal phenotype being downstream of the spatial one.

      Tighten mechanistic framing in the Discussion to more clearly distinguish what is demonstrated in this study from what is inferred from prior work, and to acknowledge the contributions of other inhibitory cell types.

      Minor changes and figure clarifications as noted by reviewers.

      Outside of the scope of this study or unable to be performed:

      There were several recommendations or points that the reviewers brought up that we do not have the resources to address. Nevertheless, we appreciate the reviewers noting these.

      SE control (as discussed above)

      Ai32 control (as discussed above)

      Behavioral consequences of NPAS4 knockout and the effects on learning and memory • Ripple analysis

      Drift observed in E4 and what this might look like over larger timescales

      Comparison between male and female mice to determine whether there are sex-dependence differences

      In conclusion, the reviewers recognized this as a well-designed and internally consistent study. We believe that many of the critiques including the request for a standard housing control, questions regarding the extent of NPAS4 expression across the pyramidal cell population, and points about the timeline of NPAS4 expression and synaptic reorganization are addressed by the clarifications provided in this response. We agree with many of the suggested analytical and textual changes and look forward to incorporating those into the revised manuscript.

      References:

      (1) Heinz, D. A., Cui, W., Cooper, K. L. & Bloodgood, B. L. Experience-induced NPAS4 reduces dendritic inhibition from CCK+ inhibitory neurons and enhances plasticity. J. Neurophysiol. 134, 361–371 (2025).

      (2) Hartzell, A. L. et al. NPAS4 recruits CCK basket cell synapses and enhances cannabinoid-sensitive inhibition in the mouse hippocampus. Elife 7, (2018).

      (3) Bloodgood, B. L., Sharma, N., Browne, H. A., Trepman, A. Z. & Greenberg, M. E. The activity dependent transcription factor NPAS4 regulates domain-specific inhibition. Nature 503, 121–125 (2013).

      (4) Sharif, F., Tayebi, B., Buzsáki, G., Royer, S. & Fernandez-Ruiz, A. Subcircuits of deep and superficial CA1 place cells support efficient spatial coding across heterogeneous environments. Neuron 109, 363–376.e6 (2021).

      (5) Quirk, C. R. et al. Precisely timed theta oscillations are selectively required during the encoding phase of memory. Nat. Neurosci. 24, 1614–1627 (2021).

      (6) Ramamoorthi, K. et al. Npas4 regulates a transcriptional program in CA3 required for contextual memory formation. Science 334, 1669–1675 (2011).

      (7) Chiaruttini, N. et al. ABBA+BraiAn, an integrated suite for whole-brain mapping, reveals brain-wide differences in immediate-early genes induction upon learning. Cell Rep. 44, 115876 (2025).

      (8) Früh, S. et al. Neuronal Dystroglycan Is Necessary for Formation and Maintenance of Functional CCK-Positive Basket Cell Terminals on Pyramidal Cells. J. Neurosci. 36, 10296–10313 (2016).

      (9) Lin, Y. et al. Activity-dependent regulation of inhibitory synapse development by Npas4. Nature 455, 1198–1204 (2008).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In this paper, the authors use a doxycycline-inducible DLD1 cell line expressing a Clover-tagged RNA-binding-defective TDP-43 2KQ mutant that forms nuclear "anisosomes" (TDP-43 shell with HSP70 core) to carry out a small-molecule screen using the LOPAC 1280 library to identify compounds that reduce anisosome number or shift their morphology and dynamics. They also conducted a genome-wide siRNA screen to identify genetic modifiers of anisosome formation and dynamics. From these screens, the authors identify pathways in RNA splicing, translation, proteostasis (proteasome and HSP90), and nuclear transport, including XPO1. They then focus on XPO1 as their primary hit. Pharmacological inhibition of XPO1 using KPT-276, Verdinexor, and Leptomycin B reduces anisosome number while enlarging remaining condensates, which retain liquid-like behavior by FRAP and fusion assays. XPO1 overexpression causes fewer, enlarged TDP-43 puncta, including cytoplasmic puncta, with little or no FRAP recovery, interpreted as gel or solid-like aggregates. Anisosome induction reduces detectable nucleoplasmic XPO1 staining. Finally, the authors examine a homozygous TDP-43 K181E iPSC-derived forebrain organoid model, showing increased cytosolic pTDP-43 in K181E/K181E organoids compared to wild-type controls. Chronic low-dose KPT-276 reduces cytoplasmic pTDP-43 without changing total TDP-43 levels. Bulk RNA-seq shows only a modest fraction of dysregulated genes in K181E/K181E organoids are rescued by KPT-276. They conclude that nuclear export, via XPO1, is a key regulator of TDP-43 liquid-to-solid phase transitions and that cytoplasmic aggregation per se may contribute only modestly to TDP-43 proteinopathy, with RNA-processing defects being dominant.

      We thank the reviewer for carefully summarizing our study.

      The study presents well-executed chemical and genome-wide siRNA screens in a DLD1 TDP-43 2KQ anisosome model and follows up on nuclear transport, particularly XPO1, as a modulator of TDP-43 phase behavior and cytoplasmic aggregation. The screens are impressive in scale, and the microscopy and fluorescence recovery after photobleaching (FRAP) work is technically strong. However, the central mechanistic and disease-relevance claims are not yet sufficiently supported. There are major concerns about the heavy reliance on non-physiological, RNA-binding-defective, and acetylation-mimetic TDP-43 (2KQ) and a homozygous TDP-43 K181E organoid model. An underdeveloped and partly contradictory mechanistic link exists between XPO1 and TDP-43 phase transitions in the context of prior work showing TDP-43 is not a canonical XPO1 cargo. The paper also appears to overinterpret organoid data to conclude that cytoplasmic TDP-43 aggregation plays only a minor role in pathology, based largely on pTDP-43 antibody staining with limited sensitivity and relatively modest rescue readouts. A deeper mechanistic analysis and additional, more physiological validation are needed for this to reach the level of rigor and impact implied by the title and abstract. The work feels screen-rich but conceptually underdeveloped, with key claims outpacing the data. A major revision with substantial new data and tempering of conclusions is warranted. I outline several problematic areas below:

      (1) The central mechanistic discoveries are derived almost entirely from a DLD1 colon cancer cell line overexpressing an RNA-binding-defective, acetylation-mimetic TDP-43 2KQ mutant and homozygous TDP-43 K181E iPSC-derived organoids. Both systems are far from physiological. The 2KQ mutation is a synthetic double lysine-to-glutamine mutant originally designed to mimic acetylation and disrupt RNA binding. In this study, essentially all cell-based mechanistic data on phase behavior, screens, and XPO1 effects rely on 2KQ. Yet there is no quantification of how much endogenous TDP-43 is acetylated in degenerating human neurons, nor whether a 2KQ-like acetylation state is ever achieved in vivo. It is not established that the phase behavior of 2KQ recapitulates the physiological or pathological phase behavior of wild-type TDP-43 or genuine disease-linked mutants, which may retain partial RNA binding and different post-translational modification patterns. As a result, it is difficult to know whether the modifiers identified here regulate a highly artificial 2KQ condensate or physiologically relevant TDP-43 condensates. To address this concern, the paper would benefit from quantifying endogenous TDP-43 acetylation at the relevant lysines in control and ALS/FTD patient tissue or more disease-proximal models such as heterozygous TARDBP mutant iPSC neurons, which would justify the focus on an acetyl-mimetic mutant. Key phenomena, including XPO1 dependence of phase behavior, effects of proteasome and HSP90 inhibition, and effects of splicing and translation inhibitors, should be tested for wild-type TDP-43 expressed at near-physiological levels and for one or more bona fide ALS/FTD-linked TARDBP mutants that are not acetyl mimetics. At a minimum, the authors should show that endogenous TDP-43 in neuronally differentiated cells exhibits qualitatively similar responses to XPO1 modulation, rather than exclusively relying on DLD1 2KQ overexpression.

      Acetylation of endogenous TDP-43 was reported by several studies. Although it occurs at low levels under normal conditions, TDP-43 acetylation is upregulated under stress conditions (e.g. oxidative stress and proteotoxic stress) (PMID: 25556531; PMID: 28724966). Importantly, Cohen et al. reported the identification of acetylated TDP-43 in ALS patient spinal cord (PMID: 25556531), while Yu et al. showed that endogenous wildtype TDP-43 undergoes demixing when neurons were treated with either a deacetylase inhibitor or proteasome inhibitor (PMID: 33335017). These studies also show that acetylated TDP-43 is defective in RNA binding and more prone to aggregation. Furthermore, ectopic expression of acetylated TDP-43 mimetics in cells and mice induces cellular defects similar to those observed in disease models (PMID: 28724966). Thus, our findings, based on previously established TDP-43 mimetics, should provide valuable information regarding the phase regulation of a disease-relevant TDP-43 mutant. We have included more background information to justify the use of TDP-43 acetylation mimetics in the introduction.

      (2) The organoid model is based on a homozygous K181E knock-in line. However, in patients, TARDBP mutations are overwhelmingly heterozygous. Homozygosity is thus a severe, arguably non-physiological sensitized background that may exaggerate nuclear RNA mis-splicing and phase defects and alter the relative contribution of cytoplasmic aggregation versus nuclear loss-of-function. In addition, it is not fully clear from this manuscript whether the structures in K181E organoids are bona fide anisosomes as defined in Yu et al. 2021, characterized by HSP70-enriched central liquid cores with TDP-43 shells and similar FRAP and fusion behavior to anisosomes in the DLD1 model. At present, the organoid section is framed as validation of "anisosome-bearing organoids," but the figures in this manuscript mainly show pTDP-43 puncta and total TDP-43 immunostaining, without detailed structural or biophysical characterization. The authors should explicitly compare heterozygous K181E/+ organoids or another heterozygous TARDBP mutant line with homozygous K181E/K181E organoids to assess whether XPO1 inhibition has similar effects in a genotype that more closely resembles patient genetics. They should provide direct evidence that the K181E condensates in organoids are anisosomes through HSP70 core immunostaining, three-dimensional reconstruction, and FRAP measurements, and clarify whether KPT-276 is acting on anisosome-like structures or more generic cytoplasmic aggregates or puncta. Without this, the leap from a DLD1 2KQ cancer cell model to human ALS/FTD-relevant neurons is not convincingly supported.

      The reviewer is correct that the use of homozygous K181E organoids generates a background that is more sensitive for detecting phospho-TDP-43. The goal was to test whether XPO1 inhibition mitigates the phosphorylation of a TDP-43 disease mutant. For this purpose, we believe that our experimental setup is suitable. We agree that we should not extrapolate the result to over emphasize on its disease connection. We have revised the paper to tone down this section. We also remove the RNAseq data as it is not essential for our conclusions.

      It is also noteworthy that TDP-43 disease mutations are usually loss-of-function alleles. Although heterozygous background is sufficient to induce disease phenotype in aged humans, heterozygous background in experimental settings is usually unable to generate severe defects. Thus, it is quite common to study TDP-43 disease-related defects in homozygous knockout or RNAi-mediated depletion conditions (e.g. PMID: 35197626; 41120751; 38277467).

      Regarding the immunostaining signals in K181E organoids, we did not report them as anisosomes. As documented in the literature, p-TPD-43 is widely used as a marker to indicate pathological TDP-43 aggregation. P-TDP-43 is enriched in pathological aggregates in human ALS and FTD patients, colocalized with other aggregation signatures such as ubiquitin and other aggregation-prone proteins in the cytoplasm (PMID: 36008843), and is being used as a diagnostic marker for neurodegeneration (PMID: 31661037). The characterization of K181E organoid is reported in a pre-print by Zhang Q. et al., 2026 (PMID: 41292965), which is currently under revision for Science Advances. In Fig. 1I of this manuscript, we confirmed the cytosolic localization of p-TDP-43 in cells that were isolated from K181E organoids. In the current manuscript, Figure 7 is to show that nuclear export inhibition mitigates the accumulation of p-TDP-43 in a brain-like tissues. We revise the subheading and the corresponding text to avoid the confusion.

      (3) The title and framing assert that "nuclear export governs TDP-43 phase transitions." However, prior studies such as Pinarbasi et al. 2018 and Duan et al. 2022 indicate that TDP-43 is not a canonical XPO1 cargo and that its export is largely passive, with active nuclear import being the dominant determinant of nuclear localization. The authors cite these studies but still position XPO1 as a central, quasi-direct regulator. The data presented are largely correlative or based on pharmacologic manipulation and overexpression in an overexpression mutant background, with no direct evidence that XPO1 engages TDP-43 in a specific, regulated manner. Even if XPO1 does not engage WT TDP-43, it could still engage the 2KQ variant, which needs to be tested.

      We did not mean to conclude or imply that the regulation of TDP-43 by XPO1 is direct. In fact, we explicatively mentioned on page 8 of the original manuscript that the regulation is likely indirect and mediated by other factors. The sentence reads as “Since XPO1 does not bind TDP-43 directly (Pinarbasi et al., 2018), additional factors might link XPO1-mediated nuclear export to TDP-43 nuclear egression.”

      We now add new data in Figure 6, showing that in an in vitro reconstitution assay using semi-permeabilized cells, LMB treatment significantly stabilizes anisosomes in an RNA dependent manner. This new data suggests that XPO1 inhibition leads to increased nuclear RNA availability, which indirectly favors anisosome assembly and maturation (see discussion). We believe that this new finding has provided significant new insight into how nuclear transport modulates TDP-43 phase behavior. We have revised the title, the abstract and changed the framing according to the reviewer’s suggestion.

      (4) The XPO1 perturbations yield somewhat confusing phenotypes. XPO1 inhibition using Leptomycin B, KPT-276, and Verdinexor reduces anisosome number and enlarges remaining anisosomes, which remain liquid-like by FRAP recovery and fusion assays and stay nuclear. XPO1 overexpression causes fewer, enlarged puncta, but these are FRAP-impaired (gel-like) and redistribute to the cytoplasm. Thus, both decreased and increased XPO1 activity reduce anisosome number and enlarge puncta, but with opposite phase behaviors and subcellular localizations. The model presented in Figure 5L is relatively qualitative and does not resolve these issues. Moreover, XPO1 inhibition globally impairs nuclear export of many cargos and profoundly alters the nuclear environment, transcription, RNA processing, and chromatin. It is therefore difficult to conclude that the observed effects are specific to TDP-43 phase regulation as opposed to secondary consequences of broad nuclear export blockade.

      The reviewer correctly summarizes our data and interpretation: XPO1 loss-of-function and gain-of-function generate opposite phenotypes regarding TDP-43 phase regulation.

      Regarding the mechanism underlying XPO1-dependent TDP-43 phase regulation, as mentioned above, we developed a semi-permeabilized cell-based assay in which we used the pore-forming toxin streptolysin O to damage the plasma membrane after anisosome induction. We noticed that upon cell permeabilization and cytosol loss, anisosomes were mostly lost (Figure 6B, C). This is probably due to a reversible partition of TDP-43 into a less fluorescent soluble fraction. Supporting this idea, when permeabilized cells were incubated with cytosol plus an energy regenerating system, small puncta containing TDP-43 2KQ could be reformed in an energy dependent manner (Figure 6D, E). Interestingly, in LMB-treated cells, anisosomes remained stable despite cell permeabilization(Figure 3F). Since LMB treatment did not increase TDP-43 nuclear concentration (Supplemental Figure 1), this data suggest that nuclear export inhibition likely alter the nuclear environment to stabilize anisosomes. Indeed, when cells were permeabilized in the presence of a small RNAase, LMB-stabilized anisosomes also collapsed (Figure 6G).

      We now add more discussions on the potential effect of RNA on TDP-43 phase behavior in XPO-1 inhibited cells considering these new findings.

      (5) The authors show that anisosome induction depletes nucleoplasmic XPO1 signal and that mCherry-XPO1 can be seen in some TDP-43 puncta. However, antibody penetration into anisosomes is limited, so XPO1 depletion from nucleoplasm could reflect sequestration in the anisosome shell or core, but this is not demonstrated. There is no demonstration of physical interaction, even indirect interaction, between XPO1 and TDP-43 or a defined adaptor, nor identification of a specific mutant of XPO1 that selectively disrupts this putative interaction while preserving other functions. The known TDP-43 NES has been shown to be weak and not a functional XPO1-dependent NES in multiple studies. If XPO1 is acting through an adaptor that recognizes 2KQ or K181E specifically, that by itself would bring into question the generality of the mechanism for wild-type TDP-43.

      We agree that our data does not demonstrate an interaction between XPO1 and TDP-43. Considering our new data (mentioned above), it is possible that the effect of anisosome induction on endogenous XPO1 localization is also mediated by RNA. We now mention more explicitly that the regulation of TDP-43 by XPO1 is likely indirect (Page 8). We have revised our paper to separate any speculative statements from the data, and also discussed the possibility of alternative interpretations.

      (6) To support a mechanistic claim that nuclear export governs TDP-43 phase transitions, more targeted evidence is needed. The authors should test whether siRNA knockdown or CRISPR interference of XPO1 in the DLD1 2KQ model reproduces the effects seen with Leptomycin B and KPT-276, including FRAP and fusion phenotypes, and verify on-target effects by rescue with an siRNA-resistant XPO1 construct. They should demonstrate that canonical XPO1 cargos behave as expected under the inhibitor conditions used, as a positive control, and that the concentrations used are not grossly toxic. They should attempt to identify or at least constrain candidate adaptors that might enable XPO1-dependent export of TDP-43 through proteomic analysis of XPO1 co-purifying with 2KQ condensates or loss-of-function studies of candidate adaptors from the siRNA screen. Finally, they should test whether a TDP-43 mutant that cannot bind the proposed adaptor still responds to XPO1 manipulation.

      The anisosome enlargement phenotype upon XPO1 depletion was seen in our siRNA screens, which was identified by machine-based image analyses using 6 different siRNAs. This, together with the chemical inhibition experiments, demonstrate that the phenotype is specifically caused by XPO1 inactivation.

      When characterizing the effect of XPO1 inhibition on anisosome dynamics, we preferred chemical inhibitor because the effect is acute, and therefore less likely to be secondary.

      Regarding the inhibitor concentration, according to the literature, Leptomycin B was commonly used at 50-200 nM. We chose 200 nM to ensure a quick and complete inhibition of XPO1-mediated nuclear export (see Figure 3 in PMID: 9628873). This dose is also well tolerated by our cells.

      We did not suggest any specific adaptor that mediates XPO1 interaction with TDP-43. Whether there is an adaptor, and if so, the identity of such adaptor is out of the scope of this study. We revise our paper on page 8-9 to clarify these points.

      (7) Even with these data, what is currently shown is that global modulation of nuclear export capacity can alter the phase behavior and localization of a highly overexpressed RNA-binding-defective TDP-43 mutant and of K181E in organoids. This is important, but it is weaker than asserting that XPO1 directly governs TDP-43 phase transitions in physiological contexts. The title, abstract, and Discussion should be tempered to reflect that nuclear export is one of several pathways, alongside RNA splicing, translation, and proteostasis, that influence TDP-43 phase states in this model, and that the specific mechanism and cargo relationship between XPO1 and TDP-43 remain unresolved and may be indirect.

      We have revised the title, abstract, and main text to temper our conclusions.

      (8) The authors conclude that cytoplasmic TDP-43 aggregation plays only a modest role in TDP-43 proteinopathies because in homozygous K181E organoids, chronic KPT-276 treatment almost abolishes cytoplasmic pTDP-43 puncta, yet bulk RNA-seq shows only a relatively small fraction of dysregulated genes are rescued. There are several issues with this inference. Relying primarily on pTDP-43 antibody staining to define cytoplasmic TDP-43 aggregation is limiting. pTDP-43 antibodies label only phosphorylated species and may miss non-phosphorylated, oligomeric, or amorphous TDP-43 species that could still be toxic. Different pTDP-43 antibodies vary in epitope accessibility depending on aggregate conformation and subcellular location. More sensitive approaches, such as high-affinity TDP-43 RNA aptamer probes developed by Gregory and colleagues, biochemical fractionation for SDS-insoluble and urea-soluble TDP-43, and filter-trap assays, would provide a more quantitative assessment of cytoplasmic aggregation and its reduction by KPT-276. Without these, it is not safe to assume that cytoplasmic aggregation has been eliminated, as opposed to one antigenic subclass.

      We agree with the reviewer that p-TDP-43 may not represent all aggregate species. However, p-TDP-43 antibodies detect the pathologically validated species tightly associated with TDP-43 proteinopatheis. In human ALS and FTD-TDP tissues, cytoplasmic inclusions are strongly immunoreactive for phosphorylated TDP-43 (typically S409/410, as detected here). Additionally, p-TDP-43 immunohistochemistry is a routine diagnostic criterion in neuropathology. For these reasons, we believe that the observation that inhibition of XPO1 significantly reduces p-TDP-43 is a significant finding, as it suggests that inhibition of nuclear transport may rescue TDP-43 proteinopathy. We revised the text on page 9 to better explain the significance of p-TDP-43 staining.

      (9) The treatment window, spanning from day 87 to 122 with 20 nanomolar KPT-276, may be too late or too mild to reverse entrenched nuclear RNA-processing defects, even if cytoplasmic inclusions are cleared. Once widespread cryptic exon inclusion and alternative polyadenylation misregulation are established, many downstream changes may become self-sustaining or only partially reversible. Moreover, XPO1 inhibition will massively rewire nucleocytoplasmic transport of many transcription factors, splicing factors, and RNA-binding proteins. Thus, the lack of full transcriptomic rescue cannot be cleanly interpreted as evidence that cytoplasmic aggregates are only modest contributors. It may instead reflect that nuclear dysfunction is primary and XPO1 inhibition does not correct, and may even exacerbate, certain nuclear defects.

      We agree with the reviewer that the lack of rescue may be caused by some technical issues. We have removed the RNAseq data and the related texts since it is not essential.

      (10) To support a causal statement about the modest contribution of cytoplasmic aggregates, one would want more direct measures of neuronal health and function, such as cell death, neurite complexity, synaptic markers, and electrophysiology before and after KPT-276, not only transcriptomics. A way to selectively reduce cytoplasmic aggregation without globally inhibiting nuclear export would allow comparison of outcomes.

      We have removed the discussion regarding the role of cytoplasmic aggregates in disease.

      (11) Given these caveats, the concluding statements that cytoplasmic TDP-43 aggregation is only a modest contributor should be substantially softened. A more defensible interpretation is that in this homozygous K181E organoid model, chronic global XPO1 inhibition reduces pTDP-43-positive cytoplasmic puncta but only partially normalizes the steady-state transcriptome, suggesting that persistent nuclear RNA-processing defects and other pathways continue to drive pathology.

      We agree with the review and have removed the RNAseq part.

      (12) The screens are a major strength but need more rigorous validation for key hits, especially nuclear transport factors. For the siRNA screen, hits are filtered by anisosome number per nucleus, but there is no direct demonstration in the main text that XPO1 or CSE1L knockdown is efficient at the messenger RNA or protein level. For the highlighted genes, Western blot or quantitative polymerase chain reaction validation and phenotypic rescue would strengthen confidence. For small-molecule hits, it is not systematically shown that anisosome modulation is independent of changes in total TDP-43 2KQ expression or gross toxicity. Translation inhibitors are tested for this, but for many other hits, including proteasome, HSP90, and kinase inhibitors, expression and general nuclear structure should be monitored. Given the reliance on anisosome count as a readout, secondary screens that specifically distinguish changes in TDP-43 expression levels, changes in nuclear morphology or cell cycle, and specific changes in anisosome phase behavior, including FRAP and fusion for top hits, would greatly increase interpretability.

      For the siRNA screen, each positive hit was confirmed by two rounds of screen with 6 independent siRNAs in total. Although we did not validate the knockdown efficiency due to the large number of hits, we routinely include a positive siRNA control in our study (Cell death siRNA), which targets several essential gene. Transfection efficiency was controlled by measuring cell viability after knocking down of these genes. In addition, the identification of XPO1 as a positive regulator of TDP-43 phase behavior was independently validated by our chemical genetic screens with three XPO-1 inhibitors. We feel confident that XPO1 is a key modulator of TDP-43 phase behavior.

      For chemical treatment experiments, the anisosome fusion phenotypes could be detected as early as 5 h post treatment. Given the relatively short treatment, we do not expect a significant change in protein level or toxicity. To alleviate this reviewer’s concern, we performed an immunoblotting experiment to measure the total TDP-43 protein levels in drug-treated cells. Except for VLX, we did not detect any significant changes in the level of TDP-43 after drug treatment (Supplemental Figure 1).

      (13) The classification of condensates as liquid versus gel-like or solid is based almost entirely on FRAP recovery or lack thereof. While FRAP is appropriate, interpretations could be made more robust by including half-region-of-interest bleach controls and assessing mobile fractions and recovery kinetics more quantitatively across conditions. Complementing FRAP with other phase-behavior assays such as sensitivity to 1,6-hexanediol, shape relaxation after deformation, and coarsening behavior over longer timescales would strengthen the analysis. At present, some assignments, such as that XPO1 overexpression drives a gel-like transition, are reasonable but somewhat qualitative.

      In this study, we used two types of FRAP assays. We either bleached TDP-43 within anisosomes or bleached the surrounding TDP-43 molecules(Figure 2). The two complementary methods yield consistent results that allow unambiguously distinguish between TDP-43 LLPS state and gel-like condensation.

      In XPO1-related experiments, the two types of condensates formed by TDP-43 2KQ can be distinguished by several features including their subcellular localization, shape, and the fluorescence recovery kinetics. We feel that these combined data clearly segregate these puncta into two distinct types of assemblies. The proposed half-region-of-interest bleach is technically challenging for small anisosomes under normal conditions. However, whenever possible, (e.g. anisosomes enlarged by Leptomycin B), we did perform both whole anisosome bleach and partial bleach (Figure 5D, I). Both assays demonstrate that TDP-43 in these enlarged anisosomes is highly mobile.

      (14) For the Leptomycin B and KPT-276 experiments in cells and organoids, it would be important to confirm that canonical XPO1 cargo proteins accumulate in the nucleus and that the concentrations used are within a range that is not overtly toxic over the experimental timeframe. Assessing nuclear morphology, chromatin condensation, and general transcriptional activity through global RNA synthesis or key reporter genes would ensure that observed effects are not secondary to severe global nuclear export collapse.

      In Leptomycin B treatment experiments, we carefully chose a dose that was previously validated (see Figure 3 in PMID: 9628873). Based on our DAPI staining, the nuclear morphology appears normal with no abnormal chromosome condensation (Figure 5A). Additionally, in cell line-based experiments, the effect of Leptomycin B on anisosomes was detected 6-8 hours post treatment. The change in global protein synthesis because of RNA changes should be relatively minor at this stage. Indeed, our new immunoblotting experiment showed that LMB treatment did not affect TDP-43 protein level (Supplemental Figure 1). Most importantly, the in vitro semi-permeabilized assay demonstrates a direct role for RNA in stabilizing anisosomes.

      (15) In the organoid section, it is not clear how many independent iPSC clones and organoid batches were used per condition, nor whether batch effects were assessed in the bulk RNA-seq analysis. This should be fully specified and ideally controlled with isogenic wild-type and K181E clones. For transcriptional rescue, it is important to know whether the changes in wild-type organoids treated with KPT-276 are negligible. A direct wild-type comparison with or without KPT-276 is important to disentangle general drug effects from K181E-specific rescue. More detailed quantification of total TDP-43 and pTDP-43 in both nuclear and cytoplasmic fractions, including biochemical fractionation if possible, would strengthen the assertion that KPT-276 specifically reduces cytosolic pTDP-43 aggregates while sparing nuclear TDP-43.

      The organoid experiment was performed with two batches per condition to reduce the effect of batch variation. The wildtype cells and K181E mutant are derived from the same genetic background. This information is now included in the method section on page 14. Given the criticisms by review 1 and 2 on the RNAseq data, we have removed this non-essential data. 

      (16) Beyond the core issues above, several additions could greatly enhance the impact. The manuscript currently emphasizes XPO1, but the genetic and chemical data clearly implicate RNA splicing, translation, and proteostasis as equally strong or stronger regulators of TDP-43 phase states. A more integrated model that explains how these pathways intersect, for example, how splicing factor availability, ribosome loading, and proteasome capacity co-govern anisosome nucleation, growth, and hardening, would be valuable.

      We now discuss a new model in discussion based on our new Figure 6, which integrates the role of RNA splicing and nuclear transport in TDP-43 phase regulation on page 10. We agree with the reviewer that other questions are also important for future studies.

      (17) A key unresolved question is whether XPO1 is acting directly on TDP-43, or instead primarily regulates anisosomes by exporting other factors that more proximally control TDP-43 phase behavior. Given that TDP-43 is not a canonical XPO1 cargo and prior work indicates that its nuclear export is largely passive, it seems at least as plausible that XPO1 inhibition alters the nuclear concentration or localization of splicing factors, RNA-binding proteins, chaperones, or other modifiers identified in the screens, and that changes in these proteins secondarily reshape anisosome dynamics. In other words, XPO1 may be exporting a more direct regulator of anisome formation and hardening, rather than exporting TDP-43 itself in a specific, regulated way. The current data do not distinguish between these possibilities. Systematic identification of XPO1-dependent cargos that colocalize with or biochemically associate with anisosomes, combined with targeted perturbation of their nuclear export, would be needed to determine whether the relevant XPO1 substrate in this system is actually TDP-43 or an upstream modulator of its phase behavior.

      As discussed above, our new data regarding the role of RNA in TDP-43 phase regulation should alleviate this concern, although we cannot exclude the possible involvement of splicing factors in this process. We also clearly state that there is no evidence to support a direct interaction between TDP-43 and XPO1 on page 8.

      (18) Testing whether identified modifiers converge on nuclear TDP-43 concentration would be informative. Since phase separation is concentration-dependent, measuring nuclear versus cytoplasmic TDP-43 levels across key perturbations, including splicing inhibition, translation inhibition, proteasome inhibition, HSP90 inhibition, and XPO1 modulation, would help determine whether modifiers mainly work by changing nuclear TDP-43 concentration or by altering interaction networks and the material properties of condensates.

      In the newly performed immunoblotting experiment, we measured the TDP-43 levels in drug-treated cells but found no effect by most drugs (Supplemental Figure 1).

      (19) Examining other ALS-relevant RNA-binding proteins would be valuable. Given the role of XPO1 and other hits, it would be informative to briefly test whether similar principles apply to FUS, hnRNPA1, or other ALS-relevant RNA-binding proteins in the same cellular context, to argue for generality versus TDP-43-specific idiosyncrasies of the 2KQ system.

      We agree that this is an important issue but we feel the proposed experiments are beyond the scope of the study.

      (20) The Introduction sometimes implies that anisosomes are common and well-established intermediates en route to pathology. It would be helpful to more clearly state that, to date, anisosomes are primarily observed in overexpression and mutant systems and have not yet been unequivocally demonstrated in human patient tissue. The link between PDGFRβ, PAK4, GSK-3β, and YAP and TDP-43 phase dynamics is intriguing but only briefly mentioned. The authors should either expand on this or tone down the emphasis in the Results section.

      We have revised the introduction and added the following sentence on page 4. “The 2KQ-containing anisosomes, observed mostly in the nucleus under overexpression conditions, have not been validated in human patient samples.”

      (21) In the organoid methods, the authors should consider clarifying whether doxycycline is continuously used, which might alter TDP-43 expression and nuclear transport in a non-negligible way.

      The organoid model does not involve protein overexpression or doxycycline treatment. We measured endogenous p-TDP-43, which is why we feel this experiment is very significant. Unlike many other p-TDP-43 detection studies that rely on TDP-43 overexpression or exposing cells to excess stressors, we could detect substantial p-TDP-43 in 3D organoids grown under normal conditions, whereas the same cells grown and differentiated in 2D culture do not show p-TDP-43 (Zhang Q. et al., BioRxiv 2025).

      (22) For statistical methods, it would be beneficial to indicate whether multiple-comparison corrections were applied for the many FRAP, anisosome count, and size comparisons beyond DESeq2 internal corrections for RNA-seq.

      We have added more statistical information to the figure legends.

      (23) Some figure legends could more clearly indicate whether the images shown are single z-planes or maximum intensity projections and how the thresholding for anisosome detection was performed.

      We revised the figure legends to include this information. As for anisosome detection, because they are so obvious, standard thresholding combined with automated counting was sufficient to identify them.

      (24) In its current form, the manuscript contains an impressive set of screens and some nicely executed imaging of TDP-43 condensates, highlighting nuclear export among other pathways as a modulator of TDP-43 phase behavior. However, the physiological relevance is undercut by heavy reliance on an acetylation-mimetic, RNA-binding-defective TDP-43 mutant and a homozygous K181E organoid model. The mechanistic link between XPO1 and TDP-43 remains largely inferential and partly at odds with prior work. The conclusion that cytoplasmic TDP-43 aggregation is only a modest contributor to disease is not firmly supported by the available data.

      We agree with the reviewer that the strength of the study is our unbiased approach that identifies pathways capable of modulating TDP-43 phase behavior. In the revised paper, we included several experiments using an in vitro semi-permeabilized cell system to further dissect the role of nuclear export in TDP-43 phase separation. We believe that these new results should provide significant mechanistic insight that links nuclear export and RNA transcription and splicing to TDP-43 phase regulation. Additionally, we have revised our paper carefully to discuss the physiological relevance and the limitation of our study.

      (25) With substantial additional mechanistic work, particularly around XPO1, rigorous validation in more physiological TDP-43 contexts, more sensitive detection of cytoplasmic TDP-43 aggregates, and a tempering of the central claims, this study could make a meaningful contribution to understanding how nucleocytoplasmic transport and other cellular pathways influence TDP-43 phase transitions and aggregation. The work should be reframed as an important screening study that identifies nuclear export as one among several cellular processes that modulate TDP-43 phase behavior in a model system, rather than as a definitive demonstration that nuclear export governs pathological TDP-43 aggregation in disease.

      We now reframe the study as an important screening study that identifies nuclear export among several other pathways as modulators of TDP-43 phase behavior. We also propose a model that links RNA splicing to nuclear export in TDP-43 phase regulation.

      Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and timely question in TDP-43 biology by systematically identifying regulators of TDP-43 anisosome formation, with a particular focus on nuclear export via XPO1. Using a combination of unbiased chemical screening, genetic perturbation, and advanced imaging approaches, the authors propose that inhibition of nuclear export modulates the abundance and biophysical properties of TDP-43 anisosomes. The study is conceptually innovative and has potential relevance for neurodegenerative diseases characterized by TDP-43 pathology. However, significant concerns regarding experimental controls, reporting transparency, and model translatability currently limit the strength of the conclusions and the interpretability of several key findings.

      We thank the reviewer for acknowledging the significance and innovation of our study.

      Strengths:

      (1) The study employs an unbiased, hypothesis-free compound screen to identify regulators of TDP-43 anisosome formation, which is a major strength and reduces confirmation bias.

      (2) The authors combine chemical and genetic screening approaches, providing orthogonal validation of key pathways and increasing confidence in the biological relevance of top hits.

      (3) The focus on biophysical properties of TDP-43 assemblies, assessed through imaging and FRAP, moves beyond simple presence/absence of aggregates and provides mechanistic insight into the biophysical states of TDP-43.

      (4) The use of multiple experimental modalities, including live-cell imaging, FRAP, pharmacological perturbation, and transcriptomic analysis, reflects a technically sophisticated and ambitious study design.

      (5) The authors attempt to extend findings beyond immortalized cancer cell lines by incorporating organoid models, demonstrating awareness of disease relevance and translational importance.

      Overall, the manuscript is clearly written and logically structured, making complex experimental workflows accessible and the central hypotheses easy to follow.

      Weaknesses:

      Despite its strengths, the manuscript has several major limitations that affect data interpretation and confidence in the conclusions.

      (1) Lack of appropriate controls for overexpression experiments:

      A central concern is the absence of proper controls for TDP-43 and XPO1 overexpression. Prior studies (including those cited by the authors, Archbold et al.2018) show that overexpression of WT TDP-43 alone is toxic to neurons. Thus, the experimental system itself may induce anisosome formation independently of the mechanisms under study. Similarly, XPO1 overexpression lacks a suitable control (e.g., mCherry alone or mCherry fused to a protein known to be independent of TDP-43). The near-complete colocalization of XPO1 with TDP-43 anisosomes upon overexpression raises the possibility that these structures reflect non-physiological protein accumulation rather than regulated assemblies.

      As mentioned in our response to reviewer 1, point 1, we have added more discussions to justify the use of acetylation mimetics in our study. We agree with the reviewer that these large puncta (both anisosomes and gel-like structures) likely resulted from TDP-43 overexpression. Nevertheless, in a titration experiment done by Yu et al. 2020 (PMID: 33335017), they showed that ectopic TDP-43 undergo demixing even at concentrations lower than endogenous TDP-43, although the demixed puncta were very small. Their result suggested that overexpression per se does not change TDP-43 phase behavior, only enlarge the demixed TDP-43 structures, which is necessary for our screen and imaging-based characterization.

      For XPO1 overexpression, we have done the mCherry alone control but due to space limit in Figure 5, we did not include it. We now include the data in Supplemental Figure 4. This figure shows that overexpression of mCherry did not change TDP-43 localization or anisosome structures.

      (2) Insufficient experimental and analytical transparency:

      The manuscript frequently lacks clear reporting of experimental details. In multiple figures, the stated number of independent experiments does not match the number of data points shown, making it difficult to assess statistical validity. Concentrations used in the compound screen are not clearly defined, nor is it stated whether multiple concentrations were tested. It is unclear how many wells, cells, or independent cultures were analyzed. The criteria used to reduce 1,533 screening hits to 211 candidates via STRING analysis are not explained. Knockdown and overexpression efficiencies are not reported.

      We apologize for these omissions. We have added more experimental details to the figure legends and the method. For the imaging experiments, data points reflect randomly selected individual cells imaged in 2-3 independent biological repeats. This is now stated in the figure legends. For chemical screens, we screened against NCATS libraries was first done at top concentration (10 mM) to ensure inhibitory efficacy for all potential hits. In the follow-up validation study, we validated the top hits using a series of concentrations, as shown in Figure 1B. Drug concentrations are provided in Figure 2A, 4A, C, E, F, 5A-D, F, Figure 6F, G, Figure 7A)

      We explain the STRING analysis in more detail now. Basically, STRING is a protein-protein interaction network that reports all potential interactions between any proteins in human proteome. Given the potential off-target effect of siRNA, we assume that if the screen identifies multiple components of a protein interaction network or pathway, the result is more likely to be real.

      We did not check XPO1 knockdown efficiency in high through-put screens (HTS) for several reasons. Firstly, the large number of positive hits makes it impossible to check knockdown efficiency for all of them. Secondly, the effect of XPO1 knockdown on anisosomes was seen with 6 different siRNAs in two rounds of screens. Thirdly, in the HTS protocol, we routinely included a transfection control (siRNAdeath) to control transfection efficiency. We would only process the data if siRNAdeath control killed > 90% of the cells. Lastly, the XPO1 knockdown result was independently validated by small molecule inhibitors. For TDP-43 overexpression, the study by Yu and colleagues suggested that the expression is more than 20-fold higher than endogenous TDP-43, but they showed that anisosome formation is not an artifact of protein overexpression. When the expression level was titrated down, they could still detect anisosomes.

      (3) RNA-seq concerns:

      The RNA-seq experiments are particularly problematic. The number of biological replicates per condition is not stated, and heatmaps suggest that only one sample per group may have been used, which would preclude statistical analysis. No baseline comparison between WT and mutant TDP-43 is shown. Given that TDP-43 is an RNA-binding protein, splicing analyses would be far more informative than gene expression alone, yet no splicing data are presented. Moreover, nuclear retention of TDP-43 does not preclude nuclear aggregation, which may still impair its splicing function.

      We apologize for the lack of clarity regarding the RNA-seq design. For each condition, organoids of two independently differentiated batches were treated in triplicate. What we showed before was averaged expression levels. We pooled the organoids of the same treatment from the two batches to reduce the impact of batch variation.

      Given the criticisms from both reviewers 1 and 2 on the limited interpretation power of the RNAseq study, we have removed this data from the revised manuscript.

      (4) Limited translatability to neuronal biology:

      All anisosome analyses are performed in a cancer cell line, raising concerns about relevance to post-mitotic neurons. While organoids are used as a secondary model, the assays performed do not overlap with those used in cancer cells, making it difficult to assess whether anisosome-related mechanisms are conserved. Neuronal toxicity, a critical outcome given known TDP-43 biology, is not assessed. Prior work has shown that WT TDP-43 overexpression alone is toxic to neurons, yet this is not addressed.

      We agree with the reviewer that the model used in this study is not directly relevant to neurodegeneration. However, as pointed out by the reviewer, neurons are much more sensitive to TDP-43-associated toxicity. By contrast, the cell line used in this study can tolerate TDP-43 overexpression with no detectable cytotoxicity. This feature makes it feasible to evaluate how different cellular processes modulate TDP-43 phase behavior without the confounding effect from cytotoxicity. Notably, the processes identified by our screens are all house-keeping pathways that are conserved in neurons. Thus, we believe that the reported findings are likely applicable to neurons. That being said, we have revised our paper to ensure that we don’t overstate the clinical relevance of our work.

      (5) Conceptual and interpretational gaps:

      The authors quantify anisosome number but also report conditions in which anisosome number decreases while size increases. The biological interpretation of larger anisosomes is not discussed, and whether this reflects improvement or worsening of pathology is unclear. Compounds targeting the same mechanism (e.g., nuclear export inhibition) are inconsistently used across experiments (KPT compounds, verdinexor, leptomycin B), raising concerns about reproducibility. In organoids, the experimental paradigm shifts to long-term treatment (35 days vs. 16 hours), further complicating interpretation.

      We thank the reviewer for these critical points. As pointed out by the reviewer 1 in point 4 above, we do not have evidence to establish a convincing correlation between the size of anisosomes and clinical phenotypes. Regarding the use of different drugs for different experiments, the initial screen identified KPT and Verdinexor because they are investigational drugs, but Leptomycin B was not in our library. In the follow-up studies, we switched to Leptomycin B because 1) it is highly potent and specific; 2) it was better characterized and more commonly used as inhibitors of XPO1 according to the literature. However, for the organoid study, we had to switch back to KPT because of the toxicity issue associated with long-term application of Leptomycin B.

      (6) Overinterpretation of rescue effects:

      Although the authors state that they aim to test whether nuclear export inhibition rescues neuronal defects, no functional neuronal readouts are provided (e.g., viability, morphology, axon outgrowth, or electrophysiological measures). RNA-seq alone is insufficient to support claims of rescue.

      Our interpretation of the RNA-seq data was that the rescue effect by nuclear export inhibition was limited and probably insignificant. Given that this negative data is not conclusive, we have removed it from the revised manuscript.

      (7) Finally, the model does not appear to exhibit cytosolic TDP-43 aggregation at baseline. It remains unclear whether longer induction would produce cytosolic gel-like assemblies and whether these would be prevented by nuclear export inhibition. Long-term data are shown only in organoids, yet anisosome formation is not assessed there.

      The expression system used in the study reaches a steady state after 24 h of induction. Prolonged expression up to 48 h did not alter the number of anisosome, nor does it change TDP-43 phase behavior. We now clarify this point on page 4.

      Reviewer #3 (Public review):

      Summary:

      TDP-43 proteinopathy is broadly found in neurodegenerative diseases. This manuscript investigates how nuclear export influences the biophysical properties of TDP-43. The authors use a combination of chemical screening and genome-wide siRNA screening to identify pathways that modulate TDP-43 liquid-to-solid transitions. Overall, the study employs a broad array of approaches and addresses an important question in TDP-43 pathobiology. The identification of nuclear export as a central regulator is compelling and conceptually aligns with the emerging view that TDP-43 nucleocytoplasmic trafficking is a major defect in neurodegeneration.

      Strengths:

      This work integrates chemical and genetic screening to identify novel modifiers. The candidates were validated in both reporter cell lines and iPS-differentiated organoids. The findings support the nucleocytoplasmic transport is important for the biophysical properties of TDP-43.

      We thank the reviewer for acknowledging the significance and strength of our study.

      Weaknesses:

      The mechanisms underlying the connection between nuclear export and phase transition need further clarification. Broader consequences of XPO1 inhibition are not addressed.

      We agree that our previous manuscript did not address how nuclear export inhibition affect TDP-43 phase behavior. As discussed in our paper, we proposed that the effect of nuclear export inhibition on TDP-43 phase separation is likely indirect. The most likely scenario is that inhibition of nuclear export changes the nuclear environment over time, which affects TDP-43 phase separation. We have tried to isolate nuclear extracts from control and LMB-treated cells and used mass spectrometry to identify proteins that are differentially present in the nucleus. However, knockdown of the identified top candidates did not abolish LMB-induced phase alteration (not shown). Considering our observation that RNA splicing is another modulator of TDP-43 phase behavior, we reasoned that it is possible that it is the combined change of RNA and protein composition in the nucleus that alters TDP-43 phase behavior. In new experiments presented in Figure 6, we now used a semi-permeabilized in vitro system to demonstrate that LMB treatment stabilized anisosomes in an RNA-dependent manner (see response to point 4 by reviewer 1). This new data allows us to propose a new model that link RNA splicing and nuclear export in TDP-43 phase regulation (Discussion).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Include appropriate controls for all overexpression experiments. In particular, overexpression of WT TDP-43 alone and suitable tag-only controls (e.g., mCherry alone or mCherry fused to a protein unrelated to TDP-43/XPO1) should be included to control for aggregation driven by non-physiological protein levels.

      In Supplemental Figure S4, we included a tag-only control, which shows that mCherry alone does not affect the localization of XPO1, neither did we see mCherry co-localizes with TDP-43.

      Since WT TDP-43 itself does not form anisosome and because the goal of the study was to test how anisosome dynamics is affected by various conditions, we did not repeat our experiments with WT TDP-43.

      (2) Address whether TDP-43 anisosomes form under endogenous or near-physiological expression levels. If possible, include experiments using lower expression systems or endogenous tagging to demonstrate that anisosome formation is not solely an overexpression artifact.

      As mentioned above, in a titration experiment done by Yu et al. 2020 (PMID: 33335017), they showed that ectopic TDP-43 undergoes demixing even at concentrations lower than endogenous TDP-43, although the demixed puncta are small. Their result suggested that overexpression per se does not change TDP-43 phase behavior. Instead, it only enlarges the demixed TDP-43 structures, which is necessary for our screen and imaging-based characterization.

      (3) Clearly define biological versus technical replicates throughout the manuscript and report exact n-numbers for all experiments in figure legends and/or methods. Resolve discrepancies between stated and displayed n-numbers (e.g., figures showing more data points than the number of independent experiments reported). Further, include how data points were defined (e.g., cells, fields of view, wells).

      We now state clearly the biological repeats in figure legends. We did not use N number to specify technical replicate. The discrepancy between the stated N number (biological repeats) and the data points is because for imaging experiments, data points usually represent single cells collected from 2-3 biological replicates (N=2 or 3). Data points are now clearly defined in the figure legends (anisosome, cell, imaging field, or independent experiment).

      (4) The authors state that they identified a list of compounds that reduced anisosomes. Please clarify how the threshold was determined: Was this a statistical analysis or a specific threshold that has been used?

      For both siRNA screen and chemical genetic screen, we calculated the Z-score and used Z-score>2 as a cutoff. This is mentioned in the method.

      (5) Provide a complete list of compounds used in the chemical screen, including concentrations tested and whether multiple doses were evaluated.

      As mentioned above, the initial screen was done with just one concentration (10 mM). Identified positive hits were re-tested with multiple doses as shown in Figure 1. The compounds are from a commercial library (LOPAC R1280, Sigma #LO4200). The list of compounds can be found at vender’s website.

      (6) Clearly explain the criteria used to reduce the initial 1,533 screening hits to 211 candidates following STRING analysis, including cutoffs and prioritization logic.

      We now explain that the Z-score was used to further narrow down the hit (page 6). Additionally, we provide an explanation on how we use STRING to further narrow down the list. The sentence reads as “To further narrow down the list, we performed a STRING protein network analysis based on the assumption that a protein interaction network bearing multiple positive hits would be more likely to be a true effector.”

      (7) Report knockdown and overexpression efficiencies for all genetic perturbations used in the study.

      For TDP-43 overexpression, the study by Yu and colleagues suggested that the stable cell line expresses 20-fold more TDP-43 than endogenous one, but they showed that anisosome formation is not an artifact of protein overexpression. When the expression level was titrated down, they could still detect anisosomes (Yu, H. et al., Science 2021). For knockdown efficiency, since the screen used 6 different siRNAs for each identified target (a few hundred), it is technically challenging to validate the knockdown efficiency of each siRNA by conventional qRT-PCR. To control knockdown efficiency, we transfected cells in parallel with siRNA-death that contains a mixture of siRNAs targeting several essential genes (Qiangen, #1027299). We would only process the data if siRNAdeath control killed > 90% of the cells, indicating good knockdown efficiency.

      (8) Clarify the biological interpretation of changes in anisosome size versus number, particularly in conditions where fewer but larger anisosomes are observed. Discuss whether larger assemblies are hypothesized to be protective, neutral, or deleterious.

      Live cell imaging was used to dissect why cells treated with certain drugs such as XPO1 inhibitors have fewer but larger anisosome. Figure 5F shows that this is caused by the fusion of small anisosomes. Our data does not suggest that the size of anisosomes can differentiate between protective or deleterious state, but rather it is the LLPS state and subcellular localization of these assemblies that may play a more critical role in determining whether TDP-43 forms deleterious protein aggregates. The discussion is on page 10.

      (9) Specify whether all anisosomes induced by XPO1 overexpression were gel-like or whether this applied only to a subset. If only a subset was affected, please provide quantifications, otherwise state clearly that all anisosomes in XPO1 overexpression were gel-like.

      All TDP-43 puncta mislocalized to the cytoplasm in XPO1-overexpressing cells are gel-like because the FRAP experiment in Figure 5I was done with randomly selected TDP-43 puncta mislocalized to the cytoplasm.

      (10) Clarify which anisosomes (nuclear vs cytosolic; gel-like vs non-gel-like) were selected for FRAP analyses in Figure 5I.

      For Figure 5I, the control anisosomes in untreated cells are nuclear while under mCh-XPO1 expressing condition, only those in the cytoplasm were randomly selected for photobleaching.

      (11) The translatability of the conclusion based on cancer cell lines to brain organoids is not convincingly shown and could be strengthened by including additional assessment of anisosomes. While this might not be feasible in 3D cultures, the authors could alternatively use 2D cultured neurons to perform the same assays as performed in the cancer cell line. Additionally, the same treatment strategy should be applied. The reasoning for increasing treatment to 35 days in the organoids is unclear.

      In another manuscript that is currently under revision, we compared 2D iNeuron culture with 3D organoids. A pre-print is available at https://www.biorxiv.org/content/10.1101/2025.11.09.687455v1.full. In this study, we found that endogenous TDP-43 K181E mutant do not undergo phosphorylation-dependent transition to aggregate in 2D cultures. Only when these cells were grown into 3-D organoids, TDP-43 phosphorylation could be detected. (see supplemental Fig. S1c, d in https://www.biorxiv.org/content/10.1101/2025.11.09.687455v1.full). Thus, it is not possible to repeat the experiments in this study in 2D iNeuron cultures. We agree with the review that there is a gap between the study using the cancer cell line and the use of K181E iPSC-derived 3D organoids. We have toned down our conclusions throughout the text.

      (12) Address neuronal vulnerability explicitly by assessing toxicity, viability, or functional neuronal readouts, particularly given prior reports that WT TDP-43 overexpression alone is neurotoxic.

      We agree that this is an important point, but the main goal of this study was to dissect the cellular pathways/mechanisms that govern TDP-43 phase separation. We feel that the requested experiments are beyond the scope of the current study.

      (13) Clearly state the number of biological replicates used for each RNA-seq condition. Establish baseline transcriptional differences between WT and mutant TDP-43 prior to assessing the effects of nuclear export inhibition. Include PCA plots and heatmaps, including all samples.

      As mentioned above, we have decided to remove the RNAseq data from the manuscript to save room for new results.

      (14) Given the role of TDP-43 as an RNA-binding protein, consider including splicing analyses to assess whether nuclear export inhibition preserves or disrupts TDP-43-dependent RNA processing.

      We thank the reviewer for this suggestion. However, we feel that the proposed experiments are beyond the scope of the current study.

      (15) Improve clarity of transcriptomic visualizations (e.g., GO-term plots) and explicitly define all group labels used (e.g., Group A vs Group B).

      We have removed the RNAseq data.

      (16) Ensure consistent use of disease terminology (ALS vs FTD) throughout the manuscript, e.g., lines 222 and 244.

      We have checked the usage of these terms to make sure they are accurately used.

      (17) Correct figure and axis labeling errors (e.g., Figure 3A x-axis range).

      Figure 3A indicates the Z score distribution of the entire human genome. As stated on page 6, 21,404 genes were targeted.

      (18) Avoid overstatements in the Discussion that are not directly supported by the presented data, particularly regarding the interpretation of proteasome inhibition and gel-like anisosome states.

      We have revised our discussion substantially to tone down our conclusions.

      (19) Clarify the rationale for switching between different nuclear export inhibitors across experiments and discuss whether results were consistent across compounds.

      In the acute experiments down with the cancer cell line, we used LMB because it is potent and well characterized. In organoid experiment, we switched to KPT-276 because it is better tolerated by organoids, especially during longer treatment.

      Reviewer #3 (Recommendations for the authors):

      Major concerns that require clarification or further strengthening:

      (1) The connection between nuclear export and liquid-solid phase transition is not clear. The 2KQ mutant forms nuclear anisosomes. The manuscript does not provide data about its nuclear-cytoplasmic distribution normally, nor how the distribution is changed upon nuclear export inhibition or enhancement. In Figure 5I, it is unclear whether the anisosomes are in the nucleus or cytoplasm. The dynamics of nuclear vs cytoplasmic anisosomes should be measured separately. What is the mechanism that promotes nuclear export and changes the dynamics, especially nuclear anisosomes?

      As mentioned by the reviewer, the 2KQ mutant forms anisosomes only in the nucleus. This was documented in Yu, H. et al., Science 371 (2021), and also shown in our Figure 4A, F, Figure 5A. Figure 5A also shows that nuclear export inhibition does not change anisosome localization, only making them bigger while reducing the numbers. For Figure 5I, the control anisosomes in untreated cells are nuclear while under mCh-XPO1 expressing condition, only those present in the cytoplasm were randomly selected for bleaching.

      (2) Figure 5J, no obvious XPO1 is sequestered to anisosomes, as described in lines 208-209.

      Unlike Figure 5G, this experiment studied the localization of endogenous XPO-1 by immunostaining. As discussed in Yu et al., Science 371 (2021), proteins inside anisosomes could not be stained by antibodies due to an accessibility problem. This explains why we could only detect reduced XPO1 after anisosome induction.

      (3) Figure 6A, the localization of phosphor-TDP-43 is not clear. And it is not clear what cell types contain the aggregates. Higher-resolution images need to be included. The mechanism by which XPO1 inhibition reduces TDP-43 aggregation requires further validation. It remains unclear whether it is directly mediated through altered nucleocytoplasmic transport of TDP-43.

      We agree that it is technically challenging to visualize the precise subcellular localization of p-TDP-43 in 3D organoids. In the manuscript that reports the characterization of the 3D organoids, we dissociated cells from the 3D organoids by trypsin digestion and plated them out in 2D before immunostaining and imaging. We could clearly see p-TDP-43 co-localizes with the neuronal marker TUJ1 and is localized outside of nucleus (see figure 1 of https://www.biorxiv.org/content/10.1101/2025.11.09.687455v1.full)

      In the newly added Figure 6, we used a semi-permeabilized cell system to dissect the phase separation dynamics of TDP-43 2KQ in cells treated with the nuclear export inhibitor LMB. Our data suggests that nuclear export inhibition alters the nuclear environment, making it more favorable for the liquid phase of TDP-43. This is dependent on nuclear RNA.

      (4) XPO1 controls the export of numerous essential proteins, and its inhibition can produce broad, potentially toxic effects unrelated to TDP-43. The manuscript should include a discussion of these off-target consequences.

      We thank the reviewer for this point. Given the new data in Figure 6, we now add some more discussion on the potential mechanism by which nuclear export inhibition modulates TDP-43 phase separation. This can be found on page 10.

      References:

      Zhang, Q. et al. A human forebrain organoid model phenocopies dysregulated RNA and protein homeostasis in ALS/FTD-associated TDP-43 proteinopathies. bioRxiv (2025). (https://www.biorxiv.org/content/10.1101/2025.11.09.687455v1.full

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      Summary:

      This preprint investigates the molecular mechanism by which warm temperature induces female-to-male sex reversal in the ricefield eel (Monopterus albus), a protogynous hermaphroditic fish of significant aquacultural value in China. The study identifies Trpv4 - a temperature-sensitive Ca²⁺ channel - as a putative thermosensor linking environmental temperature to sex determination. The authors propose that Trpv4 causes Ca²⁺influx, leading to activation of Stat3 (pStat3). pStat3 then transcriptionally upregulates the histone demethylase Kdm6b (aka Jmjd3), leading to increased dmrt1 gene expression and ovo-testes development. This work aims to bridge ecological cues with molecular and epigenetic regulators of sex change and has potential implications for sex control in aquaculture.

      Strengths:

      (1) This study proposes the first mechanistic pathway linking thermal cues to natural sex reversal in adult ricefield eel, extending the temperature-dependent sex determination paradigm beyond embryonic reptiles and saltwater fish

      (2) The findings could have applications for aquaculture, where skewed sex ratios apparently limit breeding efficiency

      Weaknesses:

      Although the revised manuscript represents an improvement over the original version, substantial weaknesses remain.

      We thank you for the critical comments. We have responded to your concerns by a point by point manner, and please see detail below.

      Scientific Concerns

      (1) Western blot normalization and exposure: The loading controls (GAPDH) in Fig. S3C appear overexposed, as do several Foxl2 blots. Because these signals are likely outside the linear range, I am not convinced that normalization is reliable. This raises concerns about the validity of the quantified results.

      We thank you for the concerns. We have repeated the experiments, and new blots were loaded in Fig.S3C.

      (2) Antibody validation and referencing (Line 776): The authors need to refer explicitly to figures demonstrating antibody validation. At present, these data are provided only as a supplementary file that is not cited in the manuscript. In addition, the Sox9a antibody appears to yield indistinguishable signals in control and RNAi conditions, suggesting that it may not recognize eel Sox9a. This issue is not addressed by the authors. Furthermore, antibody validation Western blots should be quantified.

      We thank you for the comments. We have repeated the siRNA experiments to show the specificity of the antibodies used. This file, named as the supplementary file 1, is now cited in “WB analysis” in the Materials and Method part. As required, the antibody validation of WB are uploaded in the supplementary file 1. Antibody validation for WB are now quantified, and please see the new figure 3 and supplementary Figure 3.

      (3) Unclear sample sizes (N values): Sample sizes remain unclear for several figures:

      (a) Fig. 3F - No N value is provided. Each graph shows three data points; does this indicate that only three samples were quantified? If ten samples were collected, why were all not quantified?

      We apologize for the confusion. Three data points were previously used to shown data of 3 replicates. In new figure 3F, 10 randomly selected sections were imaged, and the data are shown. In the revised manuscript, the sample numbers (the N values) are added, and all the information can be found in the figure legend.

      (b) Fig. 4 - No N values are reported.

      Now N values are added. Please see the figure legend.

      (c) Fig. 5A - Again, only three data points are shown per group, despite the apparent availability of twelve samples. The rationale for this discrepancy is not explained.

      We apologize for the wrong data representation. Now all the data points are shown in Figure 5.

      (4) qRT-PCR normalization: The manuscript does not specify the reference gene(s) used for qRT-PCR normalization. Although expression levels are reported as "relative," neither the identity of the reference gene(s) nor the justification for their selection is provided.

      We now have specify the reference gene in “Quantitative real-time PCR (qPCR) experiments” part in the Materials and Methods section.

      (5) Specificity of key antibodies: While the authors have made some effort to validate anti-Amh, anti-Sox9, and anti-Dmrt antibodies, the results remain incomplete. The Amh and Dmrt antibodies detect reduced protein levels following knockdown of their respective targets, which is encouraging. However, the Sox9a antibody shows no difference between control and RNAi conditions, suggesting it does not recognize eel Sox9. This is not acknowledged in the manuscript. In addition, no validation data are presented for Foxl2. Antibody validation data must be clearly referenced in the main text and presented in an interpretable and quantitative manner.

      The antibody specificity is very important. For that reason, we have generated at least two different antibodies for each target protein, using full-length or small peptide as antigen. We have repeated the experiments for key antibodies such as Dmrt1 and Sox9a. IF and WB results clearly showed the specificity of the antibodies.

      Author response image 1.

      Foxl2 antibody has also been reported in ricefield eel (Hu et al. SCIENTIFIC REPORTS | 4: 6884 | DOI: 10.1038/srep06884, Molecular cloning and analysis of gonadal expression of Foxl2 in the ricefield eel Monopterus albus).

      After short term warm temperature exposure, only a small portion of somatic cells in ovary may be induced to express the male markers. As different techniques have different capacity (sensitivity), some techniques were more easy to detect that change. For instance, qPCR and WB are ready to detect it, whereas IF is a little difficult in obtaining good quality data.

      (6) Immunofluorescence data quality: The immunofluorescence images remain difficult to interpret. I strongly encourage the authors to enlarge the image panels and to present monochrome images (white signal on black background). The current presentation severely limits interpretability.

      We thank you for the comments. We think that our IF images are of decent quality. Due to the limits of the Figure space (already busy for Figure 3), enlarging the image panels or presenting additional monochrome images will compromise the quality of other data. Alternatively, if you still concern its quality, we can put it in the supplementary.

      Author response image 2.

      (7) Unreferenced supplementary figure: Fig. S4 is included in the submission but is not referenced anywhere in the manuscript text.

      We now have renamed the supplementary Figures. And we have double checked the text to make sure all Figure information is correctly referenced. Figure S4 is removed, as it is not necessary.

      (8) Fig. 5B image resolution: The micrographs in Fig. 5B are too small to allow meaningful evaluation of the data.

      Now new Figure 5B images with higher resolution were shown.

      (9) Unexplained data inclusion (Fig. 5E): Fig. 5E includes a pERK blot that is not mentioned in the Results section. The rationale for including these data is unclear.

      Previous work have shown that FGF/ERK signaling may play a role in sex change of ricefield eel (in Chinese). We therefore examined the Erk activity to explore whether it is involved in sex reversal. The results showed that pErk was comparable between ovary and ovotestis. At your suggestion, we decided to remove the data.

      (10) Poor blot quality (Fig. S3C): The blots in Fig. S3C exhibit high background and overexposure. I am concerned about the reliability of the quantification shown in panel D.

      The experiments have been repeated at least three times, and similar results were obtained. We now have replaced some of the WB that were of high background or overexposure.

      (11) Poor blot quality (Fig. S5G): The Stat3 blots in Fig. S5G contain numerous white artifacts, raising concerns about their suitability for normalization in panel H.<br />

      We now have repeated the experiments, and uploaded a new representative blot with better quality.

      (12) Missing controls (Fig. 6E): Fig. 6E lacks controls for HO-3867 and Colivelin treatments alone. Without these controls, it is not possible to determine whether the reported effects are meaningful.

      We thank you for the comments. We now have added the data required (with HO-3867 and Colivelin treatments alone).

      (13) Graphical presentation: The use of a light blue-to-pink gradient in bar graphs throughout the manuscript does not aid interpretation. I recommend using more distinct colors (e.g., red, orange, green, blue, purple, gray, black) to improve clarity.

      We thank you for the comments. We now have changed the blue-to-pink gradient to more distinct color system to better present the data. Please see the detail in the revised Figures.

      In summary, the interpretation of the study remains limited by persistent issues related to data presentation, image quality, and reagent specificity.

      We thank you for the critical comments about our data, in particular for antibody specificity and image quality, and the detailed instruction for how to better present the data. Answering your questions have greatly improved the quality of the manuscript. We admit that due to the technique challenging (with different conditions and different doses of small molecules) and higher cost of animal experiments, some of the WB or IF experiments may not be of high standards.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Editorial Concerns

      (1) Overstatement of conclusions: In lines 16-18, the authors state that Trpv4 "mediates" warm temperature-driven sex reversal. This claim is too strong given the data and should be toned down.

      We agree with our editorial comment about the overstatement. Now it reads “Trpv4 links environmental temperature to testicular differentiation in ricefield eel”.

      (2) Misuse of statistical language (Line 213): The term "significant" is used where statistical significance was not measured. The wording should be revised.

      We thank you for the point, and now have replaced “significant” to “marked”.

      (3) Terminology (Line 238): The term "co-expression" is inaccurate in this context. I suggest replacing it with "co-upregulation."

      We thank you for the point, and have changed it accordingly.

      (4) Drug description errors (Lines 241-242): The manuscript incorrectly identifies which drug functions as an agonist and which as an antagonist. This caused considerable confusion and must be corrected.

      We have carefully checked the sentence, and it was correct, as RN1734 and GSK1016790A are known Trpv4 specific antagonist and agonist, respectively.

      (5) Gene examples missing (Lines 247-250): The authors should explicitly name the testis-biased and ovary-biased genes referred to in this section.

      We thank you for the point, and now it reads “warm temperature exposure increased the expression of testicular differentiation genes such as dmrt1 and gsdf, accompanied by moderately decreased expression of ovarian differentiation genes such as cyp19a1a and foxl2”.

      (6) Lack of experimental context (Lines 322-324): Rather than simply listing the drugs used, the authors should briefly explain what each compound inhibits or activates and why it was employed.

      We have described this in the manuscript. The information of pStat3 activator and inhibitor has been described in Lines 305-309, as “HO-3867, a curcumin analogue, is a selective pStat3 inhibitor, which blocks pStat3 activity by directly binding to Stat3 DNA binding domain, and Colivelin is a potent synthetic peptide activator of pStat3, which increases pStat3 levels by acting through the GP130/IL6ST complex”, and the rationale has been stated in lines 32--322 as “To functionally demonstrate that pStat3 signaling is downstream of Trpv4, rescue experiments were performed by injecting into ovaries with individual and combined small molecules”.

      (7) Discussion of evolutionary differences: The Discussion misses an important opportunity to address why Stat3 activates kdm6b in ricefield eel but represses it in turtles. It is difficult to reconcile how the same transcription factor could exert opposite effects on the same gene during sex determination without additional context. A comparison of kdm6b regulation and sequence conservation between turtles and ricefield eel would strengthen this section.

      We have downloaded the promoter sequences of red eared turtle and ricefield eel. Based on the DNA sequences (Author response image 3), the similarity (conservation) was low between the two species.

      Author response image 3.

      It was appeared that DNA around the Stat3 binding sites in turtle are GC rich (CpG island), which may be subjected to DNA methylation modification, whereas the DNA in ricefield eel are not GC rich.The observations imply that the role of pStat3 is to promote the repression of kdm6b in turtle but the activation of kdm6b in ricefield eel.

      Moreover, our unpublished data showed that Trpv4-controlled calcium signaling is required to remove the repressive histone modification H3K27me3 at the kdm6b gene. If pStat3 is downstream of Trpv4 in this case, it supports again that Trpv4-pStat3 axis activate kdm6b in ricefield eel.

      Warm temperature promotes female sex in turtle but male sex in ricefield eel. If pStat3 is mediating Trpv4, it is not surprising that it represses kdm6b in turtle but activate it in ricefield eel.

      Based on above, we have added some sentences in the discussion part, and it reads “We reasoned that a yet-unidentified co-factor may determine whether Stat3 is a transcriptional repressor or activator. A comparison of promoter sequences of kdm6b between turtle and ricefield eel supported this”.

      (8) Supplementary figure formatting: Supplementary figures should be provided in accordance with eLife formatting guidelines.

      We have now formatted the supplementary figures that are in accordance with eLife formatting requirement. Please see the new uploaded supplementary figures.

      In sum, the interpretations are still limited by the above concerns regarding data presentation and reagent specificity.

      We thank our editor for the inspiring comments. We believe we have addressed all the major concerns by our editor.

    1. Author response:

      eLife Assessment

      This study provides a valuable advance in understanding how disordered proteins interact with cell membranes by identifying the sequence rules that enable aromatic residues to penetrate deeply into the membrane interior. The integration of complementary computational approaches, including molecular simulations, large-scale sequence analysis, and the development of an online prediction server, makes the work potentially impactful for the membrane protein and intrinsically disordered protein communities. The evidence supporting the main conclusions is generally convincing, although its transferability across diverse membrane compositions and its validity as a prediction tool for real protein-membrane systems remain to be further established.

      We thank the editors for recognizing our study as a valuable advance. This work lays a solid foundation for future developments to account for diverse membrane compositions and further refinements after additional experimental tests.

      Public review:

      Reviewer #1:

      A primary limitation is the heavy reliance on computational modeling. Training for AroMIP is generated using PPM rather than direct experimental measurements, and so the model may primarily reproduce PPM behavior rather than true membrane insertion thermodynamics. Moreover, all simulations use a single lipid composition (POPC:POPS:PIP<sub>2</sub> 70:25:5), but biological membranes vary substantially in cholesterol, cardiolipin, and acidic lipid content. Whether AroMIP's predictions transfer to diverse lipid environments remains untested. The 5% PIP<sub>2</sub> concentration used in the simulations is higher than that of a normal mammalian cell and may therefore overemphasize electrostatic contributions. Applicability beyond short 9-residue motifs is unclear, as longer-range interactions or secondary structure in full-length IDRs could modulate insertion in ways the current model does not capture. This could be considered for future development.

      The reviewer’s point on our reliance on PPM for training, a single lipid composition, and potential effects beyond a 9-residue motif is well taken. Regarding PPM, we chose it as the optimal compromise for high-throughput data. However, we complemented the high-throughput PPM data with experimental data on an initial set of 10 peptides. Moreover, we validate AroMIP on an additional 12 IDRs (intrinsically disordered regions; Table S2). On membrane composition, we now acknowledge the limitation of our work based on a single composition and point to future developments of AroMIP involving membrane-specific parameterization (p. 19, 3rd paragraph). On potential effects beyond a 9-residue motif, we now add justification and note neglected factors for future developments (paragraph running from p. 19-20), as suggested by the reviewer.

      Reviewer #2:

      (1) Aromatic residues have been shown to partition preferentially to the headgroup region of the lipid bilayer. Most of the papers on this problem were published in the mid 1990s to early 2000s. Some of the most important papers in this regard are the following: von Heijne, Annu. Rev. Biophys. Biomol. Struct. 1994, 23, 167-192; Doyle et al. Science 1998, 280, 69-77; Landolt-Marticorena, et al. J. Mol. Biol. 1993, 229, 602-608; Killian & von Heijne, TIBS 2000, 25, 429-434; Marx & Fleming J. Am. Chem. Soc. 2021, 143, 764-772. Strangely enough, none of these articles is cited.

      We have now citations to the Landolt-Marticorena paper and the von Heijne reviews (refs 25-27). The Doyle paper is not particularly relevant. As for the Fleming paper, we cited a 2016 JACS paper (original ref 27; now ref 30) that specifically dealt with aromatic residues.

      (2) This is the most important point and the most serious weakness. The authors find that the PPM method is able to reproduce the results from MD simulations, and the AroMIP model is able to perform well in comparison with PPM and MD, after training AroMIP on a large set of IDR sequences (intrinsically disordered protein regions) of the human proteome. The defining feature of the AroMIP calculation is the recognition of the importance of flanking residues in the membrane-insertion propensity of a sequence containing a central aromatic residue. All this sounds good. However, this is all theoretical. There is no connection to experiment or to any method that draws from experiment. The entire approach relies on the assumption that the MD simulations produce the correct results. There is no proof of the correctness of anything. As one of the greatest physicists of our times, Richard Feynman, wrote, "The test of all knowledge is experiment. Experiment is the sole judge of scientific "truth".”

      We emphasize that we have presented substantial experimental support for AroMIP. It correctly predicts the membrane insertion status of the initial set of 10 peptides, which were characterized experimentally. In addition, we validated AroMIP on an additional set of 12 IDRs (Table S2), most of which were characterized by experimental techniques including solution and solid-state NMR, fluorescence, H/D exchange, and cryo-EM. Lastly, we now show good correlation between our insertion scores and binding free energies calculated from the scale determined experimentally by White and co-workers (new Figure S10; p. 15, second paragraph).

      (3) The drawings in Figures 2 and 3 are incorrect and misleading. The size of the Tryptophan side chain is about 5.5 Å, whereas one-half of the bilayer ("a monolayer") thickness is about 15 Å. But in the figures, the lipid length and the Trp side chain seem about the same size. This is incorrect even in a qualitative sense.

      We have now revised these figures.

      Reviewer 3:

      (1) Membrane composition and lipid shape characteristics: The authors chose to use a model membrane bilayer of a distinct lipid composition, POPC: POPS: PI4,5P2 (70:25:5 molar ratio), for their all-atom simulations of the various model peptides. While this may be pertinent for some of these peptides, it is not for many, such as sequence 2 derived from Drp1, which preferentially binds target conical lipids such as cardiolipin (CL) and phosphatidic acid (PA). The rationale behind using PI4,5P2, which can induce positive membrane curvature when sequestered, versus CL and PA, which both induce negative membrane curvature, is not explained.

      We now acknowledge the limitation of our work based on a single composition and point to future developments of AroMIP involving membrane-specific parameterization (p. 19, 3rd paragraph). In this Discussion paragraph, we also speculate that conical lipids, by promoting membrane defects, may facilitate membrane insertion.

      (2) Parallel vs. perpendicular peptide orientation of sequence 2 in peripheral Drp1-lipid interactions: On page 11, the authors state that their simulation results of sequence 2 derived from Drp1 "contrasts with a transmembrane orientation proposed by Mahajan et al." However, upon review, a transmembrane orientation for this region has never been proposed anywhere. Drp1 is a peripheral membrane protein that reversibly binds CL- and PA-containing membranes via its intrinsically disordered variable domain containing an aromatic-centered WRG motif. Indeed, the model presented in Figure 9 of Mahajan et al. displays a peripheral and parallel orientation of the transiently helical WRG-containing motif rather than a transmembrane (i.e., across the bilayer) orientation. While the authors can distinguish between a parallel vs. perpendicular orientation of this sequence relative to the plane of the membrane bilayer surface from their simulations, suggesting that previous studies indicated a transmembrane orientation for Drp1 is disingenuous and misleading. The term "transmembrane" should be removed or replaced, as it presents a wrong image.

      We have now deleted the sentence mentioning “transmembrane orientation”.

      (3) Mutational analysis of W vs. F in membrane insertion of W-centered insertion motifs and vice versa: The PPM-based workflow suggests that F-centered sequences have the highest membrane insertion properties as opposed to W-centered ones. A W552F mutation in the WRGML sequence of Drp1 was, however, found to impair function. How do the authors rationalize this? A cross-mutational analysis of W vs. F in W-centered motifs and F-centered motifs is warranted.

      AroMIP predicts a membrane insertion propensity of 0.782 for the WRGML sequence and a moderately higher propensity, 0.837, with a W552F mutation. This increase contradicts the experimental observation of a 3.6-fold increase in membrane binding affinity by Mahajan et al. We now speculate that the specific lipid, cardiolipin, as the reason for the discrepancy (p. 19, 3rd paragraph). This discrepancy provides a concrete example for the need to account for membrane composition in future developments.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This valuable study combined careful computational modeling, a large patient sample, and replication in an independent general population sample to provide a computational account of a difference in risk-taking between people who have attempted suicide and those who have not. It is proposed that this difference reflects a general change in the approach to risky (high-reward) options and a lower emotional response to certain rewards. Evidence for the specificity of the effect to suicide, however, is incomplete, which would require additional analyses.

      We thank the editors and reviewers for this important assessment. Based on clinical interviews, we included patients with and without suicidality (S<sup>+</sup> and S<sup>-</sup> groups). However, in line with suicidal-related literature (e.g., Tsypes et al., 2024), two groups also differed substantially in the severity of symptoms (see Table 1). To address the request for evidence on specificity to suicidality beyond general symptom severity, we performed separate linear regressions to explain in gambling behaviour, value-insensitive approach parameter (β<sub>gain</sub>), and mood sensitivity to certain rewards (β<sub>CR</sub>) with group as a predictor (1 for S<sup>+</sup> group and 0 for S<sup>-</sup> group) and scores for anxiety and depression as covariates. Results remained significant after controlling anxiety and depression (ps < 0.027; Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, ps < 0.001), we performed Principal Components Analysis (PCA) on the clinical questionnaire to extract the orthogonal components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. We then performed linear regressions using these components as covariates to control for anxiety and depression. Our main results remained significant (ps < 0.027; Table S9). We believe that these analyses provide evidence that the main effects on gambling and on mood were specific to suicide.

      Moreover, as Reviewer 3 pointed out, these “absence of evidence” cannot provide insights of “evidence of absence”. Although we median-split patients by the scores of general symptoms (e.g., depression and anxiety-related questionnaires) and verified no significant differences in these severities (Figure S11), we additionally conducted Bayesian statistics in gambling behavior, value-insensitive approach parameter, and mood sensitivity to certain rewards. BF<sub>01</sub> is a Bayes factor comparing the null model (M<sub>0</sub>) to the alternative model (M<sub>1</sub>), where M<sub>0</sub> assumes no group difference. BF<sub>01</sub> > 1 indicates that evidence favors M<sub>0</sub>. As can be seen in Table S7, most results supported null hypothesis, suggesting that general symptoms of anxiety and depression overall did not influence our main results. Overall, we believe that these analyses provide compelling evidence for the specificity of the effect to suicide, above and beyond depression and anxiety.

      Beyond these specific findings, this work highlights the broader utility of computational modelling and mood to better understand behavioral effect, showing how to use both mood and choice data to better comprehend a psychiatric issue.

      Please see Tables S7, S8, S9 and our revisions below:.

      Page 17:

      “Within patients, this group effect on gambling rate remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.024; also see Figure S11, Table S7 and Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, (ps < 0.001), we performed Principal Components Analysis (PCA) to extract main components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. To further control for anxiety and depression, linear regression using these components as covariates revealed that the group effect on gambling rate remained significant (p = 0.024; Table S9).”

      Pages 18-19:

      “Within patients, this group effect on the approach parameter remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.027; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on approach parameter remained significant (p = 0.027; Table S9).”

      Page 21:

      “Within patients, this group effect on βCR remained significant after controlling for gambling rate, earnings, mood-related outcome effect, mood drift effect, sex, illness duration, family history, diagnosis, and various medications use (ps < 0.032), as well as general symptoms (e.g., depression and anxiety; p = 0.001; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on this mood parameter remained significant (p = 0.001; Table S9).”

      Page 27:

      “Beyond these specific findings, this work highlights the broader utility of computational modelling and mood to better understand behavioral effect, showing how to use both mood and choice data to better comprehend a psychiatric issue.”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors use a gambling task with momentary mood ratings from Rutledge et al. and compare computational models of choice and mood to identify markers of decisional and affective impairments underlying risk-prone behavior in adolescents with suicidal thoughts and behaviors (STB). The results show that adolescents with STB show enhanced gambling behavior (choosing the gamble rather than the sure amount), and this is driven by a bias towards the largest possible win rather than insensitivity to possible losses. Moreover, this group shows a diminished effect of receiving a certain reward (in the non-gambling trials) on mood. The results were replicated in an undifferentiated online sample where participants were divided into groups with or without STB based on their self-report of suicidal ideation on one question in the Beck Depression Inventory self-report instrument. The authors suggest, therefore, that adolescents with decreased sensitivity to certain rewards may need to be monitored more closely for STB due to their increased propensity to take risky decisions aimed at (expected) gains (such as relief from an unbearable situation through suicide), regardless of the potential losses.

      Strengths:

      (1) The study uses a previously validated task design and replicates previously found results through well-explained model-free and model-based analyses.

      (2) Sampling choice is optimal, with adolescents at high risk; an ideal cohort to target early preventative diagnoses and treatments for suicide.

      (3) Replication of the results in an online cohort increases confidence in the findings.

      (4) The models considered for comparison are thorough and well-motivated. The chosen models allow for teasing apart which decision and mood sensitivity parameters relate to risky decision-making across groups based on their hypotheses.

      (5) Novel finding of mood (in)sensitivity to non-risky rewards and its relationship with risk behavior in STB.

      Weaknesses:

      (1) The sample size of 25 for the S- group was justified based on previous studies (lines 181-183); however, all three papers cited mention that their sample was low powered as a study limitation.

      We thank the Reviewer for rising this concern. We agree that the sample size for S<sup>-</sup> group (n=25) is modest, and the prior studies we cited also acknowledged limited power. We wanted to point out that we obtained a comparable sample size to a prior study. In the revision, we therefore updated the section to justify this sample size in which we acknowledge the limited power of our study in the limitation section. Please see our clarification below:

      Page 32:

      “Third, despite replicating our main results in an independent dataset (n=747), the modest S<sup>-</sup> subgroup size (n=25) has a limited statistical power.”

      (2) Modeling in the mediation analysis focused on predicting risk behavior in this task from the model-derived bias for gains and suicidal symptom scores. However, the prediction of clinical interest is of suicidal behaviors from task parameters/behavior - as a psychiatrist or psychologist, I would want to use this task to potentially determine who is at higher risk of attempting suicide and therefore needs to be more closely watched rather than the other way around (predicting behavior in the task from their symptom profile). Unfortunately, the analyses presented do not show that this prediction can be made using the current task. I was left wondering: is there a correlation between beta_gain and STB? It is also important to test for the same relationships between task parameters and behavior in the healthy control group, or to clarify that the recommendations for potential clinical relevance of these findings apply exclusively to people with a diagnosis of depression or anxiety disorder. Indeed, in line 672, the authors claim their results provide "computational markers for general suicidal tendency among adolescents", but this was not shown here, as there were no models predicting STB within patient groups or across patients and healthy controls.

      Thank you for these thoughtful comments. Our study focuses on why adolescent patients with suicidality have increased risk behavior, aiming to provide a mechanism-based target for suicide prevention. Therefore, our dependent variable in the mediation model was gambling behavior. We also agree that the clinically relevant question is whether suicidality can be predicted from task-derived behavior/parameters. We thus used risky behavior and the potential mental parameters to predict STB. Linear regressions showed that gambling behavior, as well as the value-insensitive approach parameter, can predict suicidal symptom scores among patients (former: β = 9.189, t = 2.004, p = 0.048; latter: β = 5.587, t = 2.890, p = 0.005). In healthy controls, these predictions failed (gambling behavior: β = 1.471, t = 0.825, p = 0.411; approach: β = 0.874, t = 1.178, p = 0.241). These results suggest that clinical relevance of these findings apply exclusively to people with a diagnosis of depression or anxiety disorder. We found same patterns for the mood parameter (mood sensitivity to certain rewards: patients: β = -28.706, t = -2.801, p = 0.006; healthy controls: β = -2.204, t = -0.528, p = 0.599). In sum, we believe that our statement of “computational markers for general suicidal tendency among adolescents” is reasonable now. Please see our revisions below:

      Page 17:

      “Furthermore, linear regression showed that gambling rate can predict the current suicidal ideation score (BSI-C, β = 9.189, t = 2.004, p = 0.048) among patients, but not among HC (β = 1.471, t = 0.825, p = 0.411), suggesting that gambling behavior has patient-specific predictive utility for suicidal symptoms.”

      Page 19:

      “Furthermore, linear regression showed that approach parameter can predict the current suicidal ideation score (β = 5.587, t = 2.890, p = 0.005) among patients, but not among HC (β = 0.874, t = 1.178, p = 0.241), suggesting that value-insensitive approach parameter has patient-specific predictive utility for suicidal symptoms.”

      Page 21:

      “Furthermore, linear regression showed that mood sensitivity to CR can predict the current suicidal ideation score (β = -28.706, t = -2.801, p = 0.006) among patients, but not among HC (β = -2.204, t = 0.528, p = 0.599), suggesting that mood sensitivity to CR has patient-specific predictive utility for suicidal symptoms.”

      (3) The FDR correction for multiple comparisons mentioned briefly in lines 536-538 was not clear. Which analyses were included in the FDR correction? In particular, did the correlations between gambling rate and BSI-C/BSI-W survive such correction? Were there other correlations tested here (e.g., with the TAI score or ERQ-R and ERQ-S) that should be corrected for? Did the mediation model survive FDR correction? Was there a correction for other mediation models (e.g., with BSI-W as a predictor), or was this specific model hypothesized and pre-registered, and therefore no other models were considered? Did the differences in beta_gain across groups survive FDR when including comparisons of all other parameters across groups? Because the results were replicated in the online dataset, it is ok if they did not survive FDR in the patient dataset, but it is important to be clear about this in presenting the findings in the patient dataset.

      Thank you for raising the important issue of multiple testing and for asking us to clarify exactly which tests were covered by the FDR procedure. In the clinical dataset we conducted a large number of inferential tests (χ<sup>2</sup>, t-tests, ANOVAs, regressions) spanning: (i) group differences in demographic/clinical characteristics; (ii) sanity checks (e.g., anxiety/depression questionnaires); (iii) primary hypotheses (e.g., group differences in risky behavior); (iv) model-based analyses (parameter checks and between-group contrasts); and (v) control/sensitivity analyses. Post-hoc t-tests were performed only when the three-group ANOVA was significant. This yielded >150 p-values. FDR was applied using all these p-values. Please see Supplementary Note 8.

      (4) There is a lack of explicit mention when replication analyses differ from the analyses in the patient sample. For instance, the mediation model is different in the two samples: in the patient sample, it is only tested in S+ and S- groups, but not in healthy controls, and the model relates a dimensional measure of suicidal symptoms to gambling in the task, whereas in the online sample, the model includes all participants (including those who are presumably equivalent to healthy controls) and the predictor is a binary measure of S+ versus S- rather than the response to item 9 in the BDI. Indeed, some results did not replicate at all and this needs to be emphasized more as the lack of replication can be interpreted not only as "the link between mood sensitivity to CR and gambling behavior may be specifically observable in suicidal patients" (lines 582-585) - it may also be that this link is not truly there, and without a replication it needs to be interpreted with caution.

      Thank you for these important comments. This study focused on cognitive and affective computational mechanisms underlying increased risky behavior in STB. Accordingly, we compared patients with STB (S<sup>+</sup>) with patients without STB (S<sup>-</sup>) and healthy controls (HC) to examine the effects of STB on risky behavior. Therefore, group comparison, instead of dimensional measure of suicidal symptoms by Beck Scale for Suicidal Ideation, can answer our research questions directly.

      To enhance consistency between the clinical and replication datasets, we included all participants in each dataset when performing the mediation analysis. Given that S<sup>-</sup> and HC did not differ in gambling behavior or the approach parameter in the clinical dataset, we merged these two groups. In the replication dataset, to mirror the S<sup>+</sup> vs. S<sup>-</sup> contrast used clinically, we categorized the general sample into S<sup>+</sup> and S<sup>-</sup> based on BDI item 9. The mediation results remained significant in both datasets (the clinical dataset: a×b = 0.321, 95% CI = [0.070, 0.549], p = 0.016; the replication dataset: a × b = 0.143, 95% CI = [0.016, 0.288], p = 0.031), suggesting that STB is associated with increased risk behavior via stronger approach motivation.

      We also acknowledge the non-replication of the correlation between gambling behavior and mood sensitivity to certain rewards in the online sample. While this pattern might indicate that the link is specific to suicidal patients, it may also reflect sample-specific or unstable effects; thus, we now state this explicitly and interpret the finding with caution. Please see our revisions below:

      Page 15:

      “We next verified our results in an independent dataset, including the same task and BDI questionnaire in 747 general participants (500 females; age: 20.90±2.41)[46]. One item in BDI involves the measurement of STB. In item 9 of BDI, participants chose one option that describes them best: Option 1, “I don't have any thoughts of killing myself.”; Option 2, “I have thoughts of killing myself, but I would not carry them out.”; Option 3, “I would like to kill myself.”; Option 4, “I would kill myself if I had the chance.”. In line with the current definition of S<sup>+</sup>/S<sup>-</sup> in the clinical dataset, we identified S<sup>+</sup> group as choosing Option 2, 3, or 4, while participants selecting Option 1 were categorized as S<sup>-</sup> group.”

      Page 19:

      “Given significant correlations between group, approach parameter, and gambling rate for gain trials (ps < 0.017), we further conducted a mediation analysis with the assumption of the mediating effect of approach motivation of suicidality on the risk behavior. Given that we aimed to test the effect of STB, with S<sup>-</sup> and HC as controls, and given that S<sup>-</sup> and HC did not differ in gambling behavior or in the approach parameter, we merged these two groups for the mediation analysis. Results supported our hypothesis (a×b = 0.321, 95% CI = [0.070, 0.549], p = 0.016; Figure 2C), confirming that suicidal thoughts and behavior increase risk behavior through stronger approach motivation.”

      Page 26:

      “However, we did not observe any significant correlation between mood sensitivity to CR and gambling behavior (ps > 0.389), which suggests that the link between mood sensitivity to CR and gambling behavior may be specifically observable in suicidal patients. Alternatively, this non-replicated result may also reflect sample-specific or unstable effects, which needs to be interpreted with caution.”

      (5) In interpreting their results, the authors use terms such as "motivation" (line 594) or "risk attitude" (line 606) that are not clear. In particular, how was risk attitude operationalized in this task? Is a bias for risky rewards not indicative of risk attitude? I ask because the claim is that "we did not observe a difference in risk attitude per se between STB and controls". However, it seems that participants with STB chose the risky option more often, so why is there no difference in risk attitude between the groups?

      Thank you for pointing out the ambiguity. In our manuscript, “motivation” and “risk attitude” are defined at the computational level. Following prior work with this task Rutledge et al., (2015, 2016), we decompose observed gambling into (i) value-dependent valuation parameters that capture risk attitude (e.g., risk aversion and loss aversion, which scale the subjective value of outcomes), and (ii) value-insensitive, valence-dependent biases that capture approach/avoidance motivation. Accordingly, a higher gambling rate does not imply a change in risk attitude per se: it can arise from an increased value-insensitive approach bias even when risk-attitude parameters are comparable between groups which is what we observe for S<sup>+</sup> vs. controls. We have clarified this point in the computational modeling section.

      Pages 12-13:

      “Please note that a higher gambling rate does not imply a change in risk attitude per se: it can arise from an increased value-insensitive approach bias even when risk-attitude parameters are comparable between groups. Risk attitude is indeed conceptualized in economics as the curvature of the utility function (i.e., the subjective value) of the objective outcomes, with concave curves associated with risk aversion, and convex curves associated with risk seeking [54,56]. By contrast, the approach or avoidance bias apply to all the value. A possible interpretation of the approach bias is that participant approach the option with the highest possible gain (the lottery) in the gain frame; the avoidance bias would then reflect a tendency to systematically avoid the highest potential losses (the lottery) in the loss frame.”

      Reviewer #2 (Public review):

      Summary:

      This article addresses a very pertinent question: what are the computational mechanisms underlying risky behaviour in patients who have attempted suicide? In particular, it is impressive how the authors find a broad behavioural effect whose mechanisms they can then explain and refine through computational modeling. This work is important because, currently, beyond previous suicide attempts, there has been a lack of predictive measures. This study is the first step towards that: understanding the cognition on a group level. This is before being able to include it in future predictive studies (based on the cross-sectional data, this study by itself cannot assess the predictive validity of the measure).

      Strengths:

      (1) Large sample size.

      (2) Replication of their own findings.

      (3) Well-controlled task with measures of behaviour and mood + precise and well-validated computational modeling.

      Weaknesses:

      I can't really see any major weakness, but I have a few questions:

      (1) I can see from the parameter recovery that the parameters are very well identified. Is it surprising that this is the case, given how many parameters there are for 90 trials? Could the authors show cross-correlations? I.e., make a correlation matrix with all real parameters and all fitted parameters to show that not only the diagonal (i.e., same data is the scatter plots in S3) are high, but that the off-diagonals are low.

      Thank you for raising these thoughtful concerns. The current task consisted of 90 choices and 36 mood ratings. There were 5 choice parameters and 4 mood parameters. The apparently strong identifiability is not unexpected, as 90 choice trials and 36 mood ratings are comparable to those in prior computational modeling literature (Blain & Rutledge, 2022).

      As suggested, we computed cross-scorrelations between all generating (“true”) and recovered (“fitted”) parameters. The resulting matrix showed high diagonal (choice winning model: rs > 0.91; mood winning model: rs > 0.90) and low off-diagonal (choice winning model: abs(rs) < 0.63; mood winning model: abs(rs) > 0.40) correlations, further supporting parameter recovery. Please see Supplementary Pages 2-3.

      “Parameter recovery: Figure S3 shows good parameter recovery for both choice and mood winning model (choice: rs > 0.91, ps < 0.001; intraclass coefficients > 0.78; mood: rs > 0.90, ps < 0.001; intraclass coefficients > 0.86). Moreover, we computed cross-correlations between all generating (“true”) and recovered (“fitted”) parameters. The resulting matrix showed high diagonal (choice winning model: rs > 0.91; mood winning model: rs > 0.90) and low off-diagonal (choice winning model: abs(rs) < 0.63; mood winning model: abs(rs) > 0.40) correlations, further supporting parameter recovery.”

      Page 10:

      “The numbers of choice trials and mood ratings were comparable to those in prior computational modeling studies [34,35].”

      (2) Could the authors clarify the result in Figure 2B of a correlation between gambling rate and suicidal ideation score, is that a different result than they had before with the group main effect? I.e., is your analysis like this: gambling rate ~ suicide ideation + group assignment? (or a partial correlation)? I'm asking because BSI-C is also different between the groups. [same comment for later analyses, e.g. on approach parameter].

      Thank you for pointing out the lack of clarity. We performed group difference analysis and correlation of suicidal ideation analysis, separately. We first performed group difference analysis to test our hypothesis of STB effects. We then conducted correlational analysis to further specify our findings.

      (3) The authors correlate the impact of certain rewards on mood with the % gambling variable. Could there not be a more direct analysis by including mood directly in the choice model?

      Thank you for this insightful suggestion. As suggested, we tried to integrate mood into choice models by adding mood bias component(s) in line with previous literature (Vinckier et al., 2018). The first model (mcM1) assumes that mood biases choice, building on cM3 (the winning choice model). cmM2 further separated the mood bias parameter into two components according to participants’ choices.

      However, model comparison using BIC supported cM3 (Table S6), that is, without consideration of mood in choice modeling. This can be due to the lack of block design in our experimental design unlike e.g., Vinckier et al., (2018) and Eldar & Niv, (2015). Please see Supplementary Note 6.

      (4) In the large online sample, you split all participants into S+ and S-. I would have imagined that instead, you would do analyses that control for other clinical traits. Or, for example, you have in the S- group only participants who also have high depression scores, but low suicide items.

      Thank you for this insightful suggestion. Following prior suicide-related literature (Tsypes et al., 2024), we controlled for depression by including them as covariates. Note that depression scores were derived from our established bifactor model (Wang et al., 2025), which decomposed depression from the anxiety. These results remained largely significant (ps ≤ 0.050), except a marginally significant effect of group on gambling behavior (p = 0.059). Despite a trend, this effect with covariates of depression-related questionnaires is strong in our clinical cohort (p = 0.024; Table S8). This suggests that the link between suicidality and risky behavior persists above and beyond general depressive symptoms.

      Please see our clarifications below:

      Page 26:

      “After controlling for depression severity using our established bifactor model (see ref 60 for details), these results remained significant (ps ≤ 0.050), except a marginally significant effect of group on gambling behavior (p = 0.059). Despite a trend, this effect with covariates of depression-related questionnaires is strong in our clinical cohort (p = 0.024; Table S8). This suggests that the link between suicidality and risky behavior persists above and beyond general depressive symptoms.”

      Reviewer #3 (Public review):

      This manuscript investigates computational mechanisms underlying increased risk-taking behavior in adolescent patients with suicidal thoughts and behaviors. Using a well-established gambling task that incorporates momentary mood ratings and previously established computational modeling approaches, the authors identify particular aspects of choice behavior (which they term approach bias) and mood responsivity (to certain rewards) that differ as a function of suicidality. The authors replicate their findings on both clinical and large-scale non-clinical samples.

      (1) The main problem, however, is that the results do not seem to support a specific conclusion with regard to suicidality. The S+ and S- groups differ substantially in the severity of symptoms, as can be seen by all symptom questionnaires and the baseline and mean mood, where S- is closer to HC than it is to S+. The main analyses control for illness duration and medication but not for symptom severity. The supplementary analysis in Figure S11 is insufficient as it mistakes the absence of evidence (i.e., p > 0.05) for evidence of absence. Therefore, the results do not adequately deconfound suicidality from general symptom severity.

      Thank you for this important comment. Based on clinical interviews, we included patients with and without suicidality (S<sup>+</sup> and S<sup>-</sup> groups). However, in line with suicidal-related literature (e.g., Tsypes et al., 2024), two groups also differed substantially in the severity of symptoms (see Table 1). To address the request for evidence on specificity to suicidality beyond general symptom severity, we performed separate linear regressions to explain in gambling behaviour, value-insensitive approach parameter (β<sub>gain</sub>), and mood sensitivity to certain rewards (β<sub>CR</sub>) with group as a predictor (1 for S<sup>+</sup> group and 0 for S<sup>-</sup> group) and scores for anxiety and depression as covariates. Results remained significant after controlling anxiety and depression (ps < 0.027; Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, ps < 0.001), we performed Principal Components Analysis (PCA) on the clinical questionnaire to extract the orthogonal components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. We then performed linear regressions using these components as covariates to control for anxiety and depression. Our main results remained significant (ps < 0.027; Table S9). We believe that these analyses provide evidence that the main effects on gambling and on mood were specific to suicide.

      As pointed out, these “absence of evidence” cannot provide insights of “evidence of absence”. Although we median-split patients by the scores of general symptoms (e.g., depression and anxiety-related questionnaires) and verified no significant differences in these severities (Figure S11), we additionally conducted Bayesian statistics in gambling behavior, value-insensitive approach parameter, and mood sensitivity to certain rewards. BF<sub>01</sub> is a Bayes factor comparing the null model (M<sub>0</sub>) to the alternative model (M<sub>1</sub>), where M<sub>0</sub> assumes no group difference. BF<sub>01</sub> > 1 indicates that evidence favors M<sub>0</sub>. As can be seen in Table S7, most results supported null hypothesis, suggesting that general symptoms of anxiety and depression overall did not influence our main results. Overall, we believe that these analyses provide compelling evidence for the specificity of the effect to suicide, above and beyond depression and anxiety.

      Please see Table S7, S8 &S9 and our revisions below.

      Page 17:

      “Within patients, this group effect on gambling rate remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.024; also see Figure S11, Table S7 and Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, ps < 0.001), we performed Principal Components Analysis (PCA) to extract main components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. To further control for anxiety and depression, linear regression using these components as covariates revealed that the group effect on gambling rate remained significant (p = 0.024; Table S9).”

      Pages 18-19:

      “Within patients, this group effect on the approach parameter remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.027; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on approach parameter remained significant (p = 0.027; Table S9).”

      Page 21:

      “Within patients, this group effect on βCR remained significant after controlling for gambling rate, earnings, mood-related outcome effect, mood drift effect, sex, illness duration, family history, diagnosis, and various medications use (ps < 0.032), as well as general symptoms (e.g., depression and anxiety; p = 0.001; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on this mood parameter remained significant (p = 0.001; Table S9).”

      (2) The second main issue is that the relationship between an increased approach bias and decreased mood response to CR is conceptually unclear. In this respect, it would be natural to test whether mood responses influence subsequent gambling choices. This could be done either within the model by having mood moderate the approach bias or outside the model using model-agnostic analyses.

      Thank you for this important suggestion. As suggested, one interesting question was whether mood responses influence subsequent gambling choices and how to model them. First, we median-split mood responses (except the final rating) to compare gambling rate. Results showed a trend for less gambling rate in higher mood (t = -1.971, p = 0.050). However, there was no significant group difference (F = 0.680, p = 0.507). Second, with the assumption that mood biases choice, we constructed mcM1 based on cM3 (the winning choice model). Based on our finding of the negative correlation between mood sensitivity to certain rewards and gambling rate in S<sup>+</sup>, we separated β<sub>Mood</sub> parameter into β<sub>Mood-CR</sub> and β<sub>Mood-GR</sub> (cmM2). Model comparison using BIC supported cM3 (Table S6), that is, without consideration of mood in choice modeling. This can be due to the lack of block design in our experimental design unlike e.g., Vinckier et al., (2018) and Eldar & Niv, (2015). Please see Supplementary Note 6.

      (3) Additionally, there is a conceptual inconsistency between the choice and mood findings that partly results from the analytic strategy. The approach bias is implemented in choice as a categorical value-independent effect, whereas the mood responses always scale linearly with the magnitude of outcomes. One way to make the models more conceptually related would be to include a categorical value-independent mood response to choosing to gamble/not to gamble.

      We apology for the unclear statement. The approach bias is implemented in choice as a continuous value-independent effect, ranging from -1 to 1.

      It was true that the mood responses always scale with the magnitude of outcomes, since mood ratings were request after the outcomes. Therefore, mood parameters and the approach bias were both continuous.

      We also attempted to integrate mood into choice modelling. See Response 2 for Reviewer 3 for details.

      (4) The manuscript requires editing to improve clarity and precision. The use of terms such as "mood" and "approach motivation" is often inaccurate or not sufficiently specific. There are also many grammatical errors throughout the text.

      Thank you for this important suggestion. We have now explained motivation and mood in the Introduction section and the computational modeling section. Please see our clarifications below:

      Pages 3-4:

      “A growing literature indeed shows that risky behavior can be far better explained after adding value-insensitive approach and avoidance components to prospect theory [18,19], that is by including a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference. This class of models highlights the important role of value-insensitive motivational components in decision making in addition to risk attitude-driven valuation (e.g., loss/risk aversion) [20].”

      Page 5:

      “Although mood is thought to persist for hours, days, or even weeks [30–33], momentary mood, measured over the timescale in the laboratory setting, represents the accumulation of the impact of multiple events at the scale of minutes [30,32,34–38]. Momentary mood external validity is demonstrated e.g., through its association with depression symptoms [37]. Mood is different from emotions, which reflect immediate affective reactivity and is more transient (e.g., from surprise to fear) [31–33,39].”

      We have corrected grammatical errors throughout the manuscript.

      (5) Claims of clinical relevance should be toned down, given that the findings are based on noisy parameter estimates whose clinical utility for the treatment of an individual patient is doubtful at best.

      Thank you for this comment. We agree that we did not evaluate the noise in our estimate e.g., by assessing the test-retest reliability on the task parameters, which is outside the scope of the study, and it is indeed possible that parameter estimate is somehow noisy. Therefore, we tone down the clinical relevance of our results. Please see our revision below:

      Page 32:

      “Next, we did not evaluate the noise in our estimate e.g., by assessing the test-retest reliability on the task parameters and it is indeed possible that parameter estimate is somehow noisy.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Title: I believe "aberrant mood dynamics" is both too general and overstating the results of this study, which did not measure mood dynamics longitudinally. "Aberrant" is also overly pathologizing. I would suggest sticking more directly to the results, for instance, "Insensitivity of momentary mood to non-risky rewards in adolescent suicidal patients".

      Thank you for this suggestion. We have now corrected it.

      (2) Abstract: in line 61, "Our study uncovers the cognitive and affective mechanisms" suggests that these are the only ones, and you uncovered them. Of course, there could be more mechanisms contributing to risk behavior in STB, so I would suggest removing the word "the" or adding "one of the".

      Thank you for this suggestion. We have now corrected it.

      (3) One major weakness of this study is that suicidal thoughts and behaviors were not assessed via a clinical instrument such as the Columbia Suicide Severity Rating Scale - this should be mentioned upfront.

      Thank you for this comment. According to medical records and information from family and friends by the researcher and psychiatrists, patients with suicidal thoughts and behaviors were categorized as suicidal group (S<sup>+</sup>), while patients without suicidal thoughts and behaviors were identified as control group (S<sup>-</sup>). Note that medical records and information were recorded from clinical interviews where the psychiatrists were vigilant for signs of suicidal ideation and inquired about suicidal-related thoughts and behaviors from both the patients and their families. Therefore, the current group operation was possibly comparable to Columbia Suicide Severity Rating Scale.

      (4) Table 1: female/male are sex, not gender (gender is man/woman/transgender/non-binary).

      Thank you for this suggestion. We have now corrected it.

      (5) Equation 1: It would be good to clarify what happens in gain-only or loss-only trials (the other value is then 0, but this can be clarified as it is not technically a loss or a gain).

      Thank you for this suggestion. We have now corrected it. Please see below for our revision:

      Page 12:

      “Please note that V<sub>gain</sub> is 0 in gain trials and V<sub>loss</sub> is 0 in loss trials.”

      (6) Figure 1E: The model prediction is not informative here. Given the linear regression model, there is no other option except that the mean prediction would overlap with the mean empirical measurement (unless the model was specified incorrectly). The same is true in Figure 2A.

      Thank you for this suggestion. We have now removed plots for model prediction.

      (7) Figure 1G: There was no analysis of the differences between groups in terms of earnings, given that the ANOVA was not significant. Still, if the claim is that risky behavior is sometimes suboptimal in this task, it would be good to show that there is a correlation between, say, symptoms of STB across groups and 1) risky behavior and 2) earnings.

      Thank you for this insightful comment. In the patient cohort, risky behavior (gambling rate)—but not earnings predicted the current suicidal ideation score (BSI-C, β = 9.189, t = 2.004, p = 0.048; earnings, β = 0.001, t = 0.582, p = 0.562). The lack of association for earnings is consistent with the task design, in which there is no stable optimal policy and payouts are only a coarse proxy for decision quality. Future work in learning paradigms, where optimality is well defined, may be better suited to test earning-based links to STB. We have clarified this point below:

      Page 32:

      “Second, although we assumed that increased risky behavior in STB was suboptimal, the current task was not suited to test this, given the task design of random feedback for gambling option. Future work in learning paradigms, where optimality is well defined, may be better suited to test earnings-based links to STB.”

      (8) Line 290: "beta_gain: -1-1" is unclear. I believe you meant beta_gain \in [-1,1].

      Thank you for this suggestion. We have now corrected it to make it clear.

      (9) The gain and loss biases are modeled as minimum and maximum probabilities for choosing the gamble. This is a legitimate choice for value-agnostic biases, but it is not the traditional choice (as far as I know). I wonder if the same results would hold with the more traditional formulation of the bias as an added constant to the utility of the gamble, i.e., p(gamble) = 1/(1+ exp(-mu(U_gamble + beta_gain - U_certain)). I believe in this case, you would also not have to specify different equations for positive or negative biases, or to limit the bias to the range of [-1,1] (indeed, the bias would be in reward-equivalent units).

      Thank you for this suggestion. The winning choice model we used here was consistent with previous literature (Rutledge et al., 2015 & 2016), which decomposed the decision process into risk-attitude-driven valuation (e.g., loss and risk aversion) and value-insensitive motivational components. These approach/avoidance parameters are a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference.

      As suggested, we also compared the traditional bias choice model. Model comparison did not support this. Please see Supplementary Page 4.

      (10) Also, for equations 5-8, it seems that 5-6 are identical to 7-8 except for the use of beta_gain versus beta_loss. You might want to consider simplifying by putting beta in the equations and specifying in the text that, depending on the trial type (loss or gain), the relevant beta is used.

      Thank you for this suggestion. We have now simplified it. Please see our revision below:

      (11) It is not clear what equations are applied to mixed trials in cM3.

      Sorry for the confusion. We have now clarified this point.

      Page 12:

      “Approach/avoidance parameters are not applied to in mixed trials.”

      (12) Model comparison: the mood models are nested within each other (e.g., mM3 can be derived from mM1 by setting beta_EV = beta_RPE). In this case, model comparison can use the likelihood ratio test instead of BIC, which can be too conservative (and therefore does not support the extra beta parameter for RPE, different from previous results in the literature). I wonder if a likelihood ratio test would lead to results more in line with previous findings with this task?

      Thanks for this suggestion. We agree that mM1 (CR+EV+RPE) and mM3 (CR+GR) are nested. However, our model space also included unnested models, such as mM5 (CR+GR<sub>better</sub>+GR<sub>worse</sub>). Therefore, it was not reasonable in our model space to use likelihood ratio tests.

      (13) Line 346: The replication sample is described as "healthy participants," however, their health (or mental health) status was not assessed, and they may as well have mental health concerns. I would suggest calling this a general sample or an undifferentiated sample - but not a healthy sample.

      Sorry for the confusion. We have now corrected this phrase.

      (14) Line 363: "in addition to the replication of previous findings in the validation dataset" is unclear. Are those tests not two-tailed?

      Sorry for the unclear statement. In the replication analyses, we used one-tailed t-tests because the direction of the effect was revealed on the clinical dataset. Please see our clarification below:

      Page 15:

      “For the replication of previous findings in the validation dataset, we used one-tailed tests in line with our clinically motivated directional hypothesis.”

      (15) Line 372: "validating our group manipulation" - the presented work does not have a manipulation. Maybe you meant "validating our grouping of participants"?

      Thank you for this suggestion. We have now corrected it to make it clear.

      (16) Figure 2B: It is not clear how the data were binned for illustration purposes only, and why this binning is necessary (I have not seen it in other papers) - presenting the data from each subject and the correlation line with error margins (as is done here) should be sufficient.

      Thank you for flagging this. For illustration only, we binned the data proportional to group sizes: in the patient sample (S<sup>-</sup> n = 25; S<sup>+</sup> n = 58; ≈1:2), we displayed 3 bins for S<sup>-</sup> and 6 bins for S<sup>+</sup>. We agree that binning is not necessary; all statistics were computed on raw, unbinned data. The binned panel was included solely for visualization, consistent with our prior work (Blain et al., 2023).

      (17) Table 2: delta BIC should be presented per subject (that is, divided by the number of subjects in each group), as the groups are of different sizes, so as presented now, the columns are not comparable across groups.

      Thank you for the helpful suggestion. Our goal in Table 2 is not to compare ΔBIC magnitudes across groups, but to identify the winning model within each group. The ΔBICs are aggregated at the group level solely to rank models for that group. Dividing by the number of participants would rescale each group’s column by a constant and would therefore not affect the within-group ranking or the conclusion that cM3 is the best model in all groups. For this reason, we retain the current presentation and interpret each column within group rather than across groups.

      (18) Line 640 - the effect of expectations and prediction errors on mood was not only shown in healthy people, but also in people with depression (Rutledge et al., 2007, https://pubmed.ncbi.nlm.nih.gov/28678984/)

      Thank you for this comment. Indeed, Rutledge et al., (2017) showed evidence for CR+EV+RPE mood model in adult people with depression. However, our study recruited adolescents with depression or anxiety, given that adolescent period might provide a developmental window for opportunities for early intervention of suicidality. Therefore, it is also possible that the current winning model was specific to adolescents. Please see our clarifications below:

      Page 28:

      “It is also possible that the current winning model was specific to adolescents. Given that Rutledge et al., (2017) supported the “CR-EV-RPE model” in adults with depression, our study with adolescent populations may suggest a developmental change for mood sensitivities.”

      (19) Supplemental material: Is the R2 section about R-squared? Perhaps you can use superscript on the 2 to make that clearer? For Figure S2, how was model recovery determined? Should I interpret the confusion matrix as suggesting that the winning model for each and every simulated subject was the generating model, or was the winning model determined for the whole simulated population in each of the 100 simulations? Traditionally, confusion matrices use the former measure, but the results of 100% recoverability make me suspect the latter was used here. In Figure S3, should we not be looking at simulated parameters and recovered parameters? What are "real parameters" here?

      Thank you for these important comments. We now consistently denote the coefficient of determination as R<sup>2</sup> (with a superscript 2) throughout the manuscript and Supplementary Materials.

      For the model recovery analysis in Figure S2, we have clarified that the confusion matrix is computed at the population level. Specifically, for each of the 100 simulations we generated a full dataset under each candidate model, fit all models to that dataset, and selected the winning model based on group-level model evidence (BIC). Each cell in the confusion matrix therefore reflects the proportion of simulations in which model j was selected as the best-fitting model when the data were generated by model i. This operation was reasonable because the decision of the winning model is made on the population-level dataset rather than on individual subjects.

      In Figure S3, the term “real parameters” referred to the parameters used to generate the simulated data. To avoid confusion, we now relabel these as “simulated (generating) parameters” and explicitly describe the figure as showing the relationship between simulated (generating) parameters and recovered parameters. Please see Supplementary Pages 2-3:

      “Model recovery: We generated 100 simulated datasets for each model (3 choice models and 8 mood models) using the fitted parameters of each model as the ground truth. Each dataset contained 201 trials and included 3 (or 8) sets of simulated data corresponding to the respective models. For each simulated dataset, we then fit all models and determined the winning model at the population level based on group-level BIC, yielding a confusion matrix in which each entry represents the proportion of simulations in which model j was selected as the best-fitting model when the data were generated by model i. As shown in Figure S2, all models are highly identifiable, indicating excellent recovery performance for both the choice and mood models.”

      “Parameter recovery: Figure S3 shows good parameter recovery for both choice and mood winning model (choice: rs > 0.91, ps < 0.001; intraclass coefficients > 0.78; mood: rs > 0.90, ps < 0.001; intraclass coefficients > 0.86). Moreover, we computed cross-correlations between all generating (“generating”) and recovered (“fitted”) parameters. The resulting matrix showed high diagonal (choice winning model: rs > 0.91; mood winning model: rs > 0.90) and low off-diagonal (choice winning model: abs(rs) < 0.63; mood winning model: abs(rs) > 0.40) correlations, further supporting parameter recovery.”

      Typos:

      (1) Line 90: original → originate

      (2) Line 596-598 - the same phrase is repeated twice.

      (3) Line 616: on the other word → hand.

      Sorry for the mistakes. We have now corrected them throughout the manuscript.

      Reviewer #2 (Recommendations for the authors):

      For people unfamiliar with interpersonal theory or motivational-volitional model, or three-step theory (lines 105-106), could you briefly explain the key idea of mood and suicide before going to the decision-making tasks? And from this, maybe motivate the predictions in your task? In particular, in the abstract and introduction, the phrasing could be a bit more concise and simpler. In the abstract, sentences were sometimes quite long. In the introduction, some paragraphs are somewhat repetitive. In the discussion, there were some typos.

      Thank you for these suggestions. We have now explained the key idea of mood and suicide before going to the decision-making tasks in the introduction, which can be seen below:

      Pages 4-5:

      “Contemporary theories of suicide converge on the idea that STB is initially caused by low mood experience. The interpersonal theory of suicide proposes that suicidal desire arises when people simultaneously feel socially disconnected (“thwarted belongingness”) and like a burden on others (“perceived burdensomeness”), experiences that are tightly linked to chronically low mood [25]. The motivational–volitional model [26] and the three-step theory [27,28] similarly emphasize that when negative mood and feelings of defeat or entrapment are experienced as inescapable, they can give rise to suicidal ideation, and that the progression from ideation to suicide attempts depends on additional factors such as reduced fear of death, increased pain tolerance, and a tendency to act impulsively under intense affect. Some official organizations, e.g., National Institute of Mental Health, have also listed mood problems as warning signals [8]. Interestingly, within the framework of decision making under uncertainty, gambling on lotteries with a revealed outcome has been found to induce high mood variance [29], providing an opportunity to assess the relationship between deficient mood and increased gambling decisions in STB.”

      We have also refined the wording and corrected typos throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Since many readers might only read the abstract, it is important that it is both informative and accurate. I have two suggestions in this respect. First, for the abstract to be more informative, it may be helpful to indicate already there that these are value-insensitive approach-avoidance parameters, in the sense that they favor/disfavor the gamble regardless of the potential outcomes' magnitude or probability. This issue is also present throughout the text, where the phrases "approach and avoidance motivation" are referred to as if they have established and precise computational definitions. In my view, these terms could just as easily be interpreted as parameters that multiply the value of potential gains or losses, which is not what the authors mean. It would be helpful to clarify this terminology.

      Thank you for these suggestions. In line with previous literature (Rutledge et al., 2015 & 2016), approach and avoidance motivation are indeed defined at the computational level, referring to a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference. We have cited these papers in the manuscript. We also make it clear to further clarify approach and avoidance parameters in the abstract and introduction. Please see our revisions below:

      Page 2 (Abstract):

      “Using a prospect theory model enhanced with value-insensitive approach-avoidance parameters revealed that this rise in risky behavior resulted only from a heightened approach parameter in S<sup>+</sup>.”

      “Altogether, model-based choice data analysis indicated dysfunction in the approach system in S<sup>+</sup>, leading to greater propensity for gambling in the gain domain regardless of the lottery expected value.”

      Page 3 (Introduction):

      “A growing literature indeed shows that risky behavior can be far better explained after adding value-insensitive approach and avoidance components to prospect theory [18,19], that is by including a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference. This class of models highlights the important role of value-insensitive motivational components in decision making in addition to risk attitude-driven valuation (e.g., loss/risk aversion) [20].”

      (2) The statement "our study uncovers the cognitive and affective mechanisms contributing to increased risk behavior in STB" is overstating the findings, as the study may have uncovered some contributing mechanisms, but likely not all of them. Removing the word "the" would fix this issue.

      Thank you for this suggestion. We have now corrected it.

      (3) Since mood is typically defined as lasting hours, it's inappropriate to refer to ratings that only reflect the last few trials as self-reports of mood. To be sure, I view the distinction between emotions and moods as quantitative, not qualitative, so I do not think there is a problem studying the former to understand the latter, but to avoid confusion, the terminology should follow common usage.

      Thank you for this suggestion. We follow previous work and operational definitions regarding mood (Rutledge et al., 2014, Eldar & Niv, 2015, Vinckier et al., 2018). Emotion is usually a very brief response to a specific stimulus (Emanuel & Eldar, 2023), e.g., leading to rapid changes like surprise then fear. In contrast, mood is defined as a diffuse state that is not specific to one stimulus. Here, we operationally and computationally define mood as an affective state reflecting the recent history of safe and gamble outcomes. We now clarify that point in the main text. Please see our revision below:

      Page 5:

      “Although mood is thought to persist for hours, days, or even weeks [30–33], momentary mood, measured over the timescale in the laboratory setting, represents the accumulation of the impact of multiple events at the scale of minutes [30,32,34–38]. Momentary mood external validity is demonstrated e.g., through its association with depression symptoms [37]. Mood is different from emotions, which reflect immediate affective reactivity and is more transient (e.g. from surprise to fear) [31–33,39].”

      (4) Line 78: The phrases "increase in risk attitude", "decrease in loss attitude", and "decrease in value-independent choice biases" are unclear to me in terms of their directionality. An attitude might be avoidant or embracing. If it is the former then increasing it would decrease risk-taking.

      Thank you for pointing out the ambiguity. We have now corrected them throughout the manuscript. Please see our revision below:

      Page 4:

      “We therefore hypothesized that heightened approach motivation, or weakened avoidance motivation, would account for increased risk behavior in STB.”

      (5) Line 125: I was not sure why one would expect the mood response to gamble-related quantities (EV and RPE) to be lower in STB and not higher.

      Sorry for the typo. We hypothesized that mood would respond more strongly to gambling-related quantities expected value (EV) and reward prediction error (RPE)—in adolescents with STB than in controls, given prior evidence that STB is associated with greater risk-taking.

      (6) The text could use proofreading, as there are many typos. These are from the first 100 lines alone:

      (a) Abstract: regardless the lotteries -> regardless of the lotteries'.

      (b) Line 78: it remains whether.

      (c) Line 80: can each -> each can.

      (d) Line 90: may original from.

      Sorry for the mistakes. We have now corrected them throughout the manuscript.

      (7) The rationale for focusing on the S+ group for mood model comparison is incorrect. The purpose is to identify parameters that vary as a function of suicidality, and for that, the S- group is just as important.

      Thank you for this comment. We agree that the S<sup>-</sup> group is as important as the S<sup>+</sup> group. A direct comparison was complicated because the winning mood models differed (S<sup>+</sup>: mM3; S<sup>-</sup>: mM5; Table 3). To ensure comparability, we checked results from both model specifications (mM3 and mM5). The conclusions were convergent: mood sensitivity to certain rewards (CR) was lower in S<sup>+</sup> than in S<sup>-</sup> (see Fig. 3 for mM3 and Fig. S8 for mM5).

      (8) There appears to be a contradiction between the inclusion criteria, which include having experienced suicidal thoughts and behaviors, and the definition of the S- group as not having suicidality.

      Thank you for pointing out this mistake. The corrected version of inclusion criteria can be seen on Page 7:

      “Patients were included if they met the following criteria: 1) both the researcher and psychiatrists agreed on their group classification; 2) they had a current diagnosis of major depressive disorder (MDD; unipolar depression), generalized anxiety disorder (GAD), or bipolar disorder with depressive episodes (BD), confirmed by two experienced psychiatrists using the Structured Clinical Interview for DSM-IV-TR-Patient Edition (SCID-P, 2/2001 revision; see Supplementary Note 1 for details);3) they were between 10 and 19 years of age; 4) they had no organic brain disorders, intellectual disability, or head trauma; 5) they had no history of substance abuse; 6) they had no experience of electroconvulsive therapy.”

      (9) It would be helpful to specify whether mood modeling was based on objective or subjective values, and why.

      Thank you for this helpful suggestion. We have now clarified whether mood modeling was based on objective or subjective values, and why. Specifically, we constructed two model families: one in which mood was driven by objective monetary outcomes (objective values) and one in which mood was driven by subjective values derived from each participant’s fitted choice model (subjective values). We then used the VBA_groupBMC function in the VBA toolbox to perform family-wise model comparison, with 8 candidate mood models within each family. Consistent with previous literature, the objective-value family provided a clearly superior fit to the data (exceedance probability, EP = 1.000). Based on this result and for parsimony, we report and interpret the mood modeling results from the objective-value family in the main text. We have clarified this point in Supplementary Note 9.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents an interesting approach for finding electrophysiological models that match experimental patch-clamp data. The authors develop a new method for deriving optimized current clamp protocols by training a neural network on synthetic data. This optimized current clamp is then used on both computational training data and on experimental data to predict current gating and conductance parameters that correctly reconstruct the electrical phenotype.

      Strengths:

      (1) The fitting of gating variables through an optimized patch clamp protocol is interesting.

      (2) The inclusion of experimental data is important, and the approach is shown to be effective in fitting them.

      Weaknesses:

      (1) Some clarity is necessary on the generation and selection of variable IPSC models. With such a large variation in so many parameters, I would expect some resulting parameters to generate non-realistic phenotypes, quiescent cells, etc. Are all 200,000 or 1,100,000 generated cells viable? Or are they selected somehow for realistic cell properties?

      Thank you for this important point. We agree that broad parameter variation can generate non-physiological model behavior. Indeed, with the +/-40% perturbation range, some simulated cells produced non-realistic outputs, including quiescent behavior, and failure to generate a complete action potential. These cases were excluded from the dataset. As a result, only cells exhibiting physiologically meaningful and numerically stable behavior were retained for further analysis. We have clarified this selection procedure in the Methods section. We applied a large variation to ensure that all possible combinations and morphologies were included in the training and testing data so the model would readily ingest new data and perform robustly.

      (2) The error shown in Figure 4 between different population sizes is not completely explained in the text - there seems to be a minimal difference between a population of 1,000 and 10,000, followed by a very good fit at 200,000. Is there a particular threshold that needs to be crossed where the error drops off? Related, how was the 200,000 number chosen?

      Thank you for this observation. We agree that the decrease in error shows a gradual performance improvement as the population size increases, rather than a strict cutoff. As shown in Figure 4, the difference between 1,000 and 10,000 samples is small, but as we continue to increase and get to around 200,000 samples, we see strong error minimization. This indicates how much training data is needed for optimal model performance. This improvement is due to better coverage of the high-dimensional parameter space, which helps the network learn the nonlinear relationships between the parameters and outputs.

      We tested a range of training data sets and found that above 200,000 training data sets, the model consistently produced low, stable errors and good test-training agreement. The test error decreased with the training error as the population size increased, indicating better generalization and suggesting that the model accurately predicts unseen data rather than overfitting to the training set.

      (3) Related to the point above, the 1,100,000 population for fitting experimental data also needs a more complete explanation: how was this number chosen, and how does the error compare with the other population sizes shown in Figure 4?

      Thank you for this question. We found that at a training data set size of 1,100,000 we were able to cover the large parameter space induced by +/-40% parameter perturbation. iPSC-CM measurements are known to exhibit high variability, and we wanted to capture the full range in the training data set so the model could ingest a wide range of experimental data. It is trivial to generate new training data, for example, to capture different experimental conditions like temperature differences, mutations, drugs, or ionic variability. We view this flexibility as a substantial strength of the approach. But the large perturbations we show in this study (+/-40%) allow the generation of a very broad range of cellular phenotypes while maintaining physiologically realistic ionic current properties and action potential behavior. Consistent with Figure 4, increasing population size reduces prediction error and improves generalization. The larger dataset provided more stable, accurate predictions when fitting experimental data, without evidence of overfitting.

      (4) Why are the optimized current clamp protocols different between panels A and B in Figure 5? Are they somehow informed by experimental data?

      Thank you for this question. The stimulation protocol used in panels A and B is identical. Panels A and B show whole-cell currents recorded under the same stimulation conditions as in Figure 3. The differences reflect variability in the underlying whole-cell ionic currents of the model cells rather than differences in the applied protocol. This is exactly the idea: the exact same protocol will generate different whole-cell currents in individual cells, but the model can find parameter sets for all of them.

      (5) Figure 6D: Is the EAD risk in panel D specific to cell 1, 2, or the pooled variants of both?

      Thank you for this question. We have clarified this point in the revised manuscript. The EAD risk shown in panel D is computed from the pooled variants of both Cell 1 and Cell 2, rather than being specific to either cell individually.

      (6) How sensitive is the fitting to minor parameter variation? Further, if one were to pick, let's say, the next-best-fitting value, would that fall close to the best one? Is the solution found unique, or are there multiple sets with good fits?

      Traditional optimization methods, such as Nelder–Mead, directly fit the model to the observed data by iteratively minimizing the error for each dataset. As a result, the solution can depend on the initial parameter guess and may converge to different local minima. In contrast, our approach trains a deep learning model on synthetic data generated from the baseline model, learning a mapping from whole-cell currents to the corresponding 52-parameter sets by minimizing prediction error. The mean squared error (MSE) decreases from approximately 10⁻² to below 10⁻³, with training and test errors overlapping closely, indicating stable training, good generalization, and accurate reproduction of the observed signals.

      The model achieves very low MSE and reproduces the electrophysiological outputs with high fidelity. However, accurate reproduction of the outputs does not imply a unique parameter solution. This is illustrated in Figure S1, where baseline and predicted parameter values show close agreement overall, yet small deviations persist across parameters. This indicates that different parameter combinations can yield similar whole-cell behaviors due to parameter correlations and compensatory effects. In such cases, the model learns to predict a representative parameter set that is most consistent with the training data and loss function, rather than converging to a single unique solution within a fixed numerical tolerance.

      Reviewer #2 (Public review):

      Summary:

      The authors present a computational framework for generating "cell-specific" digital twins of human iPSC-CMs from a single optimized voltage clamp recording. Using deep learning trained on > 1 million artificial cells, the authors demonstrate that the model can infer 52 biophysical parameters governing 6 major ionic currents, and the resulting digital twins can reproduce experimentally recorded action potentials.

      Strengths:

      The framework has clear potential for understanding cellular heterogeneity in iPSC-CMs, predicting individual drug responses, and reducing the experimental burden of multiple patch clamp protocols.

      Weaknesses:

      There are several concerns about the validation of the model and its clarity. First, the biological variability being modeled in this manuscript is not defined well. It is unclear whether the framework addresses cell-to-cell differences within a single differentiation batch, variability across iPSC lines, or donor-to-donor differences. This ambiguity makes it difficult to interpret what the "digital twin populations" actually represent biologically. Second, the main claim, "the digital twins enable drug testing and arrhythmia prediction that would be impractical experimentally", is not experimentally validated. For example, the E-4031 simulations predict EAD rates, but no direct experimental head-to-head comparison is provided to confirm that these predictions are accurate. Third, technical reproducibility and biological representativeness are not assessed. Single voltage clamp recordings are inherently noisy. Without knowing how much variability comes from the recording process (technical variation) vs true biological differences, it is difficult to judge whether observed "cell-specific" parameter differences are meaningful. In addition, the optimized protocol is claimed to be superior to conventional approaches, but again, no experimental comparison is shown.

      The authors should address these concerns, with particular emphasis on clarifying the biological context and providing direct experimental validation. Below are detailed specific points:

      (1) Ambiguous definition of iPSC-CM heterogeneity. The authors model "typical iPSC-CM heterogeneity" by varying 52 parameters +/- 40% around a baseline model (Figure 1), generating > 1 million synthetic cells. However, the manuscript does not clearly state what biological variability this model is intended to capture. Is this modeling within-line, cell-to-cell variability (e.g., cells from the same dish or differentiation batch that differ due to stochastic gene expression or maturation state)? Or is this modeling between-line or between-donor variability (e.g., genetic background differences, reprogramming efficiency)? This distinction is critical for interpretation. If the goal is to understand why different cells in the same dish behave differently, then training data should reflect that. If the goal is to compare patient lines or disease models, the framework needs validation across multiple donors or lines.

      For example, the experimental validation in Figure 5 uses a single iPSC line (iPS-6-9-9T.B), but how many differentiation batches or dishes were tested, or whether cells came from the same preparation are unclear. Another example is that the wide AP diversity in the training population (Figure 1A) is impressive, but there is no demonstration that real experimental cells actually fall within this assumption range of +/- 40%.

      From a biological perspective, iPSC-CMs are known to be highly heterogeneous within lines (maturation state, metabolic differences, epigenetic variation, spatial differences within the same dish, etc) and between lines (different donor/genetic background). Thus, please explicitly state whether the +/- 40% variation is intended to model within-line or between-line heterogeneity, and justify this choice with wet experiment data (or reference to experimental literature on iPSC-CM variability). Please clarify how many dishes, differentiation batches, and time points post-differentiation were used for experimental recordings (Figures 5-6). If the framework is intended to generalize across lines from different donors, please test the model on multiple independent iPSC lines (from different donors).

      Thank you for this important and insightful comment. The selected ±40% range was chosen to broadly explore all physiologically plausible electrophysiological behaviors, not to match a specific experimental distribution. Our goal was to cover enough behaviors for the model to learn a reliable mapping between responses and ionic parameters.

      We recognize that this approach does not explicitly account for variability between lines or donors. We have a current project focused on extending the framework to include multiple iPSC-CMs from patient donors, but given that the model framework successfully reproduces such a broad range of cell phenotypes, we feel confident that it will readily apply to different genetic backgrounds from patient-specific cells. This study is underway.

      We have updated the manuscript to clarify how the modeled variability is interpreted and added a discussion of these limitations. Furthermore, we clarified the experimental conditions, such as the number of differentiation batches and recording settings, in the revised Methods section.

      (2) Biological representativeness of single-cell measurements.

      The framework generates digital twins from single voltage clamp recordings. The patch clamp recordings in iPSC-CMs are subject to substantial technical variability. The manuscript does not address a fundamental question: "How representative are the measurements from a single cell on the dish (or line)?" In other words, if I measure one cell from a dish of a million cells, does that cell's digital twin tell me something about the dish as a whole, or just about that one cell? The manuscript presents Cell 1 and Cell 2 (Figures 5-6) as distinct individuals, but it's unclear whether these differences reflect true biological heterogeneity or simply sampling variability. I think the authors should perform replicate recordings on multiple cells (e.g., > 10 cells) from the same dish (same differentiation batch) and quantify how much the inferred parameters vary, and then compare between lines.

      Thank you for this important comment. We agree that the representativeness of single-cell measurements and the impact of technical variability are important considerations in interpreting the results. In this study, the framework is designed to generate digital twins that reflect the electrophysiological properties of individual recorded cells, rather than to directly represent the behavior of the entire cell population within a dish.

      As such, differences observed between Cell 1 and Cell 2 are intended to reflect variability at the single-cell level, which may arise from a combination of biological heterogeneity and experimental variability. We agree that systematic replicate recordings across multiple cells are valuable to quantify the relative contributions of biological and technical variability, and to assess the consistency of inferred parameters. However, this is beyond the scope of the current study. We have added clarification in the manuscript to explicitly state this limitation and to outline this as an important direction for future work.

      (3) No experimental validation of the main claim that in silico populations can replace wet experiments.

      The most exciting claim in the manuscript is that digital twins enable drug testing and arrhythmia prediction "at scale" without requiring hundreds of patch clamp experiments. Specifically, the authors show that in silico populations derived from two experimental cells (Figure 6C) predict dose-dependent EAD incidence for the IKr blocker E-4031 (Figure 6D), with ~3% of cells showing EADs at 50 nM.

      However, this prediction is not validated experimentally. If I actually patch 20-30 real iPSC-CMs and apply 50 nM E-4031, will ~3% of them show EADs, as the model predicts? Without this validation, I think the drug testing framework is purely hypothetical. The model may be internally consistent (e.g., Cell 1's twin behaves differently from Cell 2's twin), but there is no evidence that these in silico populations reflect real biological variability in drug response. Please provide experimental validation that justifies the prediction by digital twins.

      Thank you for this important comment. We agree that experimental validation of population-level drug response will be valuable for establishing the quantitative accuracy of the predicted EAD incidence. The E-4031 simulations are intended as a proof-of-concept illustrating how the framework can identify susceptible subpopulations and quantify relative proarrhythmic risk in silico. We agree that direct comparison with large-scale experimental datasets is a key next step, and we are working hard to get the study funded so that we can perform those experiments and bring this technology to scale.

      (4) Experimental validation and head-to-head comparison of optimized protocol.

      The authors claim that their deep learning-optimized voltage clamp protocol (Figure 3, Figure 4A) is superior to conventional approaches, but they have not validated this experimentally by doing a head-to-head comparison. The manuscript does not compare the optimized protocol to any published voltage clamp designs. If the optimized protocol is genuinely easier to implement and more informative than existing approaches, this would be a major practical advance. But without side-by-side comparison, it is impossible to judge whether the optimization made a real difference.

      Thank you for your comment. We agree that comparing directly with traditional voltage-clamp protocols through experiments would be useful. In this study, our main aim was to show that the optimized protocol enhances parameter inference within the modeling framework, not to prove experimental superiority. We have clarified this point in the revised version.

      Reviewer #3 (Public review):

      Summary:

      This work uses a convolutional neural network to optimize a voltage clamp protocol to identify features and parameters from human pluripotent stem cell-derived cardiomyocytes.

      Yang et al. introduce an innovative experimental framework that integrates computational modeling and deep learning to generate a digital twin of human pluripotent stem cell-derived cardiomyocytes (hPSC-CMs).

      Strengths:

      The major strength is the methodology used to bridge in silico prediction of cell behavior and mechanistic insights from the experimental dataset.

      The approach used in this study represents a significant step toward precision medicine by enabling in silico prediction of cellular behavior and mechanistic insight from experimental datasets. The study addresses an important and timely challenge in stem cell-based and personalized medicine, and the authors compellingly leverage state-of-the-art methods alongside strong expertise in computational modeling and cardiac electrophysiology

      Weaknesses:

      While the overall approach is highly compelling and the potential impact is substantial, there are two areas where clarification and refinement, particularly in the phrasing and framing used throughout the manuscript, would further strengthen the work.

      (1) While the overall goal of the study is compelling, the manuscript would benefit from clearer articulation of how the proposed framework is intended to be used in practice. In particular, it is not entirely clear whether the authors envision this approach as:

      (a) a method to extract population-level trends that, when paired with biological data, enhance statistical power and interpretability, or

      (b) a strategy capable of constructing a population-based model from limited single-cell recordings. If the latter is intended, additional guidance on the number of action potentials required per cell and the assumptions underlying this extrapolation would greatly clarify the scope and applicability of the method.

      Thank you for this thoughtful comment. We agree that the intended use of the framework should be more clearly articulated. In this study, we generate a large synthetic population of iPSC-CM models by varying 52 biophysical parameters governing key ionic currents. A neural network is trained on simulated whole-cell current responses to learn a mapping between current profiles and model parameters. Experimental recordings are then used as inputs to this trained model to infer ionic parameters, rather than directly fitting the model to data. This enables individual recordings to be interpreted within a large, physiologically plausible parameter space and supports population-level analysis of electrophysiological variability. The primary goal of the framework is therefore to facilitate mechanistic interpretation of variability and relate experimental observations to underlying ionic currents. But the longer-term intended goal is to develop digital twins from patient-derived cell lines and then use populations constructed from patient-specific digital twins to screen therapeutics and identify arrhythmia marker vulnerability in a very thorough and high-throughput way. We have clarified this in the revised manuscript.

      (2) The manuscript would also benefit from a clearer explanation of how electrophysiological heterogeneity observed in hPSC-CMs is linked to inter-patient variability. Although the authors state that this framework can be generalized to compare patient-specific hiPSC-CM lines, it remains unclear how this generalization is achieved, given the substantial sources of variability intrinsic to hiPSC-CMs (e.g., batch effects, reprogramming strategy, differentiation protocol, and maturation state). As acknowledged by the authors, addressing this level of variability likely requires large datasets; further clarification of how the proposed approach mitigates or accommodates these challenges would strengthen the translational claims.

      Below are my suggestions that could help strengthen the claims in the manuscript:

      (1) Adding a dedicated section describing the electrophysiological phenotype of the hPSC-CMs used in this study would help justify the choice of the underlying ionic model and the selection of the six ion currents analyzed. These currents are not only developmentally regulated but may also vary substantially across different hPSC-CM lines, which has implications for generalizability.

      Thank you for this important suggestion. We agree that providing additional context on the electrophysiological phenotype of the hPSC-CMs strengthens the rationale for both the underlying ionic model and the selection of currents analyzed.

      We have expanded the Methods section to clarify this point. Briefly, the ionic currents were selected based on the Kernik-Clancy iPSC-CM model developed in our prior work, which was specifically designed to capture the range of electrophysiological variability observed within an iPSC-CM cell line using a population-based framework. In this model, variation in key ionic conductances is sufficient to reproduce the diversity of action potential morphologies, spontaneous activity, and repolarization dynamics commonly reported experimentally, while avoiding non-physiological behaviors.

      Accordingly, we focused on six primary ionic currents that are known to play dominant roles in shaping action potential characteristics and variability in iPSC-CMs. This selection reflects a balance between model parsimony and physiological relevance, enabling the framework to capture the expected spectrum of variability within a given cell line. We also note that the framework is extensible, and additional currents or alternative parameterizations can be incorporated to account for differences across cell lines, donors, or experimental conditions in future studies. See updated discussion.

      (2) If feasible, inclusion of patch-clamp data from an additional hPSC-CM line would significantly strengthen the claim that this framework can harmonize and generalize across datasets and cell sources.

      Thank you for this helpful suggestion. We agree that adding data from more hPSC-CM lines would improve the framework's generalizability. In this work, our goal was to show that the digital twin framework is data-driven and can easily be expanded to include more hPSC-CM lines, allowing for cross-line comparisons in future studies. We have clarified this and included a discussion of this limitation in the revised manuscript. We are currently seeking funding for patient-specific lines as well to allow scalability.

      (3) The authors note that the experimental cells exhibited high variability in action potential morphology. This is an important observation that directly supports the motivation for the study and should be explicitly presented, even if only in the supplementary materials.

      Thank you for this suggestion. We agree that explicitly showing the variability in experimental action potential morphology strengthens the motivation for this study. We have now added a section in the discussion discussing this and referencing the many prior studies that focused on iPSC-CM variability, including the studies upon which our initial model (Kernik-Clancy) was based.

      (4) In the hERG-blocker experiments, further clarification is needed regarding the biological relevance of the reported 3% incidence of early after depolarizations (EADs). Additionally, an interrupted sentence in this section makes it unclear whether the goal is to demonstrate that the digital twin can capture rare arrhythmic risk events or whether the digital twin is necessary to determine whether this level of risk is clinically meaningful.

      Thank you for this important comment. We agree that more clarification is needed on the ~3% EAD incidence and the digital-twin role. This analysis aims to show that electrophysiological variability can create a small, susceptible subpopulation under drug effects, not to set a clinical risk threshold. The observed ~3% EAD incidence reflects the emergence of such a susceptible subpopulation under hERG block. While relatively small, this fraction is important because it arises from modest, physiologically plausible variation in ionic properties and would be difficult to capture using single-cell or small-sample approaches. As described in the Discussion, this variability-driven emergence of EADs provides a quantitative measure of proarrhythmic risk at the population level. The digital-twin framework enables systematic identification and quantification of these rare events, linking cell-level variability to population-level responses. We have revised the manuscript to clarify this point.

      (5) The manuscript states that some action potentials were excluded from the experimental dataset. A brief explanation of the exclusion criteria, along with guidance on how to distinguish high-quality from low-quality recordings, would improve transparency and reproducibility.

      Thank you for this comment. We agree that the definition of failed recordings should be clarified. We have now specified the exclusion criteria in the Methods section.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It would be helpful if the network cartoon in Figures 2 and 3 were replaced with a simplified sketch of the actual neural network used.

      Thank you. We now have new figures 2 and 3.

      (2) Subsection title for the Introduction has a typo.

      Thank you. We have fixed it.

      Reviewer #2 (Recommendations for the authors):

      (1) Technical quality control criteria are not specified.

      The Methods section states that "any incomplete or failed recordings were excluded," but does not define what constitutes a failed recording. The criteria could be subjective.

      Thank you for pointing this out. We agree that the definition of failed recordings should be clarified. We have now specified the exclusion criteria in the Methods section.

      “Recordings were excluded if they exhibited no spontaneous firing, abnormally slow firing rates, or failed to capture a complete action potential waveform. These criteria were applied consistently across all recordings.”

      (2) "Cell-specific" may overstate the claim.

      The term "cell-specific digital twins" (title, throughout) implies that the inferred parameters reflect the true biological state of each cell. However, parameters are derived only from curve-fitting to electrophysiological data and do not reflect other biological components (e.g., gene expression, contractility, calcium handling, metabolism, etc). Please consider rephrasing to "electrophysiology-based digital twins", "voltage clamp-matched digital twins", etc.

      Thank you for this important comment. We agree that the term “cell-specific” could be interpreted as implying a complete representation of the biological state of each cell. We have also adjusted the wording in relevant sections to avoid over-interpretation.

      Reviewer #3 (Recommendations for the authors):

      (1) I would add the list of the 52 parameters in the method section/SI and not just in the reference. Additional justification of why the perturbation was set as +/- 40% for the 52 parameter or +/- 20% for the EAD population would also help.

      Thank you for this helpful comment. We have included model equations and highlighted the 52 parameters in the Supplementary Information and provided additional justification in the Methods.

      (2) In Figure 1B, might be helpful to add the axis of the Vm instead of the dotted line indicating 0 mV to show differences in the diastolic potential.

      Thank you! We have now updated Figure 1B.

      (3) Figure 1C-I might be more impactful to show traces from the AP shown in Figure B to reinforce the impact of a single current in the AP shape.

      We have now updated Figure 1C-I to include traces from the AP shown in Figure 1B.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the manuscript by Winke et al, the authors present evidence that fear-induced analgesia is mediated by somatostatin projection cells from the vlPAG to the RVM. This study uses a mouse model of fear-induced analgesia, and incorporates optogenetic circuit manipulation with behaviour and electrophysiology to gain a meaningful insight into a novel circuit involved in fear-induced analgesia.

      Strengths:

      (1) This is a well-constructed study with appropriate controls and analyses.

      (2) Alternative interpretations of the data are systematically considered and eliminated via rational experiments. The authors are commended for a nice piece of experimental work.

      (3) The vlPAG is a known region of pain modulation, and this study adds valuable insight to the circuit involved in fear-associated analgesia.

      We are very thankful to the referee for these positive comments.

      Weaknesses:

      (1) Only male mice are included in this study.

      We thank the reviewer for this point. We used only males in this first study for practical reasons to work with a population as homogeneous as possible. However, taking sex differences in biological mechanisms into account, we included this restriction in the summary and discussion

      (2) Animals are excluded from analyses based on clearly defined criteria, but it is not clear how many mice were excluded from each group.

      We thank the reviewers for raising this point. As stated in the Methods, we applied strict inclusion criteria for mice undergoing the hot-plate test, specifically a discrimination index ≥ 0.4 and a conditioning index ≥ 0.3. Using these criteria, 23% of wild-type mice were excluded for failing to meet the discrimination criterion. In the transgenic groups, an average of 20% of mice failed to meet the learning criteria, and an additional 12% were excluded due to incorrect opsin injection or misplaced optic fiber placement.

      (3) The authors implement a pain sensitivity assay that involves a hot plate with progressively increasing temperature. The time to nociceptive responses is reported. Without reporting the actual temperature at which the mice respond, it makes it difficult to compare nociceptive responses to previously published work (which typically use a defined and static hotplate temperature).

      We thank the reviewer for this comment. We provided this information related to the actual temperature of the nociceptive response in the original manuscript in supplementary figures 1, 2 and 5.

      (4) The authors present evidence that inhibition of SST vlPAG cells enhances spinal nociceptive electrophysiological responses, but the corresponding pain sensitivity is not altered (Figure 2, CS- condition). The reason for the discrepancy between electrophysiological and behavioural responses is not clear.

      We believe this comment arises from a misunderstanding of our results. In our study, inhibiting SST+ vlPAG cells did not increase nociceptive electrophysiological responses. Instead, it decreased spinal nociceptive transmission, as evidenced by reduced nociceptive field potentials and WDR responses in Figure 4c,e. Consistent with this electrophysiological effect, photoinhibition of SST+ vlPAG cells also produced behavioral analgesia, as evidenced by increased nociceptive response latency in the hotplate test under both CS− and CS+ conditions (Figure 2f). Therefore, our electrophysiological and behavioral findings are not contradictory but instead support the conclusion that inhibiting SST+ vlPAG cells reduces pain sensitivity regardless of defensive state. We will revise the text to clarify this point.

      Reviewer #2 (Public review):

      Summary:

      Wenke et al. investigated the role of vlPAG somatostatin-expressing neurons in the mediation of analgesia during defensive states. A newly developed paradigm of cued fear-conditioned analgesia, which consists of a combination of an auditory fear retrieval session and a pain test, was used to evaluate this cell population's contribution to fear-mediated analgesia. Optogenetic manipulation of vlPAG SST+ neurons modulated the responses to a nociceptive cue (Hot Plate) presented concomitantly with an aversively conditioned tone. At the same time, alterations in the freezing levels could be observed during optogenetic activation of vlPAG SST+ neurons. In order to disentangle the impact of these cells on analgesia from their impact on the expression of defensive behaviors, the authors performed electrophysiological recordings from the dorsal horn in the spinal cord of anesthetized mice. A vlPAG-RVM-DH pathway was identified to trigger nociceptive C-fibers upon optic activation of the RVM. Finally, pathway-specific activation of SST+ vlPAG-RVM neurons could abolish CS-induced analgesia.

      Strengths:

      The study addresses a relevant topic, that is, brainstem circuits for pain-modulatory mechanisms as part of defensive states evoked by threat. This is important because the circuit mechanisms underlying pain are still not fully understood, and defining molecular markers of cellular circuit substrates may support the identification of potential pharmaceutical targets in treating pain. The authors confirm a previous study in that a somatostatin-positive cellular population presents a crucial vlPAG circuit element mediating anti-nociceptive effects. Key novelty aspects of the present study are the demonstration that these neurons seem to play a role specifically in threat-induced analgesia. This was possible by the elegant design and application of a novel fear analgesia paradigm, combined with cell- and pathway specific optogenetics.

      We thank the referee for such positive feedback.

      Weaknesses:

      Despite the convincing and rigorous experimental approach, the study leaves some interpretational room when it comes to the proposed circuit mechanism. This could either be addressed by additional experiments or by more discussion of alternative circuit layouts.

      Major Comments:

      (1) The paper by Zhang et al. (https://pubmed.ncbi.nlm.nih.gov/36641028/), which identified a role for vlPAG SOM+ neurons in mediating anti-nociception in neuropathic pain, needs to be referenced and its results discussed, if not reconciled. While functionally, both studies find an analgetic role of vlPAG SOM+ neurons projecting to the RVM, Zhang et al., using slice physiology, characterize those neurons as glutamatergic. In Figure 4E of Zhang et al. they find general (fear-independent) analgetic effects with PAG-RVM specificity by performing chemogenetic experiments.

      We thank the reviewer for highlighting this important point. We agree that the study by Zhang et al. is highly relevant and should be discussed in the revised manuscript. Their work shows that inhibiting vlPAG SST/SOM neurons with chemogenetic methods produces analgesia in a neuropathic pain model, and in our study, we similarly found that inhibiting SST+ vlPAG neurons increases hotplate response latency (Figure 2f), which aligns with an analgesic effect. Additionally, we observed that activating SST+ vlPAG neurons suppresses fear-conditioned analgesia.

      At the same time, there are important differences between the two studies that may explain the differences in interpretation. First, the behavioral paradigms are not identical. Zhang et al. used a hotplate protocol where animals were directly exposed to a nociceptive temperature, whereas in our study, we used a progressive temperature ramp and explicitly compared responses during a conditioned stimulus (CS+) and a non-conditioned control stimulus (CS−). These controls were important for us to distinguish fear-specific effects from more general effects related to stress, arousal, sensitization, or other non-associative processes.

      Second, the two studies differ in experimental context. Zhang et al. examined this circuit in a neuropathic pain model, whereas our study focused on acute nociceptive processing and fear-conditioned modulation of pain. We therefore believe that the apparent discrepancy might reflect differences in pain state and behavioral context, rather than a direct contradiction.

      Finally, Zhang et al. showed in slice recordings that SST+ vlPAG neurons provide excitatory input to RVM neurons. This is an important finding that we now address in the revised manuscript. At the same time, because the RVM contains heterogeneous neuronal populations with different projection targets and functions, these recordings alone do not prove that all recorded RVM neurons are part of the descending pathway controlling spinal nociception. Therefore, we have revised the Discussion to explicitly acknowledge Zhang et al. and to emphasize both the similarities and differences between the two studies.

      It can be argued that in addition to the two functionally distinct inhibitory SOM subtypes hypothesized by Winke et al., there is another, excitatory subpopulation. Also, the different experimental conditions (chronic vs. acute pain, non-threat vs. fearful cues/contexts may recruit different vlPAG SOM+ populations. All of this is conceivable, yet I wonder whether the contrasting findings could more parsimoniously be reconciled. The author's own results presented here in Supplementary Figure 3 suggests that SOM+ vlPAG cells are colocalizing with glutamate and thus could also be excitatory. In addition to this rather complementary piece of evidence, a more extensive characterization of vlPAG neurons using IHC and slice physiology would be needed to justify the unambiguous identification of their inhibitory nature.

      We thank the reviewer for this thoughtful comment. We agree that our current data do not support a definitive conclusion that all SST+ vlPAG neurons are inhibitory. As the reviewer notes, our Supplementary Figure 3 shows that SST+ vlPAG cells can also co-localize with glutamatergic markers, which is consistent with the possibility of cellular heterogeneity within this population. We also agree that different experimental conditions, such as chronic versus acute pain and non-threatening versus fear-related contexts, may activate different SST+ vlPAG subpopulations.

      Our intention was not to claim that SST+ vlPAG neurons constitute a uniform inhibitory population, but rather that SST+ cells are strongly represented among inhibitory neurons in the vlPAG. We agree, however, that more detailed characterization, including additional immunohistochemical analyses and slice physiology, is necessary to more definitively determine the neurotransmitter phenotype and functional connectivity of these neurons. We have therefore revised the text to temper our interpretation and to explicitly acknowledge the likely heterogeneity of SST+ vlPAG neurons, including the possibility of an excitatory subpopulation. We therefore modified the discussion accordingly:

      “Our results align with the parallel inhibition- excitation model, where inhibitory and excitatory cells form two distinct, parallel descending pathways for pain modulation.

      Indeed, previous research demonstrated the presence of an inhibitory pathway projecting throughout the PAG–RVM-spinal cord dorsal horn neuraxis. Our results complement this study by suggesting that one of these previously proposed parallel pathways is mediated by SST+ vlPAG cells and has a functional role in mediating analgesia. At the same time, our data indicate that vlPAG SST neurons are heterogeneous, with approximately one-third of these cells co-localizing with excitatory markers. Together with the recent observation that excitatory SST+ vlPAG neurons project to the RVM (Zhang et al., 2023), this raises the possibility that a subset of long-range SST+ vlPAG neurons contributes to an excitatory descending pathway within the PAG–RVM–spinal dorsal horn neuraxis. By contrast, local GABAergic SST+ vlPAG neurons may participate in local circuit mechanisms related to defensive-state expression, including freezing. Further anatomical and functional studies will be required to resolve these possibilities.”

      In the absence of a direct identification of these cells exclusively releasing GABA, an alternative explanation should be considered. What about looking at vlPAG SOM+ neurons as a putatively mixed bag of local, inhibitory interneurons and long-range, RVM-projecting excitatory cells? This model would then open up interesting questions as to the actual function of somatostatin as a modulator of vlPAG circuit activity and associated function, and from my perspective, would nicely fit into the view of PAG circuits as integrators of complex survival responses.

      We thank the reviewer for this insightful suggestion and agree that, in the absence of direct evidence that vlPAG SOM+/SST+ neurons are exclusively GABAergic, an alternative interpretation should be considered. In particular, we agree that this population may be heterogeneous and could include both local inhibitory interneurons and long-range excitatory neurons projecting to the RVM. We believe this is an important and constructive framework for interpreting our data, and we have revised the Discussion accordingly. In the revised text, we now explicitly acknowledge the likely heterogeneity of vlPAG SST+ neurons and discuss the possibility that distinct local and long-range SST+ subpopulations may contribute differently to defensive-state regulation and descending pain modulation. We agree with the reviewer on this point and have modified the discussion accordingly (see point above).

      (2) "Our data indicate that the optogenetic inhibition of SST+ vlPAG cells promotes analgesia irrespective of the animal's defensive state. In contrast, the optogenetic activation of long-range SST+ vlPAG cells that project to the rostral ventromedial medulla (RVM) abolishes the analgesia mediated by fear behavior." (lines 32-35). Consider toning down these conclusions, as contrasting activation with inhibition of two different (though overlapping) populations cannot be fully conclusive. Alternatively, a pathway-specific (vlPAG-RVM) inhibitory experiment could help to fully understand the circuit mechanism and verify the necessity of these neurons.

      We thank the reviewer for raising this point. We agree that inhibition of the entire SST+ vlPAG population and activation of the long-range SST+ vlPAG neurons projecting to the RVM population are not directly equivalent manipulations. Our conclusion was intended at the level of observed functional effects: inhibition of SST+ vlPAG neurons promotes analgesia regardless of the defensive state, while activating long-range SST+ vlPAG neurons projecting to the RVM suppresses fear-conditioned analgesia. This occurs regardless of whether the SST vlPAG neurons are excitatory or inhibitory. To address the excitatory or inhibitory nature of SST vlPAG neurons, we have revised the discussion to include a reference to the Zhang et al study.

      (3) Despite an overall very thorough reporting style, some information is missing from the manuscript:

      (a) In Figures 2d and f, what are the freezing levels during optogenetic manipulation? From Figure 3d, one can expect that freezing is inhibited during the hot plate test, which could bias the NC response towards shorter latencies.

      We thank the reviewer for this important comment. As shown in Figure 1e, we previously quantified freezing both at CS onset and at the time of the nociceptive response in the hot plate test. These analyses indicate that freezing levels at the time of the nociceptive response do not differ between the CS+ and CS− conditions. Therefore, the variation in hot plate response latency is unlikely to be due to differences in freezing at the time of response.

      We acknowledge, however, that freezing was not directly measured during optogenetic manipulation in this experiment. Based on the temporal profile of freezing shown in Figure 1e, we still consider it unlikely that the effect of optogenetic manipulation on nociceptive latency is mainly caused by a change in freezing behavior.

      (b) In Figure 5, the histological experiment showing the vlPAG-to-RVM pathway is presented by a qualitative image only. Here, some quantification would strengthen the finding.

      We thank the reviewer for this comment. The aim of the histological experiment in Figure 5 was to provide qualitative anatomical evidence that vlPAG projections reach the RVM and are positioned in close apposition to spinally projecting RVM neurons. We did not intend this experiment to serve as a quantitative characterization of connectivity. We agree that a more systematic quantification would be informative, but this would require additional dedicated experiments beyond the scope of the present manuscript.

      (c) In Figures 6 c and d "Consistently, activation of the SST+ vlPAG-RVM pathway during CFCA had no impact on CS-presentation, whereas the same manipulation performed during CS+ blocked the increase in NC response latency compared to GFP controls." (line 194-196). Is it possible that the NC response cannot be any lower than the one during CS-, thus constituting a floor effect?

      We are thankful to the reviewer for this important point. We agree with the reviewer that this is indeed a possibility. We have added a sentence in the discussion to acknowledge this limitation.“Another possibility is that our nociceptive test with a slow ramp of temperature induces a floor effect on nociceptive response latency, which may limit the detection of further decreases in latency under certain conditions.”

      (c) Connected to major point 1- this experiment is important for defining the circuit mode and therefore should be as convincing as possible. However, for the colocalization experiment in Supplementary Figure 3, the methodological description is missing and thus makes it hard to comprehend how this data set was generated (how many data points, etc.). The visual depiction of the results is non-standard and not easily graspable. Consider e.g., a Venn diagram.

      We apologize for this omission in the original manuscript. We have now provided this methodological information in the method section. We have now expanded the description of these data in the figure legend to ease the comprehension of the figure.

      Reviewer #3 (Public review):

      Summary:

      Conditioned analgesia refers to the ability of a learned fear cue to suppress pain-related behavior and neural activity. Understudied, the authors developed a novel conditioned analgesia procedure in which a cue that had been paired or unpaired with shock was played while a hot plate increased temperature. Compared to several control conditions, the authors found increased latency to a nociceptive response (paw licking). The authors identified somatostatin neurons in the periaqueductal gray as a likely mediator of the behavior. They then showed that: (1) stimulating vlPAG-SST neurons blocked nociceptive response latency increases to the CS+, (2) stimulating vlPAG-SST neurons suppressed fear retrieval freezing, (3) stimulating vs. inhibiting vlPAG-SST neurons drove opposing modulation of c-fibers and Aδfibers, (4) direct-projecting vlPAG SST neurons modulate freezing while RVM-projecting vlPAG SST neurons modulate conditioned analgesia.

      Strengths:

      These experiments have many strengths. The behavioral assay is chief among them. The assay is robust and controls for confounding factors to reveal a repeatable effect of a shock-paired cue to delay nociceptive responding. The optogenetic experiments provide the correct level of temporal precision, given the authors' time-specific interest in cued responding. Combining neuronal manipulations with spinal recordings is particularly innovative, especially in the context of more behavioral neuroscience-based assays. All-in-all, I found this to be an exceptionally strong set of experiments.

      Weaknesses:

      No obvious weaknesses were identified by this Reviewer.

      Recommendations for the authors:

      Comments from Reviewing Editor:

      Summary

      Three reviewers have assessed your manuscript on vlPAG somatostatin pathways contributing to conditioned analgesia. Conditioned analgesia refers to the ability of a learned fear cue to suppress pain-related behavior and neural activity. Understudied, the authors developed a novel conditioned analgesia procedure in which a cue that had been paired or unpaired with shock was played while a hot plate increased temperature. Compared to several control conditions, the authors found increased latency to a nociceptive response (paw licking). The authors identified somatostatin neurons in the periaqueductal gray as a likely mediator of the behavior. They then showed that: (1) stimulating vlPAG-SST neurons blocked nociceptive response latency increases to the CS+, (2) stimulating vlPAG-SST neurons suppressed fear retrieval freezing, (3) stimulating vs. inhibiting vlPAG-SST neurons drove opposing modulation of c-fibers and Aδ-fibers, (4) direct-projecting vlPAG SST neurons modulate freezing while RVM-projecting vlPAG SST neurons modulate conditioned analgesia.

      Strengths

      All three reviewers converged on multiple strengths. The assay developed was seen to be novel, rigorous, and included a variety of controls that convincingly demonstrated conditioned analgesia. Focusing on the ventrolateral periaqueductal gray, and more specifically on somatostatin-expressing cells, made prior sense, and the results more than justified this selection. Approaching the vlPAG and circuits with many converging methods provided further, compelling evidence for a role in conditioned analgesia.

      Weaknesses

      Specific weaknesses are described in the individual reviews. Generally, the following weaknesses were identified. The study only used male mice, a choice that should be better justified. Animals were reasonably excluded from analysis, but the final group ns for analyses were not always clear. Some statistical results lacked clarity. The relevance of these findings to prior work (particularly Zhang et al. 2023, Journal of Pain) was not always described. Relatedly, the results would be better contextualized by appreciating and describing the likely diversity of somatostatin functional types and projection types.

      Recommendations

      (1) Provide rationale for only using male mice, discuss the limitation of the exclusion of females, and note that male mice were the subjects in the abstract.

      Thank you for this recommendation, we have mentioned this information in the abstract and in the discussion. We have also mentioned the limitations of not including female mice in the abstract and the discussion of the revised manuscript.

      (2) Complete final report ns for each statistical analysis. If you have not already done so, please include full statistical reporting including exact p-values wherever possible alongside the summary statistics (test statistic and df) and, where appropriate, 95% confidence intervals. These should be reported for all key questions and not only when the p-value is less than 0.05 in the main manuscript.

      An extended table with all statistical tests and analysis for all figures has been provided in sup Table 1.

      (3) Include example videos of CFCA sessions, demonstrating optogenetic effects.

      We understand the editor’s request to include video material illustrating the behavioral responses. However, we would prefer not to include such videos in the manuscript, in accordance with our institution's guidelines and recommendations on the dissemination of animal experimentation footage. Importantly, all behavioral sessions were systematically video-recorded from both sides of the apparatus, allowing detailed offline analysis of the animals’ responses. These recordings were carefully examined by an experienced experimenter to assess nociceptive behaviors, including jumping responses and licking of the stimulated hindpaw. This procedure ensured a reliable and accurate evaluation of pain-related behavioral reactivity. While the videos themselves cannot be included in the manuscript for the reasons mentioned above, we believe that the behavioral scoring procedures described in the Methods section provide a clear and rigorous description of how these responses were assessed. In addition, Figure 1 includes an example image illustrating hindpaw licking behaviour, which is typically more subtle and more difficult to identify than jumping responses. We therefore believe that this visual example, together with the detailed description of the scoring procedure and the quantitative data provided, adequately supports the interpretation of the behavioural results.

      (4) Provide summary expression and ferrule placement figures.

      We thank the editor for this comment. We have now included schematic summaries of fiber placements for both SST and VIP mice used in this study, based on histological verification (Supplementary Figures 10 and 11). Representative images of viral expression are also provided (Figure 2a, Supplementary Figure 7b and f).

      (5) Detail how behavior judgments were made.

      We thank the editor for emphasizing this important methodological point. During all behavioral sessions, mice were video-recorded simultaneously from both sides of the apparatus, allowing a comprehensive and unobstructed view of the animals’ posture and movements throughout the experiment. These recordings were subsequently analyzed offline by an experienced experimenter trained to evaluate nociceptive behaviors. Pain-related behavioral responses were assessed based on well-established indicators of nociceptive reactivity. In particular, we quantified overt escape-like reactions such as jumping, which reflects a strong aversive response to the stimulus. In addition, we evaluated more localized nociceptive behaviors directed toward the stimulated limb, including licking of the hindpaw. These measures are commonly used in rodent pain assays and provide reliable behavioral readouts of nociceptive sensitivity. The combination of bilateral video recordings and expert behavioral scoring ensured that both subtle and robust nociceptive responses could be accurately detected and categorized during the analysis.

      (6) Provide the temperature at which nociceptive responses were initiated. Check grammar and references.

      The temperature at which nociceptive responses were initiated were originally reported in Supplementary Figure 1, 2 and 5.

      Reviewer #1 (Recommendations for the authors):

      (1) The authors use optogenetic manipulation of SST activity in the vlPAG to show that this cell type is involved in fear-induced analgesia. They include a valuable control to show that manipulation of another inhibitory cell type (VIP) also does not impact analgesia. It would be helpful to know the expression level of VIP cells in the vlPAG. Is this a predominant inhibitory projection cell in the vlPAG (besides SST)?

      We thank the reviewer for pointing this. While we did not quantify the expression level of VIP+ cells in the vlPAG in the present study, available data suggest that this population is relatively sparse compared to other inhibitory cell types. In particular, reference to the Allen brain atlas indicates that VIP gene expression in the vlPAG is limited and primarily localized around the fourth ventricle, within the lateral and ventrolateral PAG, rather than broadly distributed across the region. Consistent with this, we provide an example of viral expression in VIP-Cre mice in Supplementary Figure 7f, illustrating the restricted distribution of VIP+ neurons in the vlPAG. We have also provided a summary of ferrules placement for SST and VIP mice used in our study in Supplementary Figures 11 and 10, respectively.

      (2) The numbers of animals dropped from each experiment should be indicated - perhaps on the statistics table?

      We thank the reviewer for pointing this.

      As stated in the Methods, we applied strict inclusion criteria for mice undergoing the hot-plate test, specifically a discrimination index ≥ 0.4 and a conditioning index ≥ 0.3. Using these criteria, 23% of wild-type mice were excluded for failing to meet the discrimination criterion. In the transgenic groups, an average of 20% of mice failed to meet the learning criteria, and an additional 12% were excluded due to incorrect opsin injection or misplaced optic fiber placement.

      (3) Line 105: "...,which activity..." change to "..., whose activity..."

      Done

      Reviewer #2 (Recommendations for the authors):

      (1) Please also provide absolute temperature values of the nociceptive response threshold.

      The temperature at which nociceptive responses were initiated was originally reported in Supplementary Figure 1, 2 and 5.

      (2) It would be nice to see an example video of a CFCA session (with and without optogenetic manipulation).

      We understand the editor’s and reviewer’s request to include video material illustrating the behavioral responses. However, we would prefer not to include such videos in the manuscript, in accordance with our institution's guidelines and recommendations on the dissemination of animal experimentation footage. Importantly, all behavioral sessions were systematically video-recorded from both sides of the apparatus, allowing detailed offline analysis of the animals’ responses. These recordings were carefully examined by an experienced experimenter to assess nociceptive behaviors, including jumping responses and licking of the stimulated hindpaw. This procedure ensured a reliable and accurate evaluation of pain-related behavioral reactivity. While the videos themselves cannot be included in the manuscript for the reasons mentioned above, we believe that the behavioral scoring procedures described in the Methods section provide a clear and rigorous description of how these responses were assessed. In addition, Figure 1 includes an example image illustrating hindpaw licking behaviour, which is typically more subtle and more difficult to identify than jumping responses. We therefore believe that this visual example, together with the detailed description of the scoring procedure and the quantitative data provided, adequately supports the interpretation of the behavioural results.

      (3) Please provide a schematic summary of fiber placements and opsin expressions confirmed by histological examinations.

      We thank the reviewer for this comment. We have now included schematic summaries of fiber placements for both SST and VIP mice used in this study, based on histological verification (Supplementary Figures 10 and 11). Representative images of viral expression are also provided (Figure 2a, Supplementary Figure 7b and f).

      (4) "Valid nociception readout responses included jumping or licking the hindpaw." (Line 453). How was this evaluated- manually or automated, blinded etc.?

      We thank the reviewer for emphasizing this important methodological point. During all behavioral sessions, mice were video-recorded simultaneously from both sides of the apparatus, allowing a comprehensive and unobstructed view of the animals’ posture and movements throughout the experiment. These recordings were subsequently analyzed offline by an experienced experimenter trained to evaluate nociceptive behaviors. Pain-related behavioral responses were assessed based on well-established indicators of nociceptive reactivity. In particular, we quantified overt escape-like reactions such as jumping, which reflects a strong aversive response to the stimulus. In addition, we evaluated more localized nocifensive behaviors directed toward the stimulated limb, including licking of the hindpaw. These measures are commonly used in rodent pain assays and provide reliable behavioral readouts of nociceptive sensitivity.The combination of bilateral video recordings and expert behavioral scoring ensured that both subtle and robust nociceptive responses could be accurately detected and categorized during the analysis.

      (5) Line 226 REF33 doesn't seem to fit.

      The reference list has been updated. Related to this section in which we discuss the disinhibition mechanisms inducing nociception in chronic stress mice. We have cited the work of Samineni et al., 2015 (reference 15) and Tovote el al., (reference 23) both related to these disinhibition mechanisms.

      Full sentence for reference 33 (now 35): “Two independent previous studies found that long-range inhibitory inputs from the central medial amygdala contact inhibitory cells within the vlPAG, implicated in different roles: the modulation of fear behavior (23) and nociceptive transmission (35)”.

      Ref 35 - Yin, W. et al. A Central Amygdala–Ventrolateral Periaqueductal Gray Matter Pathway for Pain in a Mouse Model of Depression-like Behavior. Anesthesiology 132,1175–119 (2020)

      (6) Some minor language, semantic, and grammatical flaws.

      The manuscript has been evaluated for language, semantic and grammatical flaws

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      (1) There are certainly some areas of the manuscript that would benefit from deeper exploration, such as electron microscopy/other imaging approaches to explore whether deletion of PfMSP2 has a visible impact on merozoite surface structure.

      We in principle agree with the reviewer that applying enhanced resolution microscopy approaches to understand structural and functional changes with loss of PfMSP2 could be of interest. However, based on our ongoing work, this represents a significant body of work in terms of experimental optimisation in an effort to gain the detail required to make meaningful insights. Therefore, this will remain outside the scope of this manuscript and we hope to provide these insights in future studies.

      (2) Further replicates of the video microscopy assays to see whether trends in the data could reach significance (although these are very time-consuming and technically difficult assays).

      Conclusions we have drawn from live-cell imaging data for MSP2 knock-out parasites encompass some 43 invading merozoites from 21 schizont ruptures for PfDd2 WT and 35 invading merozoites from 18 schizont ruptures for PfDd2 DMSP2 parasites. One of the leading studies to apply live-cell microscopy to film invading merozoites based conclusions of invasion kinetics on: 3D7 (number of merozoite invasion =63, number of schizont ruptures =23), D10 (invasions =33, ruptures =20) and W2mef (invasions =39, ruptures = 15; this line is of the same lineage as Dd2) (Weiss et al. PLoS Pathogens, 2015). Although there are variations within and between lines from this gold-standard study, our dataset is mostly comparable in terms of the number of schizont ruptures and merozoite invasions filmed and analysed to look at changes in kinetics. What we can say definitively is that there is no strong phenotype in the absence of inhibitory antibodies against other antigens for either live-cell or growth inhibition assays. Therefore, we have focussed the data interpretation in the manuscript to highlight the lack of statistical significance and limited phenotype seen, which given the previously believed importance of MSP2 to P. falciparum invasion of red blood cells is somewhat surprising.

      In order to address this suggestion, we have modified the discussion to better represent any non-significant changes in invasion and growth seen.

      “Despite the abundance of PfMSP2 on the merozoite surface and previous work suggesting a role in RBC invasion, we found merozoites invade and grow with similar kinetics to wildtype parasites in the absence of PfMSP2. This does not exclude a role for PfMSP2 in vivo where there are additional pressures, such as immune-effector mechanisms and flow dynamics, on merozoite invasion. However, given we have knocked-out PfMSP2 from two different P. falciparum isolates, our findings do not currently support a major role for PfMSP2 in the mechanics of merozoite invasion. Thus, it appears that the function of the two most abundant proteins on the merozoite surface, PfMSP1 (Das et al., 2015; Kals et al., 2024) and PfMSP2, are not obviously linked to merozoite binding to the RBC and subsequent invasion.”

      (3) Follow up of some of the genes where expression is changed by PfMSP2 knockout (as the authors point out, there are no candidates that have a very obvious link to invasion suggesting that they may be compensating for PfMSP2 function, although several are expressed in schizont stages).

      A thorough investigation of the genes where expression changes with PfMSP2 knock-out would require a substantial body of additional work, not least because they would all have to be investigated as there is no single likely candidate based on stage of expression, membrane binding properties or previous links to merozoite surface architecture. Given this, potential follow up of these proteins will be left for future studies.

      We also thank the reviewer for the recognition of the work provided in the manuscript and the modifications made that have improved the manuscript from version 1. The reviewer also recognises the value in our detailed characterisation, including data where phenotyping changes with MSP2 knock-out could not be seen, in defining the function of PfMSP2 as commented below:

      However, there is already a substantial amount of data in the manuscript, and more detailed follow-up is reasonable to leave to future work. Overall, with the modifications made through the review process, including the addition of new controls for key experiments, the claims and conclusions are justified by the data, and the manuscript generates important new information about a highly studied Plasmodium falciparum merozoite surface protein.

      Reviewer #3 (Public review):

      Major points:

      (1) Much of the manuscript describes negative results and this reviewer found it arduous to get through many negative or nonsignificant results before finally getting to the significant effect on AMA1 inhibitory antibodies, not presented until Figure 6! Computational studies in Fig. 1 could be a supplementary figure. Figs. 2 and 3. demonstrate knockout in 3D7 and Dd2, respectively and could be assembled into a single figure. (Notably Fig. 2A and 3A are almost identical with use of some different primers.) Fig. 2E, 2F, 3D-H, all of Fig. 4, most of Fig. 5 are all negative or insignificant results that could also be moved to supplementary data. As MSP4, MSP5, and SUB1 are presumably included in the whole genome RNA-seq experiments shown in Fig. 4C, it makes sense to remove Fig. 4A data from the paper fully. These consolidating changes would help highlight the key finding of improved binding and block of AMA1's role in invasion.

      We have chosen to not take the approach proposed by Reviewer 3 as it would leave the manuscript with only around 2.5 Figure panels and undersells the very significant amount of work that has been done to characterise PfMSP2 knock-out lines. Although, as noted by the reviewer, piggyBac mutagenesis studies predict PfMSP2 is dispensable, much of the field likely expect PfMSP2 to be essential to P. falciparum blood stage parasite growth due to the results of earlier reverse genetics approaches and many years of publications that have speculated on the importance of the protein. Therefore, we are also conscious of providing very clear and comprehensive evidence to support our findings. While this may delay highlighting the findings in Figure 6, we also note that the lengths we have gone to in characterising an important antigen with a difficult phenotype is still valued as evidenced by Reviewer 2 (Public Review Comments on the original manuscript):

      “PfMSP2 knockouts are made in two different strains, which is important as it is known that invasion pathways can vary between strains, but is a level of comprehensiveness that is not always delivered in P. falciparum genetic studies. The knockout strains are characterised very thoroughly using multiple different assays, and the authors should be commended for publishing a good deal of negative data, where no phenotype was detected.”

      (2) The potentiating effects on anti-AMA1 antibodies are shown with rabbit sera and purified antibodies, mouse monoclonal antibodies, and smaller i-bodies inspired by shark antibody-like receptors but not with human monoclonal antibodies (hmAbs). As naturally acquired hmAbs targeting AMA1 have been identified and characterized (PMIDs: 39632799, 40020675), would it not be important to test these antibodies in the ∆MSP2, especially as the authors emphasize the importance of their model in designing better human malaria vaccines?

      As the reviewer noted, we demonstrated enhanced inhibitory activities of antibodies to AMA1 using rabbit polyclonal antibodies, mouse mAbs, and i-bodies. We note that the WD34 i-Body we used was humanised to be IgG-like with a human Fc-region (IgG1 backbone). Rabbit IgG is very similar to human IgG1. Therefore, we have provided evidence of the enhancing effect using different types and sources of antibodies relevant to human immunity to support our conclusions. Our findings open new avenues for future research and we agree with the reviewer that future studies using panels of human mAbs to defined epitopes would be interesting and may further inform vaccine design; however this is beyond the scope of the current paper. We do not have the mAb mentioned by the reviewer to test in our system. To perform studies with human mAbs would take a substantial amount of time (many months), requiring the generation of different human mAbs and quantification of their activity and testing them for potentiation effects. While this would be an interesting future endeavour, we do not feel that such studies are needed at this stage to support our conclusions, and instead would be a future extension from our current paper. To acknowledge the reviewer's comment, we have extended our comment in the discussion about future studies with different panels of invasion inhibitory antibodies to include huMabs targeting AMA1 as follows:

      “Further investigation using the parasite lines developed in this study and a wider panel of antibodies that target different stages of the merozoite invasion process, including human monoclonal antibodies against AMA1 (Patel et al., 2025), could shed more light on this potentially novel mechanism of vaccine derived antibody efficacy.”

      (3) Fig. 7 presents quantitative fluorescence microscopy to measure anti-AMA1 binding and support a model where MSP2 serves to sterically hinder antibody access to AMA1 on individual merozoites. I understand that the negative WD33 control is useful to contrast to the positive WD34 antibody (both bind AMA1 but only WD34 exhibits parasite growth inhibitory effects), but it seems that use of smaller i-bodies rather than conventional larger mouse or ideally human monoclonal antibodies may compromise demonstration of steric hindrance by MSP2 because smaller i-bodies may be less hinder.

      The antibodies used in this experiment have fluorescent tags attached. So while the untagged WD33 and WD34 i-bodies are approximately 14 kDa, when fused to GFP or mCherry their expected size increases to approximately 42 kDa, approaching that of the Fc-tagged WD34 i-body (78 kDa) that shows increased growth inhibitory activity in the absence of MSP2. Therefore, we expect steric hindrance to be a significant factor with these fluorescently tagged antibodies.

      (4) Some explanation for why WD33 fails to inhibit growth despite targeting the same antigen as WD34 is needed. Are the epitopes known? Does one bind further from the RON2 binding pocket?

      As reported in Angage et al., Nature Communications 15, 7206 (2024). WD34 has been identified to bind to, and block, a site within the hydrophobic AMA1 and RON2 binding pocket found on Domain II of AMA1. In contrast, WD33 recognises a distinct conserved epitope in Domain II of AMA1 near to, but not overlapping with, the hydrophobic AMA1 and RON2 binding pocket. We have clarified this by including additional description when first describing the i-bodies as follows:

      “When we tested the i-body WD34 (Angage et al., 2024) which binds a highly conserved epitope that includes the PfRON2-binding pocket on PfAMA1 domain II, we observed a small potentiation of PfAMA1 specific activity with knock-out of PfMSP2 in Pf3D7 (1.3-fold; IC<sub>50</sub> PfD7 WT 0.012 mg/mL; IC<sub>50</sub> Pf3D7 DMSP2 0.009 mg/mL; p=0.08 Figure 6F).”

      Then

      “A second i-body, WD33 (Angage et al., 2024), which binds AMA1 between domain II and domain III but does not appear to overlap with the PfRON2-binding pocket on PfAMA1, had very limited invasion inhibitory activity against Pf3D7 parasites and did not show improved potency with knock-out of Pf3D7 MSP2 (0.9-fold; IC<sub>50</sub> Pf3D7 WT 1.02 mg/mL; IC<sub>50</sub> Pf3D7 DMSP2 1.1 mg/mL; p=0.8; Figure 6I).”

      Recommendations for the authors:

      Reviewing Editor Recommendations:

      Although providing microscopic images might require a lengthy process, including results based on human mAbs (if available) might enhance the strength of evidence. The reorganization of the figures and the presentation of results usually falls into the realm of personal preferences, however, if the comments/suggestions are useful, it might highlight your message.

      As covered in the Response to Public Reviewer Comments for Reviewer 2 and indicated by the editor, investigations of phenotypes found in this study using high-resolution imaging techniques (e.g. electron microscopy) will require very significant additional work and will be attempted in future studies. We also provide a response to Reviewer 3 in regards to the potential to test human monoclonal antibodies and believe this is best done more thoroughly in future studies. We have elected to not make substantial changes to the data presented as suggested by Reviewer 3. We have addressed additional comments as covered below.

      Reviewer #3 (Recommendations for the authors):

      Minor Comments

      (1) Scale bar in Fig. 7A is not resolved well. The image is too pixelated to resolve merozoites or the actual dimensions of the scale bar.

      We have updated this figure to provide improved clarity of the scale bar.

      (2) Lines 69, 216, 221, 253, 628-629, 648 all suggest that MSP2 was heretofore assumed to be essential. However, piggyBac insertional mutagenesis revealed that MSP2 is highly dispensable (MIS of 0.988, per PlasmoDb.org; PMID: 29724925). I would suggest to tone down this claim as it does not detract from the authors' production of useful ∆MSP2 clones.

      We agree with the reviewer that the piggyBac insertional mutagenesis study results should also be acknowledged and apologise for this oversight. To address this, we have reviewed the sentences highlighted by the reviewer and, where appropriate for the historical interpretation of PfMSP2 function, have added the following modified information through the text:

      P. falciparum merozoite surface protein 2 (PfMSP2), an antigen reported to be refractory to gene knock-out in P. falciparum (Sanders et al., 2006) but that has also been reported to be dispensable in a piggyBac mutagenesis study (Zhang et al., 2018), has been of long-term interest as a vaccine candidate.”

      “Given previous unsuccessful attempts to disrupt pfmsp2 (Sanders et al., 2006), and its high abundance on the merozoite surface (Gilson et al., 2006), PfMSP2 has been traditionally viewed as an essential P. falciparum protein with an essential function in merozoite invasion, although more recent piggyBac mutagenesis studies have called this understanding into question (Zhang et al., 2018).”

      We have chosen not to modify this text and it remains the same as below. The reason for not changing this text is the result that we could knock-out MSP2 from 3D7 was still unexpected given the published reverse genetics studies and results from piggyBac mutagenesis studies are also sometimes not reliable indicators of what happens when reverse genetics is performed. Therefore, the following text we believe is a reasonable description.

      “Unexpectedly, we confirmed successful disruption of pfmsp2 by replacing the coding sequence between 132 bp and 819 bp of the gene with a hDHFR drug selection cassette in the 3D7 P. falciparum laboratory-adapted line (Figure 2A and B), resulting in Pf3D7 DMSP2 parasites.”

      “As a previous reverse genetics study in 3D7 reported that PfMSP2 was essential for P. falciparum growth in vitro (Sanders et al., 2006), we investigated whether PfMSP2 could also be removed from PfDd2, an isolate of P. falciparum that differs from 3D7 in geographical origin, RBC receptor usage and allelic type of pfmsp2.”

      “However, CRISPR-Cas9 gene editing used in this work has shown that, in contrast to previous attempts to knock-out PfMSP2 (Sanders et al., 2006), PfMSP2 is not essential for P. falciparum blood stage parasite growth in vitro.”

      “Advancements in gene-editing techniques in P. falciparum have allowed us to directly demonstrate using reverse genetics in two different parasite lines that PfMSP2 is not essential for P. falciparum growth in vitro.”

      (3) Figs. 2B, 2C, 2D show PCR, immunoblots, and IFA with a ∆MSP2 clone but two clones (termed clone 1 and clone 2) are show in panels 2E and 2F. Which clone is used in each panel? Without clarification, readers may wonder if one clone was used for PCR but another clone gave a desired result in immunoblots? By convention, validation studies (PCR and immunoblots) should be performed and shown (in Supplementary figures) for all clones used for phenotype studies; alternatively, a single clone can be used throughout if all clones are presumed identical. Which of these clones was used for the RNA-seq experiments in Fig. 4C? Similar questions arise for the two knockout clones made in the Dd2 line (Fig. 3D).

      We agree with the reviewer that it would be helpful to have this information provided more clearly through the Results. To this end, we have updated the Figure legends across Figures 2, 3, 4, 5, 6, 7 and Supplementary Figure 5 as appropriate to specifically indicate the clones used for the downstream experiments. All clones were validated by PCR and, after growth characteristics were found to be the same, a single clone was used for all downstream experiments for PfMSP2 knock-outs in both 3D7 and Dd2.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      This study aims to understand how cell fusion contributes to wound healing using a laser-induced injury in the notum epithelium of a developing fruit fly. The authors meticulously characterize the epithelial fusion events using a live imaging approach and report that syncytia arise by 'border breakdown' and 'cell shrinking'. The syncytial epithelial cells also appear to outcompete mononucleated cells and preferentially dissolve their tangential borders, which correlates with the accumulation of actin at the leading edge.

      Strengths:

      The strength of this study is the authors' live imaging approach to capture these dynamic fusion events that are a fundamental, yet poorly understood biological process.

      Weaknesses:

      A major weakness is that all the authors' conclusions are based on descriptive studies, in which the role of cell fusion is not directly tested. This is particularly important because other models of wound induced polyploidization have demonstrated that another cytoskeletal protein, myosin, was upregulated and dependent on endoreplication, and not cell fusion. Therefore it remains unclear to what extent cell fusion, endoreplication, or both are required to outcompete mononucleated cells as well as pool actin as described in this study.

      We thank the reviewer for appreciating our live imaging and meticulous approach. In this revision we have identified that the gene Atg1 is required for wound-induced fusion in the pupal notum: when Atg1 is knocked down, there is a reduction in wound-induced cell fusions, both border breakdown and cell shrinking. Analysis of Atg1 knockdown shows that the wounds close more slowly. This is a direct test of the role of cell fusion in speeding wound closure, presented in new Fig. 4.

      Reviewer #2 (Public Review):

      Summary:

      Overall, this study provides a thorough description of the formation of syncytia following wounding of the proliferation-competent diploid epithelium of the pupal notum. While this phenomenon has already been described briefly for this particular tissue by the Galko lab in Wang et al 2015, the authors provide a much more detailed description and characterisation of the process providing some novel insights (radial versus tangential border breakdown, cell shrinkage, timings, syncytia outcompeting mononucleated cells, etc.).

      Strengths:

      This paper provides an elegant, thorough, descriptive characterisation of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors show that laserinduced wounding of this diploid, proliferation-competent epithelium results in the formation of syncytia of various sizes in the first few cell rows around the wound edge, which progressively become bigger as healing proceeds. This results in ~50% of cells becoming part of these syncytia. The cell fusion events were convincingly demonstrated by showing the disappearance of p120ctnRFP and E-Cadherin-GFP from cell-cell borders as well as cytoplasmic GFP mixing of GFPpositive cells with a GFP-negative cell.

      Apart from cell-cell fusion by border breakdown that mostly happens in the first 2h following wounding, the authors also found that at later stages of wound healing cell shrinkage following cytoplasmic mixing contributed to sycytia formation.

      Next, the authors provided some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound.

      The authors then show that radial border breakdown occurs much less frequently than tangential border breakdown. They suggest that radial border breakdown reduces the requirement for cell-cell intercalations. They also hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. To test this, the authors generate single-cell clones that overexpress Actin-GFP. They then show convincingly how a single Actin-GFP-positive cell in the second cell row fuses with one GFP-negative cell in the first cell row. The Actin-GFP signal then spreads in the fused cell and labels some previously unlabelled actin-rich structure near the wound edge which most likely is the actomyosin cable. This provides some evidence for resource sharing by cytoplasmic mixing following fusion.

      Weaknesses:

      The authors provide some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound. The authors suggest that the syncytial cells might be better able to close the wound. However, some genetic studies would need to be done to establish this more convincingly. E.g. Could the authors genetically block syncytia formation and then show that these wounds now heal slower?

      We now present such data in new Fig. 4, which describes knocking down Atg1, previously shown by the Leptin lab to promote wound-induced fusions in larval epidermis. We quantify the resulting reduction in fusion in the pupal notum and show that the leading edge advances more slowly to heal the wound.

      The authors suggest that radial border breakdown reduces the requirement for cell intercalation. While this might be true it also raises the question of how the various syncytia facing the wound border change shape to allow the shrinkage of the first cell row over time to allow wound closure. None of the four movies included in the study shows the whole wound healing process until the later stages, making it hard to assess this. It would be good to include one such movie showing the syncytia in the whole wound and comment on this point.

      In response to the reviewer's request, we now extend Supplemental Video S1 out through 8 hours after wounding (same video as included previously but extended longer). In this video, as in many of the wounds, it is hard to determine the exact moment of closure because a syncytium extends across the wound whereas the nuclei do not. However, during the process of closure, one can clearly observe the large syncytia becoming more wedge-shaped – drastically reducing the section of their perimeter remaining in contact with the wound’s leading edge.

      In addition, we now explore how syncytia reduce the need for intercalation in a computational model, presented in new Fig. 7 and Supplemental Videos S5 and S6. One can observe the modeled syncytia becoming similarly wedge-shaped. The modeling shows that the presence of syncytia and their ability to reshape can speed closure by about 1/3 even if the syncytia have no special properties aside from their relative size.

      In both the experiments and models, some syncytia are also removed from the leading edge by intercalation, but the presence of syncytia reduces the total number of intercalations needed.

      The authors hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. They show convincingly through the fusion of a single Actin-GFP-positive cell in the second cell row with a GFP-negative cell in the first cell row that Actin-GFP spreads in the fused cell and labels the previously unlabelled actomyosin cable. While the hypothesis of resource sharing to improve healing is intriguing and makes sense, this experiment doesn't necessarily prove the benefit of resource sharing. It does show cytoplasmic mixing following fusion, now allowing the GFPlabelled actin to diffuse and be incorporated into the actomyosin cable. In a wild-type condition, fusion would not increase the total concentration of resources, although it would increase the total amount of resources within this bigger fused cell. The question is whether resource sharing without increasing the protein concentration is beneficial and increases the efficiency of certain wound healing mechanisms. There might be a benefit of cell fusion, if for example certain resources were only present in limited amounts or if protein transport could increase the concentration locally. To provide better evidence for the hypothesis that resource sharing improves wound healing, maybe the authors could look at the actomyosin cable in a wounded epithelium (such as in Figure 4E, F), in which all cells express MyoII-GFP. The authors could compare the average intensity of the actomyosin cable at the wound edge in mononucleated cells versus in syncytia. If resource sharing is indeed beneficial, it might be that the actomyosin cable is stronger/brighter in syncytia or it forms quicker.

      We agree with the reviewer that we have not "proved the benefit of resource sharing". Because we cannot inhibit resource sharing while still allowing cell fusion, we can think of no rigorous way to test this hypothesis. We appreciate the reviewer's suggestion of quantifying the myosin at the leading edge cable, but we can imagine too many caveats to the interpretation to make it worthwhile. Rather, we accept the limitation that this is an untested, perhaps untestable, hypothesis -- but nevertheless intriguing.

      We do want to clarify ideas about the concentration of resources after fusion. We agree that the overall concentration of a given resource (mass/volume) throughout a syncytium would be the same as the overall concentration in the unfused progenitor cells; however, a syncytium would have a larger total resource mass to direct subcellularly, allowing for local subcellular concentration to be greater in a syncytium vs. an unfused cell. We demonstrate this subcellular localization of actin in a syncytium twice, in Fig. 7C and E (previously Fig. 6C,E), which we think is evidence for increased local concentration.

      The biggest limitation of this study is that the authors don't address how the formation of these syncytia is regulated. While the manuscript in its current form provides some valuable new insights into syncytial-driven wound closure, it would be much more informative if it also provided some mechanistic details. The authors could test if some of the mechanisms shown to regulate syncytial formation in other types of syncytia-driven wound healing are also involved here. E.g. Yorkie was shown to negatively regulate cell fusion in adult syncytial-driven wound closure (Losick et al 2013). The authors could test for the effect of Yorkie-RNAi in the epithelium on wound closure and syncytia formation. Expression of the dominant negative RacN17 also blocked cell fusion in adult syncytial-driven wound closure (Losick et al 2013).

      Moreover, JNK activation was shown to be needed in larval syncytial-driven wound closure (Galko and Krasnow 2004). The authors could test JNK pathway reporters to assess pathway activation or test if the JNK pathway is needed for syncytial-driven wound closure by expressing a dominantnegative form of Basket JNK in the epithelium.

      Or could syncytia formation be regulated by changes in Integrin-mediated adhesion as shown by the Galko lab in Wang et al 2015? They show that wounding provoked a striking relocalization of PINCH and ILK, indicating the disassembly of functional FA complexes concomitant with syncytium formation. Maybe the authors could investigate some of these.

      We investigated the role of JNK in fusion by expressing bsk<sup>DN</sup> on one side of the wound. Comparing the numbers of border-loss fusion on each side, we did not find a significant difference in our seven-sample cohort (see Author response image 1). If we had increased the sample size, we may have found a significant difference with a small effect size, but because of the small difference in fusions on each side we did not think this was worth pursuing. Instead, we include data that the autophagy gene Atg1 is required for cell fusion in new Fig. 4, which begins to address mechanism, and relates the wound-induced fusion described here in pupae to wound-induced fusion shown in larvae. A complete mechanism for wound-induced fusion is outside the scope of this paper, as we focus on the function of syncytia in healing wounds.

      Author response image 1.

      Another general question that the authors raise but don't address enough is whether syncytia-driven wound closure in proliferation-competent epithelia is any different from the one in post-mitotic, polyploid epithelia. Since the mechanism regulating the former is not known, this remains unclear.

      We now include a paragraph on this question in the discussion.

      Finally, it is not clear, whether syncytia in these proliferation-competent epithelia get resolved after wound healing. Do they get removed and replaced by mononucleated proliferation-competent cells or do the syncytia stay in the epithelium like a scar? The authors should provide some images of wound areas a few hours after wound closure is complete and comment on this.

      To answer the reviewer’s question: some but not all syncytia do get removed during wound closure by remarkable apoptotic/extrusion events. This will be the subject of a future manuscript, as it is outside the scope of this paper focusing on the function of syncytia in promoting wound healing.

      Minor points:

      Figure 3: It would be better to have the microcopy images alongside the quantifications.

      The images in Figs. 1 and 2 show the border breakdown and shrinking cells, and we do not see benefit in adding them in Fig. 3.

      Figure 4A: The syncytium at the wound edge here doesn't look straight but wavy. Does it not form an actomyosin cable that straightens the front? Or are there lamellipodia/filopodia?

      We assume the reviewer is asking about the wavy edge outlined at 400 min after wounding (now Fig. 5A). As shown by Jacinto and colleagues in the first pupal wounding paper (JCB 2013), the actin cable forms quickly, within 15 minutes; much later actin protrusions extend from the leading edge to close the wound. This result is consistent with the wavy edge 400 min after wounding.

      248: The authors suggest an interesting hypothesis that mitochondria or ER could be pooled in fused cells. It would be nice to see some evidence: e.g. by labeling mitochondria and assessing where they are in syncytia versus mononucleated cells and whether they are concentrated around the wound edge.

      Although we don't think that exploring mitochondria or ER is central to this manuscript, we agree it would be an interesting question for the future.

      141-145 (Figure 4B and C) This example is not completely convincing. First, it is hard to see where the wound edge is. Second, it would be good to include an even later time point when the cell is clearly no longer at the wound edge.

      We have revised this figure, now Fig. 5B,C, to include a later image at 360 min after wounding healing, and this additional panel clarifies that the smaller cell leaves the wound edge. As noted in the text, the wound edge is indicated by the cell borders lacking p120ctn.

      Reviewer #3 (Public Review):

      Summary:

      White et al. described laser-induced wound healing of the Drosophila pupal notum. They found that the epithelial monolayer is dynamically induced to form syncytia by cell-cell fusion as an important part of repair. They reveal two processes: cell shrinking and border breakage that occur as part of syncytia formation. Expression of GFP in the cytoplasms of some epithelial cells reveals that cytoplasmic contents mix following injury and the GFP rapidly diffuses between cells. Using live imaging they observe that syncytia expand towards the wound, maintain their positions close to the leading edge, and apparently displace smaller cells. They propose that syncytia redistribute cellular components towards the wound facilitating repair and show that labelled actin becomes concentrated at the leading edge.

      Strengths:

      The manuscript is interesting and on an important and emerging topic of wound healing in a genetically tractable organism. The manuscript is very well written.

      Weaknesses:

      There are three major issues that the authors must address: 1. Is cell-cell fusion sufficient to enhance/facilitate wound healing? 2. Characterization of "border breakdown"; Is this phenomenon disassembly of apical junctions following membrane fusion? 3. Are cells really shrinking or is it only the apical domains that "shrink" as the cells join the syncytium.

      We thank the reviewer for recognizing the importance of this topic. Our responses to the specific weaknesses are below.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Major Components:

      (1) For syncytia measurements the nuclei are labeled with histone-GFP which is expressed in all cell types. How do you know the nuclei within the cell junctions are epithelial and not another cell type, such as immune cells recruited to the injury site? It would be helpful to verify the number of nuclei per cell using an epithelial-specific nuclear marker as well. This could be via epithelial Gal4-specific expression of a UAS-nls-GFP.

      This is an interesting point. In response to the reviewer's question, we investigated by doing the converse experiment, labeling immune cells with hml-Gal4, UAS-GFP, and observing what they do after wounding (analyzing six wounded pupae). They do get recruited to the wound, but they remain either in the wound center or at the basal side of the leading edge. Because they are labeled with cytoplasmic GFP, we would be able to ascertain whether they fused with epithelial cells because they would share their GFP with epithelial cells in the epithelial plane, and they did not. Thus we are confident that the many syncytial nuclei are not derived from immune cells. Our live tracking throughout the manuscript, and specifically of GFP-labeled clones, also supports our interpretation that syncytial nuclei derive from epithelial cells.

      (2) The manuscript focuses on cell fusion, but other mechanisms of cell enlargement have been observed to occur during wound healing via endoreplication. To what extent do epithelial cells in pupae notum endocycle or endomitosis post injury? It is unclear if the increase in syncytia size during a 1-2hr period could also be due to endomitosis, which would also increase nuclear number.

      Since the first submission of this manuscript, we published our results demonstrating limited wound-induced endoreplication after this type of explosive laser injury to the pupal notum (White et al, 2024, PMID: 38495588). We chose to publish this work separately because we could not offer the same degree of depth for endoreplication as we could for fusion: our pupal notum injury model is extremely well-suited to analyzing cell fusion and wound closure by live imaging; however, it is not particularly well-suited for analyzing endoreplication in fixed tissue. With respect to reviewer's question about endomitosis -- i.e. nuclear divisions that are not accompanied by cell divisions -- even after many years we have not observed an endomitosis event, which would be visible by live imaging, whereas we frequently and easily observe mitosis of diploid cells.

      (3) One of the major conclusions of this study is that cell fusion is necessary to pool resources at the leading edge. Therefore it is critical that authors identify a mechanism to inhibit cell fusion to test this assumption.

      We now include new Fig. 4, an analysis of the role of Atg1 in promoting wound-induced fusion and wound closure. These results build on the finding of the Leptin lab (Kakanj et al, 2022) that autophagy genes are required for fusion. Our results are consistent with the model that syncytia speed wound closure.

      (4) There is evidence that myosin increases in endoreplicating cells during wound healing hence it is, maybe equally - if not more - probable that the increase in resources (here actin-GFP) at the leading edge is dependent on endoreplication instead of cell fusion.

      Some of the new data we provide for this manuscript is a correlation between cell size and distance traveled, showing that larger cells travel more within the wound (Fig. 4F,G). Endoreplication would certainly be expected to contribute to increasing cell size, and our published 2024 data indicates that there can be one extra S-phase induced by these types of wounds. Doubling the genome is not a significant contribution to cell size compared to the 10s of nuclei we observe in syncytia from fusion. Nevertheless, we do not claim that actin is the only important resource that can be pooled subcelluarly for the benefit of the cell; we use it only as a proof-of-principle. Finally, we discuss the work on myosin in wound-induced endoreplicating cells (Losick and Duhaime, 2021).

      Reviewer #3 (Recommendations For The Authors):

      Major comments

      (1) Can induction of epithelial fusion enhance wound healing?

      Different epithelial cell-cell fusion processes have been well-characterized: i) Trophoblast fusion in the placenta mediated by Syncytins. ii) Viral induced cell-cell fusion mediated by diverse viral glycoproteins (e.g. gp41 from HIV, Hemaglutinin from Influenza, GP from Ebola, and G glycoprotein from VSV). iii) Epidermal, myoepithelial, and other epithelial cell-cell fusion in C. elegans mediated by EFF-1 and AFF-1. iv) Cell-cell fusion in the eye lens (unknown fusogens). The authors may want to compare and discuss the temporal dynamics and intermediates observed in the diverse processes of epithelial cell-cell fusion with the characterization of syncytia formation during wound healing of the Drosophila pupal notum. Since some of these characterized cell-cell fusogens can fuse heterologous cells, including Drosophila S2 cells (Shilagardi et al., 2013; https://pubmed.ncbi.nlm.nih.gov/23470732/), the authors may consider expressing these fusogens in Drosophila pupal notum before, during and after injury. This could determine whether syncytia formation is sufficient to stimulate efficient wound healing.

      We thank the reviewer for the suggestion of comparing and discussing temporal dynamics and intermediates observed in the many types of epithelial fusion that are well understood. Regretfully, we do not think this article is the right venue for such a complex discussion, especially since we have little by way of comparison in our own wound-induced fusion data. As for overexpression of fusogens, it is an intriguing idea to force cell fusion with a heterologous fusogen such as EFF-1 and then investigate any resulting changes in wound healing. However, since half the cells within 70 µm of the wound already fuse even without a heterologous fusogen, it seems unlikely we could meaningfully increase the level of cell fusion unless we expressed the fusogen universally, forcing the fusion of nearly all the epithelial cells as well as other cells throughout the body that express pnr-Gal4. Because the overexpression of EFF-1 in C .elegans results in lethality (PMID: 26854231), a widespread induction of fusion would be expected to cause other types of physiological problems that would interfere with the interpretation of wound closure rates. Further, the conditional expression tools in Drosophila allow excellent spatial control, but temporal control is still somewhat low-resolution, so that we would have difficulty expressing EFF-1 before, during, and after wounding at times that would be relevant to understanding wound healing.

      (2) The phenomenon of "border breakdowns" described here is not clear. The authors are probably studying the disassembly of the apical junctions following the initiation of membrane fusion and pore expansion. This should be clarified by using membrane labels to directly observe membrane fusion. Researchers have used electron microscopy and membrane fluorescent probes to follow cell-cell fusion. For example, GPI-mCherry, FM4-64, lipid-modified-GFPs (e.g. PH-domain fluorescently labeled proteins) DiO, DiI, and many others. See for example: Markosyan et al., 2016; https://pubmed.ncbi.nlm.nih.gov/26730950/; Mohler et al., 1998; https://pubmed.ncbi.nlm.nih.gov/9768364/; Meng et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32668210/.

      We agree completely with the reviewer, that border breakdowns represent the disassembly of apical junctions following initiation of membrane fusion and pore expansion. Direct evidence for this order of events is found in the video stills of Figure 1 panel I and video S2, which show that cytoplasmic GFP is transferred to the fusion partner 14 minutes before there is a visible decrease in the apical adherens junction marker p120ctn. The reproducibility of this order of events is documented in Fig. 3: among 107 GFP-labeled cells, 30 of them first visibly shared GFP with a fusion partner, and then 11/30 displayed border breakdown, 16/30 displayed cell shrinking, and 3/30 did not fuse. This last category is consistent with a fusion pore that closed rather than expanded productively. Although we have obtained TEM images of wound-induced fusion pores, these are included in another manuscript currently in revision and so cannot be included here, and further these EM images do not shed light on border breakdown per se, as only live imaging can establish the relationship between border breakdown and pore formation (GFP-sharing).

      (3) The observation of cell shrinking may be misleading. The process the authors describe as "cell shrinking" may involve shrinking of the apical domain, maintaining the cell volume. To clarify this process, the authors may simultaneously label the apical and basolateral domains. It is possible that fusion pore formation occurs in the basolateral, apical, or both domains. The apical shrinking could reflect the migration of the apical junctions following fusion. A similar process has been described in epidermal and vulval cells of C. elegans and other nematodes (Mohler et al., 1998; https://pubmed.ncbi.nlm.nih.gov/9768364/; Sharma-Kishore et al., 1999; https://pubmed.ncbi.nlm.nih.gov/9895317/; Kolotuev and Podbilewicz 2008; https://pubmed.ncbi.nlm.nih.gov/18031720/).

      We thank the reviewer for pointing out these examples of cell fusion in nematodes, and we now compare our findings to Mohler et al, 1998. In Fig. 2D, we specifically investigated what happened to the cell volume of these shrinking cells, and we hope we have now clarified both the text and the annotations on the figure to make our findings more clear. In the X-Z plane, the entire cell volume of two shrinking cells is visible from cytoplasmic GFP labeling. For both cells, the cytoplasmic volume moves laterally into the neighboring syncytia, appearing to initiate the movement from the basal-most area of the cell so that 150 minutes after wounding, both cells have a reduced apical footprint and only a whisp of apically-oriented cytoplasm, with the remainder of the cytoplasm having moved into the syncytia. These images make it clear that fusion is occuring, and that when the apical area disappears the corresponding cytoplasm has also moved into the territory of the neighboring syncytium. In response to the reviewer's suggestion, we did try labeling basolateral domains, but the fluorescent proteins we examined are not restricted to the basolateral domain and are difficult to interpret.

      Minor comments

      (1) Lines 40-43. Repair of injuries has also been observed in non-proliferative syncytial epidermal cells and involves cell-cell fusogens. The authors may want to include this reference: Meng et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32668210/.

      We thank the reviewer for the suggestion, and we have included this reference in the Discussion paragraph about fusogens.

      (2) Lines 128-130. Is "Shrinking fusion" an "artefact"?

      The apical junction shrinks not the cell. I suggest following basolateral membranes to see whether the cell is indeed shrinking as it fuses. The authors may want to share whether the cell volume is maintained but spills into an existing syncytium; the apical junction shrinks because it disappears/disassembles (see also Major comment 3).

      As discussed in Major comment 3, we do provide evidence that the cell cytoplasm spills into an existing syncytium. Perhaps the reviewer finds the term "shrinking cell" to be misleading, as we all agree that the cell contents do not disappear. We have updated the manuscript to use the term "apical shrinking" throughout.

      (3) Lines 157-159. Are these small cells or instead they are small apical junctions? The interpretation should include basolateral domains of the small cells to determine their size! It is also possible that some small cells have fused with the syncytia but on the basolateral domain without apical junction disassembly.

      We appreciate the reviewer's rigor. As noted above, we were not able to analyze the basolateral domains of these cells. Because our all analyses are live-imaging videos, we are able to identify the cells are undergoing apical shrinking and clearly delineate those from stable diploid cells. We now realize that the term "small cells" is confusing and can be mixed up with apical shrinking. These cells are not "small" but normal sized, small only in comparison with the gigantic syncytia around them. We have removed the term "small" from this description.

      (4) Lines 204-206. Many genes required for myoblast fusion in Drosophila have been shown to play a role in different stages of cell-cell fusion. Do they play roles in epithelia fusion during wound closure in the pupal notum?. For example, actin polymerization? Dynamin? Ig-domain and integrin cell adhesion machineries?

      We now provide a new Fig. 4 that shows that the autophagy gene Atg1 reduces wound-induced cell fusion, as it does in larvae (Kakanj et al, 2022), and importantly these wounds close more slowly. We have not analyzed mutants in actin polymerization because we are confident they would interrupt many aspects of wound healing. The Galko lab has identified that integrins suppress wound-induced cell fusion in larval epidermis, but we have not tested these. We have a manuscript in revision demonstrating a requirement for Dynamin and other endocytosis genes in wound-induced fusion, and without dynamin-mediated fusion, these wounds close more slowly.

    1. Author response:

      We sincerely thank the editors and reviewers for their time and thoughtful feedback on our manuscript. The reviewers' constructive comments have been very helpful in guiding our revision plan. Below, we outline our plan.

      In response to Reviewer #1's comments on clarifying the factors that affect image difficulty and categorization rules, we will implement several revisions. First, to clarify what drives image difficulty, we will test whether image typicality within categories, quantified using methods such as Kramer et al. (2023; Sci Adv 9.17: eadd2981), can explain monkey categorization performance. Second, we will also examine whether performance on generalization images depended on their similarity to specific repeated images and on their category typicality. Third, to address whether monkeys and humans apply similar category rules, we will focus on images for which monkeys consistently made errors and examine whether these same images also yielded lower performance (i.e., longer reaction times) in humans.

      Reviewer #1 also raised an important question about how well macaque IT representations and behavior align. The IT categorization performance estimated in our manuscript is currently lower than monkey behavior, but this may reflect the limited number of recorded neurons. We will estimate ceiling IT performance as a function of neuron count and compare it with monkey and human behavior.

      In response to Reviewer #2's suggestion to enhance narrative flow, we will reorganize the text and adjust the ordering of certain figures and sections to ensure smoother transitions between findings and analyses. Specifically, we will more clearly state which parts of the manuscript establish monkeys' categorization ability and which parts compare their behavior with models or humans before performing a triangular comparison across all three.

      Regarding Reviewer #2's suggestion to test DNN performance on control experiments (non-natural stimuli, arbitrary categorization), we agree this is an excellent addition. We will perform these analyses and plan to report the results in the revised manuscript.

      We believe these revisions will substantially strengthen the manuscript and fully address the reviewers' feedback.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study constructed engineered NK-92 cell extracellular vesicles displaying CD19 single-chain variable fragment and evaluated their therapeutic efficacy in MRL/lpr mouse models of systemic lupus erythematosus, demonstrating that these vesicles could deplete B cells, alleviate lupus nephritis, and improve mouse survival. However, this strategy lacks significant innovation compared to existing research. The current results are not sufficient to provide strong support for the experimental hypotheses.

      Weaknesses:

      (1) This study proposes using engineered EVs displaying CD19 scFv to target B cells for SLE treatment. However, similar core therapeutic strategies have been reported in previous studies. For instance, recently, studies have reported engineered EVs for SLE therapy (J Control Release. 2025, 384:113886; Ann Rheum Dis. 2025, 84(11):1811-1821; J Nanobiotechnology. 2026, 24(1):203). Another research team from China also constructed engineered EVs displaying anti-CD19 scFv for SLE treatment, which is highly consistent with the present work in targeting strategy, delivery vehicle, and disease model (Mol Ther. 2026:S1525-0016(26)00080-8). Moreover, the human trial of allogeneic CD19-targeted CAR-NK therapy for SLE has been published (Lancet. 2026, 406(10522):2968-2979). This study has not made original improvements in therapeutic vectors, targeting modules, therapeutic mechanisms, and indications, and thus finds it difficult to meet the requirements of high-level journals for originality and novelty.

      J Control Release. 2025, 384:113886; Ann Rheum Dis. 2025, 84(11):1811-1821; J Nanobiotechnology. 2026, 24(1):203). Another research team from China also constructed engineered EVs displaying anti-CD19 scFv for SLE treatment, which is highly consistent with the present work in targeting strategy, delivery vehicle, and disease model (Mol Ther. 2026:S1525-0016(26)00080-8). Moreover, the human trial of allogeneic CD19-targeted CAR-NK therapy for SLE has been published (Lancet. 2026, 406(10522):2968-2979).

      Reviewer 1 mentioned 4 publications

      (1) J Control Release. 2025, 384:113886; Genetically engineered extracellular vesicles expressing decoy protein TACI provide a therapeutic effect in systemic lupus erythematosus mouse model

      (2) Ann Rheum Dis. 2025, 84(11):1811-1821; J Nanobiotechnology. 2026, 24(1):203)Genetically modified CD19-targeting IL-15 secreting NK cells for the treatment of systemic lupus erythematosus. –but not Evs

      (3) Lancet. 2026, 406(10522):2968-2979) Efficacy and safety of allogeneic CD19 CAR NK-cell therapy in systemic lupus erythematosus: a case series in China。

      (4) Anti-CD19 engineered exosomes enable B-cell targeted anti-BAFF mRNA delivery to alleviate lupus progression”, 

      We sincerely thank the reviewers for their valuable and constructive feedback. We fully acknowledge the important contributions made by the publications cited, and we respectfully submit that they do not invalidate our findings. A critical point to emphasize is that our study employed engineered NK-92 cell extracellular vesicles (EVs) not the cells themselves and we would like to respectfully reiterate the fundamental differences between whole cells and non-cellular EVs, particularly in terms of safety and efficiency profiles. Our safety hypothesis is further supported by the clinical use of inactivated NK-92 cells (as demonstrated in this study: [URL]), which we believe provides a strong and relevant precedent. We are also very grateful that the originality and novelty of our approach have been favorably recognized by Reviewers 2 and 3, which we take as an encouraging validation of our work.

      (2) Numerous core experiments are missing, including the validation of CD19 scFv fusion protein expression on EVs, systematic characterization of engineered EVs, verification of EVs functions and therapeutic mechanisms, and in vitro and in vivo safety assessments. The available data are insufficient to support complete conclusions.

      (3) The stable expression of CD19 scFv on EVs should be further verified by Western blot or flow cytometry. The anchoring of CD19 scFv on the outer membrane surface of EVs must be confirmed. In addition, the loading capacity of CD19 scFv on exosomes should be quantified for the dosage selection in SLE treatment.

      We sincerely thank the reviewers for raising these important points. We note that points (2) and (3) address essentially the same concern, and we fully agree that further validation of CD19 scFv fusion protein expression on EVs is necessary. We are pleased to confirm that we will present additional data on this in due course. Furthermore, we respectfully acknowledge that several other aspects—including the EVs' functions, therapeutic mechanisms, in vitro and in vivo safety profiles, and CD19 scFv loading capacity—remain to be thoroughly investigated. We are committed to addressing these important questions in our follow-up studies, and we hope to provide more comprehensive insights in future work.

      (4) In vitro experiments are required to confirm the specific targeting ability of CD19 scFv-EVs to B cells and clarify the precise mechanism of B cell depletion, particularly whether it is mediated by effector molecules carried by exosomes such as perforin and granzyme B.

      We are most grateful to the reviewer for raising this important point. We are happy to report that we have successfully obtained data demonstrating the specific targeting of CD19 scFv-EVs to B cells, and we will be pleased to include these findings in our revision. With regard to the mechanism of action, we respectfully acknowledge that perforin and granzyme B are recognized as key mediators of NK cell targeting. Nevertheless, we are not aware of any published evidence to date that supports the presence of this same machinery in NK exosomes. We consider this a valuable question for future exploration, and while it lies beyond the scope of the current work, we are diligently investigating it in related ongoing studies.

      (5) The key quality control parameters, such as the stability, purity, buoyant density, and particle/protein ratio of engineered exosomes, should be characterized and identified.

      Agreed, We will provide additional characterization data for the engineered EVs in our revision.

      (6) For the in vivo treatment experiments, the author needs to explain how the treatment dose of CD19scFv-EVs was determined in order to clarify the dose-effect relationship.

      We sincerely thank the reviewer for this valuable suggestion. We fully agree and will be happy to revise the dose calculation accordingly in the updated manuscript.

      (7) It is necessary to supplement with in vivo imaging and tissue distribution data to prove that the CD19 scFv-EVs can specifically accumulate in B-cell organs such as the spleen or lymph nodes. 

      We sincerely thank the reviewer for this valuable suggestion. We fully acknowledge that this is a challenging experiment for several reasons: (1) EV internalization is a rapid process and is therefore difficult to capture; and (2) currently, there is no reliable method available for labeling EVs. Nevertheless, we respectfully assure the reviewer that we will make every effort to attempt this experiment and will report our findings in due course.

      (8) The author needs to clarify the mechanism by which CD19 scFv-EVs reduce B cells in vivo and verify the caspase apoptosis pathway.

      We sincerely thank the reviewer for these valuable comments. We are pleased to confirm that we have successfully demonstrated the specific targeting ability of CD19 scFv-EVs to B cells, and we will gladly incorporate these results in our revised manuscript.

      Regarding the mechanism of action, we fully acknowledge that perforin and granzyme B are well-established mediators of NK cell targeting according to textbook knowledge. However, to the best of our knowledge, there is currently no evidence indicating that NK-derived exosomes are equipped with the same machinery. We respectfully recognize that this is an interesting and important question; while it lies beyond the scope of the present study, we are actively pursuing it in our ongoing parallel work.

      We also appreciate the reviewer's comment regarding the apoptosis pathway. We respectfully note that this aspect was not assessed in any of the publications mentioned by Reviewer 1, which suggests that such analysis may be considered optional rather than mandatory. Nevertheless, we fully agree that this is a worthwhile avenue for further investigation, and we are committed to exploring it in our future studies."

      (9) For the in vivo therapeutic experiments, the clinical first-line drugs and the free CD19scFv should be used to supplement the control group to highlight the advantages of the engineered EVs.

      We sincerely thank the reviewer for this thoughtful and constructive advice. We fully agree that if we were developing this approach for clinical trials, regulatory agencies such as the FDA would require it to demonstrate superiority over current first-line clinical drugs. However, we respectfully wish to clarify that the primary objective of the present study is to provide a proof-of-concept that this strategy is feasible. We fully acknowledge that efficacy and safety will need to be investigated more intensively in future studies before any clinical translation can be considered. We are grateful for this valuable perspective and will be sure to discuss these considerations more explicitly in the revised manuscript.

      (10) Safety assessment in this manuscript is completely absent. Routine toxicity examinations, including hepatic and renal function tests, routine blood tests, and histopathological analysis of major organs in mice, must be supplemented. In addition, the systemic inflammatory cytokine profile and anti-drug antibody levels should be determined to rule out critical safety risks such as cytokine release syndrome and immunogenicity. The authors only focused on alterations in B cells; the impacts of the treatment on T cell subsets, NK cells, and monocytes/macrophages should be further investigated.

      We sincerely thank the reviewer for this valuable advice. We fully agree and will be happy to provide additional data to address this point in our revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      Sun and colleagues report the development of an engineered extracellular vesicle platform derived from NK-92 cells that display an anti-CD19 single-chain variable fragment (scFv) on their surface via fusion with LAMP-2B (V-CD19-Exo). In an MRL/lpr mouse model of SLE, the authors demonstrate that intraperitoneal administration of V-CD19-Exo reduces splenic CD19+CD20+ B cells, attenuates proteinuria and lupus nephritis pathology, downregulates pro-inflammatory cytokines (IL-17A, IFN-γ) and autoantibodies (anti-dsDNA, ANA), and improves survival from approximately 25% to 80%. The authors propose that this "cell-free" targeted extracellular vesicle strategy offers advantages over conventional cell therapies, including lower immunogenicity, scalable production, and no requirement for lymphodepletion.

      The study addresses an important question in autoimmune disease therapeutics: how to achieve targeted B cell depletion while avoiding the complexities and safety risks associated with CAR-T/CAR-NK cell therapies. The concept is novel, and the initial in vivo efficacy data are encouraging. However, several significant limitations in experimental design, mechanistic depth, and evidence rigor temper the strength of the conclusions.

      Strengths:

      (1) Novel conceptual approach.

      The adaptation of CAR targeting principles to extracellular vesicles represents a creative and potentially impactful strategy. By displaying CD19 scFv on NK-92-derived vesicles, the authors successfully confer B cell-targeting capability while retaining the cytotoxic effector functions of the parental NK cells. This "cell-free" concept addresses genuine limitations of live cell therapies, including the need for lymphodepletion, risks of cytokine release syndrome, and manufacturing complexity.

      (2) Comprehensive in vivo efficacy readouts.

      The study evaluates therapeutic effects across multiple clinically relevant endpoints: B cell depletion (flow cytometry), renal function (proteinuria, UPCR), renal histopathology (HE staining with semi-quantitative scoring), systemic inflammation (IgE, IL-17A, IFN-γ), autoantibody production (anti-dsDNA, ANA), and survival. This multi-dimensional characterization strengthens the phenotypic evidence for efficacy.

      (3) Appropriate control groups.

      The inclusion of non-targeted NK92-Exo as a control allows attribution of the observed effects to CD19-mediated targeting rather than non-specific vesicle-associated activities.

      (4) Significant survival benefit.

      The improvement in survival from 25% to approximately 80% in V-CD19-Exo-treated mice is substantial and represents arguably the most compelling evidence for therapeutic potential in this model.

      Weaknesses:

      (1) Mechanism of B-cell reduction remains unclear.

      The manuscript reports a dramatic reduction in splenic CD19+CD20+ B cells (from 10.53% to 1.51%) following V-CD19-Exo treatment. However, the authors do not establish whether this results from direct cytotoxicity (e.g., perforin/granzyme-mediated killing, apoptosis induction) or from functional suppression/downregulation of CD19 expression. The authors speculate that the effect is likely mediated by cytotoxic proteins carried by NK-92-derived vesicles, but no data are provided to support this mechanism. Essential experiments would include the detection of apoptosis markers (Annexin V, activated caspase-3/7) in B cells, assessment of perforin/granzyme B content within V-CD19-Exo, or in vitro co-culture assays demonstrating direct B cell killing.

      We sincerely thank the reviewer for raising this excellent question. We fully agree that it is an important point that truly needs to be addressed. We are pleased to confirm that we have already begun investigating this and hope to obtain meaningful results in due course.

      (2) Small sample sizes.

      Most experimental endpoints were assessed with n=5 per group, which is marginal for detecting modest effect sizes and may amplify the influence of individual biological variation. While the survival study had n=10 per group, the main mechanistic and endpoint analyses would benefit from larger cohorts (n=8-10) to increase statistical power and robustness.

      We are most grateful to the reviewer for this thoughtful and constructive comment. We completely agree that the sample size in our current analysis is somewhat limited for robust statistical evaluation. We are pleased to report that we have since collected additional data, which we will incorporate into our revised manuscript to strengthen the statistical power. If further data become available, we will gladly update them in subsequent revisions.

      (3) No dose-response or dosing optimization studies.

      All experiments used a single dose (10<sup>9</sup> particles per injection) and a fixed schedule (twice weekly for three weeks). The absence of dose-response data leaves unclear whether the observed effects represent maximal efficacy or could be achieved with lower doses, and whether alternative dosing regimens could improve outcomes or reduce potential off-target effects.

      We appreciate the reviewer's thoughtful and important question. We completely agree that this needs to be addressed, and we have already started working on it. We will be pleased to update our data in later comments once further results are obtained.

      (4) Lack of safety assessment.

      The authors emphasize the theoretical safety advantages of extracellular vesicles over cell therapies, but no systematic safety evaluation is presented. Key missing data include: histopathological examination of non-target organs (liver, lung, heart, gastrointestinal tract), assessment of off-target immune activation (T cell responses, cytokine profiles beyond those measured), and evaluation of potential accumulation or toxicity with repeated dosing.

      We appreciate the reviewer's careful and important observations. We fully agree that a systematic safety assessment is necessary.We are actively conducting these experiments and will update our manuscript with the findings as soon as possible.

      (5) Incomplete characterization of the engineered vesicles beyond targeting.

      While the manuscript successfully demonstrates CD19scFv display and vesicle enrichment of exosomal markers, it does not characterize whether V-CD19-Exo retains the full spectrum of NK-92 effector molecules (perforin, granzymes, FasL, TRAIL, cytokines such as IFN-γ) at functional levels. Quantitative or semi-quantitative comparison of cargo between V-CD19-Exo and parental NK-92 cells or non-engineered NK92-Exo would help contextualize the observed in vivo effects.

      We thank the reviewer for this valuable comment. We fully agree that further characterization of the engineered vesicles including NK-92 effector molecules and cargo comparison is needed. We are actively working on this and will update the manuscript as soon as the data become available.

      (6) Sex as a biological variable is not systematically addressed.

      The authors note in the Discussion that the same treatment showed more significant efficacy in male mice compared to females (data not shown), yet all main experiments were conducted exclusively in female mice. Given the strong sex bias in SLE epidemiology (approximately 9:1 female-to-male ratio) and potential differences in immune responses between sexes, this observation warrants systematic investigation rather than a footnote. Presenting the sex-differential data or alternatively, conducting adequately powered sex-stratified analyses would substantially strengthen the manuscript.

      We appreciate the reviewer's important comment. We agree that sex is a relevant biological variable, but a systematic analysis is beyond the current scope. We will consider this for future studies and will acknowledge this limitation in the Discussion.

      (7) Translational claims are premature.

      The manuscript repeatedly emphasizes advantages over cell therapy (low immunogenicity, scalable production, no requirement for lymphodepletion) as if these are established properties of V-CD19-Exo. However, no experiments directly compare V-CD19-Exo to CAR-NK or CAR-T cells in terms of efficacy, immunogenicity, or safety. Similarly, claims of "scalable production" and "high batch-to-batch consistency" are not supported by any manufacturing or quality control data. These statements should be toned down or supported with empirical evidence.

      We thank the reviewer for this important observation. We fully agree that our therapeutic claims are premature without direct comparative and manufacturing data. We will revise the manuscript to temper these statements and present them as potential advantages that warrant future investigation.

      Reviewer #3 (Public review):

      Summary:

      This manuscript describes the development of engineered NK-92-derived extracellular vesicles (EVs) displaying CD19scFv for targeted treatment of systemic lupus erythematosus (SLE). Using a CD19scFv-LAMP2B fusion strategy, the authors generated EVs intended to selectively target pathogenic B cells in the MRL/lpr lupus mouse model. The study reports reductions in CD19⁺CD20⁺ B-cell populations, improvements in proteinuria and renal histopathology, decreased inflammatory cytokines and autoantibody levels, reduced splenomegaly, and improved survival outcomes following treatment. The work aims to position engineered EVs as a cell-free alternative to CAR-T/CAR-NK therapies for autoimmune disease treatment. While the concept is interesting and potentially translational, the study currently lacks sufficient methodological rigor, EV purification standards, mechanistic validation, and comprehensive characterization to fully support many of the claims presented.

      Strengths:

      (1) The study addresses an important unmet clinical need in systemic lupus erythematosus and explores an innovative cell-free therapeutic strategy.

      (2) The concept of combining CAR-like targeting approaches with engineered EVs is interesting and potentially translational.

      (3) The manuscript includes both in vitro and in vivo experiments, including functional renal assessments, immune profiling, histopathology, and survival studies.

      (4) The authors attempt to evaluate multiple disease-associated readouts, including proteinuria, cytokines, autoantibodies, splenomegaly, and survival outcomes, which strengthens the overall biological relevance of the work.

      (5) The use of engineered NK92-derived vesicles as a scalable alternative to CAR-NK therapy represents a potentially attractive therapeutic platform.

      (6) The in vivo therapeutic observations in the MRL/lpr lupus model are encouraging and warrant further mechanistic investigation.

      Weaknesses:

      (1) The EV isolation strategy is not sufficiently rigorous for defining the isolated particles as "exosomes" according to current International Society for Extracellular Vesicles/MISEV guidelines. The precipitation-based workflow without density gradient purification or SEC raises major concerns regarding EV purity and identity.

      We thank the reviewer for this valuable and timely comment. We fully agree that our precipitation-based isolation does not meet MISEV guidelines for defining particles specifically as 'exosomes.' Since our characterization is based on shape, protein markers, and size, we will replace 'exosome' with 'extracellular vesicles' throughout the manuscript to more accurately reflect our methodology.

      (2) No direct validation was provided demonstrating successful surface localization or functional accessibility of CD19scFv on EV membranes.

      We thank the reviewer for this valuable point. We agree, and we are happy to confirm that we have obtained data on surface localization and functional accessibility of CD19 scFv, which we will include in the revision.

      (3) The characterization of EVs is incomplete and insufficient. Additional positive/negative EV markers, purity metrics, and orthogonal characterization methods are required.

      We thank the reviewer for this important point. We fully agree that more comprehensive EV characterization is needed. We are pleased to confirm that we have obtained data on CD19 scFv surface localization and accessibility, which we will include in the revision. We also acknowledge the need for additional markers and purity metrics, and will address this as a limitation in the Discussion.

      (4) The absence of density gradient ultracentrifugation is particularly concerning, given the systemic injection of EV preparations into mice, as contaminating soluble factors and non-vesicular particles may contribute to the observed therapeutic effects.

      We sincerely thank the reviewer for raising this important technical concern. We fully agree that density gradient ultracentrifugation is a more rigorous method for EV purification and that contaminating soluble factors or non-vesicular particles cannot be completely ruled out in our current preparation. We also acknowledge that even with gradient ultracentrifugation, absolute purity is not guaranteed. Nevertheless, we respectfully note that the therapeutic effect of CD19 scFv from EVs was evident when compared to appropriate controls, suggesting that the observed efficacy is attributable at least in part to the EVs themselves. We will add a clear statement of this limitation in the Discussion and will consider more stringent purification methods in our future studies.

      (5) The manuscript lacks adequate mechanistic studies explaining how engineered EVs mediate B-cell depletion or immune modulation.

      We thank the reviewer for this important point. We agree that mechanistic studies would be valuable, but we respectfully note that our current paper focuses on establishing a proof-of-concept. We plan to investigate the mechanisms of B-cell reduction and immune modulation in our future work.

      (6) The in vitro functional assays are weakly designed, particularly the use of A549 cells for evaluating CD19-targeted vesicle function.

      We thank the reviewer for this comment. We wish to clarify that the A549 experiment was intended to confirm that the engineered EVs retain their native function, not to validate CD19 targeting (which will be addressed in point (2). We will revise the manuscript to make this distinction clearer.

      (7) Important methodological details are missing, including EV normalization strategies, flow cytometry gating controls, blinding procedures, and randomization approaches.

      We thank the reviewer for this important observation. We agree that several methodological details were missing. We will reorganize and expand the Methods section to include EV normalization, flow cytometry gating controls, blinding, and randomization procedures.

      (8) Several figures, particularly TEM and western blot images, are of low quality and difficult to interpret.

      We thank the reviewer for this comment. We agree that the TEM and Western blot images are of low quality. We will provide improved, higher-resolution images in the revision

      (9) The study does not sufficiently exclude the possibility that observed therapeutic effects result from contaminating soluble immune mediators rather than EV-specific activity.

      We appreciate this concern. Based on our data, we believe the effects are EV-specific. We will acknowledge this limitation and plan additional controls in future work.

      (10) Broader immune profiling is lacking despite the systemic immune complexity of SLE.

      We thank the reviewer for this important point. We agree that broader immune profiling would be valuable, especially for clinical translation. However, our current study is designed as a proof-of-concept to establish feasibility. We will acknowledge this limitation in the Discussion and plan to address immune profiling in our future work.

      (11) The statistical analysis section includes tests that are not reflected in the Results section, creating concerns regarding data presentation and consistency.

      We thank the reviewer for pointing this out. We agree that the statistical tests in the Methods do not match those in the Results. We will revise both sections to ensure consistency throughout.

      (12) Overall, while the concept is interesting, the manuscript currently falls short of the experimental rigor expected for high-impact translational EV studies.

      We sincerely thank the reviewer for this thoughtful comment. We fully agree that this is a very early-stage translational study, and we acknowledge that considerable work remains before any clinical application can be envisioned. Nevertheless, we respectfully believe that our findings provide a valuable conceptual framework and an initial proof-of-concept that may inform and guide future translational development."

    1. Author response:

      We appreciate the reviewers’ positive assessment of the overall concept and the strength of the wild-type mouse data. We also agree with the main concern raised by the reviewers and editors: the Alzheimer’s disease model findings are more preliminary and should be distinguished more clearly from the stronger conclusions supported by the wild-type data. In the revised manuscript, we will soften the abstract, and discussion to avoid overstating disease-model efficacy, and will frame the AD-model results as suggestive and hypothesis-generating rather than definitive.

      We also plan to address the major methodological and interpretive issues raised in the reviews. We will add sex breakdowns to the figure legends and, where feasible, include sex in the analyses. We will further examine the existing EEG/EMG data to determine which additional sleep bout or spectral analyses can be included, while also clarifying the interpretation of increased dark-phase sleep as a redistribution of sleep and activity rather than a generalized improvement in sleep. We will also clarify PER2::LUC SCN phase analyses and better define the limits of our conclusions regarding central clock strengthening.

      In addition, we will improve the Methods and reporting throughout the manuscript, including clearer information about light conditions, behavioral testing timing, pathology quantification, sample sizes, exclusions or missing data, exact p values, and sex balance. We will also revise the discussion to acknowledge the limitations of the sequential design, the incomplete dissection of individual LiFE components, and the possibility that control wheel access may have reduced the dynamic range for detecting disease-model effects.

      Finally, we will correct and update the references noted by the reviewers and make the requested figure and terminology clarifications.

      Overall, we are encouraged that the reviewers found the study creative, interesting, and potentially important. We believe these revisions will sharpen the claims, improve statistical transparency, and more clearly separate the robust wild-type findings from the preliminary AD-model observations.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This useful study presents an improved protocol for long-term in vitro culture of Schistosoma mansoni that enables progression toward sexually dimorphic stages, representing a meaningful advance for studying parasite development and reducing reliance on animal models. The findings show that host-specific culture conditions support essential developmental and metabolic functions required for parasite maturation, although development remains delayed compared to in vivo conditions. The evidence is solid overall, but limited pairing efficiency and the absence of egg production indicate that the system does not yet fully recapitulate complete reproductive development.

      On behalf of the co-authors, we thank the three reviewers and the editors for their complimentary remarks as well as the major and minor comments/ concerns. Addressing these concerns have led to revisions that improved the manuscript. In particular, further analyses have generated an updated Figures 3 and 4, and Supplementary Tables S1, and S4-S6.

      Public Reviews:

      Reviewer #1 (Public review):

      Pichon, Rémi et al. describe an in vitro method for transforming Schistosoma cercariae into mature adult worms. The authors show that human serum (HS) supports parasite growth and differentiation more effectively than fetal bovine serum (FBS). They also observed differences in parasite growth and activity, with worms cultured in HS efficiently digesting human red blood cells (hRBC). Cultured worms were able to pair with ex vivo adult worms and produce eggs, indicating functional maturation suitable for downstream applications such as drug screening. While the experimental approach is comprehensive and supports the advantage of HS culture conditions, the pairing efficiency was low (≈7%) and required long culture periods (70-80 days), highlighting limitations that may affect reproducibility.

      We acknowledge the reviewer for the positive highlights. Regarding the low in vitro pairing efficiency, we have now edited the manuscript to clarify a misleading statement related to 7%. We decided to remove the value of 7% — which corresponds to the percentage of experiments in which couples were observed, as it does not accurately represent the actual number of observed worm pairs and it is probably misleading. We have updated the text as follows:

      Results, lines 230 ff.:

      “While the establishment of sexual dimorphism was robust and reproducible across more than 15 independent experiments, pairing between male and female parasites was rare. Pairing was observed only in experiments lasting more than 80 days in which we were only able to observe a few couples. In addition, these pairings were temporary (Figures 6A, B; Supplementary Video S4).”

      We also agree with the reviewer that the extended culture periods required to obtain fully sexually dimorphic parasites remain a limitation. As elaborated in Discussion (see below), key factors, probably derived from the host, are missing in the in vitro system explaining both the slow in vitro development and low rate of spontaneous pairing between in vitro developed, sexually dimorphic male and female worms. This was discussed as follows (lines 340-343): “That said, while our system was highly efficient in producing sexually dimorphic worms, spontaneous pairing between male and female parasites was extremely rare, mainly in aged in vitro cultures (from 80 to 100 days in culture) indicating that other factors, e.g., cholesterol, may be missing [35].”

      A major strength of the study, in particular, is that the authors clearly differentiate the effects of FBS versus HS on developmental progression. The conversion rate observed in HS cultures is significant and consistent with previously published data.

      While the study has several strengths, some aspects of the work are not fully explored. In particular, the role of hRBC supplementation requires further clarification. Although HScultured worms were shown to digest hRBC more readily, the implications of this observation remain unclear. Specifically, it would be useful to understand whether hRBC supplementation influences (1) long-term culture stability, (2) molecular pathways associated with development and differentiation, or (3) the pairing capacity of the worms. While addressing these questions may not be the main objective of the study, further discussion of these points would strengthen the manuscript.

      We agree that deciphering the role of the human Red Blood Cells (hRBCs) supplementation is critical. Regarding the influence of hRBCs on the long-term culture stability in parasite development it has been well established for more than four decades that schistosomes do need red blood cells to grow in culture [Basch, P. F. Cultivation of Schistosoma mansoni in vitro. II. production of infertile eggs by worm pairs cultured from cercariae. J Parasitol 67, 186-190 (1981); Basch, P. F. Cultivation of Schistosoma mansoni in vitro. I. Establishment of cultures from cercariae and development until pairing. J. Parasitol. 67, 179-185 (1981)]. The molecular pathways underlying development, sexual differentiation and pairing and modulated by hRBCs in culture is currently being investigated by our team. We decided not to include these data and analyses in the current manuscript, as they fall outside its scope.

      The manuscript is clearly written and represents a valuable contribution to the field. Overall, the experimental approach is sound, and the results support a useful methodological framework for the in vitro culture of Schistosoma worms and the attainment of sexual maturity, particularly for adult male worms.

      We thank the reviewer for highlighting the manuscript’s strengths.

      Reviewer #2 (Public review):

      Summary:

      The authors perform confirmation studies of Paul Basch's seminal schistosome work from 1981, demonstrating the development of transformed schistosomules into sexually dimorphic adult parasites, albeit without successful egg production. In addition to the findings from Basch's earlier work, the authors add some new molecular data in the form of an analysis of proliferative cells in in-vitro-derived animals.

      Strengths:

      The authors successfully confirm experimental results from earlier schistosome researchers, providing a potential new tool for studying schistosome biology without the need for vertebrate hosts.

      We thank the reviewer for highlighting the manuscript’s strengths.

      Weaknesses:

      The display of data from the authors is sometimes difficult to follow/understand where it comes from. For example:

      (1) Line 136: The authors claim that parasites in HS and FBS conditions have substantially different mortality rates (11.3 +/- 2.7 vs 5 +/- 2.3) but a quite high p-value (0.8). Analyzing the raw data myself, I obtained a mean of 8.2 +/- 1.7% vs 4.8% +/- 4.3% with a p-value of 0.15. Either the data are not clearly presented, and I did not follow them, or the data presented in the text do not match the raw data in the supplemental files.

      We thank the reviewer for pointing this out; we have now edited Supplementary Tables S1 and S6 by turning them into a long format for the sake of clarity. Accordingly, Results, Methods sections, and indicated supplementary tables were edited as follows:

      Results, lines 142 ff.:

      “No morphological differences were observed between parasites cultured either in FBS or HS within the first week in culture; in both conditions most parasites were classified as early schistosomula [category 1: 76% ± 30 (average ± SD) in FBS and 73% ± 29 (average ± SD) in HS] with few lung (category 2) and early liver schistosomula (category 3) (Figure 1B, week 1; Supplementary Figure S1). The mean mortality (category 0) at week 1 was slightly higher, but not statistically significant (P= 0.42), in worms cultured in HS [9.75% ± 2.76 (average ± SD)] compared to the mortality registered in FBS-cultured parasites [5.52% ± 5.18 (average ± SD), Supplementary Table S6], consistent with previous findings [39].”

      Methods, lines 463-465:

      “To evaluate differences in mortality between HS- and FBS-cultured parasites, data from 5 experiments were combined and analysed using a Shapiro-Wilk normality test to test normality of the data and a non-parametric Wilcoxon rank sum exact test (Supplementary Tables S1 and S6).”

      Supplementary Tables:

      Supplementary Table S1. “Raw counts of parasites within each developmental stage category. Each row corresponds to a picture of parasites in culture medium containing FBS or HS. Each column corresponds to the raw parasite counts at indicated stage development (categories 0 to 5), time in culture (Time in days - D), and experimental condition.”

      Supplementary Table S6. “Summary of all statistical tests employed in this study. 1. Statistical tests of parasite mortality and the raw data table used for this test. 2. Statistical tests for worm size comparisons (correspond to Figure 2). 3. Statistical tests for worm black gut comparisons (correspond to Figure 3). BG: Black gut. 4. Statistical tests for EdU positive cells comparisons (correspond to Figure 4). Replicate code: E, M and L correspond to day 2, 8 and 15 respectively; R and W correspond to the presence (R) or absence (W) of RBCs added 13 days after transformation.”

      For clarity, below we provide the R script used to perform the statistical tests on the data shown in Supplementary Table S6 (column ‘Raw count of parasite developmental category per image and experiment’)

      Author response image 1.

      (2) Line 187/Figure 4: Though it is not clearly stated, it appears that the authors treat their EdU counts as an ordinal data set of 61 steps (from 0 to >60) rather than a continuous measure of EdU+ cells per animal. In this author's opinion, the graph strongly suggests a continuous data set, and the fact that this reviewer had to dig through poorly-labeled raw data to discover the nature of the data is problematic. The authors should either switch to a continuous data set or make it explicit that the data shown are ordinal. If counting EdU+ cells is too arduous, the authors could consider comparing the amount of EdU+ area to the amount of DAPI+ area in maximum intensity projections of their confocal images, as this would roughly approximate the amount of proliferative cells in the animals.

      As the reviewer correctly pointed out, the data were treated as ordinal because counting worms with more than 60 Edu+ cells became extremely difficult and highly inaccurate. Therefore, we decided to group in a single category, “60 EdU+ cells”, all worms showing more than 60 EdU+ cells. We have now updated Figure 4 where medians are shown instead of media values, Supplementary Table S5 to provide more comprehensive access to the raw counts, and Supplementary Table S6 to indicate the data for EdU+ cells per worm were considered ordinal. Accordingly, we have revised the corresponding sections as follows:

      Results, lines 211 ff:

      “HS-cultured schistosomula showed higher numbers of proliferating stem cells, with a median of >48 and >60 EdU+ cells per worm at days 8 and 15, respectively (Figure 4). On the other hand, most FBS-cultured parasites displayed no more than an average of 20 EdU+ cells per worm (Figure 4).”

      Methods, lines 520 ff:

      “EdU+ cells per parasite were counted for an average of 100 parasites across three independent experiments (Supplementary Table S5). Worms were grouped based on the number of cells per individual, but all those showing ⪰ 60 EdU+ cells were counted in the same group named ‘60 EdU+ cells'. Therefore, the data were considered ordinal data. Statistical analysis was performed by Kruskal-Wallis test with Dunn multiple comparison post-hoc test, with P≤0.05 considered significant (Supplementary Table S6).”

      Figure 4 legend, lines 830 ff:

      “A. Violin plots showing the number of Edu+ cells per worm at indicated time points (2, 8, and 15 days post cercarial transformation) in parasites cultured either in Foetal Bovine Serum (FBS, blue) or Human Serum (HS, light brown). Human Red Blood Cells (hRBCs) were added in the culture at day 13 post cercarial transformation. The small black dots indicate individual worms, and the big black point indicates the median of EdU+ cells per worm. All worms showing ⪰ 60 EdU+ cells were counted and clustered together in the group named ‘60 EdU+ cells’. Hence, the data were treated as ordinal and statistical analysis performed by Kruskal-Wallis test with Dunn multiple comparison post-hoc test, with P≤0.05 (*) considered significant (Supplementary Tables S5 and S6).”

      We thank the reviewer for the very interesting suggestion to quantify cell proliferation by calculating the ratio between EdU+ area to DAPI+ area in maximum intensity projections images. Measuring the fluorescence area for each worm in maximum projection is an excellent idea; however, due to the number of EdU+ cells present in some samples, we think this technique would not provide additional information or produce more detailed data compared with our analysis when the number of Edu+ cells exceeds 60 per worm. We will certainly consider this approximation for future studies.

      There are some minor issues as well:

      (1) Line 122: It is perhaps incorrect to refer to humans as "the" definitive host of schistosomes, as S. japonicum is primarily considered a zoonotic infection with water buffalo/cows being the primary definitive host.

      We thank the reviewer for pointing this out; we have now replaced ‘schistosomes’ with ‘Schistosoma mansoni’ (current line 131)

      (2) Line 185/298: The authors refer to EdU pulse-chase experiments, but the experiments described here are EdU pulse experiments.

      This is a very good point, we thank the reviewer for bringing this up and have accordingly edited by replacing ‘EdU pulse-chase’ with ‘EdU pulse’ experiments in lines 37, 204, and 321.

      Reviewer #3 (Public review):

      Summary:

      This study is significant as it established a protocol for the long-term culture of Schistosoma mansoni newly transformed cercariae, which developed in vitro into sexually dimorphic forms. The impact of two different sera, Fetal Bovine Serum (FBS) and Human Serum (HS), added to the culture medium supplemented with human red blood cells was evaluated. The authors demonstrated that HS-cultured parasites were able to digest red blood cells, a critical step for long-term parasite development. Furthermore, while most FBS-cultured parasites did not progress beyond an early liver stage, sexual dimorphism was clearly evident in the HS-cultured worms, albeit delayed compared to in vivo development.

      Strengths:

      This study could contribute to further in vitro studies for a better understanding of the unique sexual biology of Schistosoma mansoni and for screening novel schistosomicidal compounds. By increasing parasite development in in vitro studies, this protocol could have a positive impact on the principles of the 3Rs (Replacement, Reduction and Refinement) for animal research.

      We thank the reviewer for highlighting the manuscript’s strengths.

      Weaknesses:

      As the authors mentioned, "pairing between male and female parasites was rare. Pairing was observed in approximately ~7% of the experiments, usually after day ~ 80 in culture. Egg production was also not achieved with this protocol.

      Following the reviewer’s point and to clarify a misleading point, we have now decided to remove the value of 7% - which corresponds to the percentage of experiments in which couples were observed. However, this value does not accurately reflect the actual number of observed worm pairs, and it is probably misleading. We have updated the text as follows:

      Results, lines 230 ff:

      “While the establishment of sexual dimorphism was robust and reproducible across more than 15 independent experiments, pairing between male and female parasites was rare. Pairing was observed only in experiments lasting more than 80 days in which we were only able to observe a few couples. In addition, these pairings were temporary (Figures 6A, B; Supplementary Video S4).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The manuscript is well-written overall. However, there are some minor revisions that would further improve the clarity and presentation of the data.

      (1) At the beginning of the manuscript, it would be helpful to clearly state three to four specific aims or objectives. This would help readers better understand the expected outcomes and the broader methodological contribution of the study.

      We agree with the reviewer and accordingly have stated the overall goals of the study, as follows:

      Introduction, lines 106 ff:

      “We aimed at optimising a platform to study intra-mammalian schistosomes that supports in vitro sexual dimorphism establishment, consequently leading to an overall positive impact in the 3Rs (Reduction, Replacement, Refinement) for animal research (https://nc3rs.org.uk/) [42]”.

      (2) In the abstract, you highlighted the relevance of the work according to the 3R principles of reduction in animal experimentation. However, this point is not clearly introduced in the Introduction section. Including a short discussion of this aspect would improve continuity and context.

      Following this and previous item raised by the reviewer, we have now clarified the potential impact in the 3Rs by our research outcomes and included that link to the NC3Rs website and a representative reference [Louis-Maerten E, Rodriguez Perez C, Cajiga RM, Persson K and Elger BS (2024). Conceptual foundations for a clarified meaning of the 3Rs principles in animal experimentation. Animal Welfare, 33, e37, 1–11)].

      (3) In line 43, please italicize Schistosoma spp.

      Edited accordingly.

      (4) When discussing the importance of "interfering with sexual development," in line 52, please specify the life cycle stages being referred to.

      Revised accordingly as follows:

      Introduction, lines 54-56:

      “This suggests that interfering with the sexual development of schistosome intra-mammalian stages could potentially restrict human pathology.”

      (5) Between lines 56-58, please rephrase this sentence for clarity.

      We thank the reviewer for this editorial suggestion. The text has been revised as follows:

      Introduction, lines 58 ff :

      “Therefore, novel control strategies are urgently needed, and new targets for drug/ vaccine development became a priority. A better understanding of the mechanisms underlying schistosome development, including sexual dimorphism establishment, will pave the wave to achieve this goal.”

      (6) In lines 66-68 & line 88, please clarify whether the transcriptomic studies cited were performed in vivo, in vitro, or ex vivo, and indicate the developmental stages analyzed.

      We have now included the information suggested by the reviewer as follows:

      Introduction, lines 69-70:

      “Transcriptomic studies, at both bulk [7-11] and single cell [12-1]4 levels for intra mammalian stages in vivo and ex vivo,...”

      (7) Please indicate, in line 110, the day of culture for reference. Without this information, the conversion rates per life cycle stage are difficult to interpret and reproduce. Overall, please try to give an overview in the text of these rates of conversion for context, wherever possible.

      Following the reviewer’s question, we have clearly indicated the in vitro and in vivo timings for ‘conversion’ (understood as sexual dimorphism establishment.) We have written:

      Introduction, lines 117-120:

      “Finally, while most of the FBS-cultured parasites did not progress beyond lung and early liver stage, HS-cultured parasites reached sexually dimorphic stages by week 6, albeit at a slightly delayed rate compared to in vivo development. In the mouse model, parasites become dimorphic by day 21 post-infection (~3 weeks) [12].”

      (8) The section beginning with "Furthermore, phenotypic...cell proliferation" (line 110) may be easier to follow if moved earlier in the Introduction.

      Following the reviewer’s suggestion, we have moved and slightly rewritten the sentence to current line 112, as follows: “First, phenotypic differences between FBS- and HS- cultured parasites became evident as early as 48 hours in culture, with HS-cultured parasites exhibiting higher rates of cell proliferation resulting in larger worms in the HS condition.”

      (9) In line 126, please remove the DOI and add the citation.

      Edited accordingly.

      (10) When referring to 10-week-old parasites, in line 130, please indicate the developmental stage at which they stalled and relate this to the phenotypic scoring shown in Figure 1.

      Based on this suggestion, we have now revised the third paragraph of Results section (‘Sexually dimorphic schistosomes developed entirely in vitro from cercariae’), as follows:

      Results, lines 137 ff.:

      “The development of schistosomula derived from mechanically transformed cercariae was assessed in at least 15 independent experiments, five of which were maintained over a period of at least 10 weeks to assess parasite survival and ability to mate and produce fertile eggs (Figure 1A; Supplementary Table S1).”

      Lines 151 ff.:

      “Differences in parasite development between the two conditions became apparent by week 2 (Figure 1B). At this time point, 14.8% ± 24.9 (average ± SD, excluding dead worms) or 36% ± 33.6 (average ± SD, excluding dead worms) of the parasites cultured in FBS or HS, respectively, have reached category 3, i.e., early liver schistosomulum. Parasites in FBS rarely progressed beyond this stage during the 10-week experiment, with very few parasites (<0.1% ± 0.2, average ± SD) reaching category 4, i.e., late liver schistosomulum. In contrast, worms cultured in HS developed over time across all categories, achieving marked sexual dimorphism by week 6 (13.4% ± 18.6, average ± SD) (Figure 1B; Supplementary Figure S3A), as confirmed by PCR (Supplementary Figure S3B; Supplementary Table S2). No differences in the timing for sexual dimorphism establishment were observed between male and female parasites. The mortality rate of FBS-cultured parasites reached an average of 76.24% ± 23.46 (average ± SD) by week 10, after which the experiments under this condition were stopped as most parasites were dead (Supplementary Figure S2). From that time point onwards only parasites in HS were kept in culture. As previously described for the in vivo development of schistosomes [12], in vitro cultured parasites showed developmental asynchrony in agreement with Basch’s observations [33]; however, by week 10 most of the worms in HS (73.7% ± 25.4, average ± SD) acquired an evident sexual dimorphism (Figure 1B).”

      (11) In line 142, please provide a standard deviation value for the reported average of 14.8%, if available. As well as the absolute numbers of these parasites or indicate them in the supplementary. Otherwise, it is difficult to understand the true conversion rate.

      We followed the reviewer’s suggestions and have now rewritten the text (see above, item 10). In addition, Supplementary Table S1 was edited in long format (see answer for item 1, reviewer #2)

      (12) Please explain, IN line 144, why all cultures were maintained for 10 weeks and provide the rationale for this experimental design.

      We thank the reviewer for this opportunity to clarify this point and hence improve the manuscript. The experimental condition stopped at week 10 included only FBS-cultured worms, not HS-cultured parasites. This is relevant as most of the parasites in FBS were dead by this time, unlike the HS-developed schistosomes. Indeed, some experimental groups consisting of parasites cultured in HS were maintained for up to 22 weeks. We have now updated the text to clarify this point, as follows:

      Results, lines 160 ff.:

      “The mortality rate of FBS-cultured parasites reached an average of 76.24% ± 23.46 (average ± SD) by week 10, after which the experiments under this condition were stopped as most parasites were dead (Supplementary Figure S2). From that time point onwards only parasites in HS were kept in culture.”

      (13) In lines 146-151, please streamline the timelines of culture conditions and observed outcomes in FBS versus HS media. As the current wording makes interpretation difficult.

      Following the reviewer’s suggestion we have streamlined the culture timelines and observed outcomes, as follows:

      Results, lines 137 ff.:

      “The development of schistosomula derived from mechanically transformed cercariae was assessed in at least 15 independent experiments, five of which were maintained over a period of at least 10 weeks to assess parasite survival and ability to mate and produce fertile eggs (Figure 1A; Supplementary Table S1).”

      Results, lines 151 ff.:

      “Differences in parasite development between the two conditions became apparent by week 2 (Figure 1B). At this time point, 14.8% ± 24.9 (average ± SD, excluding dead worms) or 36% ± 33.6 (average ± SD, excluding dead worms) of the parasites cultured in FBS or HS, respectively, have reached category 3, i.e., early liver schistosomulum. Parasites in FBS rarely progressed beyond this stage during the 10-week experiment, with very few parasites (<0.1% ± 0.2, average ± SD) reaching category 4, i.e., late liver schistosomulum. In contrast, worms cultured in HS developed over time across all categories, achieving marked sexual dimorphism by week 6 (13.4% ± 18.6, average ± SD) (Figure 1B; Supplementary Figure S3A), as confirmed by PCR (Supplementary Figure S3B; Supplementary Table S2). No differences in the timing for sexual dimorphism establishment were observed between male and female parasites. The mortality rate of FBS-cultured parasites reached an average of 76.24% ± 23.46 (average ± SD) by week 10, after which the experiments under this condition were stopped as most parasites were dead (Supplementary Figure S2). From that time point onwards only parasites in HS were kept in culture. As previously described for the in vivo development of schistosomes [12], in vitro cultured parasites showed developmental asynchrony in agreement with Basch’s observations [33]; however, by week 10 most of the worms in HS (73.7% ± 25.4, average ± SD) acquired an evident sexual dimorphism (Figure 1B).”

      (14) In lines 153-159, please clarify comparisons between worms cultured in FBS and HS at equivalent time points (e.g., 2 weeks FBS vs 2 weeks HS), rather than comparing only 10 week cultures.

      Following the reviewer’s comment, we have now rewritten the whole third paragraph in Results, under the heading “Sexually dimorphic schistosomes developed entirely in vitro from cercariae” - changes detailed in answers to items 10 and 13 (above).

      (15) It would also be helpful to include information on male versus female development in the context of sexual dimorphism.

      This is a relevant point that we have not clarified in the original submission - we have now indicated in the text that no differences were detected in the timing for male and female dimorphism establishment. New text included as follows:

      Results, lines 159-160:

      “No differences in the timing for sexual dimorphism establishment were observed between male and female parasites.”

      (16) In line 163, please resolve the editing marks and punctuation.

      Resolved accordingly.

      (17) In lines 169 and 172, when referring to stages such as "early liver stage," please indicate the corresponding time in culture (e.g., 3 weeks, 7 weeks + 3 days), or define these stage classifications earlier in the manuscript.

      Following the reviewer’s suggestion we have now included the developmental category after stating ‘early liver stage’, as follows:

      Results, line 187:

      “Even though few parasites in FBS reached the early liver stage (category 3)…”

      (18) Please indicate, in line 173, the developmental stage of worms used when assessing hRBC digestion in HS and FBS cultures. Additionally, here, it would be useful to discuss how hRBC supplementation may influence worm development beyond culture conditions, including possible molecular mechanisms. As a revision, that way maybe you can include data, if already performed or conduct it, to show the effect of adding or not adding hRBC even in HS cultured worms.

      We thank the reviewer for highlighting this important item that warrants further clarification. As stated in Results washed human red blood cells (hRBCs) were added to the culture at day 13. Pilot experiments in which hRBCs were added at different time points had been previously performed; no hemoglobin digestion was apparent when hRBCs were added at days 4, 5 and 6 consistent with previous findings (Correnti JM, Jung E, Freitas TC, Pearce EJ. Transfection of Schistosoma mansoni by electroporation and the description of a new promoter sequence for transgene expression. Int J Parasitol. 2007 Aug;37(10):1107-15. doi: 10.1016/j.ijpara.2007.02.011. Epub 2007 Mar 18. PMID: 17482194.).

      Following this observation, we have added a line to clarify this point, as follows (lines 181187): “Based on both previous reports [45], and pilot experiments in which adding human Red Blood Cells (hRBCs) to the culture before day ~10 did not show obvious haemoglobin digestion, we decided to supplement the culture media with hRBCs at day 13. The addition of hRBCs allowed the parasites to feed and thus continue their development [19]. At this point, they began to swallow and degrade erythrocytes, producing hemozoin, a black pigment derived from host haemoglobin degradation and visible in the worms' intestines.”

      Regarding the specific effect of adding hRBCs in the culture, this is a very good point. First, it has been well established for more than four decades that schistosomes need red blood cells in culture to grow, as example see (Basch, P. F. Cultivation of Schistosoma mansoni in vitro. II. production of infertile eggs by worm pairs cultured from cercariae. J Parasitol 67, 186-190 (1981); Basch, P. F. Cultivation of Schistosoma mansoni in vitro. I. Establishment of cultures from cercariae and development until pairing. J. Parasitol. 67, 179-185 (1981). Second, we are currently analysing transcriptomic data from parasites cultured in different conditions, including in the presence or absence of hRBCs. We decided not to include these data and analyses in the current manuscript, as they fall outside its scope.

      (19) In line 183, please clarify whether the referenced single-cell transcriptomic data were obtained from adult worms.

      We have now clarified this point in the manuscript as follows:

      Results, lines 199 ff:

      “In schistosomes, a complex stem cell system consisting of both somatic and germline stem cells has been described by leveraging recent single cell transcriptomic data across different developmental stages, including schistosomula and adult worms [47].”

      (20) In lines 210 and 213, please indicate the absolute number of worms used for these observations, rather than only percentages. If possible, also report any sex bias in pairing.

      Following this and a similar item raised by reviewer #3 (public review), we decided to remove the mention of 7% given it is misleading. This percentage corresponds to the percentage of experiments in which couples were observed. However, this value does not accurately reflect the actual number of observed worm pairs, and it is probably misleading. We have updated the text as follows:

      Results, lines 230 ff.:

      “While the establishment of sexual dimorphism was robust and reproducible across more than 15 independent experiments, pairing between male and female parasites was rare. Pairing was observed only in experiments lasting more than 80 days in which we were only able to observe a few couples. In addition, these pairings were temporary (Figures 6A, B; Supplementary Video S4).”

      (21) In the final results section, please clarify whether pairing enhances sexual maturation of already mature worms or whether maturation occurs primarily after pairing.

      This is a very relevant point, and we thank the reviewer for giving us the opportunity to clarify it in the manuscript. As described in the manuscript the parasite sexual dimorphism was established in vitro and developed male and female parasites were capable of pairing. Moreover, enlarged oocytes in the ovary’s posterior section of in vitro developed female parasites became apparent after pairing. This observation (Figure 6E, F and Supplementary Video S6) suggests that these female parasites, fully developed in HS-supplemented culture media, were not only capable of pairing, but of starting to fully maturate. We have clarified this aspect in the manuscript as follows:

      Results, lines 243 ff.:

      “Moreover, in vitro developed females coupled with ex vivo collected mature males displayed signs of primordial ovary maturation with larger oocytes towards the posterior region of the ovary (Figure 6E, F; Supplementary Video S6). On the other hand, females developed in vitro but not paired with ex vivo collected males remained immature.”

      (22) Further in the Materials and methods sections, please clarify, isn't 8000 schistosomula/well of a 6-well plate really a confluent culture condition, and does it contribute to NTS mortality in that way, as shown in previous in vitro transformation publications? Please clarify, at least with relative values, percentages of parasite transformation in such a concentrated system.

      No formal titration experiments were carried out but based on empirical observations during pilot experiments we decided to add no more than 8,000 schistosomula per well. This is something to further investigate in the future. We have now added the following sentence in Methods:

      Methods, lines 423-426:

      “The number of parasites cultured per well (~8,000 schistosomula) was determined empirically, as no formal titration experiments were performed. At higher densities (>10,000 per well), more frequent media changes were required, and parasite development appeared to be impaired.”

      (23) Also, what was the rationale of adding hRBCs as early as 13 days post-transformation, when the parasites are in the lung and early liver stage, just forming the guts? Therefore, is it possible that this would have contributed to the observation of lesser parasites disgesting hRBCs? Also, were the hRBC supplemented each time with the media change? This was not clear.

      We thank the reviewer for these questions. The rationale of adding hRBCs at day 13 has been elaborated above (question 18). In addition, in the mouse model, parasites have already migrated through and left the lungs by day 13 post-infection, as described by Nation et al [Nation CS, Da’dara AA, Marchant JK, Skelly PJ (2020) Schistosome migration in the definitive host. PLoS Negl Trop Dis 14(4): e0007951] as follows: “In the mouse, S. mansoni schistosomula begin to arrive in the lungs between 2 and 3 days post-infection, peaking at around day 7 and lasting until around day 11”. Hence, we do not think that adding hRBCs at day 13 contributed to the observation of fewer parasites digesting hemoglobin, because this was only seen in parasites cultured in FBS, not in HS.

      The hRBCs were replaced every two weeks, or sooner if their numbers decreased due to consumption. We have now clarified this point in Methods as follows (lines 427-430): “LTC medium was replaced twice a week and washed human red blood cells (hRBCs) added to a final concentration of 0.02% v/v at 13 days after transformation. Washed hRBCs were replaced every two weeks, or sooner if their numbers decreased due to consumption.”

      (24) In the Discussion, please address the limitations related to the relatively late onset and low frequency of pairing in vitro.

      Following the reviewer’s suggestion and comments from reviewer #1, we have now included a section in Discussion highlighting the limitations of the study and avenues to overcome these in the future.

      Discussion, line 360 ff.:

      “Considering these elements in future experiments will help overcome the limitations encountered in this study, including the low rate of spontaneous pairing between in vitro– developed male and female worms and the requirement for extended culture periods (>70 days). In addition, further research is needed to assess the role of host- and parasite-derived cues in schistosome development.”

      (25) Figure 1: Please consider adding arrows or markers indicating which parasites correspond to the representative developmental stages used for classification.

      We acknowledge the reviewer for the suggestion; however, we respectfully consider this may not be necessary as (1) the images shown in Figure are representative pictures of each time point included for illustrative purposes; (2) Supplementary Figure S1 clearly depicts representative images of worms in each developmental category associated with specific morphological descriptions. For greater clarity we have now added the following text at the end of Figure 1 legend:

      Figure 1 legend, line 810-811:

      “A detailed description of the developmental categories and representative images are provided in Supplementary Figure S1.”

      (26) Figure 2: This plot is somewhat misleading in showing that the HS cultured worms grew significantly more than the FBS worms, where the latter did not grow at all, as also shown by the blue bars all over the plot.

      We appreciate the reviewer’s observation; critically, the data shown in Figure 2 represent measurements of the worm's area, which means that some worms may have become longer but thinner maintaining the same area. Most of the FBS-cultured worms did not develop beyond lung or early liver stages, in which the parasites were long/ thin or shorter/wide, respectively. Therefore, the overall area of these FBS-cultured worms almost did not change (please see the raw data and statistical analyses in Supplementary Tables S3 and S6. We believe that, as presented, Figure 2 is sufficiently clear and self-explanatory. However, we would be happy to consider any suggestions to further clarify this point in the manuscript.

      (27) Figure 3: For panel A, what is the worm percentage corresponding to? The context is missing. Please clarify in the text.

      Following the reviewer’s question and for clarity, we have now (1) modified the axis-legend in Figure 3 as “Percentage of worms displaying or not Black Guts - BG (%)”, and (2) slightly edited the legend as follows:

      Figure 3 legend, lines 820-823:

      “Bar Plot representing the percentage of Human Serum (HS)- or Foetal Bovine Serum (FBS)-cultured schistosomula with (blue bar) or without (light brown bar) black guts (BG) due to the presence of intestinal hemozoin.”

      Reviewer #2 (Recommendations for the authors):

      The authors need to clarify their presentation of data. The raw data needs to be more clearly labeled/explained, and the representation of the data in Figure 4A needs to be explicitly described or changed.

      We acknowledge the reviewer for highlighting this issue related with the data presentation and have decided to follow their advice by editing Figures 3 and 4, and improving the data presentation in Supplementary Tables S1, and S4-S6. In particular:

      Figure 3. We have now modified the axis-legend as “Percentage of worms displaying or not Black Gut - BG (%)”, and slightly edited the legend as follows:

      Figure 3 legend, lines 820-823:

      “Bar Plot representing the percentage of Human Serum (HS)- or Foetal Bovine Serum (FBS)-cultured schistosomula with (blue bar) or without (light brown bar) black guts (BG) due to the presence of intestinal hemozoin.”

      Figure 4. We have edited this figure to show medians instead of media values, and updated the legend as follows: lines 830 ff.:

      “A. Violin plots showing the number of Edu+ cells per worm at indicated time points (2, 8, and 15 days post cercarial transformation) in parasites cultured either in Foetal Bovine Serum (FBS, blue) or Human Serum (HS, light brown). Human Red Blood Cells (hRBCs) were added in the culture at day 13 post cercarial transformation. The small black dots indicate individual worms, and the big black point indicates the median of EdU+ cells per worm. All worms showing ⪰ 60 EdU+ cells were counted and clustered together in the group named ‘60 EdU+ cells’. Hence, the data were treated as ordinal and statistical analysis performed by Kruskal-Wallis test with Dunn multiple comparison post-hoc test, with P≤0.05 (*) considered significant (Supplementary Tables S5 and S6).”

      Supplementary Table S1. We have clarified the data presentation by turning it into a long format and updated the legend accordingly as follows (lines 864-867): “Raw counts of parasites within each developmental stage category. Each row corresponds to a picture of parasites in culture medium containing FBS or HS. Each column corresponds to the raw parasite counts at indicated stage development (categories 0 to 5), time in culture (Time in days - D), and experimental condition.”

      Supplementary Table S4. We have clarified the table by turning it into a long format, simplified the data presentation, and updated the legend accordingly as follows (lines 873874): “Percentage of parasites displaying either black positive (hemozoin) or black negative (no hemozoin) intestine.”

      Supplementary Table S5. We have simplified the table by turning it into a long format, and explained the naming for elements in columns C (‘Group’) and D (‘Replicate’). We have updated the legend accordingly as follows (line 876 ff.): “Raw counting of EdU positive cells per parasite for indicated experimental group, replicate and experiment in long format. The worms were classified by group (column C) and replicate (column D), using the following code: E (‘early’), M (‘medium’) and L (‘late’), corresponding to days 2, 8 and 15, respectively. R and W correspond to conditions with (R) or without (W) human red blood cells, and HS and FBS to culture medium employed.”

      Supplementary Table S6. We have incorporated a new section with the statistical analyses for parasite mortality estimation and updated the legend accordingly as follows (lines 882887): “Summary of all statistical tests employed in this study. 1. Statistical tests of parasite mortality and the raw data table used for this test. 2. Statistical tests for worm size comparisons (correspond to Figure 2). 3. Statistical tests for worm black gut comparisons (correspond to Figure 3). BG: Black gut. 4. Statistical tests for EdU positive cells comparisons (correspond to Figure 4). Replicate code: E, M and L correspond to day 2, 8 and 15 respectively; R and W correspond to the presence (R) or absence (W) of RBCs added 13 days after transformation.”

      Reviewer #3 (Recommendations for the authors):

      The study was well conducted, and the data presented clearly support the conclusions. The protocol is well described, making it reproducible. The pairing experiments could be improved.

      Specific Questions.

      (1) "Male and female adult worms that developed in vivo and recovered from mice by portal perfusion on day 42 post-infection were sorted by sex and placed in culture with worms of the opposite sex developed in vitro (>70 days). Within 24 hours of initiating the co-culturing of in vitro developed worms with ex vivo collected worms, couples were observed".

      In the interest of clarity, and considering that stating ‘worms developed in vivo were collected from infected mice’ is redundant, we have now shortened and edited these lines as follows (lines 238- 242): “Male and female adult worms were recovered from mice by portal perfusion on day 42 post-infection, sorted by sex and placed in culture with worms of the opposite sex developed in vitro. Within 24 hours of initiating the co-culturing of in vitrodeveloped worms with ex vivo collected worms, couples were observed (Figures 6C, D; Supplementary Video S5).”

      (2) Have the authors conducted experiments with in vitro female and male parasites under the same experimental conditions as the in vitro/ex vivo pairing experiments? Is it possible that the tissue culture medium used for the development of sexually dimorphic forms is inhibiting pairing?

      The reviewer raises an interesting point that warrants clarification. First, the experimental conditions tested for in vitro developed parasites were the same as for the pairing experiments, as the ex vivo collected worms were washed and placed in HS-supplemented media. Second, as the culture conditions were the same (same culture protocol and medium) between in vitro pairing and in vitro / ex vivo pairing experiments, we do not think that the tissue culture medium used for developing sexually dimorphic parasites inhibited the pairing. As elaborated in Discussion (see below), key factors, probably derived from the host, are missing in the in vitro system explaining the low rate of spontaneous pairing between in vitro developed, sexually dimorphic male and female worms. This was discussed as follows (lines 340-343): “That said, while our system was highly efficient in producing sexually dimorphic worms, spontaneous pairing between male and female parasites was extremely rare, mainly in aged in vitro cultures (from 80 to 100 days in culture) indicating that other factors, e.g., cholesterol, may be missing [35].”

    1. Author response:

      Reviewer #1 (Public Review):

      Zeng et al.’s work links several key issues in Cryo Electron Tomography in ways that reinforce each other, inspired by the cycleGAN model, leading to very positive results across several benchmark datasets. The related topics include tomogram cleaning and simulations (two crucial areas in the field), with ”spin-off” outcomes in automatic annotation and the completion of the missing wedge. The manuscript covers nearly all essential topics in Tomography, making it very comprehensive and potentially critical in the field. The generalization capabilities on the SHREC 2021 data set are very interesting, although difficult to quantify. I appreciate the approach, but I have serious concerns about some of the limitations of the results presented by the authors.

      We thank the reviewer for the encouraging assessment of our work and for recognizing the potential importance of integrating tomogram denoising and simulation within a unified unsupervised framework. We appreciate the reviewer’s thoughtful evaluation and the concerns raised regarding the limitations of the current results. We address these concerns in detail below and have revised the manuscript to clarify the scope, evaluation strategy, and practical applicability of DUAL.

      (1) Simplified data versus nowadays challenging tomography data. It is acknowledged the difficulty inmaking general tests. In this work, the method shows excellent results on potentially simple data sets (the SHREC 2021, which was used for a benchmark in ET several years ago, but not much used since then) and, even more, the old Relion data set for picking).

      We appreciate the reviewer raising this important point regarding dataset difficulty and relevance. The SHREC 2021 dataset was selected because it is currently the most widely used benchmark simulated dataset for cryo-electron tomography and originates from the last SHREC contest specifically designed for evaluating cryo-ET analysis methods. It provides standardized simulated tomograms with known ground truth structures, which enables objective and reproducible quantitative comparison between different methods. The RELION ribosome dataset is also a commonly used experimental benchmark for evaluating particle detection performance. Nevertheless, we agree that demonstrating performance on additional recent and challenging datasets will further strengthen the evaluation of the method. In response to this comment, we have expanded the experimental evaluation in the revised manuscript by applying DUAL to additional recent cryo-ET datasets to further demonstrate its effectiveness on recent tomograms with more complex biological structures and imaging conditions.

      Specifically, we added an evaluation on the CZII Cryo-ET Object Identification dataset, a popular competition in 2025 with more than 1,000 participants. This experiment complements the original SHREC 2021 and RELION ribosome benchmark results and shows that DUAL can also be successfully applied to more recent cryo-ET data. The quantitative results and representative visual comparisons (shown above in Figure 1 and 2) are provided in the new section 2.6.

      (2) Reproducibility by the average user. I have found many cases in which a specific software producesexcellent results when run by the authors. Still, the average user is lost with the parameters and cannot reproduce these promising results. I propose that the authors address this issue by involving some experimental colleagues and ask them to repeat the work. This is a general concern that applies not only to this work but to many others. I think this consideration is crucial for a field that is growing very quickly and where method development happens at an extraordinary pace... but are all of them generally useful?

      We fully agree with the reviewer that reproducibility and usability are critically important for computational methods in cryo-ET. In response to this concern, we substantially improved the accessibility and reproducibility of the DUAL framework and revised the accompanying documentation to make the implementation easier to inspect and use, as two experimental colleagues have used and reproduced the results. The updated software repository now includes improved documentation, a clearer README, practical tutorials, a method-to-implementation description, a code reference, and example workflows demonstrating how to reproduce the experiments described in the manuscript. We also provide pretrained models together with the configuration files used to generate the results reported in the paper. In addition, the revised documentation clarifies the data interface, domain convention, training workflow, model outputs, and the interpretation of the trained translators. We believe that these improvements will significantly facilitate reproducibility and make it easier for users to apply the method to their own datasets.

      Reviewer #2 (Public Review):

      This study introduces DUAL (Deep Unsupervised simultAneous denoising and simuLation), an unsupervised deep learning framework that jointly addresses denoising and realistic data simulation for cryo-electron tomography (cryo-ET). By leveraging a cyclic, unpaired learning strategy, DUAL avoids reliance on paired clean ground-truth tomograms, which represents a practical advantage over many existing supervised approaches.

      We thank the reviewer for the positive summary of our work and for recognizing the advantages of the unsupervised framework in avoiding reliance on paired ground-truth data.

      Through extensive quantitative evaluations on benchmark datasets, together with qualitative and downstream analyses on diverse experimental tomograms, the authors show that DUAL performs robustly across both denoising and simulation tasks.

      We appreciate the reviewer’s recognition of the robustness of the framework and the evaluation strategy presented in the manuscript.

      If feasible, a limited quantitative or qualitative comparison with one or more recently published deep learning approaches for cryo-ET denoising or simulation, such as CryoSamba, or DeepDeWedge, would further strengthen the evaluation and help contextualize DUAL’s performance.

      We thank the reviewer for this helpful suggestion. As also recommended by the editor, we extended the experiments to include comparisons with recently proposed methods CryoSamba and DeepDeWedge. These comparisons were performed using the same evaluation metrics used in the current experiments so that the results remain directly comparable. The additional comparisons are added into section 2.6.

      Specifically, DUAL was compared with CryoSamba for denoising and with DeepDeWedge for missing wedge compensation on the CZII Cryo-ET Object Identification dataset, a popular competition in 2025 with more than 1,000 participants. The results are shown above in Figure 1 and 2.

      Reviewer #3 (Public Review):

      The paper is titled “DUAL: Deep Unsupervised Simultaneous Simulation and Denoising for Cryo-Electron Tomography.” The authors provided two closely related code branches: one for denoising and one for missingwedge correction. However, I did not find the simulation component. This is important, as the authors state that “the simulation branch provides learning-based cryo-ET simulation to generate synthetic tomograms indistinguishable from experimental ones.”

      We thank the reviewer for carefully examining the released code and for pointing out this source of confusion. We would like to clarify that, in the DUAL framework, simulation and denoising are the two simultaneous branches that are trained jointly, rather than separate sequential modules. The simulation branch learns the transformation from clean/simulated tomograms to realistic experimental cryo-ET tomograms, while the denoising branch learns the reverse transformation from experimental tomograms to the clean domain. Together, these two translators form the cyclic unsupervised learning framework described in the manuscript.

      In the original repository release, the organization of the code may not have made this relationship sufficiently clear, which likely led to the impression that only denoising and missing-wedge correction components were provided. To address this issue, we have substantially revised the repository structure and documentation. The updated repository now explicitly documents the two simultaneous branches of DUAL, explains how the simulation and denoising translators interact during training, and provides clear instructions for reproducing both functionalities. We have also added a dedicated method-to-implementation guide, code reference, and tutorial examples that describe the usage of the simulation component and its role in generating realistic synthetic tomograms that are statistically and visually consistent with experimental cryo-ET data.

      We believe these revisions clarify the implementation of the simulation branch and make the correspondence between the manuscript and the released code substantially easier to understand and reproduce.

      In addition, no pre-trained models were provided. Given that the authors indicate that all training data are publicly available, sharing trained models together with references to the corresponding datasets would significantly facilitate evaluation of the reported performance.

      We agree with the reviewer that providing pretrained models will greatly facilitate reproducibility and evaluation by other researchers. In the revised release of the repository, we have provided pretrained models corresponding to the experiments described in the manuscript together with clear references to the datasets used for training.

      The provided instructions are quite minimal and do not currently support reproduction of the reported findings.

      We appreciate the reviewer highlighting this issue. We have expanded the documentation substantially and provided detailed instructions describing the full workflow required to reproduce the experiments presented in the manuscript. In the revised repository, we added documentation that more explicitly connects the method described in the manuscript with the released implementation. The README summarizes the repository scope and data interface, the tutorial describes the practical workflow for preparing data and running training, and the method and code reference documents describe the mapping between the DUAL formulation and the main implementation files. We believe these additions will make the workflow clearer for users who wish to reproduce or adapt the experiments.

      After many hours of trial, debugging, and experimentation, I was able to train a model for missing-wedge correction using the default parameters, although the process was slow and memory-intensive.

      We thank the reviewer for investing significant effort to test the software and for reporting this observation. Training large 3D deep learning models on cryo-ET volumes can indeed be computationally demanding. We have clarified the computational requirements in the revised manuscript and provide guidance for efficient training and inference.

      Once these points are addressed, I would return to my original request that the authors provide: 3. A fully solved and functional tutorial based on their updated notebooks with all the intermediate results.

      We agree that a comprehensive tutorial will be extremely helpful for users. In the revised repository we have provided a complete end-to-end tutorial demonstrating the workflow from raw tomograms to the final outputs including simulated tomograms, denoised tomograms, and missing-wedge-corrected tomograms.

      We once again thank the editor and reviewers for their insightful comments and suggestions, which have helped us significantly improve the manuscript and the accompanying software.

  3. Jun 2026
    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors provide a resource to the systems neuroscience community by offering their Python-based CLoPy platform for closed-loop feedback training. In addition to using neural feedback, as is common in these experiments, they include a capability to use real-time movement extracted from DeepLabCut as the control signal. The methods and repository are detailed for those who wish to use this resource. Furthermore, they demonstrate the efficacy of their system through a series of mesoscale calcium imaging experiments. These experiments use a large number of cortical regions for the control signal in the neural feedback setup, while the movement feedback experiments are analyzed more extensively. The revised preprint has improved substantially upon the previous submission.

      Strengths:

      The primary strength of the paper is the availability of their CLoPy platform. Currently, most closed-loop operant conditioning experiments are custom built by each lab, and carry a relatively large startup cost to get running. This platform lowers the barrier to entry for closed-loop operant conditioning experiments, in addition to making the experiments more accessible to those with less technical expertise.

      Another strength of the paper is the use of many different cortical regions as control signals for the neurofeedback experiments. Rodent operant conditioning experiments typically record from the motor cortex, and maybe one other region. Here, the authors demonstrate that mice can volitionally control many different cortical regions not limited to those previously studied, recording across many regions in the same experiment. This demonstrates the relative flexibility of modulating neural dynamics, including in non-motor regions.

      Finally, adapting the closed-loop platform to use real-time movement as a control signal is a nice addition. Incorporating movement kinematics into operant conditioning experiments has been a challenge due to the increased technical difficulties of extracting real-time kinematic data from video data at a latency where it can be used as a control signal for operant conditioning. In this paper, they demonstrate that the mice can learn the task using their forelimb position, at a rate that is quicker than the neurofeedback experiments.

      Weaknesses:

      Many of the original weaknesses have been addressed in the revised preprint.

      While the dataset contains an impressive amount of animals and cortical regions for the neurofeedback experiment, my excitement for these experiments is tempered by the relative incompleteness of the dataset.

      As we have responded earlier, we acknowledge that some of the neurofeedback experiments include data from only a single mouse for some cortical regions, while for some cortical regions, there are several animals. This was due to practical constraints during the study, and we understand the limitations this poses for drawing broad conclusions. We felt it was still important to include these data sets with smaller sample sizes, as they might be useful for others pursuing this direction in the future. To address this, we have revised the text to explicitly acknowledge these limitations and clarify that the results for some regions are exploratory in nature. We believe our flexible tool will provide a means for our lab and others to include more animals representing additional cortical regions in future studies. Importantly, we have included all raw and processed data as well as code for future analysis.

      Additionally, adoption of the platform may be hindered by the absence of a tutorial on how to run a session.

      We thank the reviewer for this valuable suggestion. We agree that the absence of clear documentation and tutorials could limit the accessibility and broader adoption of the platform. In response, we have significantly improved the available resources by adding a comprehensive tutorial. Specifically, we have created a dedicated “Wiki” section on the GitHub repository, along with detailed documentation hosted on ReadTheDocs (https://clopy-docs.readthedocs.io). These resources now provide step-by-step guidance on setting up and running a session, along with additional usage examples to facilitate ease of use for new users.

      Reviewer #2 (Public review):

      Summary:

      In this work, Gupta & Murphy present several parallel efforts. On one side, they present the hardware and software they use to build a head-fixed mouse experimental setup that they use to track in "real-time" the calcium activity in one or two spots at the surface of the cortex. On the other side, they present another setup that they use to take advantage of the "real-time" version of DeepLabCut with their mice. The hardware and software that they used/develop is described at length, both in the article and in a companion GitHub repository. Next, they present experimental work that they have done with these two setups, training mice to max out a virtual cursor to obtain a reward, by taking advantage of auditory tone feedback that is provided to the mice as they modulate either (1) their local cortical calcium activity, or (2) their limb position.

      Strengths:

      This work illustrates the fact that thanks to readily available experimental building blocks, body movement and calcium imaging can be carried out using readily available components, including imaging the brain using an incredibly cheap consumer electronics RGB camera (RGB Raspberry Pi Camera). It is a useful source of information for researchers that may be interested in building a similar setup, given the highly detailed overview of the system. Finally, it further confirms previous findings regarding the operant conditioning of the calcium dynamics at the surface of the cortex (Clancy et al. 2020) and suggests an alternative based on deeplabcut to the motor tasks that aim to image the brain at the mesoscale during forelimb movements (Quarta et al. 2022).

      Weaknesses:

      This work covers 3 separate research endeavors: (1) The development of two separate setups, their corresponding software. (2) A study that is highly inspired from the Clancy et al. 2021 paper on the modulation of the local cortical activity measured through a mesoscale calcium imaging setup. (3) A study of the mesoscale dynamics of the cortex during forelimb movements learning. Sadly, the analyses of the physiological data appears incomplete, and more generally, the paper shows weaknesses regarding several points:

      The behavioral setups that are presented are representative of the state of the art in the field of mesoscale imaging/head fixed behavior community, rather than a highly innovative design. Still, they definitely have value as a starting point for laboratories interested in implementing such approaches.

      We agree with the reviewer that the behavioral setup presented here reflects current state-of-the-art approaches in the mesoscale imaging and head-fixed behavior community, and that similar systems have been implemented in other laboratories. However, the primary contribution of our work lies not in introducing a fundamentally new design but in providing a fully open-source, modular, and accessible implementation of such a system. By detailing both the hardware and software components, along with protocols for assembly and use, we aim to lower the barrier to entry for laboratories that may lack the specialized expertise or resources required to develop these systems independently. We hope this accessibility and ease of adoption will facilitate broader use of closed-loop and mesoscale imaging approaches across the field.

      Throughout the paper, there are several statements that point out how important it is to carry out this work in a closed-loop setting with an auditory feedback. Still, sadly there is no "no feedback" control in cortical conditioning experiments. At the same time, there is a no-feedback condition in the forelimb movement study, which shows that learning of the task can be achieved in the absence of feedback.

      We appreciate the reviewer’s insightful comment. We acknowledge that a no-feedback control group was not included in the neurofeedback experiments. This was due in part to the extensive exploration of multiple ROI combinations, as well as preliminary pilot experiments with a no-feedback condition that did not show consistent evidence of learning. Based on these initial results, we chose to prioritize conditions with feedback and did not pursue the no-feedback experiments further. We agree that including such a control would strengthen the study and consider this an important direction for future work.

      The analysis of the closed-loop neuronal data behavior lacks controls. Increased performance can be achieved by modulating actively only one of the two ROIs, this is not really analyzed, while this finding which does not match previous reports (Clancy et al. 2020) would be important to further examine.

      We agree that further analysis of this aspect would strengthen the interpretation of the dataset, and we encourage the community to explore this question using the publicly released data. In our 2-ROI paradigm, we observed that mice often adopt a strategy of predominantly modulating a single ROI to achieve task success, rather than dynamically balancing both regions. This behavior is noted in the manuscript. Importantly, our task design did not impose explicit constraints on the directionality of modulation across ROIs (i.e., increasing one while decreasing the other), in contrast to the paradigm used in Clancy et al. (2020). This difference in task structure may account for the observed divergence in strategies and outcomes.

      Reviewer #3 (Public review):

      Summary:

      The study demonstrates the effectiveness of a cost-effective closed-loop feedback system for modulating brain activity and behavior in head-fixed mice. Authors have tested real-time closed-loop feedback system in head-fixed mice two types of graded feedback: 1) Closed-loop neurofeedback (CLNF), where feedback is derived from neuronal activity (calcium imaging), and 2) Closed-loop movement feedback (CLMF), where feedback is based on observed body movement. It is a python based opensource system, and the authors call it CLoPy. Authors also claim to provide all software, hardware schematics, and protocols to adapt it to various experimental scenarios. This system is capable and can be adapted for a wide use case scenarios.

      Authors have shown that their system can control both positive (water drop) and negative reinforcement (buzzer-vibrator). This study also shows that using the closed-loop system, mice have shown to better performance, learnt arbitrary tasks and can adapt to changes in the rules as well. By integrating real-time feedback based on cortical GCaMP imaging and behavior tracking authors have provided strong evidence that such closed-loop systems can be instrumental in exploring the dynamic interplay between brain activity and behavior.

      Strengths:

      Simplicity of feedback systems design. Simplicity of implementation and potential adoption.

      Weaknesses:

      Long latencies, due to slow Ca2+ dynamics and slow imaging (15 FPS), may limit the application of the system.

      We agree that the latency introduced by calcium dynamics and imaging frame rates is an inherent limitation of calcium imaging–based approaches. Future improvements, including faster calcium indicators, higher frame-rate imaging systems, and more efficient computational pipelines, are expected to mitigate these constraints and enhance temporal precision.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This version is a substantial improvement from the previous version. My main recommendation is to add a tutorial, with visualizations of some sort, to show how to run a session with the platform. The tutorials for the probe trajectory planner PinPoint is a good example for reference (https://virtualbrainlab.org/pinpoint/tutorial.html).

      We thank the reviewer for this valuable suggestion. We agree that the absence of clear documentation and tutorials could limit the accessibility and broader adoption of the platform. In response, we have significantly improved the available resources by adding a comprehensive tutorial. Specifically, we have created a dedicated “Wiki” section on the GitHub repository, along with detailed documentation hosted on ReadTheDocs (https://clopy-docs.readthedocs.io). These resources now provide step-by-step guidance on setting up and running a session, along with additional usage examples to facilitate ease of use for new users.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Strengths:

      This is an ambitious study that provides a quantitative dissociation of the roles of phasic and tonic pain in adaptive behavior, by integrating ecological neuroscience, motivational theory, and computational modeling. The use of immersive VR combined with a freeoperant foraging task offers a more ecologically valid context to study pain-related behavior compared to traditional paradigms. Furthermore, the study employs a multimodal approach by combining behavioral data, computational frameworks, physiological signals, and EEG. In particular, one of the main strengths of the study is the use of sophisticated computational modeling to capture phasic and tonic pain effects. The experiment codes are available on GitHub, increasing reproducibility.

      We appreciate the reviewers’ recognition of the study’s ambition, the integration of ecological and computational approaches, and our efforts to support reproducibility through open code.

      Weaknesses:

      The main limitations of this article are that it provides insufficient detail on VR implementation. The design of the VR environment is, at this stage, under-described. Crucial information is missing, such as the number of pineapples per block, timing precision, details on how motion is mapped to the virtual movement, etc. This aspect strongly limits the reproducibility of the experiments.

      We thank the reviewer for highlighting the importance of detailed reporting to ensure reproducibility. In response to this valuable feedback, we have taken the following steps:

      (1) Open Access to Software and Data: We have now uploaded the full software and hardware specifications used in our study to a public GitHub repository: https://github.com/ShuangyiTong/PineappleStudy2025ReplicationSoftware. This includes the complete VR implementation, allowing readers to directly experience the task using a commercially available VR headset. The repository also contains the raw data and analysis scripts to facilitate full replication of our results. These links have been updated in “Data and Code Availability” section.

      (2) Expanded Methodological Details: We have revised the Methods section to include the specific details requested, such as:

      (a) The number of pineapples presented per block,

      (b) The temporal resolution and precision of the data collection,

      (c) The mapping between physical motion and virtual movement within the VR environment.

      Specifically, the paragraph containing the changes is following: “At the beginning of each one-minute block, a total number of 150 virtual pineapples of varying heights from 0.33 to 1 m were randomly generated in a circle centred around the participant with a diameter of 6.67 m. Five identical baskets were placed within the space. Spatial locations of trees and vegetation were generated using the game engine's default tree painting tool (Unity Technologies, San Francisco, US).”

      We hope these updates address the reviewer’s concerns and significantly improve the transparency and reproducibility of our experimental design.

      A second limitation lies in the lack of clarity regarding the study hypotheses. Although two overarching hypotheses can be inferred, they are not explicitly formulated. To this end, it is unclear which analyses were merely exploratory, especially for physiological and EEG outcomes.

      We thank the reviewer for this constructive feedback. We agree that making the hypotheses more explicit—particularly regarding the computational framework and the role of physiological measures—strengthens the manuscript. We have significantly revised the final section of the Introduction to explicitly formulate our two primary hypotheses and operationalise the associated behavioural and neurophysiological measures.

      (1) Phasic Pain Hypothesis: We hypothesised that phasic pain serves as a discrete valuation signal that updates the state-action value of specific actions. We predicted this would be evidenced behaviourally by reduced choice probability and increased ‘distance bias’ for pain-associated targets. Neurally and physiologically, we predicted that these aversive values would be tracked by skin conductance responses (SCRs) and the amplitude of pain event-related potentials (ERPs), which serve as established markers for the encoding of aversive magnitude and salience.

      (2) Tonic Pain Hypothesis: We hypothesised that tonic pain acts as a coefficient modulating the trade-off between opportunity cost and vigour cost. This was tested by applying tonic pain to the non-dominant (non-task) limb to ensure that any observed changes were motivational rather than mechanical. We predicted a global reduction in motivational vigour, operationalised as decreased movement velocities and foraging rates.

      By framing the study this way, we clarify that the physiological and EEG outcomes were used to quantitatively test whether the brain and body implement the computations (valuation and vigour-regulation) defined by our model. We have updated the text in the Introduction (see below) to reflect these explicit formulations.

      Updated paragraphs: “Our first hypothesis was that phasic pain provides a distinct valuation signal that updates the value of specific actions within complex environments. In our task, this was implemented by associating specific fruit (distinguishable by colour) with a brief electrical stimulus to the grasping hand, emulating thorns. In our computational model, this was defined as an aversive utility term incorporated into the state-action value evaluation process. We predicted that this computational mechanism would manifest behaviourally as a reduction in choice probability for pain-associated targets and an increase in ‘choice distance bias’ (the willingness to travel further for pain-free options). Neurally and physiologically, we predicted that these aversive values would be tracked by skin conductance responses (SCRs) and the amplitude of nociceptive event-related potentials (ERPs), specifically the N1-P2 complex (Favero et al., 2023).

      Second, we hypothesised that tonic pain acts as a coefficient modulating the tradeoff between opportunity cost and vigour cost, thereby serving a recuperative function. To test this in Experiment 2, we delivered continuous tonic pressure to the non-dominant arm via an inflated cuff to emulate a background state of injury. Within our free-operant framework, tonic pain was modelled as a weighting factor that shifts the optimal balance toward reduced energy expenditure. Because the stimulus was applied to the non-task limb, we specifically predicted a global reduction in motivational vigour—operationalised as decreased movement velocities and foraging rates—rather than a direct mechanical impairment. By applying this formal computational approach, we move beyond exploratory observations to provide a rigorous, mechanism-based explanation for how distinct pain states adaptively govern choice and action.”

      In Experiment 2, the reduction in vigor during tonic pain could plausibly reflect attentional load rather than pain per se. As recognized by the authors, there is no control condition involving an innocuous salient stimulus to rule out non-specific effects of distraction. Perhaps a tonic non-painful but salient somatosensory stimulus (e.g., a strong vibrotactile stimulus applied on the same arm) could have been used as a control stimulus.

      We agree that examining the potential role of attentional load on the interaction between tonic and phasic pain is an important area of future investigation. The inclusion of additional control conditions matched for attentional salience with additional experiments is possible but introduces other confounds related to their different qualities (e.g. a salient vibrotactile stimulus might invigorate behaviour). More fundamentally, attentional processes are a core part of pain function, and should not necessarily be viewed as a confound (i.e. the way that pain mediates some of its core functional effects may directly be through its salient attentional nature). This view is formalised in Wall and Melzack’s classical tripartite model of pain, and distinguishes pain from purely sensory systems such as somatosensation, vision and so on.

      Reviewer #1 (Recommendations for the authors):

      (1) Computational models may be difficult to follow without prior familiarity. Including simplified explanations could make the approach more accessible.

      We thank the reviewer for this constructive suggestion. To make the computational framework more accessible to a broader audience, we have added two new schematic diagrams (Figure 2 and Figure 8) that provide a visual overview of the models used in Experiment 1 and Experiment 2, respectively. These figures illustrate the state-action transitions and provide a clear decomposition of the payoff components—including reward, pain, and temporal costs. We believe these additions significantly clarify the modelling logic and help ground the mathematical descriptions in a more intuitive visual context.

      (2) Lines 220-222: I don't think it is possible to talk about "objective measures of pain" as pain is, by definition, subjective. I suggest rephrasing the sentence.

      We thank the reviewer for this thoughtful observation regarding our terminology. We recognise that the phrase ‘objective measures of pain’ may be misintepreted. Our intention was to highlight the distinction between the internal, reported experience and the behavioural manifestations of pain that our computational method reveals.

      To avoid ambiguity and to better align the text with the core focus of our study, which is the motivational function of pain, we have rephrased the sentence as suggested. We have shifted the emphasis from ‘measuring pain’ to quantifying its specific impact on behaviour.

      Original lines 220-222 have been revised as follows:

      "Taken together, this indicates the composite nature of overall aversiveness and highlights the benefit of combining subjective ratings with model-based measures of its motivational impact on behaviour."

      We believe this revision more accurately reflects our approach of using choice and movement as objective indices of the motivational value of pain.

      (3) The explanation for choosing the foraging task is very interesting, but should be provided in the Introduction rather than in the Methods section. In contrast, the Methods section should include the details of the VR implementation.

      We thank the reviewer for these constructive suggestions regarding the manuscript structure.

      Regarding the rationale for the foraging task: We agree that providing the theoretical justification for the task earlier in the manuscript improves the narrative flow. We have revised the Introduction to explicitly outline why a foraging paradigm was chosen by added the following sentences:

      “A foraging paradigm provides a robust, free-operant framework that captures the core components of adaptive behaviour: it is goal-directed, involves complex movement, and requires the learning of an optimal strategy to maximise rewards. This allows us to computationally dissociate how different types of pain influence the control of action.”

      We believe this addition clarifies the link between our computational hypotheses and the experimental design.

      Regarding the VR implementation: We have updated the Methods section to include the specific experimental parameters requested in the reviewer's previous comments (e.g., timing precision, stimulus counts, and motion mapping) to ensure full reproducibility. However, we have opted not to include the exhaustive engineering details of the underlying software architecture and communication protocols. To ensure complete transparency, the full software and firmware source code, which allows for the exact replication of the environment, is available in our public GitHub repository shown in the code and data availability section.

      (4) It is unclear how the sample size was determined. This information should be included.

      We thank the Reviewer for this comment. For the present study, an a priori power analysis was not conducted due to the novelty of the investigation and the complexity of the analyses. Standard power analyses are not commonly conducted for studies where computational modelling is the primary focus, as results would be potentially misleading. Instead, we based our sample size estimate of N ≈ 30 participants on previous studies using computational modelling of neurophysiological data [6], as well as EEG, SCR and pain studies [7, 8] and studies in our group using combined neurophysiological recordings and VR [9]. This approach represented a pragmatic balance which ensured the credibility of our results and the stability of our model estimates while accounting for the high persubject cost and the depth of the data collected from each individual. This has now been described more accurately in the Method section:

      “An a priori power analysis was not conducted due to the novelty of the investigation and the complexity of the analyses. Instead, we based our target sample size (N ≈ 30 per experiment) on previous studies using computational modelling of neurophysiological data (Mahajan et al., 2025), as well as EEG, SCR, and pain studies (Schulz351 et al., 2015; Zhang et al., 2018), and studies from our group using combined neurophysiological recordings and VR (Hewitt et al., 2026). This approach represents a pragmatic balance that ensures the credibility of the results and the stability of model estimates while accounting for the high per-subject cost and depth of data collected from each individual.”

      (5) Please clarify how / when the monetary performance incentive was provided.

      We thank the reviewer for the opportunity to clarify the incentive structure. The monetary performance incentive is detailed below:

      Participants were informed at the start of the study that they would earn a performance-based bonus of up to £10, determined by the points they collected during the foraging task. To ensure that motivation remained consistent across the entire session for all individuals—regardless of their baseline foraging speed—the specific exchange rate between points and currency was not disclosed. This prevented potential 'ceiling effects', where a high-performing subject might stop exertive effort after reaching the maximum bonus early, or 'floor effects', where a subject might perceive the reward for an individual action as too small to be motivating.

      Following the completion of the experimental session, all participants were compensated with the full £10 bonus in addition to their base payment for participation.

      We have updated the Methods section to reflect these details:

      “Participants were informed at the start of the experiment that their total points would be rewarded with a monetary incentive of up to £10. To maintain a constant level of motivation throughout the task, the exact point-to-currency exchange rate was not specified. Upon completion of the session, all participants were awarded the maximum bonus of £10.”

      Reviewer #2 (Public review):

      Strengths:

      Overall, this study aims to address an important topic and is generally well written.

      We thank the Reviewer for the generally positive evaluation of our work.

      Weaknesses:

      First, phasic pain was induced using electrical stimulation, which typically elicits somatosensory evoked potentials (SEPs). These responses may not reflect pain-specific processes and thus complicate interpretation. This issue bears directly on the study's conclusions, especially when discussing interactions between phasic and tonic pain. For example, tonic pain is known to reduce perceived intensity or cortical responses to phasic pain stimuli delivered elsewhere on the body - an effect not expected for SEPs elicited by electrical stimuli.

      We acknowledge the reviewer’s concern regarding the specificity of evoked potentials elicited by electrical stimulation. We agree that traditional SEPs— particularly those evoked by large surface electrodes—primarily reflect activation of non-nociceptive A-beta fibres and thus may not reliably index pain-specific processes or be modulated by tonic pain via descending nociceptive control. However, we would like to clarify that phasic pain was administered in the present study using small-diameter concentric ‘Wasp’ electrodes. These are comparable to intraepidermal electrodes shown to preferentially activate nociceptive A-delta fibres, thereby eliciting ERPs more closely associated with nociceptive processing rather than mixed somatosensory input [1, 2]. Accordingly, our ERP results demonstrated a reliable increase in N1-P2 amplitude with higher phasic pain intensity, suggesting that the evoked responses captured stimulus-evoked nociceptive processing.

      We acknowledge that these ERPs may still reflect mixed sensory processing and thus may not be fully modulated by tonic pain. Previous studies have shown that ERPs elicited by nociceptive electrical stimulation can be attenuated during tonic pain using cold-water immersion in CPM paradigms [3, 4]. However, these studies typically employ passive tasks, whereas our paradigm involved continuous voluntary behaviour during sustained tonic pressure pain. This difference in task context may engage distinct modulatory systems, possibly prioritising behavioural adaptation over sensory gating.

      We have revised the Discussion and Methods sections to explicitly clarify the electrode design and address the lack of ERP modulation by tonic pain in the context of active behaviour:

      Discussion: “Although we utilised concentric ‘Wasp’ electrodes designed to selectively activate nociceptive A-delta fibres, and confirmed that the resulting ERPs (N1-P2) were significantly modulated by phasic intensity (Figure 6E, F), we observed no such attenuation by tonic pain (Fig. 6G, H).”

      Methods: “These electrodes preferentially activate nociceptive A-delta fibres, thereby eliciting ERPs that more accurately reflect nociceptive processing compared to standard bipolar stimulation (Inui et al., 2002; Mørch et al., 2011).”

      Second, additional control experiments are necessary to rule out alternative explanations. For instance, the authors are suggested to deliver phasic pain to the contralateral arm (e.g., at 1-2 Hz), which might also reduce action velocity. Similarly, tonic pain applied to the grasping hand should be tested to disentangle hand-specific effects.

      We thank the reviewer for these suggestions regarding the spatial configuration of stimuli. The decision to deliver phasic pain to the grasping hand and tonic pain to the contralateral arm was a deliberate feature of our experimental design.

      First, delivering phasic pain to the grasping hand ensured spatial congruency between the virtual stimulus (the fruit) and the physical consequence (the pain). This congruency is essential for subjects to form a coherent representation of the 'painful' object; a contralateral delivery would have introduced a sensory-motor mismatch that could complicate the interpretation of the learning and choice data.

      Second, tonic pain was applied to the contralateral arm specifically to avoid mechanical interference with the grasping action. Applying sustained pressure to the ipsilateral limb would likely have impeded the manual dexterity and fine motor control required to operate the controller buttons. This would have introduced a physical confound, making it difficult to determine if changes in behaviour were due to motivational vigour or simply the mechanical difficulty of performing the grasp while the arm was under pressure.

      We agree that exploring the spatial generalisation of these effects is an important future direction, and we have added a paragraph to the Discussion to clarify these design choices:

      “It is also important to consider the spatial configuration of the stimuli used in this study. Phasic pain was delivered to the grasping hand to maintain spatial congruency with the virtual fruit, ensuring a coherent nociceptive feedback signal for the interactive task. Additionally, tonic pain was applied to the contralateral arm to prevent mechanical interference with motor execution, which would have occurred if pressure were applied to the ipsilateral limb used for grasping the controller. Whilst this design promotes spatial congruency and avoids mechanical confounds, future studies might explore how these effects generalise across different body parts, for which VR experiments serve as a promising tool to test relevant hypotheses (Hewitt et al., 2026).”

      Reviewer #2 (Recommendations for the authors):

      (1) First, the abstract mentions only EEG, yet Experiment 1 employed skin conductance response (SCR) measures while Experiment 2 utilized EEG. Also, the rationale for using SCR in Experiment 1 and EEG in Experiment 2 is not provided and should be explicitly stated.

      We thank the reviewer for identifying the discrepancy between the physiological signals reported in Experiment 1 and Experiment 2. We have revised the Abstract and Methods section to clarify the rationale for these measures.

      In Abstract, the following sentence has been revised: This could be explained by a free-operant computational framework that formalises and quantifies the function of tonic and phasic pain in terms of motivational vigour and decision value, and model parameters correlated with EEG “physiological and neural responses.”

      Regarding the rationale for the measurements, the following sentences were inserted into the Methods section: “Experiment 1 was designed to establish the robust behavioural effects of the foraging task while ensuring the collection of reliable physiological data. We chose SCR as it is a well-validated index of autonomic arousal that we were confident would provide a clear peripheral measure of pain-related processing in this novel VR paradigm.”

      For Experiment 2, we aimed to build on these findings by adding EEG. This was intended as a complementary piece of neural evidence to provide insights into the underlying central neural mechanisms of phasic and tonic pain interactions.

      (2) Second, the quality of both SCR (Figure 3A) and EEG/ERP data (Figure 5A-D) appears compromised by low SNR. For instance, ERP signals show baseline drift at low frequencies, potentially due to movement-related artifacts. The authors are encouraged to enhance data quality and provide cleaner, more interpretable results.

      We thank the reviewer for this observation. We acknowledge that our recordings exhibit a lower SNR compared to conventional, stationary EEG studies. This is a recognized characteristic of Mobile Brain-Body Imaging (MoBI), particularly in immersive VR experiments where participants are physically active [10]. However, previous research has demonstrated that it is possible to recover valid, interpretable neural signals in active settings using modern cleaning methods including trained ICA labels which we have adopted for artefacts cleaning [11]. We also believe we should be restrained from over cleaning the EEG data as pointed out by Delorme in the paper ‘EEG is better left alone’ [12]. Therefore, we have added a new paragraph in the Discussion:

      “It is important to acknowledge that the signal-to-noise ratio in both our physiological and neural recordings is lower than that typically observed in conventional, stationary laboratory experiments (Gramann et al., 2011). This is primarily due to the motion artefacts inherent in an immersive and active virtual reality environment. Whilst we utilised robust cleaning and artefact-correction methods (Klug and Gramann, 2021), the elevated noise floor may limit our capacity to detect more subtle neural effects or interactions. These challenges highlight a critical area for future methodological research, particularly in the development of hardware and signal-processing tools designed to isolate neural signals during complex, mobile behavioural tasks.”

      Another factor contributing to the appearance of the raw signal is the "free-operant" nature of our task. Unlike conventional neurophysiological study paradigms with fixed, sufficient intervals between trials, our participants were free to move and interact with fruit at their own pace. This means that neurophysiological signals from successive actions (e.g., picking up one fruit followed quickly by another) can overlap. For the SCR analysis, we addressed this by using a canonical response function (CRF) to model and "unfold" the overlapping signals with GLM to produce our final results [13]. While we did not perform a similar deconvolution for the EEG data, we focused our analysis on the early, salient components (N1-P2 and early time-frequency changes < 500ms) which are less susceptible to overlap from subsequent actions than the much slower SCR.

      In summary, while significant efforts representing the state-of-the-art approach for MoBI analyses have been taken to minimise the contributions of noise to the dataset, residual noise does remain in the final data. We have employed a combination of robust preprocessing and model-based analytical methods to account for the complexities of a free-operant task. We believe these results represent the best possible balance between signal clarity and the ecological validity of an active foraging task, and we have called for future research to continue improving these tools for immersive VR environments.

      (3) Third, although the authors state that time-frequency analysis was conducted on the EEG data, no corresponding results are presented in Figure 8 or elsewhere. Furthermore, the statistical maps shown appear noisy and require further clarification and possible denoising.

      We thank the reviewer for pointing this out. The time-frequency results are indeed presented in Figure 8 (now Figure 10); however, they are depicted as topographic maps of the t-statistics derived from our LMM rather than raw power change plots.

      The application of EEG to a novel, free-operant task represents a significant methodological development in this study. Unlike conventional EEG experiments where variables are strictly controlled and a "clean" pre-stimulus baseline is easily obtained, our task involves continuous participant engagement and movement. In this context, for the decision-making event, a stable baseline is unattainable as multiple variables, most notably head movements, are constantly in effect.

      Therefore, we believe that presenting the LMM statistical maps in the main text is the most appropriate and rigorous interpretation of the time-frequency results, as these maps represent the signal after accounting for these complex fixed and random effects. This approach was also adopted in previous pain studies [7]. We also updated the figure legend and caption specifically saying that the figure represented correlation between band power and variables we were investigating to improve clarity.

      Second, for more salient stimuli like phasic pain stimulation, we can indeed obtain a highly interpretable time-frequency analysis without further LMM analysis. We have added induced oscillatory responses to phasic pain stimuli to the Supplementary Material (section: Induced oscillatory responses to phasic pain stimuli). The results showed that, consistent with our ERP findings, the intensity of phasic pain significantly modulated induced responses, while the background tonic pain state did not significantly alter the induced oscillatory response to the phasic pain stimulus.

      Regarding the SNR and Denoising Strategy, we acknowledge that the statistical maps appear noisier than those from stationary studies. This is a direct consequence of the lower signal-to-noise ratio (SNR) inherent in mobile VR. Moving EEG from strictly controlled laboratory settings to ecologically valid, "real-world" VR scenarios introduces higher levels of noise, which we believe represents a key frontier for future methodology research. Regarding the denoising process, the maps in the main text represent the data after our full pipeline (including ICA-based artifact rejection and high-pass filtering). Regarding further denoising, we have deliberately chosen not to apply excessive spatial or temporal smoothing [12]. Also, it is important to note that the LMM framework itself serves as a powerful statistical "filter." By including head movement velocity as a regressor and accounting for random intercepts across subjects, the model effectively "cleans" the signal by partitioning out noise components not related to the task conditions.

      Reviewer #3 (Public review):

      Strengths:

      The experimental paradigm is highly innovative. Assessing human behaviour in a naturalistic yet highly controlled setting represents a promising approach to pain research. Notably, assessing pain magnitude implicitly, via its motivational value, offers insights about the overall pain experience that are not usually accessible via common pain ratings.

      Weaknesses:

      Despite these strengths, the manuscript would benefit significantly from more precise definitions of key concepts and an overall clearer, more coherent presentation of its main arguments. The writing, in its current form, often presents claims that are too vague or insufficiently connected with the experimental findings. Moreover, certain aspects of the computational modeling and statistical analysis appear flawed or inadequately justified.

      We thank the Reviewer for the generally positive evaluation of the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) The analyses presented in the section

      "Results/Additional cost of effort associated with movement" require clearer explanations. The intention here appears to be to assess the association between moving distances and pain intensity to test the hypothesis that the higher the average pain ratings within blocks, the longer the distances moved (i.e., the higher the effort to avoid pain). It is unclear why and how exactly "egocentric distance differences between painful and non-painful fruits" were computed.

      We thank the reviewer for pointing out the need for a clearer definition of the egocentric distance calculation. As the reviewer correctly identified, this analysis tests the hypothesis that subjects would trade off physical effort (distance) for pain avoidance. To compute this, we used a blockwise approach: for each one-minute block, we calculated the average egocentric distance travelled to pick up non-painful fruits and subtracted the average distance travelled to pick up painful fruits. This difference (labelled as "Choice Distance Bias" in Figure 3B) represents the additional effort subjects were willing to exert to reach a pain-free option. We have clarified the computation method and our motivation for using it in the revised text:

      “As shown in Figure 3B, the vertical axis represents the 'choice distance bias', calculated as the difference between the average egocentric distance to non-painful fruits and the average egocentric distance to painful fruits within each block. The egocentric distance is the fruit distance relative to the participant. This metric was computed to test whether subjects would trade off physical effort for pain avoidance; specifically, a positive bias indicates that subjects were willing to bypass closer painful fruits to reach more distant pain-free ones. As hypothesised, we found that as the pain intensity (VAS) of the aversive fruits increased, this distance bias grew significantly, confirming that subjects exerted greater movement effort to avoid higher levels of pain.”

      We have also updated the text in the beginning of " Avoidance increases with increasing phasic pain intensity" section to emphasize the calculation is analysed at the block level to clarify the computation procedure:

      “For this analysis, both aversive choice probabilities and subjective pain ratings were estimated at the block level.”

      (2) In its current form, the explanation of the first optimality equation lacks precision and transparency. Consider the following improvements:

      (a) Precisely define the features that characterize a state/decision point: e.g., i) memory of available options (= set of 7 fruits that were seen but not picked up) and ii) subject's current position, iii) pain intensity associated with green fruit in the current block.

      (b) Precisely define the set of values the action variable a can assume.

      (c) Precisely define the function u(a) in mathematical notation, including its hyperparameters. The fact that a is likely a categorical variable, while u(a) is later described as a sigmoid function (i.e., as a function of a continuous variable), is confusing. In my understanding (see Figure 2F), u is actually a function of the stimulus intensity associated with a given fruit. Since the stimulus intensity depends on the current state s (and varies from block to block), the phasic pain utility function technically also depends on s.

      (d) Precisely define the function d(a) in mathematical notation, including its hyperparameters.

      (e) Precisely describe how the separate horizontal and vertical components of C_m enter the equation.

      (f) Provide a summary of all parameters and hyperparameters being optimized. Are parameters and hyperparameters optimized jointly? What distinguishes parameters and hyperparameters practically?

      We thank the reviewer for this insightful critique. We agree that the original presentation of the optimality equation was insufficiently formal. We have now added a dedicated subsection, "Experiment 1 model summary", which includes a comprehensive table (Table 2) and supporting text to address these points with mathematical precision.

      Specifically, we have implemented the following clarifications in the revised manuscript:

      State and Action Space (a, b): We have formally defined the state s as an ordered memory list M_s of up to 7 items, governed by a FIFO principle. The action a is now explicitly defined as a one-to-one mapping from these memory items to physical reach trajectories.

      Utility and Cost Functions (c, d, e): We have provided the full mathematical notation for the phasic pain utility u(a) and the effort cost d(a). We have clarified that while the choice of fruit (a) is categorical, it serves as an indicator variable that determines the application of a continuous sigmoid utility function based on the block-level pain intensity (x_stim). We have also explicitly decomposed the effort cost into its horizontal (C_h) and vertical (C_v) egocentric components.

      Parameters and Hyperparameters (f): We have clarified that because our model focuses on steady-state motivational trade-offs rather than online learning, the hyperparameters listed are the only variables subject to optimisation. These are fixed for each subject across the duration of the experiment.

      We believe these additions, centred around the new Table 2, provide the transparency and precision requested.

      Furthermore, we would like to clarify a subtle caveat regarding the assumption of a fixed x_stim for the entirety of a block. While participants were aware that green pineapples were aversive, the specific stimulation intensity for a given block was only fully revealed upon picking up the first green pineapple.

      To ensure our model-fitting remains robust despite this 'information lag', we considered several computational alternatives:

      (1) Prior Estimation Modelling: Modelling a participant’s prior estimation of pain stimulation based on previous blocks. We found this unsuitable due to the independent block design and the limited number of trials available to establish a stable prior.

      (2) Data Trimming: Excluding all decisions made before the first green pineapple pickup. While theoretically 'cleaner', this approach introduces significant data imbalance and ignores blocks where a participant—dissuaded by high pain— only picked up a single green fruit before ceasing (approx. 8.75% of blocks).

      Crucially, we performed a sensitivity analysis by re-running the model-fitting procedure using only the data collected after the first green pineapple was harvested in each block. This analysis yielded the same qualitative statistical results as the full-block model presented in the main text. We have added a detailed discussion of this caveat and the alternative study designs we explored (such as pre-block stimulation or stochastic choice paradigms) to the Supplementary Material (Section Discussion of pain intensity information and model robustness). We believe this confirms that our current approach provides a faithful representation of the underlying motivational trade-offs.

      (3) The statistical method selected for assessing the association between decision values and pain ratings is problematic (Figure 2G): Since there are multiple data points from multiple subjects, which introduces dependence between data points, a multilevel instead of a single-level linear regression should be employed.

      We appreciate the reviewer’s suggestion to utilise a multilevel modelling approach. We agree that a single-level regression does not fully account for the nested structure of our data.

      In response, we re-analysed the association using a linear mixed-effects model with a maximal random effects structure. Specifically, we included both random intercepts and random slopes for Ratings grouped by Subject (in R syntax: PainFunc ~ Ratings + (1 + Ratings | Subject)).

      The results of this mixed effect model are consistent with our original findings, showing a significant relationship between decision values and pain ratings (p = .001). We have updated the Figure caption (now Figure 3G) to reflect these multilevel model statistics. We believe this addition addresses the concern regarding data dependence and provides a more rigorous validation of our conclusions.

      (4) The statistical method selected for assessing how decision values/pain ratings relate to SCR coefficients is problematic (Figures 3B and C): Again, a multilevel regression method should be used.

      We thank the reviewer for this important point. We agree that a multilevel approach is more appropriate for our nested data structure, and that the interpretation of the SCR data required more explicit justification in the context of the divergence between decision values and ratings.

      We have now re-analysed the relationship between SCR coefficients (both fixationevoked and shock-evoked), decision values, and subjective ratings using a multilevel (mixed-effects) regression model. This model included random intercepts and random slopes for each participant to account for individual variability. We have updated Figure 4 (previously Figure 3) caption and the corresponding Results and Discussion sections to reflect these findings (revised text are copied to the response to next comment (5) below. This more rigorous approach provided a clearer and more nuanced picture of the data. Specifically, while the simple regression previously suggested that both measures correlated with fixation-evoked SCR, the multilevel model reveals a dissociation: fixationevoked SCR is significantly associated with decision values, but not with subjective ratings.

      (5) The interpretation of the skin conductance analysis results as evidence of "dissociation between expected and experienced utility" is vague and not well-supported given the presented data and statistical shortcomings. The low R2 in Figure 2G already indicates divergence between decision values and pain ratings. It is unclear what the decision values' differential association with shock-evoked SCR coefficients adds to this insight.

      The reviewer correctly notes that the low R^2 in the correlation between decision values and pain ratings (Figure 3G) already suggests a divergence between these two measures. We agree that this is one of the key findings, as it highlights that decision values provide a dimension of pain assessment that is not fully captured by subjective report. However, we believe the SCR results add crucial physiological evidence to explain why and how these measures diverge. The updated multilevel results provide a more concrete double dissociation that aligns with the distinction between decision utility and experienced utility:

      Experienced Utility (Shock-evoked SCR): This measure of physiological arousal during the painful event was significantly predicted by subjective pain ratings (beta = 0.0154, p = .006) but not by decision values (p = .672). This suggests that ratings are more closely tied to the immediate, experienced aversiveness of the stimulus.

      Decision Utility (Fixation-evoked SCR): In contrast, arousal during the period of evaluation/fixation was a significant predictor of decision values (beta = -0.0739, p = .009) but was not significantly associated with subjective ratings (p = .105).

      By using a more rigorous statistical method, we found that decision values are actually a more robust predictor of anticipatory/evaluative arousal (fixation) than subjective ratings are. This supports our interpretation that decision values and ratings capture different temporal and functional aspects of pain processing— specifically, the evaluation of potential outcomes (decision utility) versus the reaction to the outcome itself (experienced utility). We have revised the Discussion to be more conservative regarding the strength of this evidence while clearly articulating how these physiological results provide a mechanistic grounding for the divergence observed in the behavioural data.

      Summary of changes in the manuscript:

      Figure 4 Caption: Updated to report multilevel regression statistics (beta, 95% CI, t, and p-values) instead of R^2 from simple linear regression.

      Results Section: Updated the text to describe the mixed-effects model results, highlighting the dissociation between fixation-evoked and shock-evoked SCRs. Revised text:

      “Analysis using a multilevel linear mixed-effects model revealed a clear dissociation in the relationship between physiological responses and motivational parameters. Fixation-evoked SCR coefficients were significantly associated with decision values, but not with subjective pain ratings (Fig. 4B). Conversely, shock-evoked SCR coefficients showed a significant association with subjective pain ratings, while the association with decision values was not significant (Fig. 4C). This double dissociation suggests a notable divergence between the physiological correlates of expected utility (at the decision level) and experienced utility (the actual pain experience). Taken together, these findings highlight the composite nature of the overall aversiveness of pain and underscore the benefit of combining subjective ratings with model-based measures to capture its distinct impacts on behaviour.”

      Discussion Section: Revised the paragraph discussing decision versus experienced utility to include the "further hint" provided by the divergent SCR correlations.

      Revised text:

      “In our task we get a further hint of this in the SCR measures in experiment 1, whereby a discrepancy exists between decision values and pain ratings in their respective associations with fixation-evoked SCRs and phasic pain-evoked (shock) SCRs. Taken together, this indicates the composite nature of overall aversiveness of pain, and highlights the benefit of combining subjective ratings with model-based measures of its motivational impact on behaviour.”

      (6) When investigating the effects of tonic pain on the neural processing of phasic pain (Figure 5), why were only ERPs analyzed and not induced oscillatory responses?

      We thank the reviewer for this insightful suggestion. We initially focused our analysis on Event-Related Potentials (ERPs) because the N1-P2 amplitude is an established and robust marker in pain research, providing a clear and reliable metric for comparing phasic pain processing across conditions.

      However, we agree that induced oscillatory responses provide a more comprehensive view of cortical dynamics. Following your suggestion, we have performed a Time-Frequency Representation (TFR) analysis at electrode Cz. These results, now included in the Supplementary Material (Figure S4, S5), are entirely consistent with our ERP findings. Specifically:

      Phasic Modulation: Both ERP amplitudes and induced oscillatory power (notably in the theta and gamma bands) were significantly modulated by the intensity of the phasic pain stimulus.

      Tonic Independence: Consistent with the ERP results, the presence of background tonic pain did not significantly modulate the induced oscillatory responses to phasic stimuli.

      We believe this additional analysis significantly strengthens the manuscript by demonstrating that the observed effects are consistent across both phase-locked and non-phase-locked neural domains. We have amended the ERP results section to reflect the addition of induced oscillatory responses in supplementary materials: “We focused our neural analysis of phasic pain on ERPs as phasic stimuli are well characterised by these time-locked evoked potentials. Nevertheless, to ensure a comprehensive assessment of the neural response, we also examined induced oscillatory responses. These results were consistent with the ERP findings and are detailed in the Supplementary Materials (Fig. S4, S5).”

      (7) The explanation of the second optimality equation (involving motivational vigour) requires substantial clarification. Besides the points mentioned for the previous optimality equation, specific opportunities to improve the explanations include the following:

      - In the provided formula, C_v and C_m appear indistinguishable given they are multiplied together, rendering this an ill-posed optimization problem. This should be clarified.

      - In my understanding, d(a)/V_speed corresponds to the temporal delay associated with picking fruit a. Then, what is tau, and why compute the sum tau + d(a)/V_speed?

      - V* is not introduced properly. Is V*(s') = Q*(s', a, tau)? If so, why introduce V*? Moreover, the notational similarity between V_speed and V* is confusing.

      - Gamma = 0 still holds?

      - Summarize all parameters and hyperparameters that are optimized to model the data and more precisely describe the method used for optimization.

      We thank the reviewer for these insightful comments. We agree that the transition from a standard reinforcement learning framework to one incorporating motivational vigour requires precise definitions to ensure the model is well-posed and interpretable. We have addressed these points as follows:

      (1) Clarification of C_v and C_m: We have clarified C_m and d(a) in the newly added Experiment 1 model summary table. Specifically, C_v is the scalar vigour constant and C_m is a unit vector representing the horizontal and vertical components. Because C_m is a unit vector, the optimization does not suffer from a collinearity issue from the scalar multiplication between C_v and C_m.

      (2) Bridging Theory to Practice (tau and Total Delay): In the theoretical framework of Niv et al. (2007), "delay" is an abstract sum encompassing both waiting and execution. In practice, when fitting to real-world VR data with variable execution times , we must distinguish between the waiting time tau (time spent stationary or searching) and the execution time (||d(a)|| / V_speed). This is necessary because participants take time to look around the forest to search for fruits before deciding to commit to an action. The sum tau + ||d(a)|| / V_speed represents the total delay between two actions, which directly aligns with the notion of opportunity cost of time. We have added a table (Table 3) and added a new Figure 8 to clarify these distinctions.

      (3) V*, Q*, and gamma: The reviewer is correct that V*(s') = max_{a’, tau’} Q*(s', a', tau'). We previously used V* for simplicity. Since the notation of V* and V_speed was confusing, we have updated the term to max_{a’, tau’} Q*(s', a', tau') in the optimality equation. We confirm that gamma = 0 (a greedy policy) still holds for the Experiment 2 framework to maintain focus on steady-state motivational trade-offs. We have added this statement to the method section.

      (4) Summary of Parameters and Optimization: We have summarized the hyperparameters {k, x_0, C_p, C_v, h, v} in the new summary table for Experiment 2.

      (8) It is not clear what the results of the modelling approach presented in Figure 7a+b concretely add to the comparison of movement velocities and collection rates in Figure 6.

      We appreciate the reviewer's comment regarding the relationship between the raw behavioral metrics and the computational results. While both sets of findings support the argument for reduced motivational vigour in the tonic pain condition, we believe the modeling approach provides distinct and essential value:

      (1) Finer-Grained Analysis Tool: The computational model acts as a more sophisticated analysis tool than simple velocity or rate averages. Unlike Figure 9a+b (in the revised manuscript, previously Figure 7), which summarizes overall performance, the model accounts for the trial-by-trial trade-off between opportunity costs, movement effort, and choice values. This allows us to isolate vigour from other confounding components.

      (2) Direct vs. Indirect Measurement: If we assume that motivational vigour in a free-operant task can be quantified through an RL framework, as established in animal studies, then the model's vigour constant (C_v) serves as a direct, concrete estimate of that internal state. In contrast, overall speed and collection rates are indirect markers that can be influenced by multiple factors, such as different choice sets available to the participants as the fruits locations are randomly generated.

      In summary, the computational approach provides a rigorous, parameterized bridge between observable behavior and the underlying neuro-computational mechanisms of recuperative pain. We have updated the Discussion section to more explicitly state how the computational approach provides a controlled measure that is isolated from the other confounders of the task. Added text to the Discussion:

      “Compared to overall speed and collection rate, which can be influenced by multiple factors, such as different choice sets available to participants as the fruit locations are randomly generated, the model's fitted parameters (e.g. vigour constant C_v) in theory serves as a direct, concrete estimate of that internal state.”

      (9) Claims made in the discussion should be more thoroughly and closely linked to the results presented previously. Specifically, experimental outcomes supporting the following claims should be directly referenced:

      - "tonic and phasic pain serve different motivational functions".

      - "phasic pain provides a punishment teaching signal that directs avoidance".

      - "tonic pain reduces motivational vigour".

      - "these two functions [punishment teaching signals and reduction of motivational vigour?] can be formally distinguished and quantified".

      - "We did not see interactions between tonic and phasic pain".

      We have revised the Discussion to more explicitly link these claims to our experimental results. Revised text:

      “The experiments show that tonic and phasic pain serve different motivational functions during adaptive behaviour, in line with ecological and evolutionary theories of pain (Bolles and Fanselow, 1980; Walters and Williams, 2019). Specifically, our findings point towards phasic pain providing a punishment teaching signal that directs avoidance through value-based learning, balancing the cost of future harm alongside potential reward. This is supported by the observation that increasing phasic pain intensity significantly reduced choice probability and increased distance bias between choices, whereby participants were willing to travel further to reach a pain-free fruit. In contrast, we found that tonic pain reduces motivational vigour, which supports energy conservation and recuperation in the context of bodily damage. This claim is directly evidenced by the reduction in taskrelated movement velocities and fruit collection rates during tonic pain blocks. The experiments are the first to show that these two functions can be formally distinguished and quantified during ongoing behaviour. By utilising a free-operant RL computational framework, we were able to dissociate these roles phasic pain was quantified as a generally negative utility term affecting choice values, while tonic pain was formalised as a change in vigour constants that were significantly higher (increasing delays between actions) in tonic pain condition. This illustrates how pain simultaneously acts in different ways to serve self-protection.”

      “One notable aspect of our results is that we did not see interactions between tonic and phasic pain at either the behavioural or neural level. Behaviourally, we observed that average aversive choice probabilities remained similar regardless of the presence of tonic pain, with no significant interaction effect on punishment sensitivity. Furthermore, our model-fitting confirmed that tonic pain did not significantly modulate the fitted phasic pain utility values. There are two contexts in which these might be predicted. First, in `conditioned pain modulation' paradigms (Kennedy et al., 2016), a tonic pain stimulus is sometimes seen to reduce both the perceived intensity and the cortical evoked responses to phasic pain stimuli delivered somewhere else on the body (Hoffken et al., 2017; Enax-Krumova et al., 2020). Although we utilised concentric ‘Wasp’ electrodes designed to selectively activate nociceptive A-delta fibres (Inui et al., 2002), and confirmed that the resulting ERPs (N1-P2) were significantly modulated by phasic intensity, we observed no such attenuation by tonic pain. Indeed, neither subjective pain ratings nor the N1-P2 amplitude showed a significant modulation by the tonic pressure pain stimulus. In contrast, our results were more compatible with a trend in the other direction.”

      (10) The paragraph in the discussion "A concern that is sometimes raised..." (lines 243 - 254) raises interesting points, but its particular relevance to the study at hand is unclear.

      We appreciate the reviewer's feedback. The motivation for including this discussion is to address a common critique we received for the study: whether the observed reduction in vigour under tonic pain is "simply" due to distraction or cognitive load, rather than being a specific functional output of the pain system. We have revised this paragraph to link the concern to our paper’s specific finding.

      Our central argument is that for tonic pain, distraction is not a confounding "sideeffect" but rather the primary mechanism of action. By being inherently "distracting," tonic pain successfully withdraws resources from ongoing tasks (like foraging) to promote the energy conservation required for recuperation.

      (11) The clinical perspective of the methodological framework presented at the end of the discussion is interesting and could be expanded.

      We thank the reviewer for this encouraging comment. We have expanded the final paragraph of the Discussion to more explicitly state the clinical utility of our framework. Specifically, we now contrast our approach with standard clinical assessments such as Quantitative Sensory Testing (QST). We highlight that while QST is a valuable tool, it can lack ecological validity; in contrast, our VR-based task allows for a more realistic, behaviourally sensitive assessment of how pain impacts a patient’s daily functional activities and motivational state. We believe this represents a significant step towards more objective and "real-world" clinical pain phenotyping.

      (12) The statistical analyses part in the methods section should provide a clear definition of dependent and independent variables and clearly state which test was used for which analysis, e.g., by referencing the corresponding subfigure in the main text.

      We agree that a more structured summary of the statistical approach would improve the clarity of the Methods section. We have now included a comprehensive summary table (Table 1) in the Statistical Analysis subsection. This table explicitly defines the dependent and independent variables for each analysis, identifies the specific statistical model used (e.g. Linear Mixed Models or repeated measures ANOVA), and directly maps these to the corresponding figures in the results section.

      Minor comments:

      (1) Introduction:

      (a) The introduction should elaborate more on the advantages of employing an "ecologically meaningful context".

      We thank the reviewer for suggesting further elaboration on the advantages of employing an "ecologically meaningful context". We have updated the introduction to provide additional reasoning of choosing an ecologically valid context for the study:

      “One of the challenges in studying adaptive functions of pain is the difficulty of embedding experiments within ecologically meaningful contexts. To solve this, we designed an immersive foraging task using virtual reality (VR), in which humans search a forest to collect fruits from the low-lying bushes at varying heights. A foraging paradigm provides a robust, free-operant framework that captures the core components of adaptive behaviour: it is goal-directed, involves complex movement, and requires the learning of an optimal strategy to maximise rewards. This allows us to computationally dissociate how different types of pain influence the control of action.”

      (b) It would be helpful to clarify why tonic pain applied to a limb not involved in the task is expected to influence the motivational vigour with respect to the task.

      We thank the reviewer for pointing out additional clarification for applying tonic pain to the non-dominant arm. We have added the following text to the introduction clarifying our hypothesis and why it was applied to the non-task limb:

      “Second, we hypothesised that tonic pain acts as a coefficient modulating the tradeoff between opportunity cost and vigour cost, thereby serving a recuperative function. To test this in Experiment 2, we delivered continuous tonic pressure to the non-dominant arm via an inflated cuff to emulate a background state of injury. Within our free-operant framework, tonic pain was modelled as a weighting factor that shifts the optimal balance toward reduced energy expenditure. Because the stimulus was applied to the non-task limb, we specifically predicted a global reduction in motivational vigour—operationalised as decreased movement velocities and foraging rates—rather than a direct mechanical impairment.”

      (2) Results/Experiment 1:

      (a) How were monetary rewards implemented exactly? How much money per fruit?

      We thank the reviewer for the opportunity to clarify the incentive structure. Participants were informed at the start of the study that they would earn a performance-based bonus of up to £10, determined by the points they collected during the foraging task. To ensure that motivation remained consistent across the entire session for all individuals—regardless of their baseline foraging speed—the specific exchange rate between points and currency was not disclosed. This prevented potential 'ceiling effects', where a high-performing subject might stop exertive effort after reaching the maximum bonus early, or 'floor effects', where a subject might perceive the reward for an individual action as too small to be motivating.

      Following the completion of the experimental session, all participants were compensated with the full £10 bonus in addition to their base payment for participation. We have updated the Methods section to reflect these details:

      “Participants were informed at the start of the experiment that their total points would be rewarded with a monetary incentive of up to £10. To maintain a constant level of motivation throughout the task, the exact point-to-currency exchange rate was not specified. Upon completion of the session, all participants were awarded the maximum bonus of £10.”

      (b) A green pine apple is not ripe and, in a naturalistic context, possesses some aversive value, even in the absence of phasic pain stimuli. Why was the color coding not counterbalanced across individuals? To what degree could this have confounded the results?

      We thank the reviewer for this insightful point. We acknowledge that the lack of counter-balancing for fruit colour (green vs. yellow) is a limitation of the current study design. However, we believe the potential confounding effect of "unripe" green pineapples on the final analysed data is minimal due to the principles of associative learning.

      While a naturalistic heuristic (green = unripe) might establish a weak prior bias, fundamental associative learning [14] and reinforcement learning models [15] demonstrate that extensive training with a highly salient unconditioned stimulus (such as pain) rapidly overrides mild initial priors. The task objective focused strictly on maximizing reward points, and participants underwent extensive training (10 blocks in Experiment 1; 6 blocks in Experiment 2) before the analysed sessions began. During this time, the strong, explicit contingencies (green = pain, yellow = safe) were learned and verbally verified. Therefore, by the time the main experimental data was collected, any weak baseline aversion to green had been overshadowed by the explicit task contingencies, making the learned associative value the primary driver of behaviour. We have added a statement acknowledging this limitation and outlining this theoretical rationale in the Methods section.

      “While the colour association (green for painful, yellow for pain-free) was not counter-balanced across subjects, any inherent aversive value of green pineapples (e.g., as 'unripe' fruit) is expected to have a minimal confounding effect on the analysed data. In associative learning frameworks, while mild prior biases may influence initial value estimations, extensive training with a highly salient unconditioned stimulus (e.g. phasic pain) rapidly updates these values, driving them toward an asymptote determined entirely by the explicit task contingencies (Rescorla & Wagner, 1972; Sutton & Barto, 2018). Because participants underwent extensive training (10 blocks in Experiment 1 and 6 blocks in Experiment 2) to establish the explicit pain associations prior to the analysed sessions, the observed avoidance behaviour was predominantly driven by the learned phasic pain contingencies rather than baseline colour preferences.”

      (c) In the "Avoidance increases with increasing phasic pain intensity" section, clarify upfront that pain ratings and choice probabilities were estimated at the block level. This information is provided only in a later section.

      We agree with the reviewer that this information should be stated earlier for clarity. We have updated the beginning of the "Avoidance increases with increasing phasic pain intensity" section to specify that these metrics were estimated at the block level:

      “For this analysis, both aversive choice probabilities and subjective pain ratings were estimated at the block level.”

      (3) Results/Experiment 2:

      (a) ERP visualizations (Figure 5) should include standard error indicators.

      We have updated Figure 5 (now Figure 6) to include 95% confidence intervals for standard error of the mean across subjects for all ERP traces. This provides a clearer visualization of the variance in the neural response.

      (b) In the section "A unified model...", clarify what is meant by saying that the unified model is "validated by the behavioural data", since behavioral data is what is being modeled in the first place.

      We clarify that "validation" in this context refers to the consistency between the parameters estimated by our generative unified model and the results obtained from the independent, model-free regression analysis of the raw behavioural data. While both approaches use the same source data, the unified model provides a finer-grained analysis of latent internal states (like motivational vigour), whereas the regression provides a direct empirical benchmark (more details were discussed in the response to major comment (8)). We have rephrased this section to better describe this as a consistency check against empirical regression results.

      (c) In the context of Figure 8a, the term "correlations" is misleading if referring to pairwise comparisons.

      We appreciate the opportunity to clarify our terminology. The results presented in Figure 8a (and the associated text) are derived from a Linear Mixed Model (LMM) where the tonic pain condition was treated as a binary independent variable. The term "correlation" was used to describe the statistical association (represented by the t-values) between the presence of tonic pain and EEG band power, accounting for subject-level random effects. It does not refer to simple pairwise comparisons (like t-tests). However, we agree that "correlation" can be ambiguous when applied to a binary predictor. We have revised the text and figure legends to use the terms "associated with" or "predicted by" to more accurately reflect the LMM framework.

      (d) Based on the presented data, there is no evidence for the section headings claim "Neural activities link to vigour".

      We agree with the reviewer that our results primarily provide evidence for a significant neural association with the tonic pain condition rather than a direct, statistically robust correlation with the vigour parameter itself (after Bonferroni correction). While tonic pain is associated with reduced vigour behaviourally, the EEG markers we identified are more accurately described as signatures of the pain state. We have revised the section heading and the corresponding text to focus on the characterisation of the tonic pain state to ensure our claims are strictly supported by the statistical evidence.

      (4) Methods:

      In the supplementary materials, the headings pertaining to different LMMs are confusing and not consistent with the Figure labeling in the manuscript (e.g., 4(ii)b likely corresponds to Figure 4d).

      We thank the reviewer for identifying these inconsistencies in the supplementary material. We apologize for the confusion caused by the labelling errors during reformatting the manuscript. We have now thoroughly audited the supplementary headings and updated them to ensure they correspond directly and consistently with the figure labels in the main manuscript.

      References

      (1) Inui, K., Tran, T. D., Hoshiyama, M., & Kakigi, R. (2002). Preferential stimulation of Adelta fibers by intra-epidermal needle electrode in humans. Pain, 96(3), 247–252. https://doi.org/10.1016/S0304-3959(01)00453-5

      (2) Mørch, C.D., Hennings, K. & Andersen, O.K. Estimating nerve excitation thresholds to cutaneous electrical stimulation by finite element modeling combined with a stochastic branching nerve fiber model. Med Biol Eng Comput 49, 385–395 (2011). https://doi.org/10.1007/s11517-010-0725-8

      (3) Höffken, O., Özgül, Ö.S., Enax-Krumova, E.K. et al. Evoked potentials after painful cutaneous electrical stimulation depict pain relief during a conditioned pain modulation. BMC Neurol 17, 167 (2017). https://doi.org/10.1186/s12883-017-0946-7

      (4) Enax-Krumova, E., Plaga, A.-C., Schmidt, K., Özgül, Ö. S., Eitner, L. B., Tegenthoff, M., & Höffken, O. (2020). Painful Cutaneous Electrical Stimulation vs. Heat Pain as Test Stimuli in Conditioned Pain Modulation . Brain Sciences, 10(10), 684. https://doi.org/10.3390/brainsci10100684

      (5) Enrico Schulz, Elisabeth S. May, Martina Postorino, Laura Tiemann, Moritz M. Nickel, Viktor Witkovsky, Paul Schmidt, Joachim Gross, Markus Ploner, Prefrontal Gamma Oscillations Encode Tonic Pain in Humans, Cerebral Cortex, Volume 25, Issue 11, November 2015, Pages 4407–4414, https://doi.org/10.1093/cercor/bhv043

      (6) Mahajan Pranav, Tong Shuangyi, Lee Sang Wan, Seymour Ben (2024) Balancing safety and efficiency in human decision making eLife 13:RP101371 https://doi.org/10.7554/eLife.101371.2

      (7) Enrico Schulz, Elisabeth S. May, Martina Postorino, Laura Tiemann, Moritz M. Nickel, Viktor Witkovsky, Paul Schmidt, Joachim Gross, Markus Ploner, Prefrontal Gamma Oscillations Encode Tonic Pain in Humans, Cerebral Cortex, Volume 25, Issue 11, November 2015, Pages 4407–4414

      (8) Suyi Zhang, Hiroaki Mano, Michael Lee, Wako Yoshida, Mitsuo Kawato, Trevor W Robbins, Ben Seymour (2018) The control of tonic pain by active relief learning eLife 7:e31949

      (9) Hewitt, D., Tong, S., Schreiber, S., & Seymour, B. (2026). Tonic pain modulates neural correlates of associative phasic pain memories. PAIN. DOI: 10.1097/j.pain.0000000000003917

      (10) Gramann, K., Gwin, J. T., Ferris, D. P., Oie, K., Jung, T.-P., Lin, C.-T., Liao, L.-D., and Makeig, S. (2011). Cognition in action: imaging brain/body dynamics in mobile humans. Reviews in the Neurosciences, 22(6):593–582.

      (11) Klug, M. and Gramann, K. (2021). Identifying key factors for improving ica-based decomposition of eeg data in mobile and stationary experiments. European Journal of Neuroscience, 54(12):8406–8420.

      (12) Delorme, A. EEG is better left alone. Sci Rep 13, 2372 (2023). https://doi.org/10.1038/s41598-023-27528-0

      (13) Bach, D. R., Flandin, G., Friston, K. J., and Dolan, R. J. (2010). Modelling event-related skin conductance responses. International Journal of Psychophysiology, 75(3):349–356.

      (14) Rescorla, R. and Wagner, A. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement, volume Vol. 2

      (15) Sutton, R. S. and Barto, A. G. (2018). Reinforcement learning: An introduction, 2nd ed. Adaptive computation and machine learning. The MIT Press, Cambridge, MA, US.

    1. Author response:

      (1) Clarification of the scope of the present study and future mechanistic analyses

      We agree that the downstream molecular mechanisms by which SOX17 regulates Sertoli valve formation remain to be elucidated. Our findings are consistent with a model in which SOX17 regulates Sertoli valve formation through paracrine signaling; however, the downstream effectors have not yet been identified. Despite extensive analyses of Sox17 conditional knockout and wild-type mice, including single-cell RNA sequencing, identifying the downstream molecular targets of SOX17 has remained challenging (Uchida et al., 2022). The transgenic mouse model generated in the present study now provides a valuable experimental platform for investigating SOX17-dependent molecular pathways. We are currently performing transcriptomic analyses using this model to identify candidate downstream pathways and genes regulated by SOX17. However, further investigation will be required to determine whether these candidates represent direct transcriptional targets of SOX17 and whether they function specifically within the rete testis during Sertoli valve formation.

      Accordingly, we will avoid overinterpreting the molecular mechanisms in the present study and will revise the Discussion to more clearly acknowledge these limitations while emphasizing that elucidation of these mechanisms represents an important direction for future research. We therefore believe that a comprehensive mechanistic analysis is beyond the scope of the present study.

      (2) Clarification of the quantitative methodology

      We will provide a more detailed description of the methodology used for Sertoli cell quantification. Specifically, Sertoli cells were counted within the SV region extending 100 μm from the rete testis (RT) boundary, and Sertoli cells protruding into the RT lumen were also included in the analysis. The sampling procedure for sagittal RT-SV-seminiferous tubule (ST) sections will be described more explicitly in the revised Methods to improve reproducibility.

      (3) Clarification regarding expression levels

      We appreciate the reviewer's comment regarding the quantitative assessment of SOX17 and other SV-associated molecules.

      The Sertoli valve (SV) is an extremely small transitional structure, with only approximately 20 SVs present in each mouse testis. In addition, Sertoli cells within the SV are tightly interconnected. Consequently, selectively isolating the SV without contamination from adjacent tissues while obtaining sufficient material for quantitative molecular analyses, such as quantitative PCR, remains technically challenging.These technical limitations partly explain why the Sertoli valve has remained an understudied structure in testicular biology. Therefore, in the present study, the expression of SV-associated molecules was primarily evaluated by histological and immunohistochemical analyses. We will clarify these technical limitations in the revised manuscript and revise the relevant text accordingly.

      (4) Additional revisions

      We will address the remaining comments, including clarification of the phenotypic differences between Tg26 (established line) and Tg27 (F0), standardization of gene nomenclature, correction of methodological descriptions, and improvements to the Discussion and figure presentation where appropriate.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Li and colleagues examines how defensive responses to visual threats during foraging are modulated by both reward level and social hierarchy. Using a naturalistic paradigm, the authors test how the availability of water or sucrose, with sucrose being more rewarding than water, shapes escape behavior in mice exposed to looming stimuli of different intensities, which are used to probe perceived threat level and defensive responses. In parallel, the study compares dominant and subordinate animals to assess how social rank biases the trade off between reward seeking and threat avoidance. By combining detailed behavioral analyses with computational modeling, the work addresses how reward level and social context jointly influence escape decisions in an ethologically relevant setting.

      Across the different experimental conditions, perceived threat level is the main determinant of behavior. The authors show that looming stimuli associated with higher threat (contrast) consistently elicit faster and more robust escape responses than lower threat stimuli. This effect is particularly evident during early exposures, when animals are highly vigilant and have not yet habituated to the looming stimulus (learned that it is not dangerous). Later they described that as animals gain experience and habituate, behavior becomes more flexible, and reward level begins to exert a graded modulation of the escape response. Importantly, the authors show that under high threat conditions increasing reward value leads to more frequent and faster escape rather than greater reward pursuit. This finding is particularly relevant, as it suggests that highly valued rewards can heighten vigilance and thereby enhance responsiveness to threat, highlighting that reward does not simply compete with defensive behavior but can also reshape it depending on the perceived level of danger, in contrast to low threat conditions, where threat can be more easily outweighed by reward. Thus, an important conceptual contribution of the study is the introduction of vigilance as a useful framework to interpret these effects. Vigilance is treated as a behavioral state reflecting heightened attention to potential danger. In line with what is known from natural foraging, mice initially maintain high vigilance when confronted with an innate threat. This perspective helps clarify a finding that might otherwise appear counterintuitive. One might expect higher rewards to motivate animals to tolerate risk, explore more, and habituate faster in any scenario. Instead, the data suggest that highly rewarding outcomes can elevate vigilance, making animals more responsive to threat and leading to faster or more frequent escape under high threat conditions. In this sense, reward does not simply compete with threat but can also amplify sensitivity to it, depending on the internal state of the animal.

      The social results are particularly interesting in this context as well. Dominant mice consistently prioritize avoidance over reward, showing stronger escape responses and slower habituation than subordinates. This behavior is well captured by the vigilance framework proposed by the authors: dominant animals appear to maintain higher vigilance, which biases decisions toward threat avoidance. The authors further suggest that stable social relationships sustain high vigilance and slow habituation, framing this as an evolutionarily conserved strategy that may enhance survival. This interpretation provides a valuable perspective on how social structure shapes defensive behavior beyond immediate physical interactions. At the same time, there are important limitations to this interpretation. All experiments were conducted in male mice, and it is possible that the relationship between social hierarchy, vigilance, and defensive behavior would differ substantially in females. In addition, the idea that stable social relationships maintain elevated vigilance does not straightforwardly align with broader views of social stability as protective for mental health and as a buffer against anxiety and stress. These points do not undermine the findings but suggest that the social effects described here should be interpreted with caution and within the specific context of the task and sex studied.

      We thank the reviewer for raising this important point. In the context of repeated looming exposure, slower habituation reflects more sustained vigilance over time. Compared to individually housed mice, group-housed mice exhibit slower habituation (Lenz et al., 2022), and pair-housed mice showed even slower habituation in our current work. Importantly, this pattern does not indicate that pair-housed mice have higher overall vigilance than individually housed animals. Although individually housed mice habituate more quickly, they display higher initial vigilance, as reflected by their increased probability of escaping in response to looming stimuli (Lenz et al., 2022). Thus, pairhoused mice exhibited reduced defensive responses compared to individually housed animals, consistent with a social buffering effect.

      Furthermore, in a separate study (Rank- and Threat-Dependent Social Modulation of Innate Defensive Behaviors; Li, Gao, Li, 2026, eLife 15:RP109571), we directly compared responses to looming stimuli when mice were tested alone versus in the presence of a social partner and observed clear evidence of social buffering.

      Another important limitation is that the neural mechanisms underlying these effects remain speculative. The manuscript includes an extensive discussion of candidate circuits, particularly involving the superior colliculus and downstream structures, but this section is necessarily based on prior literature rather than on data presented in the study. Given the complexity of the circuits involved in integrating internal state, reward, social context, and vigilance, the current work should be viewed as providing a strong behavioral and conceptual framework rather than direct insight into underlying neural mechanisms.

      We fully agree that the proposed neural mechanisms remain speculative and that the circuits involved in integrating internal state, reward, and social context are likely far more complex. We have revised the manuscript to acknowledge this limitation.

      Methodologically, the behavioral paradigm is well suited for studying escape decisions in socially housed animals, and the machine learning based classification of defensive responses is a clear strength. The computational model provides a useful formalization of how threat level, reward level, and vigilance interact and may be valuable for other laboratories studying escape, approach avoidance, or conflict situations, particularly as a way to classify behavioral outcomes after pose estimation. More generally, the work will be of interest to the neuroethology community for its detailed characterization of escape behavior under naturalistic conditions.

      Given the ethological nature of the study and the high inter individual variability reported by the authors, clarity and precision in the methods are especially important for reproducibility. While the revised manuscript addresses many earlier concerns, some aspects remain slightly difficult to follow. For example, the main text states that animals were not water deprived to avoid differences in internal state, whereas parts of the methods describe conditions in which animals were water deprived, suggesting that internal state manipulation may differ across experiments. Clearer separation and explanation of these conditions would further strengthen confidence in the work.

      To improve clarity, we have revised the Methods section to clearly distinguish between experimental conditions that involved water deprivation and those that did not.

      Overall, this study provides a rich and thoughtful analysis of how reward level and social hierarchy modulate defensive behavior through changes in vigilance. It offers a useful conceptual advance for thinking about escape behavior in naturalistic settings and lays a solid foundation for future work aimed at linking these behavioral states to underlying neural circuits.

      Reviewer #2 (Public review):

      Zhe Li and colleagues investigate how mice exposed to visual threats and rewards balance their decisions in favour of consuming rewards or engaging in defensive actions. By varying threat intensity and reward value, they first confirm previous findings showing that defensive responses increase with threat intensity and that there is habituation to the threat stimulus. They then find that water-deprived mice have a reduced probability of escaping from low contrast visual looming stimuli when water or sucrose are offered in the environment, but that when the stimulus contrast is high, the presence of sucrose or water increases the probability of escape. By analysing behaviour metrics such as the latency to flee from the threat stimulus, they suggest that this increase in threat sensitivity is due to increased vigilance. Analysis of this behaviour as a function of social hierarchy shows that dominant mice have higher threat sensitivity, which is also interpreted as being due to increased vigilance. These results are captured by a drift diffusion model variant that incorporates threat intensity and reward value.

      The main contribution of this work is quantifying how the presence of water or sucrose in water-deprived mice affects escape behaviour. The differential effects of reward between the low and high contrast conditions are intriguing, but I find the interpretation that vigilance plays a major in this process not supported by the data. The idea that reward value exerts some form of graded modulation of the escape response is also not supported by the data. In addition, there is very limited methodological information, which makes assessing the quality of some of the analyses difficult, and there is no quantification on the quality of the model fits.

      (1) The main measure of vigilance in this work is reaction time. While reaction time can indeed be affected by vigilance, reaction times can vary as a function of many variables, and be different for the same level of vigilance. For example, a primate performing the random dot motion task exhibits differences in reaction times that can be explained entirely by the stimulus strength. Reaction time is therefore not a sound measure of vigilance, and if a goal of this work is to investigate this parameter, then it should be measured. There is some attempt at doing this for a subset of the data in Figure 3H, by looking at differences in the action of monitoring the visual field (presumably a rearing motion, though this is not described) between the first and second trials in the presence of sucrose. I find this an extremely contrived measure. What is the rationale for analysing only the difference between the first and second trials? Also, the results are only statistically significant because the first trial in the sucrose condition happens to have zero up action bouts, in contrast to all other conditions. I am afraid that the statistics are not solid here. When analysing the effects of dominance, a vigilance metric is the time spent in the reward zone. Why is this a measure of vigilance? More generally, measuring vigilance of threats in mice requires monitoring the position of the eyes, which previous work has shown is biased to the upper visual field, consistent with the threat ecology of rodents.

      (2) In both low and high contrast conditions, there are differences in escape behaviour between no reward and water or sucrose presence, but no statistically significant differences between water and sucrose (eg: Figure 3B). I therefore find that statements about reward value are not supported by the data, which only show differences between the presence or absence of reward. Furthermore, there is a confound in these experiments, because according to the methods, mice in the no-reward condition were not water-deprived. It is thus possible that the differences in behaviour arise from differences in the underlying state.

      (3) There is very little methodological information on behavioural quantification. For example, what is hiding latency?

      Is this the same are reaction time? Time to reach the safe zone? What exactly is distance fled? I don't understand how this can vary between 20 and 100cm. Presumably, the 20cm flights don't reach the safe place, since the threat is roughly at the same location for each trial? How is the end of a flight determined? How is duration measured in reward zone measures, e.g., from when to when? How is fleeing onset determined?

      (4) There is little methodological information on how the model was fit (for example, it is surprising that in the no reward condition, the r parameter is exactly 0. What this constrained in any way), and none of the fit parameters have uncertainty measures so it is not possible to assess whether there are actually any differences in parameters that are statistically significant.

      These are the public reviews for the original submission. The corresponding authors responses are provided below.

      (1) We agree that reaction time can be influenced by multiple factors, including stimulus strength. Consistent with this, reaction times (i.e. latencies to flee) were substantially shorter under high-contrast conditions (Figure 3E). However, even under the same high-contrast condition, reaction times were significantly shorter in the water condition compared to the no-reward condition, suggesting that other factors such as vigilance may contribute.

      Upward-directed attention includes rearing, up-stretching, and upward head orientation, which will be clarified in the Method section. To address concerns about statistical validity, we will quantify these behaviors across the first 10 trials rather than limiting the analysis to the first two.

      As for the dominance-related results, we interpret them as reflecting both enhanced vigilance and reduced reward-seeking behavior. Time spent in the reward zone is not a measure of vigilance but an indicator of reward-seeking motivation. We will clarify this in the revised manuscript.

      (2) In Figure 3B, the difference between water and sucrose conditions did not reach statistical significance (p = 0.08). We plan to collect additional data to determine whether this is due to limited statistical power. It is also possible that some behavioral readouts are more sensitive to the differences between water and sucrose conditions. For example, Figure 3F shows that escape speed was significantly higher in the sucrose than in the water condition under high-contrast stimulation.

      Thank you for pointing this out. To control for the potential confounds related to internal state, mice were not water-deprived under any of the three conditions in Figures 3A-3H. We will clarify this in the main text and Methods. For Figures 3I-3M, which compare decision-making under no-reward and water conditions, we will conduct additional experiments using non-deprived mice in the water condition.

      (3) Hiding latency was defined as the time from stimulus onset to the animal’s arrival at the safe zone. Reaction time was quantified as the latency to flee, measured from stimulus onset to the initiation of the first flight state. The flight state was defined as locomotion exceeding 10 cm at a speed greater than 10 cm/s. Distance fled was defined as the distance covered between stimulus onset and offset for all trials. However, in trials classified as no reaction or freezing, this measure does not accurately reflect escape behavior. We will therefore rename it as distance under threat to better capture its meaning. The reward zone was defined as the region within 15 cm of the reward port at the end of the arena. Duration in the reward zone was measured as the time spent within this region during the 20 seconds following stimulus onset. In Figure 4E, the percentage of time spent in the reward zone was calculated relative to the total time the mouse remained in the arena during the 2-hour social session.

      All definitions and additional details on behavioral quantification will be included in the revised Methods section.

      (4) We appreciate the comment and agree that further clarification is needed. We will provide a more detailed description of the model fitting procedure in the revised Methods section. Specifically, the drift rate parameter (r), which reflects the perceived reward value, was constrained to zero in the no-reward condition. To enable statistical comparison across conditions, we will report uncertainty measures for all fit parameters.

      Comments on the revised manuscript:

      The manuscript has been revised and improved significantly by the addition of methodological details and new analysis. I remain, however, unconvinced by the argument that increased vigilance in the presence of reward leads to heightened escape behaviour.

      In response to my criticism that the work does not measure vigilance directly, the authors have included measures of foraging interval and foraging speed, which they state are "two direct behavioral analyses of vigilance". I disagree - like reaction time, foraging speed and foraging interval can be modulated, for example, by changes in threat sensitivity. Increased threat sensitivity comes with diverse behavioral changes that may well include increased vigilance, but foraging interval and foraging speed can certainly change without the animal expressing increased vigilance behaviors. A bigger issue I still have though, is with the conclusion that the presence of reward increases "direct escape behaviors". Comparing the no reward, water and sucrose groups indeed shows a difference (which is now clear after the split into early and late phases), but the issue is that these are different mice. As the text is written, is sounds like introducing reward will acutely increase escape. But if we look at the raw data show in Figure 2C, what I think is happening is that the presence of reward is decreasing habituation to the stimulus. The data for trials 1 and 10 in the three conditions show this - there is habituation with no reward (reaction times are all shifting to the right), a bit less with water and very little with sucrose. This is interesting in its own right and we can speculate why it might be happening, but I think this is conceptually different from what the authors are proposing.

      We agree that vigilance is not directly observable as a single variable. Our intent was not to claim that foraging speed and foraging interval provide a direct measure of vigilance, but rather to suggest that they may serve as indirect behavioral correlates.

      We also considered an alternative interpretation: these two measures could reflect perceived reward value under high-threat conditions across distinct reward types. If that were the case, animals would be expected to exhibit shorter intervals and faster speeds across no reward, water, and sucrose conditions. However, our data do not support this interpretation (Figures 3L and 3M), suggesting that these measures are more likely correlated with vigilance.

      Furthermore, it is unlikely that changes in foraging interval and speed are driven by altered threat sensitivity, as animals could not see the threat during most of the foraging bout and only encountered it at the end.

      Regarding the conclusion that the presence of reward increases direct escape behaviors, our interpretation is that increased reward value reduces habituation, thereby maintaining higher vigilance during the late phase. This was discussed in the second-to-last paragraph of the "Economic and social modulations of innate decision-making under threat" subsection in the Discussion.

      Reviewer #3 (Public review):

      Male mice were tested in a classic behavioral "flee the looming stimulus" paradigm. This is a purely behavioral study; no neural analyses were done. Mice were housed socially, but faced the looming stimulus individually, using an elegant automated tunnel (see videos for clarity).

      The additional changes made to the paper clarify the work done. While there are some limitations (male mice, weird stimulus), the general results are interesting and a valuable addition to the experimental literature. The main claim of the paper is that the different rewards (none, water, sucrose) did not change the escape properties early in learning, but did late, particularly that in the late (already experienced) conditions, reward value (assuming sucrose > water > no reward) interacted with the salience of the looming stimulus (light gray, dark gray). (Panels 3D, 3G, 3K, 3N).

      For readers, I want to note that one of the most interesting results is actually in Figure S2, where they find that a looming stimulus behind the mouse still makes a mouse run to the nest. In these conditions, the mouse runs past the looming stimulus to get to safety! (I also do love the video of the mouse running around the barriers like a snake to get home.)

      I have a few minor clarification questions and a few notes that I think would be useful additions for authors and readers to think about.

      Dominance: What does the mouse social science literature say about the "test tube" test? What can we conclude from this test? This would be useful when trying to understand what is causing the dominance/submissive difference in responses. Figure 4 shows that the dominant mice are more risk-averse than the submissive mice. Is "dominance" in the test-tube actually a measure of risk-seeking? Is the issue that the submissive mice don't think they can get back to the food-site easily, so they are less willing to sacrifice the current (if dangerous) foraging opportunity? Is the issue that the submissive mice can't get back to the nest? As I understand it, the nest was always available to all the mice, so I suspect inability to get to the nest is an unlikely hypotheses. Is the issue that the submissive mice also don't feel safe in the nest?

      The tube test is a widely used assay in the rodent social behavior literature to assess dominance hierarchies, operationally defined by the ability of one animal to force its opponent to retreat from a narrow tube. Importantly, this assay does not directly measure risk-seeking or anxiety-related traits, but rather competitive outcomes during social conflict. Furthermore, our data indicate that the behavioral responses of subordinate mice to looming stimuli are primarily driven by the visual threat itself rather than by social avoidance. This point was elaborated in the second paragraph of the “Social modulation of innate decision-making” subsection in the Results section.

      Limitations of the study: There is an acknowledged limitation to male mice, and the limitations of the small data sets that are typical of such experiments. In addition, however, it is also worth noting the strangeness of the looming stimulus, which is revealed clearly in the videos. The stimulus is a repeating growing circle, growing in a single location within the environment. The stimulus repeats 10 times, once per second. This is not what an attacking hawk or owl would look like. (I now have this image of an owl diving down, and then teleporting up and diving down again.) Note - I am fine with this stimulus. It produces an interesting experiment and interesting results. I do not think the authors need to change anything in their paper, but readers need to recognize that this is not a "looming predator".

      These "limitations" are better seen as "caveats" when folding these results in with the rest of the literature that has gone before and the literature to come. (Generally, I do not believe that science works by studies making discoveries that change how we think about problems - instead, science works by studies adding to the literature that we integrate in with the rest of the literature.) Thus, these caveats should not be taken as problems with the study or as fixes that need to be done. Instead, they are notes for future researchers to notice if differences are found in any future studies.

      Thus, my only suggestion is that I think authors could write a more careful paper by using the past and subjunctive tense appropriately. Experimental observations should be in past tense, as in "the influence of reward was contextdependent and emerged in the late phase" instead of "the influence of reward is context-dependent and emerges in the late phase" - it emerged in the late phase this once - it might not in future experiments, not due to any fault in this experiment nor due to replicability problems, but rather due to unexpected differences between this and those future experiments. At which point, it will be up to those future experiments to determine the difference. Similarly, large conclusions should be in the subjunctive tense, as in "these data suggest that threat intensity is likely to be the primary determinant of decision making" rather than "threat intensity is the primary determinant of decision making", because those are hypotheses not facts.

      We thank the reviewer for the helpful suggestions and have revised the Abstract accordingly.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Figure 5: The points in panel 5G and 5H are unreadable. What are these stars and symbols supposed to mean? They are also too small to see without zooming way in.

      We have increased the symbol size.

      Figure 5: What is the final panel of 5J? I did not understand this panel at all. The first three panels of 5J (threat-based detection, reward-based detection, vigilance-based detection) are, I believe, three patterns we should look for in the data. But then what is the "experimental results" section? It contains all three, but they don't overlap? Shouldn't we have an experimental results section for each condition?

      Panel 5J was to compare three hypothesized decision patterns with the experimentally observed data. To make this distinction explicit, we have revised the panel titles to: “H1: Threat-based decisions,” “H2: Reward-based decisions,” “H3: Vigilance-based decisions,” and “Experimental results.”

      Thank you for including the videos. They made the task construction and the stimulus much clearer.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      While the evidence in favour of the two gradients largely supports the claims, the evidence for a new visual field map cluster in the anterior temporal lobe falls short of the level used historically when identifying visual field maps in the visual cortex and is, at present, not convincing. More specifically, the progressions of polar angle within the putative anterior lobe cluster are highly variable across subjects. Few subjects have convincing polar angle reversals at either the horizontal or vertical meridians. In other cases, a putative border is shown that spans different polar angles, which does not align with the accepted definitions for visual field maps in the cortex.

      We agree with the reviewer that more evidence could be provided in support of retinotopic representations within the anterior temporal lobe. We have performed a number of new analyses to further explicate the receptive field properties of this anterior temporal lobe visual representation. We have pasted updated Figure 2e-i. We have added additional participants, increasing the total number from N=12 to N=21. In panel g, we show that in this larger group, we can still observe pRFs that are about 3x larger than those in early visual cortex, and that the relationship between their size and eccentricity shows the expected steeper slope compared to these early representations. In this new participant group, we also illustrate the visual field coverage of the left and right anterior temporal lobe representations (panel h). As expected, the left hemisphere pRFs largely sample the right visual field, and right hemisphere pRFs largely sample left visual space. One can also see that both the upper and lower visual fields are sample quite evenly, consistent with the hemi-field representation of visual field maps observed in earlier visual cortex. To quantify whether there is a left-right contralateral bias in the sampling of visual space (and to test whether such a bias is significantly different in each hemisphere), we calculated for each pRF a laterality index as previously defined by Sheremata and Silver (2015) according to the equation below:

      Where resulting values of 1 mean the pRF is contralateral, 0.5 is no laterality bias, and 0 is ipsilateral bias. Additionally, we input pRF sigma values that were adjusted for the non-linearity exponent as defined by Kay et al. (2013). For the purposes of visual comparison, we subtracted 0.5 from index values so that resulting laterality scores were relative to 0 to represent the center of the visual field, and then values were inverted with a -1 scalar so that left hemisphere pRF laterality index values are plotted on the right side of space, and the right hemisphere on the left as shown in panel i. The laterality index was calculated for each pRF for a given participant and then averaged within that participant to result in a single mean laterality index for the left hemisphere pRFs and a single index for their right hemisphere pRFs. The histograms illustrated in panel i depict density of participants (kernel smoothed). We find a significant difference between laterality indices with left AT pRFs showing significantly rightward index values compared to right AT pRFs (paired-samples t-test, t(20) = 7.6, p = 2.7 x10<sup>-7</sup>). These data thus offer stronger evidence of a hemifield representation with a contralateral bias, and it should also be noted that there is stronger ipsilateral coverage in these high-level visual pRFs compared to earlier visual field maps like V1, which is consistent visual field maps in latera stages of the visual processing hierarchy as quantified by Mackey et al. (2017).

      Lastly, we note that the progression of polar angle values on the cortical surface is certainly not as strikingly topographic as in visual field maps V1 through hV4. This is perhaps a result of the strong ipsilateral visual field coverage in which pRFs whose centers were near or within the ipsilateral field (especially those near the fovea) are not visualized appropriately when using a contralateral colormap. It is also possible that at this very late stage of visually-responsive cortex within entorhinal cortex that retinotopic topography becomes less clear as is the case in higher stages of the dorsal visual stream. To improve visualization, we have created a new Supplemental Figure 6 using a binary color map that colors lower and upper visual field in separate colors and extends into the ipsilateral visual field (pasted below for convenience). We hope that this color map helps to show the upper and lower visual field coverage. While there is a clear radial eccentricity gradient within these AT pRF clusters, and while most participants do show a polar angle gradient that runs perpendicular to this radial eccentricity gradient as expected for a visual field map, we do agree that it is difficult to observe polar angle traversals as clearly as in earlier visual cortex. Nonetheless, the presence of these pRF clusters which show their own distinct eccentricity representation (i.e., a foveal confluence) and a full sampling of the contralateral visual space is still consistent with our anatomical model’s prediction in which PC2 anchor points predict foveal representations shared by visual field map clusters. While the topographic clarity of these representations on the cortical surface is less than earlier visual cortex, the existence of contralateral representations of visual space with a full eccentricity gradient that spans the upper and lower visual field is strongly supported by the data and consistent with our anatomical model’s prediction that there should have been a distinct eccentricity gradient. These findings are also consistent with work showing that the human hippocampus also shows sensitivity to contralateral visual space (Silson et al., 2021) and suggests the hippocampus may inherit this contralateral bias from this entorhinal visual representation. We have updated the manuscript to incorporate these new findings, and refer to these AT clusters as contralateral visual representations, remaining agnostic to whether or not they can be fully defined as topographic maps which can be the focus of future work using smaller voxel sizes to better capture small topographic gradients.

      We have revised the manuscript to incorporate these points in the following sections.

      Line 466: “We performed pRF mapping on 21 participants with high-contrast, …”

      Line 601-625: “To produce maps of visual field coverage (Figure 2h) similar to previous work, … The histograms illustrated in Figure 2i depict density of participants (kernel smoothed).”

      Line 236-246: “We find that consistent with its high position within the processing hierarchy, … We find a significant difference in laterality indices between left and right AT pRF’s (pairedsamples t-test, t(20) = 7.6, p = 2.7 × 10-7).”

      Line 373-383: “The organization of polar angle in anterior temporal cortex was not as orderly as earlier visual cortex, … in more posterior portions of ventral occipitotemporal cortex.”

      Reviewer #2 (Public review):

      Weaknesses:

      (1) The neurobiological model does not take into consideration present knowledge about the microstructural organization of the visual system. This limits the way the results are interpreted correctly. Critical information on the layer-specific myeloarchitecture and cytoarchitecture (and their relation to cortical thickness), as explored for example by Sereno et al. 2013 Cereb Cortex, is missing. There is no information given with respect to how different visual areas differ in their microstructural profile. It is also not mentioned that cortical parcellation is indeed characterized by sharp boundaries between areas, rather than structural gradients, so it remains unclear why focusing on a gradient is of interest. The authors cite the parcellation atlas by Glasser et al. 2016, but do not discuss the rationale of this publication, which was not the definition of gradients, but the definition of sharp boundaries for cortex parcellation. Indeed (as explained below), the results of the authors seem to a large extent to be driven by cortex parcellation, but instead of acknowledging this fact, the authors write (line 179) that "we hypothesize that these local deviations from the canonical thickness and density of cortex underlie the finer-scale division of visual cortex into categorically distinct regions. That is, does the realization of the cortex into distinct regions involve these regions becoming more distinct from a prototypical cortical sheet (i.e., gradient 1)?" - While the first sentence is reasonable, the second sentence is pure speculation ignoring present knowledge on cortical parcellation of this area according to which there is no "prototypical cortical sheet", but each area has its distinct microstructural profile.

      We thank the reviewer for this important comment. We first want to point out that we believe there is a conceptual misunderstanding on the part of the reviewer, as we address in our lengthy response below. In this response, we explain that our findings capture what we believe is a novel finding—that variation across participants in the cortical sheet is not random across the spatial expanse of cortex but respects its functional boundaries—which we view as a finding that is complimentary to the current knowledge about the microstructure of visual cortex. It was not our intention to ignore or gloss over this present knowledge, but instead show that variation in these cortical microstructures across brains is not random.

      We agree that incorporating current knowledge about the microstructural organization of visual cortex, including its laminar architecture and sharp areal boundaries, is critical for situating our findings within the broader literature. In response, we have added key background information on the relationships among cytoarchitecture, myeloarchitecture, and cortical thickness, as described in previous studies (for example, Maingault et al., 2021; Sereno et al., 2013; Shafee et al., 2015). While our study does not aim to capture layer-specific properties per se, which would require different imaging modalities and higher-resolution data, we focus on spatial properties tangential to the cortical surface.

      We first address a concern that the particular parcellation might be driving effects with an analysis showing that we believe our finding is robust to this concern. As suggested by the overall negative covariance observed between cortical thickness and tissue density, we further confirmed this relationship not only across larger visual ROIs, which could potentially reflect effects of arealization, but also within individual ROIs at a finer spatial scale. To avoid potential circularity in ROI definition, we used a visual ROI atlas derived from population-level retinotopy based on independent datasets (Abdollahi et al., 2014). We found that at the global level, cortical thickness and T1w/T2w ratio showed a strong negative correlation across visual ROIs (Fig. 3, revised Supp. Fig. 3a & b). Although only a portion of the visual cortex is clearly delineated in this atlas, we replicated similar results across the entire visual cortex using the MMP atlas (Glasser et al., 2016). At the within-ROI level, we found robust negative correlations between cortical thickness and T1w/T2w ratio across most visual ROIs in both hemispheres, with the notable exception of V1, V2 and VO1, which exhibited a positive relationship, consistent with prior work (for example, Maingault et al., 2021; Sereno et al., 2013; Shafee et al., 2015). These results highlight both common and distinct microstructural profiles across the visual cortex and provide important context for interpreting our data-driven findings.

      We also want to address what we think is a conceptual misunderstanding by the reviewer, which likely resulted from a lack of clarity on our part. The reviewer’s confusion likely results from the fact that we theoretically “transposed” the typical PCA analysis such that we get a subject-wise contribution (PC loadings) per participant (also see response to next point), which is how we’re able to relate inter-participant variability in their loadings to behavior in Figure 3. This is also why we refer to a “typical” cortex/cortical sheet because the surface maps being visualized for PC2 can be thought of as a map explaining variance of deviation orthogonal to PC1 (which captures the primary relationship between thickness and T1/T2). Thus, because PC2 is orthogonal to PC1, it captures the spatial pattern in which participants deviate from the primary relationship (e.g., the typical relationship). Therefore, if a given participant is far from the PC1 vector and has high PC2 loading, their cortical sheet is either thicker or more myelinated than predicted by the PC1 relationship and is therefore more distinct from the “typical” or “average” cortical sheet values captured by PC1. We want to emphasize that PCA is agnostic to spatial structure across the cortex. Thus, the fact that deviation from the primary thickness-myelination relationship (i.e. PC2) captured by PC1 had any spatial structure at all is interesting. Furthermore, the fact that the spatial structure of PC2 across the cortical sheet seems to separate visual cortex into its constituent processing streams is also interesting. Therefore, we are not speculating but rather describing the PCA model itself whereby a participant’s loading on PC2 describes their deviation or distinctness from the PC1 relationship. The fact that PC2 has spatial structure on the cortical sheet (which did not have to be true) and the fact that this structure seems to capture broad borders between visual processing streams and field maps is what we find interesting and quantify within the paper. We hope this additional explanation clarifies the broader theoretical thrust of the paper. We view these findings as complimentary to the present knowledge of the microstructural organization of the visual system. Our findings suggest that variability in these microstructural features across participants (PC2) don’t occur randomly across cortex but seem to respect the functional borders of the neural populations of the underlying cortical sheet.

      Regarding the concern that our gradient approach may contradict established knowledge of cortical arealization, we would like to clarify that the primary goal of our gradient analysis is not to redefine visual areas, or to go against cortical arealization, but to explore the continuous variation in cortical architecture across brains that may co-exist alongside sharp boundaries which is phenomenon complementary to the arealization. In our study, cortical thickness maps were regressed for curvature before entering any analyses, given the covariance between cortical folding and area borders (Fischl et al., 2008). We acknowledge that cortical parcellation is traditionally characterized by discrete transitions between areas. However, our results suggest that gradients of cortical properties—particularly those shared across participants—may capture supra-areal organizing principles that reflect how distinct regions relate to one another within a broader cortical sheet.

      Finally, we agree with the reviewer that the phrase “prototypical cortical sheet” was speculative and potentially misleading. We have removed this language from the manuscript and revised the corresponding discussion.

      We have revised the manuscript to incorporate these points in the following sections.

      Line 92-94: “Thickness and density maps showed a robust anti-correlation both at the coarse across-area level based on an independent parcellation and at the finer within-area level, except in primary regions (Figure S3a, b).”

      Line 350-353: “The convergence pattern, arising from the negative correlation between thickness and density, is consistent with previous findings and may support the balloon model, whereby cortical thinning is associated with tangential stretching due to myelination.”

      Line 188-189: “That is, does the arealization of cortex into distinct regions involve these regions becoming more distinct from a typical cortical sheet (i.e., gradient 1)?”

      (2) Instead of building on present, detailed knowledge of brain anatomy and in-vivo cortex parcellation of the visual system and its known relation to visual maps, the authors focus on two metrics of cortex architecture (mean T1/T1 over depth and cortical thickness), and conduct a PCA to explore their shared variance. It needs to be clarified if the PCA was conducted correctly. There is no mention of standardizing the variables, which could bias the results. In addition, in a PCA, all possible features are categorized as vector components, and those are scanned through the samples, hence, one such analysis per vertex. But the authors write "in which participants are features and cortical vertices are samples" and "the thickness and tissue density maps were concatenated". This needs clarification. The architecture of the PCA should be visualized better.

      We thank the reviewer for pointing out the need to clarify the PCA methodology. In response, we have revised the Methods section to provide a clearer and more accurate description of our approach.

      We also would like to point the reviewer’s attention to Figure 1a, in which the PCA was illustrated graphically. The reviewer’s confusion likely results from the fact that we theoretically “transposed” the typical PCA analysis such that we get a subject-wise contributions (PC loadings) per participant, which is how we’re able to relate inter-participant variability in their loadings to behavior in Figure 3. This is also why we refer to a “typical” cortex/cortical sheet because the surface maps being visualized for PC2 can be thought of as a map explaining variance of deviation orthogonal to PC1 (which captures the primary relationship between thickness and T1/T2). Thus, because PC2 is orthogonal to PC1, it captures the spatial pattern in which participants deviate from the primary relationship (e.g., the typical relationship).

      We have revised the manuscript in the following sections.

      Line 493-502: “For each hemisphere, individual cortical thickness and T1/T2-weighted ratio maps from all HCP-YA participants—each represented as an M × N matrix, … corresponding participant-wise contributions (i.e., PC loading or individual weights) in pairs.”

      (3) Because the PCA only contains two features, PC1 is driven by the positive relationship between cortical thickness and mean T1/T2, whereas PC2 is driven by their negative relationship. Because in the early visual cortex, cortical thickness and mean T1/T2 correlate positively, it naturally follows that PC1 relates to pRF size (but mediated by the actual cortex parcellation). However, it is unclear why this insight is interesting. I also do not share the view that "these findings demonstrate that gradient 1 acts as a global gradient enveloping the entire visual cortex (...) while gradient 2 acts as a local gradient specific to individual visual streams". I think this relationship between cortical thickness and T1/T2 ratio does not have much to do with local and global gradients. But if so, stronger arguments as to why this should be the case should be presented. What the authors make of this result (particularly the discussion starting line 366) is not clear to me. I cannot follow the line of argumentation, which in my view is too far away from the data.

      We appreciate the reviewer’s thoughtful comments and agree that, in general, cortical thickness and T1w/T2w ratio tend to be negatively correlated, with early visual areas (i.e., V1 and V2) representing a notable exception—an observation we highlight and support with evidence in R2. Given this overall pattern of correlation, it may seem intuitive to interpret PC1 as capturing a convergent relationship across the two metrics, and PC2 as reflecting their divergence. Alternatively, one can think of PC2 as the orthogonal residuals from the linear relationship between thickness and myelin captured by PC1. In this framework, PC2 is not necessarily the inverse correlation, but instead what is left unexplained through a simple linear model. However, it is important to note that PCA is inherently agnostic to spatial structure, as our PCA operates solely on inter-subject variance. As such, the spatial patterns observed in the resulting component maps are not direct or trivial consequences of the input correlations.

      Upon examining the spatial properties of the PCA-derived maps (Fig. 1d), we found that PC1 manifests as a large-scale, low-frequency gradient spanning broad portions of the visual cortex, whereas PC2 exhibits a fine-scale, high-frequency pattern confined to subregions of the visual cortex (quantified in Fig. 1f, g). Our initial use of the terms “global” and “local” may have inadvertently implied functional interpretations beyond our intent. We have revised the manuscript to clarify that these descriptors were intended purely to convey differences in spatial scale based on the observed frequency content of the gradients.

      Motivated by the reviewer’s comment, we performed additional analyses to explicitly test whether the PCA components reflect consistent (i.e., global) or variable (i.e., local) relationships across visual ROIs. Specifically, we examined whether the direction and magnitude of PC1 and PC2 scores within each ROI align with the global relationships between cortical thickness and tissue density. As shown in the revised Supp. Fig. 3e, we found that in most ROIs, vertices with high PC1 scores consistently exhibit high cortical thickness and low T1w/T2w ratios, while those with low PC1 scores show the opposite pattern. This within-ROI consistency mirrors the largescale cross-ROI correlation structure (see Supp. Fig. 3a), supporting the interpretation of PC1 as reflecting a large-scale, cortex-wide organizational principle. In contrast, PC2 shows more heterogeneous profiles across ROIs, with peaks and troughs that differ in the two metrics. This variability suggests that PC2 captures more localized, region-specific features.

      We have incorporated the results of these new analyses into the Results section to strengthen our argument regarding the spatial scale and cross-regional consistency of the PCA-derived gradients:

      Line 102-107: “Within-area analyses further confirmed that PC1/2 represent the consistent/deviating components … while PC2 represents the spatial divergence from this commonality.”

      Recommendations for the authors:

      Reviewing Editor Comments:

      Through collaborative discussions among the reviewers, we first summarised the key recommendations for enhancing the significance and strengthening the evidence of the work - integrating public reviews and recommendations to authors by each reviewer individually. The individual reviewer recommendations can be found below this.

      (1) Modelling component 2

      The geodesic model for component 2 is interesting but we can recommend ways to improve the evidence and interpretation (see Reviewer 1 comments). As the polar angle reversals are inconsistent and boundaries ambiguous, the OTS maps do not meet the standard of evidence required for showing a new map. The 181 pRF maps available for these HCP data would provide an independent more powerful test of the OTS map cluster. To further strengthen the evidence for the proposed correspondence of foveal confluences and gradient 2, why not define the geodesic model anchoring points based on retinotopic measures, e.g., using HCP pRF data? About the current anchoring points for the geodesic model, what were the criteria - were they objective to avoid circularity?

      We appreciate the reviewer’s suggestion to incorporate the HCP 7T retinotopy dataset as an independent test of the proposed geodesic model and its relation to foveal confluences and gradient 2. We agree in principle that such data could provide a valuable validation resource. However, as detailed in the publication accompanying the HCP 7T retinotopy dataset (Benson et al., 2018), the authors recommend a threshold of 9.8% variance explained to distinguish reliable pRF estimates from noise. As illustrated in their Figure 4, this thresholded pRF data shows poor signal coverage in higher-order visual regions, particularly those along the occipitotemporal sulcus (OTS), where gradient 2 effects are most prominent in our data. This lack of reliable pRF signal in these regions limits the utility of the HCP retinotopy data for anchoring the geodesic model or validating the observed spatial gradients.

      To address this limitation, we relied on our in-house data collected using high-contrast, naturalistic images designed to robustly activate high-level visual areas. This approach allowed us to define more complete and consistent topographic patterns in the regions of interest. We have thus expanded the size of this in-house dataset to N=21. We also point the editor’s attention to the response to Reviewer 1’s first comment regarding the visual field maps for a more detailed response to this point. For convenience, we have pasted the Figure 2 e-i panels in which we conduct additional analyses showing that these anterior temporal pRF clusters tile contralateral visual space as one might expect (Fig 2h), and significantly differ across hemispheres in their laterality bias (Fig 2i). We have revised the manuscript accordingly.

      To mitigate the concern of circularity in defining the geodesic model’s anchor points, we conducted a split-half cross-validation. Anchors were defined on one half of the participants and used to predict the PC2 map in the other half. The PC2 maps across the two halves were highly similar (r = 1.00, p < 0.001), indicating strong reliability. Importantly, the cross-predicted geodesic model accounted for a significant portion of variance (r<sup>²</sup> = 0.23) in the held-out PC2 map, suggesting that the geodesic organization is not an artifact of overfitting or circular reasoning. We have revised the manuscript accordingly:

      Line 139-142: “A split-half cross-validation yielded similar results, … underlying the spatial organization of PC2.”

      (2) Speculation about prototypical cortical sheet

      You hypothesise that gradient 1 characterises a global "prototypical cortical sheet" characteristic, with gradient 2 reflecting that regions become more distinct from this prototype. There is an alternative simpler possibility: the data can be explained by the stronger relationship between cortical thickness and T1/T2 ratio in early compared to late sensory areas, as can for example be seen in Glasser et al. 2016 Nature, Figure 4. We recommend omitting or balancing the statement about a "prototypical" cortex, and integrating findings on cortex parcellation and the view that sharp boundaries characterize transitions between high and low T1/T2 and cortical thickness areas.

      Please see R2 for reviewer #2

      (3) Confounds

      We'd like to see more data to understand the contributions of data quality to these results. For the component 1 gradient specifically, could its features be influenced by spatial SNR inhomogeneities? Could the developmental effects for both gradients be explained by lower SNR and other data quality markers in younger and older participant data? We missed appropriate tests that gradients develop differently across age, controlling for such confounds (Reviewer 1 comments).

      Regarding the reviewer’s concern about the component 1 gradient, we believe it is unlikely to be merely a consequence of uneven spatial SNR. Our findings are consistent with previous histological studies demonstrating systematic variations in cortical architecture—specifically, thinner cortex (Wagstyl et al., 2020) and higher myelin content (Dinse et al., 2015) in occipital compared to ventral visual regions. This correspondence between in vivo MRI-derived measures and postmortem histology suggests that the large-scale organization captured by PC1 is grounded in biologically meaningful cortical architecture, and not an artifact of SNR variability.

      To statistically assess whether the two PCs show different developmental trajectories across age, we performed an ANOVA with age, LC, and their interaction as factors on LC’s similarity to PC (i.e., r ~ age + LC + age × LC). Significant age × LC interactions were observed in the developmental (HCPD: F<sub>1,118</sub> = 257.01, p < .001) and aging (HCPA: F<sub>1,132</sub> = 263.85, p < .001) cohorts, but not in the young adult cohort (HCPYA: F<sub>1,202</sub> = 0.02, p = 0.80). These findings indicate that the two gradients show distinct age-related changes during development and aging but remain stable in young adulthood. We have revised the manuscript accordingly:

      Line 313-327: “Examining the correlation between the young adult gradient and LC … F<sub>1,132</sub> = 263.85, p < 0.001).”

      (4) Implementation of PCA

      The manuscript raises questions about the correct implementation of the PCA - please clarify that the variables were first standardised to enable fair weightings, and visualise the PCA matrix in more detail than in Figure 1a to ensure the samples and features are correctly defined (Reviewer 2).

      Please see R3 for reviewer #2

      References

      Abdollahi, R. O., Kolster, H., Glasser, M. F., Robinson, E. C., Coalson, T. S., Dierker, D., Jenkinson, M., Van Essen, D. C., & Orban, G. A. (2014). Correspondences between retinotopic areas and myelin maps in human visual cortex. NeuroImage, 99, 509–524. https://doi.org/10.1016/j.neuroimage.2014.06.042

      Benson, N. C., Jamison, K. W., Arcaro, M. J., Vu, A., Glasser, M. F., Coalson, T. S., Van Essen, D. C., Yacoub, E., Ugurbil, K., Winawer, J., & Kay, K. (2018). The HCP 7T Retinotopy Dataset: Description and pRF Analysis. https://doi.org/10.1101/308247

      Dinse, J., Härtwich, N., Waehnert, M. D., Tardif, C. L., Schäfer, A., Geyer, S., Preim, B., Turner, R., & Bazin, P.-L. (2015). A cytoarchitecture-driven myelin model reveals area-specific signatures in human primary and secondary areas using ultra-high resolution in-vivo brain MRI. NeuroImage, 114, 71–87. https://doi.org/10.1016/j.neuroimage.2015.04.023

      Fischl, B., Rajendran, N., Busa, E., Augustinack, J., Hinds, O., Yeo, B. T. T., Mohlberg, H., Amunts, K., & Zilles, K. (2008). Cortical Folding Patterns and Predicting Cytoarchitecture. Cerebral Cortex, 18(8), 1973–1980. https://doi.org/10.1093/cercor/bhm225

      Glasser, M. F., Coalson, T. S., Robinson, E. C., Hacker, C. D., Harwell, J., Yacoub, E., Ugurbil, K., Andersson, J., Beckmann, C. F., Jenkinson, M., Smith, S. M., & Van Essen, D. C. (2016). A multimodal parcellation of human cerebral cortex. Nature, 536(7615), 171–178. https://doi.org/10.1038/nature18933

      Kay, K. N., Winawer, J., Mezer, A., & Wandell, B. A. (2013). Compressive spatial summation in human visual cortex. Journal of Neurophysiology, 110(2), 481–494. https://doi.org/10.1152/jn.00105.2013

      Mackey, W. E., Winawer, J., & Curtis, C. E. (2017). Visual field map clusters in human frontoparietal cortex. eLife, 6, e22974. https://doi.org/10.7554/eLife.22974

      Maingault, S., Pepe, A., Mazoyer, B., Tzourio-Mazoyer, N., & Crivello, F. (2021). Characterization of late structural maturation with a neuroanatomical marker that considers both cortical thickness and intracortical myelination. https://doi.org/10.1101/2021.02.24.432645

      Sereno, M. I., Lutti, A., Weiskopf, N., & Dick, F. (2013). Mapping the Human Cortical Surface by Combining Quantitative T1 with Retinotopy†. Cerebral Cortex, 23(9), 2261–2268. https://doi.org/10.1093/cercor/bhs213

      Shafee, R., Buckner, R. L., & Fischl, B. (2015). Gray matter myelination of 1555 human brains using partial volume corrected MRI images. NeuroImage, 105, 473–485. https://doi.org/10.1016/j.neuroimage.2014.10.054

      Sheremata, S. L., & Silver, M. A. (2015). Hemisphere-Dependent Attentional Modulation of Human Parietal Visual Field Representations. The Journal of Neuroscience, 35(2), 508–517. https://doi.org/10.1523/JNEUROSCI.2378-14.2015

      Silson, E. H., Zeidman, P., Knapen, T., & Baker, C. I. (2021). Representation of Contralateral Visual Space in the Human Hippocampus. The Journal of Neuroscience, 41(11), 2382–2392. https://doi.org/10.1523/JNEUROSCI.1990-20.2020

      Wagstyl, K., Larocque, S., Cucurull, G., Lepage, C., Cohen, J. P., Bludau, S., Palomero-Gallagher, N., Lewis, L. B., Funck, T., Spitzer, H., Dickscheid, T., Fletcher, P. C., Romero, A., Zilles, K., Amunts, K., Bengio, Y., & Evans, A. C. (2020). BigBrain 3D atlas of cortical layers: Cortical and laminar thickness gradients diverge in sensory and motor cortices. PLOS Biology, 18(4), e3000678. https://doi.org/10.1371/journal.pbio.3000678

    1. Author response:

      eLife Assessment

      In this valuable manuscript, the authors tackle a highly relevant question in biology: how cells integrate attractive and repulsive cues to achieve directed migration. They present solid data demonstrating that two wunen genes act as negative regulators of Hedgehog signalling, thereby enabling efficient primordial germ cell (PGC) migration in Drosophila embryos. Beyond its immediate scope, this work has broader implications, particularly for understanding key mechanisms underlying complex processes such as cancer metastasis, where the coordinated interpretation of guidance cues is critical.

      Thank you for the reviews and the overall assessment of our manuscript. It is our impression that both the reviewers and the senior editor find the study interesting and potentially of general relevance. The reviewers have made specific suggestions to improve the manuscript. They have also recommended ways to uncover the mechanistic basis to add to the broad appeal of the findings.

      To begin with, we would like to point out that since the discovery of Wunen in 1996 by Ken Howard and colleagues, a number of genetic and molecular studies have attempted to identify and characterize the putative target(s) of the two lipid phosphate phosphatase(s). We and others have shown that Hh acts as a guidance signal for the migrating PGCs. Our data demonstrating the ability of Wunen(s) to attenuate Hh signaling constitutes an important step in elucidating the molecular underpinnings of the repulsive activity of Wun(s) during PGC migration.

      Thus, we feel the need to share these findings with the scientific community at this juncture. In the following, we will summarize our response to the relevant points included in the individual public critiques of the reviewers without going into specific details.

      Public Reviews:

      Reviewer #1 (Public review):

      This manuscript addresses how PGCs migrate towards SGPs in the Drosophila embryo. It's been shown that Hh produced by SGPs acts as an attractive cue, and that Wunnen(s) act as repulsive cues. In this work, the authors propose that Wun and Wun2 refine PGC guidance by attenuating Hedgehog signalling coming from other tissues.

      Overall, the study is potentially interesting and could make an important contribution to the field. The data shown support the idea that Wun/Wun2 negatively regulate Hh signalling and produce PGC migration phenotypes associated with Hh. However, in my opinion, there are two major questions that should be addressed.

      (1) Which is the mechanism by which Wun/Wun2 attenuates Hh signalling? The authors propose that Wun/Wun2 block Hh ligand transmission, but their data could also be explained by other possibilities, such as altered Hh production, uptake, retention or degradation, among others. The authors should either show the effect of Wun/Wun2 in Hh transmission mechanistically or attenuate their claim.

      (2) How do Wun/Wun2 attenuate Hh signalling in PGCs? The authors propose that Wun/Wun2 function both in somatic tissues and in PGCs, but these two sites of action may have very different mechanistic implications. In the soma, Wun/Wun2 could affect Hh transmission, but a PGC-autonomous role cannot be explained simply by reduced Hh ligand transmission from producing cells; it would more likely involve ligand uptake, receptor trafficking, intracellular degradation or altered PGC responsiveness. This distinction should be central to the interpretation of the data.

      We thank the reviewer for recognizing the importance of the problem and we are sensitive to both the points of criticism regarding the mechanism(s) Wunen(s) may employ to downregulate Hh signalling.

      The reviewer correctly pointed out that we singled out Hh transmission as the putative target of Wunen(s) which need not be the case. We agree with this assessment and would like to thank the reviewer for pointing us in the right direction(s). Indeed, Wunen(s) could act at several different levels to regulate Hh signalling including “Hh production, uptake, retention or degradation”. We will modify the text to incorporate these possibilities in the appropriate sections of the manuscript.

      The only reason for the emphasis on the ‘Hh transmission’ in the text was to contrast it with Hmgcr which acts in a qualitatively opposite manner. Hmgcr potentiates Hh signalling by altering the range/strength of the Hh ligand in the embryonic context. This was also confirmed in the wing discs and adult wings as hmgcr mutants could dominantly suppress the wing duplications and abnormalities induced by the ‘gain of function’ allele of hh (hh<sup>MRT</sup>). Upon compromising hmgcr, Hh ligand was shown to be sequestered in the Hh producing cells in the ectoderm. However, we have not carried out similar experiments to either rule in or rule out the different possibilities suggested by the reviewer. We will ensure that the claims made in the manuscript will appropriately reflect the scope of the analysis and the related arguments will be suitably modified.

      The reviewer also makes a very critical point regarding cell autonomous v/s cell non autonomous activities of Wun(s). We have briefly mentioned the possible role of individual Wun(s) in the SGPs/mesoderm as well as within the PGCs. It has not escaped our notice that Wun(s) could regulate Hh internalization within the PGCs or its subcellular compartmentalization (within the ER, golgi or lysosomes). Wunen(s) could also act at the level of Hh reception by changing the activity/localization of Hh receptors, either Smoothened or Patched and could influence the outcome of the signaling pathway in a multi-pronged manner.

      We appreciate the thoughtful suggestions and as recommended, future analysis will focus on these aspects. In our view, data included in the present version of the manuscript are novel and sufficient to argue a functional relationship between Wun(s) and Hh signalling which is qualitatively antagonistic to Hmgcr.

      Reviewer #2 (Public review):

      Summary:

      In this submission, Roy et al. examine the process of Drosophila PGC migration. Directed cell migration requires the concerted activities of chemoattractants and repellents to guide cells to the correct locale. In their submission, the authors describe a role for regulated Hedgehog (Hh) signaling to inform PGC migration. In prior work, the authors reported that Hmgcr potentiates Hh signaling, providing a permissive axis. A gap in the field, however, was the identification of the repulsive cues that guide PGCs out of the midgut and toward the future gonad. In the current work, the authors report that two wunen genes (wunen and wunen 2) inhibit Hh signaling, thereby repressing Hh activity. The model is that Hmgcr and wunen(s) balance the transmission of Hh signals to enable effective PGC migration.

      Strengths:

      A strength of this work is the comprehensive genetic analysis performed by the authors. The authors examine zygotic versus maternal contributions, autonomous versus non-autonomous requirements, and use a variety of RNAi and mutant allele combinations to examine genetic requirements and interactions. Another strength is that the data presented are generally clear and well quantified. Insets are provided to enhance visualization, and relevant data are quantified through replicated experiments.

      Weaknesses:

      Weaknesses of the work include a lack of biochemical data to validate some of the proposed interactions. Although the authors do report lipidomics data, little is done with these findings to validate or place the results in the context of a mechanistic model. Despite these issues, the conclusions stated are generally well supported by the results.

      We would like to thank the reviewer for their positive feedback and a succinct description of the findings reported in the manuscript.

      We agree that the mechanistic basis of DAG accumulation was not explored in this manuscript. Prior work in the Ratnaparkhi and Kamat labs identified a Serine hydrolase that functions as a phospholipase C in biochemical assays (Kumar et al., 2024, Biochemistry 63:3000-3010). We have since conducted several genetic experiments, and preliminary data indicate that, in the embryonic context, mutations in the specific Phospholipase C display phenotypes analogous to wun(s). We hope to present these data along with the comparative molecular and biochemical analysis in the near future.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors present a method to detect natural selection on transcription factor binding sites (TFBSs), which is an upgraded version of a previously published method (Liu and Robinson-Rechavi, 2020). This upgraded version of the test implements more explicit models of evolution and is shown to outperform its predecessor in terms of both power and false positive rate. I think this method can be a valuable resource for the community and can be helpful not only to studies of TFBSs but also broader evolutionary questions related to genotype-phenotype maps or fitness landscapes.

      Major comments:

      (1) Questions related to Figure 1

      Figure 1, along with the first section of the Results, shows that the SVM score and its sensitivity to mutations are generally correlated with the strength of ChIP-seq signals. It is not very clear to me, however, what the motivation is behind this part of the paper. It seems that the model used to predict binding strength is a pre-existing one, and it is unclear what is new in this section. Was the prediction model retrained using different data? Was its validity confirmed using new data? I would appreciate some more elaboration on how these results differ from what was presented in the previous study of Liu and Robinson-Rechavi (2020).

      We agree that the current manuscript does not clearly distinguish which parts of Figure 1 are novel and which are foundational. The SVM itself is not new and is the same as in Lee et al. (2016), as used in Liu & Robinson-Rechavi (2020). In the revision, we will explicitly state that the SVM used in Figure 1 is the standard gapped-kmer SVM (ls-gkm) approach. We retrained all gkm-SVM models de novo for each species-TF dataset, ensuring consistency across all analysed ChIP-seq peaks. For this, we recalled all ChIP-seq peaks in a homogeneous and robust manner using the nf-core ChIP-seq pipeline v2.0 (Ewels et al. 2022). Figure 1A confirms that the predicted binding affinity from the SVM correlates with experimental ChIP peak height. In addition, examining scores per site rather than per peak is new compared with Liu and Robinson-Rechavi (2020). The correlations between the SVM-derived scores and other features had not been shown before to the best of our knowledge, thus Figure 1B-C is entirely novel. In other words, this analysis is meant to show that our phenotypic metric (SVM score per site) indeed tracks binding intensity, i.e. molecular phenotype.

      The existence of weak or negative correlations between SVM and coverage, which reportedly reflects low-quality peaks, seems applicable not only to this paper, but also to previous ones, so I would like to have it confirmed whether the question and the authors' answers apply to previous studies as well.

      Yes, this is a well-known issue in ChIP-seq studies. Low coverage often matches weak predicted binding affinity scores because noisy or unreliable peaks naturally have weaker signals. This is not specific to our work, and it has been observed in many other studies (e.g., Bailey et al. 2013 doi:10.1371/journal.pcbi.1003326; Nakato and Shirahige 2017 doi:10.1093/bib/bbw023). It is simply an expected property of the data.

      It is reported that SVM scores capture TF binding signals better than conservation-based statistics do. My intuitive interpretation is that both ChIP-seq peaks and SVM scores are supposed to reflect binding strength, whereas conservation is supposed to reflect selection (i.e., different definitions of "function" as mentioned above). It is not explicitly explained in the Results, however, what the difference indicates, leaving only an impression that the SVM score is "better" than the conservation statistics.

      While the reviewer is correct that there are different definitions of function, both conservation-based statistics and RegEvol seek to capture selected function. The difference is that RegEvol aims to measure functional change, whereas conservation-based statistics aim to detect sequences that retain the same function across species. In both cases, we expect a correlation with causal function (i.e., binding). We will clarify these concepts and how they apply to our results in the revised manuscript.

      (2) Lack of directional selection for low binding affinity

      In the analysis of Drosophila melanogaster ChIP-seq peaks, there were more cases of directional selection for higher binding affinity than directional selection for lower binding affinity. The authors suggested that this observation is "likely biological" because the same pattern was not seen in simulations (line 412-413). I wonder if this could have resulted from a difference in the distribution of ancestral binding affinity across TFBSs between real and simulated data. If binding affinity was generally low in the common ancestor of D. melanogaster and D. simulans, selection for low binding affinity would manifest mainly as purifying selection against mutations that increase affinity instead of directional selection. Ancestral sequences for simulations, if I understood correctly, are observed peaks in D. melanogaster (line 715-719), which would include high fraction sequences that could be rarer in the real ancestral sequences.

      The description of this particular result does not refer to a figure or table, nor is it revisited in the Discussion. Figure 5 treats peaks under directional selection as a single category. Taken together, it is hard to tell how this observation should be interpreted. If the authors consider this result as biologically meaningful, I would suggest adding more details (e.g., the number of each side).

      We appreciate this insight. We agree that the text was not clear, but in fact, the simulations were performed using the reconstructed ancestral sequences of ChIP-seq peaks themselves. Thus, simulated and empirical results should be directly comparable, and different results should be due to biology. We will revise the Manuscript to explicitly state that simulations are performed from reconstructed ancestral sequences and why. We will also add more descriptive statistics of the simulated and real data.

      (3) Selection in non-focal lineages

      Regarding the detected signals of directional selection for stronger binding in certain tissues (Figure 6), I wonder if it is the focal species or those very tissues that are "special": did the human lineage undergo more adaptive regulatory evolution than the chimpanzee lineage, or do nervous and male reproductive systems have a high "propensity" for adaptive regulatory evolution? Assuming that the binding preference of the same TF did not undergo a significant change since human-chimpanzee split (which, I believe, is a built-in assumption in both RegEvo and the permutation test), it should be possible to perform the same test using chimpanzee sequences that are homologous to the human ChIP-seq peak regions. In the case of coding sequences, for example, Bakewell et al. (2007) found that it was the chimpanzee that had more genes under positive selection than humans; I wonder if TFBSs show the same or a different pattern.

      This is an excellent suggestion. To compare in an unbiased manner, we would need transcription factor ChIP-seq from the same organs in chimpanzees and humans. We are not aware of such a dataset. If one is identified, we would be very interested in analysing it, and thus answer this question. As suggested by the reviewer, we will analyse the human homologous sequences. Although it should be clear that this will provide a biased estimate for comparing adaptation between the two species, as we will lack newly acquired binding sites in the chimpanzee.

      (4) Comments on terminology

      (a) Meaning of "function"

      The word "function" has had different meanings in the biology literature, with some authors using "functional" to refer to anything with a phenotypic effect and some using it only for targets of selection. A (putative) TFBS would be considered "functional" as long as it has TF binding affinity if we follow the effect-based definition, but only if its binding affinity is under selection if we follow the selection-based definition. In this manuscript, the term "function" appears to have been used to refer to TF binding but not selection, most notably in the first Results section. There are also places where it is less clear what "function" means exactly (e.g., "deeply conserved elements that are likely to be functionally important" of line 61). Since this paper is about evolution, it is likely that many readers prefer the selection-based definition or assume that the selection-based definition would be used. Thus, using "function" to refer to just TF binding could be confusing. To this end, I would suggest that the authors drop the word "function" or give an explicit definition early in this paper.

      We thank the reviewer for this precision and fully agree, we will revise our terminology for clarity. We will clarify the distinction between selected function and causal function, and we will pay attention to their use throughout the manuscript.

      (b) Directional selection in different directions

      In this paper, selection for increased TF binding affinity is referred to as "positive directional selection", and selection in the opposite direction is called "negative directional selection" (as exemplified in Figure 2). I understand that using such shorthand names would make the text less clumsy, but these two terms could potentially be confusing, as "positive selection" and "negative (purifying) selection" are also terms referring to specific types of selection and have some connection to directional and stabilizing selection. Therefore, I suggest that the authors use something like "selection for increased/decreased binding affinity" instead, or note explicitly in the text that "positive/negative directional selection" would be used as shorthand.

      We agree with this ambiguity in the current terminologies. We will replace the phrases “positive directional selection” and “negative directional selection” with, e.g., “selection for increased binding affinity” and “selection for decreased binding affinity” as suggested when presenting our biological result on ChIP-seq peaks. However, we will still use “positive/negative directional” for the general framework (genotype → phenotype →fitness map) and insert a note that we use “positive/negative directional” as shorthand to mean increasing/decreasing affinity in the case of CHIP-seq peaks.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Laverre et al. provides an interesting new test of selection on TF binding. Rather than focusing on sequence changes, this test is specifically for changes in predicted TF binding affinity. The authors report directional selection on 5.1% of tested regions in Drosophila, as well as a signal of selection on CTCF binding in the human CNS and male reproductive system.

      Strengths:

      Overall, I think this represents an important direction for the field of molecular evolution: now that TF binding can be predicted fairly well from sequence, it can be a very useful focus for tests of selection.

      Weaknesses:

      As mentioned several times in the manuscript, Jiang and Zhang (2024) pointed out some issues with a previous permutation-based version of this test. Foremost among these was the issue of ascertainment bias: when testing only experimentally supported TF binding sites from a focal species, and then asking what type of selection (or lack of selection) led to those sites, one is guaranteed to find more substitutions that increase affinity, simply because the sites were selected in the first place as those with maximum (empirically measured) affinity.

      To address this issue, the authors simulated Drosophila CTCF peaks evolving neutrally and then tested different ascertainment cutoffs in Figure 4D. It was not entirely clear to me what is shown in Figure 4D: the text says the bins were stratified by derived delta-SVM, whereas the figure says SVM, and the legend says derived SVM (both without the delta). I was unable to find any clarification of this in the Methods section. In any case, I am not really convinced by his, for two main reasons. First, when analyzing empirical ChIP-seq data, I would guess that only a tiny fraction of the genome is bound (far less than 1%, especially in mammalian genomes). However, the most extreme bin in Figure 4D is taking the top 10% of (delta?) SVM values. What would Figure 4D look like at bins of the highest 0.1%, 0.001%, etc? My guess is there would be a strong uptick in the FPR.

      We apologise for the confusion in Figure 4D, we will clarify the caption and text and specify that bins are stratified by derived SVM (post-simulation binding affinity proxy), not genome % or ΔSVM.

      We want to note that we used the same subsampling approach as Jiang and Zhang (2024) to evaluate ascertainment bias, and that Figure 4 both confirms the issue that they identified with Liu and Robinson-Rechavi (2020), and shows very clearly that RegEvol does not have the same issue (flat red lines). Following the reviewer's suggestion, we can extend the figure to 1% or 0.1% bins. We note that the % of the total genome is different from the % of peaks: while actual peaks cover a very small proportion of the genome, the subsampling in Figure 4 (and in Jiang and Zhang 2024) aims to estimate the impact of detecting only the strongest peaks.

      One difference between Jiang and Zhang (2024) and our study is that we simulated using whole empirical peaks, whereas they simulated 10-nucleotide transcription-binding sites, meaning that each substitution represented a 10% change. We will clarify these differences in the revised text.

      The second reason is actually more important and fundamental than the first. As long as this method is working as described, I cannot see any way that it would ‘not’ be impacted by ascertainment bias. As an extreme case, imagine that all TF binding sites tested had the maximum possible SVM scores; then none of them would have any chance of showing directional selection against binding, while even those that evolved neutrally would appear to have directional selection in favor of binding. Of course, real empirical data are not as extreme as this, but the same concept applies in less extreme scenarios.

      This bias could explain patterns observed in the real data. For example: "We observe much more positive than negative directional selection, a pattern likely biological rather than methodological, since it is absent from simulations." This is exactly the pattern predicted under ascertainment bias (in the extreme-scenario thought experiment above). I suspect it is absent from simulations simply because the authors did not properly account for this bias in their simulations.

      If the main result reported by the authors had been a lack of any directional selection in favor of binding, and instead only neutrality or directional selection against binding, then this ascertainment bias would not be an issue- it would only have made their results conservative. Unfortunately, this is not the case, and the directional selection in favor of binding, which is the main result emphasized from the empirical analysis, could be inflated by this bias.

      There is indeed a possible ascertainment bias, although we believe it concerns only the detection of negative directional selection, as long as we have only empirical peaks in the focal species and not the sister species. This is not so much a limitation of our method as an intrinsic limitation of asymmetrical sampling of species: to study both gain and loss of function, function must be studied experimentally in several species. We will revise the manuscript to highlight this limitation.

      Concerning positive directional selection, the mathematical foundation of RegEvol makes it inherently robust to ascertainment bias for positive directional selection. RegEvol calculates the likelihood of the entire sequence of observed substitutions accounting for the starting ancestral state and the mutational landscape. In other words, the model does not assume a uniform probability of phenotypic change; instead, it models the probability of each nucleotide mutation to result in a substitution (i.e., go to fixation) depending on its phenotype.

      In an extreme case where all tested TF binding sites had the maximum SVM score, detecting negative directional selection would indeed be impossible, as ancestral states would have had equivalent or lower scores. However, positive directional selection would be inferred only if the likelihood of observing the substitution pattern’s deltaSVM distribution significantly exceeded that expected under the mutational landscape. If a sequence evolved neutrally but reached a maximum SVM score, the likelihood of detecting directional selection would depend on: either the ancestral state being close to maximum with few substitutions increasing SVM (resulting in low statistical power), or the ancestral state being distant with many neutral substitutions and rare chance shifts to maximum (where the substitution distribution would be indistinguishable from neutrality). Then, even in such an extreme dataset, neutral evolution remains detectable, demonstrating RegEvol's strength beyond deltaSVM comparisons between two states.

      Minor point:

      The following statement: "In contrast, phastCons and phyloP scores lack such enrichment and have a lower dynamic range, suggesting that the conservation scores are less sensitive to fine-scale variation of TF occupancy and thus regulatory region function" is only true if one assumes that TF binding is the only function of this region. One could even turn this around and say the fact that the sites affecting TF binding are not the most conserved is actually evidence that TF binding is not a good indicator of these regions' entire function. I suggest the authors soften this claim that conservation scores are less sensitive to regulatory region function.

      We thank the reviewer for this comment, the text will be revised to soften this claim. We will explicitly state that sequence conservation reflects general functional constraints, whereas sequence-to-phenotype predictions capture highly specific and lineage-specific TF-DNA interactions.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      The authors aim to understand, in the context of leaf shape, how the constraints imposed by development inform evolution. Leaf shape is a good place to study the influence of development on evolution because it is a trait that exhibits a lot of diversity, and the developmental mechanisms that give rise to leaf shapes are apparently rather conserved across angiosperms.

      As part of the motivation for their work, the authors cite a previous study (Geeta et al), which found that in angiosperm phylogenies, transitions from complex to simple leaf shapes occur through evolution more often than transitions in the opposite direction. Is this due to developmental constraints or adaptation?

      The authors undertake two parallel lines of work:

      (1) Extending the study of Geeta et al with more data, consisting of both phylogenies and a shape classification dataset. The conclusion from this line of inquiry is that transitions from lobed to unlobed leaves are more common than transitions away from unlobed leaves.

      (2) The authors conduct evolution simulations in a computational model of leaf development. Here, they look at {\it neutral} mutations and whether simply neutral evolution is sufficient to drive the observed trend.

      The conclusion of the second part of the work is that the driver of the evolution toward simple leaf shape is entropy: there are more ways to make unlobed leaves than to make lobed leaves (at least in terms of gene regulation parameters that will produce the two leaf types). The argument is that random gene regulatory networks are more likely to produce unlobed leaves than lobed leaves; therefore, neutral evolution drives this trend.

      Data Analysis

      Roughly $9000$ images of leaves were classified into 4 categories: unlobed, lobed, dissected, and compound. These labels were applied to the tips of 5 phylogenetic trees of angiosperms (3 resolved at the genus level and 2 at the species level). By fitting a continuous-time Markov chain to the labelled trees, the authors claim that there is a significantly higher rate of transition to the unlobed leaf shape compared to transitions to more complex shapes.

      Simulation

      First, the authors validate a computational model (Runions et al) for leaf growth on an experimental dataset. By changing parameters in the model, they can recapitulate the morphological changes in the shapes of Arabidopsis leaves engendered by expression of two particular genes.

      Then the authors run an evolutionary model (without selection, just random mutations) on top of the computational leaf development model. As the random walk in parameter space reaches a stationary distribution, they look at both the proportions of the leaf categories in the steady state as well as the transition rates between different categories. The result is that transitions to unlobed leaves are more common than from unlobed leaves.

      We thank the reviewer for the helpful and clear summary of our work.

      General Comments

      The authors use angiosperm phylogenies from other works as the basis for the data analysis part of their work. Given the centrality of these phylogenies for their conclusions, more information is needed about how these phylogenies were constructed and what they mean. What is the timescale that they span? What method is used to infer them? What regions of DNA were sequenced in order to build the phylogenies? Also, maybe some more discussion of angiosperm evolution (e.g., when was the most recent common ancestor of all angiosperms?) would help put the study in context.

      We also need a more in-depth discussion of the computational model. What are all the $>100$ parameters doing, and what informs the seemingly strange mutational model that changes parameters by 3 orders of magnitude?

      I am confused about how the rates of transitions were inferred from the phylogeny. Here, one has a phylogeny inferred by some method (which needs to be described in more detail), and just the leaves are labelled. It is stated in the methods that BayesTraits was used to infer the transition rates. I realize this method is probably documented elsewhere, but a bit of a summary of how it works and how to interpret its results would (1) make the paper more selfcontained and (2) if the algorithm is credible, make the results firmer.

      We thank the referee for the suggestion to make the paper more accessible. The tool we use to infer transition rates from the phylogenies, BayesTraits, is standard in the field. However, the referee is right that for an interdisciplinary journal, it may be helpful to more fully flesh out how these methods work. To that end, we have added an additional section "Phylogenetic rate inference" in the supplementary information that includes a longer description of how BayesTraits works, and how we used it to infer transition rates from phylogenies.

      All trees are shown in the supplementary information section "Phylogenetic trees" with scale-bars showing the amount of time or genetic change that the trees span. For a broader discussion of angiosperm evolution, there is supplementary information section "The adaptive significance of leaf shape review".

      Regarding the more in-depth discussion of the computational model, we have added supplementary information section S1 "Leaf model details" to give a more detailed description of the leaf model.

      I am a bit skeptical of the authors' interpretation of the biological trend (of complex to simple leaf shapes) as being driven by neutral evolution. Why does one expect that the mutations generated by the random walk models described in the work are in fact neutral mutations?

      A random walk is a well-established way of modelling the dynamics of neutral evolution in the monomorphic regime, where the population has a narrow diversity of different genotypes. In the higher mutation rate polymorphic regime, where the diversity of genotypes in the population is larger, we also expect that a random walk should still recapitulate the correct average transition rates. The purpose of the simulations is not to model every aspect of population genetics, but to ask whether developmental bias alone is sufficient to generate the observed directional asymmetry. By assigning equal fitness to all viable leaves, we isolate the contribution of development from that of selection. The agreement with the phylogenetic transition rates therefore demonstrates sufficiency rather than exclusivity: selection may also contribute, but it is not required to explain the observed bias We discuss the evidence for the role adaptation in leaf shape further in supplementary information section "The adaptive significance of leaf shape review".

      If the entropy of simple leaf shapes is higher than that of complex leaf shapes, why did we have complex leaves at all? I suspect the authors might argue that this is due to selection. In that case, what allows these complex shapes to become simpler? Wouldn't they be losing the selective advantage that drove them to be more complex in the first place? Or maybe the idea is that the rates are inferred assuming some steady state that generates the phylogeny? I did not understand this point.

      The entropy language is a useful framing. Within that framework, one can view our study as showing that the entropy (defined here as the logarithm of the volume of parameter space mapping to a phenotype) of simple leaf shapes is higher than that of complex leaf shapes. If this entropy were to be ignored, then all states would be equally likely in our simulations, where we do not take fitness differences into account. What we show is that the differences in entropy -- related to differences in volumes of the parameter space that map to different phenotypes -- also affects the rates. The inferred transition rates for both simulation and phylogeny from unlobed to more complex shapes are lower than vice versa but not zero. Therefore, complex leaf shapes arise stochastically through mutation and in this model would eventually reach a steady state proportion, even in the absence of selection.

      Are the rates of transitions between leaf types inferred for the phylogeny assuming that the phylogeny is generated by the steady state of some Markov process? (I think the answer is no: in that case, how does one explain the initial condition?)

      The tool we use to infer transition rates from phylogenies—BayesTraits—allows the initial state at the root of the tree to vary during the numerical optimisation (Pagel, 1994). Therefore, it is not assumed that the initial state is generated by the steady state of the Markov process.

      If I take the mutation model (random walk) seriously, then shouldn't I expect that this steady state obeys detailed balance? In that case I should have $p_i r_{i\to j} = p_j r_{j\to i}$ for each of the occupancies $\{ p_i\}$ and transition rates $r_{i\to j}$ for the shape categories. How close are the rates inferred from the phylogenies to obeying detailed balance? Presumably, the Markov chain fitted to the simulation data obeys detailed balance because the mutation model itself does?

      BayesTraits allows off-diagonal transition rates of the rate matrix to vary freely during numerical optimisation (Pagel, 1994). Therefore, there is no requirement for the detailed balance to hold for the inferred rate matrix. For our simulations, the mutations are symmetric at the parameter level, therefore at this level, the process would be expected to obey the detailed balance.

      I find it hard to take the discussion of development seriously without some consideration of mechanics. Presumably, the mechanics are hidden in the computational leaf development model, but this model is not discussed in enough detail for the reader to know. It seems to me that the interesting question is: what are the {\it mechanical} constraints on development that drive the apparent trend in evolution towards simpler leaf shapes? Maybe it is something about the type of differential growth needed to make complex leaf shapes less robust to mutation. But in this case, I would assume that selection plays a role in the complexity of shape. In any case, a better understanding (or explanation) of the computational model is needed to make this interpretation.

      We thank the referee for the suggestion to make the paper more accessible. We have added a more detailed and pedagogical description of the model from (Runions, Tsiantis and Prusinkiewicz, 2017) in the supplementary information section S1 "Leaf model details". We also note that Fig. 5 in the methods that gives an overview of how the model works, including some mechanical aspects of development and growth.

      More generally, mechanics is one component of the developmental map that determines which parameter combinations produce viable leaf morphologies. Our analysis concerns the geometry of this complete developmental map, irrespective of whether its constraints arise from gene regulation, tissue mechanics, or their interaction.

      On the interesting question of what is causal, perhaps the example in figure 2 is helpful. We focus on two parameters, a morphogen repression strength, and a duration of growth. A key physical process here is called webbing, where cellular growth fills in the gaps between branching veins. This process flattens the leaf structure and creates a continuous, solid leaf blade (lamina). Strong webbing, characterized by a significant resistance to stretching and bending, results in a smoother margin (Runions, Tsiantis and Prusinkiewicz, 2017). The morphogen repression strength affects the physical parameters that determine how strong the webbing is. The duration of growth determines how long the leaf has to grow. Varying these two parameters varies the physical processes that determine leaf shape. The mechanics of growth operate downstream of these parameters that we vary in our evolutionary simulations according to the details of the leaf developmental model.

      Some discussion of timescales is needed, especially when invoking neutral evolutionary arguments. If a neutral mutation occurs, its time to fix in a population of size $N$ is $\sim N$ generations. What are the relevant angiosperm population sizes and the number of mutations that separate branches on the tree? Are timescales remotely consistent with e.g., the age of angiosperms on Earth?

      Neutral processes have a well-established role in key aspects of angiosperm evolution, for example genome complexity (Lynch and Conery, 2003). This would suggest that the relevant time scales and generation times are not completely prohibitive of neutral processes also playing a role in the evolution of angiosperm leaf shape. Effective population sizes in plants are highly variable but estimates span 10^3-10^6. Assuming diploidy (and therefore average fixation time of 4Ne) and generation times of 1-10 years, this gives fixation timescales of 10^3-10^7 years. This is within the timescales of the trees we analyse, which span >150 million years.

      Reviewer #2 (Public review):

      Strengths:

      The paper's underlying question is interesting, extending the authors' prior work on RNA along similar conceptual lines. The paper combines both image analysis of leaves and a computational analysis of a simple model of leaf development.

      Weaknesses:

      The entire paper is based on the Runion model. More intuition about the Runion model would be useful for a broader readership that cares about the evolutionary aspect of this, but may not know the developmental model in question. Obviously, this is prior well-established work, but 2 - 3 sentences highlighting the key structural aspects of such a model would be great. Currently, that intuition is found implicitly in a sentence on page 2 ("complex leaf shapes need more specificity in their GRNs than their simpler unlobed leaf shape"), but the reader is left wondering - is the Runion model a detailed mechanistic one with multiple interacting genes/proteins? If so, how many? Or is it just 2 - 3 genes but with complexity entirely in how long they are each expressed/when they are turned off, etc.

      We thank the referee for the suggestion to make the paper more useful for a broader readership. To that end, we have added a more detailed description of the (Runions, Tsiantis and Prusinkiewicz, 2017) model in supplementary information section S1 "Leaf model details".

      The Runions model has nearly 100 free parameters. Random walks in 100dimensional spaces have generic properties like a tendency to move toward regions of larger volume that have nothing to do with leaf biology. How do you disentangle the geometry of high-dimensional random walks from genuinely biological developmental bias? Would a toy model with 100 parameters and arbitrary phenotype categories also show "bias toward simplicity" if "simple" phenotypes occupy more volume?

      Our argument is largely independent of the number of parameters. While it is true that most of the volume is near the surface in a high-dimensional space, our argument is about the relative volumes of the sets of parameters that map to each of the four phenotypes, an entropic argument if you wish. The basic intuition is that a simple phenotype needs fewer parameters to be fine-tuned, and so a larger volume of parameter space will map to a simpler phenotype.

      The question about a toy-model with arbitrary phenotypes is helpful, because it allows us to clarify that what we are illustrating here with the biologically realistic example of leaf shapes is a much more generic principle. We can say with confidence that if the toy-model generates a many to one set of outputs (phenotypes) through an algorithmic process whose description length does not grow faster than logarithmically with the size of the genotype space, then it should produce a bias towards simplicity regardless of the number of dimensions, see for example Johnston et al. (2022) and Dingle, Camargo and Louis (2018) for a longer discussion of this more general point which is based on arguments from algorithmic information theory (AIT). We don’t use that framing in the current paper because the basic intuition for GRNs that more complex phenotypes need more parameters fine-tuned, and so have relatively smaller volumes, is more straightforward to understand that the more abstract AIT arguments. Our general prediction that this principle should hold more widely for GRNs can be made both by the more formal AIT route, or via the more heuristic fine-tuned parameter route.

      The discussion of Figure 4 (PCA of parameter space) uses "area" loosely when what's actually being measured is bin count in a 2D projection of a highdimensional space. I would think that, in general, PCA projections can be misleading about volume in the full parameter space, but I can't tell if that's an issue in this case. Some comments/thoughts here would be useful.

      The quantitative estimate of phenotype frequencies is computed directly in the full parameter space and does not depend on PCA. Ie. We estimate that the total volume of viable leaves maps to simple unlobed leaves about 80% of the time. However, the volume is extremely high-dimensional, and so hard to visualise. PCA is used solely to provide an interpretable visualization of this otherwise high-dimensional structure. The PCA plots in Fig 4 and Fig S16 are there to be illustrative, not quantitative. Because the volume differences are large, we do not think that the projections of the main PCA components would be misleading on at least the ordering of the sizes of the parameter space components that map to each leaf shape. We provided a similar analysis for other projections -- PC1-PC6 (supplementary information section "PCA occupancy for higher dimensions"), finding the same trend. To make this point clearer, we have now changed the sentence in the Fig. 4 caption slightly “This (reveals that --> illustrates how) unlobed leaves occupy a larger region of model parameter space than more complex shapes and that this larger space also contains the majority of more complex leaves.”

      The classifier validation section is in the Methods section, but it seems critical to the whole story. The < 80% agreement with manual classification could propagate to the rest of the estimates in the paper. Again, some comments/thoughts here would be useful.

      We have repeated the analysis of the agreement between by-eye and automatic morphometric classification. Generating a confusion matrix for the two classification methods shows that the agreement is high for unlobed, dissected and compound, with the main source of disagreement being leaves that were classified as lobed by-eye being classified as either unlobed or dissected by the automatic-morphometric method. The proportion of by-eye lobed leaves classified by the automatic morphometric method as either unlobed (27%) or dissected (23%) is relatively balanced, which we think will help cancel out some error as well. Moreover, we find that the agreement between the automatic-morphometric method and by-eye classification increases to 90.0% when using the categories unlobed and all other categories grouped into one. This is the most important classification for our finding that development and phylogeny are both biased towards unlobed.

      The authors should explain Mut2 and Mut5 in the main paper with a sentence or two, at least schematically, because how you mutate is obviously very relevant to interpreting a paper about biases in variation.

      In the results section we have added a sentence for more detail on the random walk.

      "[We mutated the initial sample using a random walk algorithm with two different mutational schemes, MUT2 (alg. 1) and MUT5 (alg. S2).] These algorithms work by iterating through model parameters one by one and perturbing the value by a small amount. We then [automatically classified the resulting shapes...]"

      Moreover, in methods section C there is already a more detailed description of both algorithms.

      “MUT2 (alg. 1) iterates through the parameters in a random order, and attempts to change the parameter by a value selected at random from an array of numbers randomly generated at 3 different orders of magnitude. MUT5 (alg. S2) is the same as MUT2 except the value each parameter is multiplied by 10% of the range of that parameter within the initial leaves (fig. S1). The aim here was to provide some way of accounting for the biologically relevant sampling range. "

      Moreover, the MUT2 algorithm is described in pseudocode in Algorithm 1 in the main text, and the pseudocode for MUT5 is in supplementary information section S1 C, as algorithm S2.

      The two mutational schemes use additive perturbations to individual parameters. Real mutations presumably affect regulatory networks in more structured ways (e.g., changing binding affinities that affect multiple parameters simultaneously). How sensitive are the results to the assumption of independent single-parameter mutations?

      The referee raises an interesting and well-known issue concerning this widely studied class of GRN models. Without a detailed understanding of how individual genetic mutations map onto model parameters, it is difficult to determine with confidence whether a mutation would produce correlated changes in certain sets of parameters. Our main argument, however, is that the primary source of the observed bias is geometric: the volume of parameter space (or equivalently, the entropy) corresponding to simple leaf morphologies is substantially larger than that corresponding to complex morphologies. As long as mutations explore parameter space approximately symmetrically, even if they involve correlated changes in multiple parameters, larger phenotype regions will tend to be encountered more frequently and retained for longer than smaller regions. We therefore expect the observed bias to be robust to many alternative mutation models, although quantifying this robustness is an interesting direction for future work.

      The connectedness argument is made using a 2D PCA projection. Is there a way to check this statement in the full parameter space or perhaps in higher dimensional projections to test the robustness of this result? Connected components can merge/split under different projections.

      Constructing the nearest neighbour graph for the full dimensional data results in the following no. connected components: unlobed-146, lobed-274, dissected-255, compound-315. This follows the same pattern identified for the PC1-PC2 projection, that unlobed splits into fewer connected components than other leaf shape categories.

      References:

      Dingle, K., Camargo, C.Q. and Louis, A.A. (2018) ‘Input–output maps are strongly biased towards simple outputs’, Nature Communications, 9(1), p. 761. Available at: https://doi.org/10.1038/s41467-018-03101-6.

      Johnston, I.G. et al. (2022) ‘Symmetry and simplicity spontaneously emerge from the algorithmic nature of evolution’, Proceedings of the National Academy of Sciences, 119(11), p. e2113883119. Available at: https://doi.org/10.1073/pnas.2113883119.

      Lynch, M. and Conery, J.S. (2003) ‘The Origins of Genome Complexity’, Science, 302(5649), pp. 1401–1404. Available at: https://doi.org/10.1126/science.1089370.

      Pagel, M. (1994) ‘Detecting correlated evolution on phylogenies: a general method for the comparative analysis of discrete characters’, Proceedings of the Royal Society of London. Series B: Biological Sciences, 255(1342), pp. 37–45. Available at: https://doi.org/10.1098/rspb.1994.0006.

      Runions, A., Tsiantis, M. and Prusinkiewicz, P. (2017) ‘A common developmental program can produce diverse leaf shapes’, New Phytologist, 216(2), pp. 401–418. Available at: https://doi.org/10.1111/nph.14449.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer 1 (Public review):

      Summary:

      This study presents a systematic investigation of parent-of-origin effects on gene expression using trio-based data from the Framingham Heart Study, which is notable for its relatively large number of trios. By combining whole-genome and RNA sequencing data, the authors examined the extent to which gene expression is influenced by whether genetic variants are inherited maternally or paternally.

      The authors report that parent-of-origin eQTLs are widespread, identifying 15,893 eQTLs from 14,733 variants and 1,824 genes that were significant in paternal, maternal, or joint tests but not detected by traditional eQTL approaches. They further classified these associations based on the relative strength and direction of paternal and maternal effects, highlighting a subset with opposing directions. The study also highlighted eGenes linked to known imprinted genes as well as those with opposing parent-specific effects, and observed that paternal eGenes are enriched for drug targets. Finally, the work revisits previous findings in which eQTL studies were used to interpret disease-associated loci, emphasizing that conventional eQTL analyses without testing the parent-of-origin may mislead gene prioritization efforts. The study recommends that future downstream analyses, such as Mendelian randomization, take into account the provided lists of SNPs and eGenes and exclude those with strong parent-of-origin effects when linking genetic regulation to disease risk.

      Strengths:

      The major strength of the study lies in the scale and quality of the dataset, the trio-based design, and the systematic application of statistical tests for parent-of-origin effects. The strengths thoughtfully employed Bayes factors rather than p-values to provide stronger evidence of association, which adds rigor to their analyses. These design choices provide compelling evidence that parent-of-origin effects are widespread and that conventional eQTL analyses miss a substantial fraction of regulatory variation. The results are clearly presented and supported by robust analyses, including the identification of opposing parental effects and the enrichment of paternal eGenes for drug targets. Notably, the two examples demonstrating how these findings can reshape disease gene prioritization highlight the broader impact of the study and encourage further work in the community to incorporate parent-of-origin effects.

      Weaknesses:

      The main limitations of the study are threefold.

      First, there is a lack of replication in independent cohorts, which is understandable given the difficulty of identifying datasets with a comparable number of trios, but replication would help establish the generalizability of the findings.

      We fully agree with the reviewer that replication in an independent cohort is a crucial step for establishing generalizability. As the reviewer notes, the Framingham Heart Study, with its 1,477 trios possessing both WGS and RNA-seq data, represents a uniquely powerful and, to our knowledge, currently unmatched resource for this specific type of parent-of-origin eQTL analysis.

      In the absence of an external cohort of comparable size and data richness, we have taken several steps to ensure the internal validity and robustness of our findings within the current study, which we will clarify and expand upon in the revised manuscript:

      Positive Control Validation: We explicitly used well-established, bona fide imprinted genes (e.g., MEG3, NDN, SNURF, as listed in Table 1 and Figure 1) as positive controls. The fact that our analysis correctly identifies their known parent-of-origin expression patterns (e.g., maternal eQTL for MEG3, paternal eQTL for NDN) serves as a powerful internal validation of our phasing methodology, statistical models, and significance thresholds. This demonstrates that our approach has the power to detect true POE signals.

      Conservative Calling Criteria: As the reviewer suggests, we prioritized specificity. Our definition of eQTL sets (Section 4.6) uses stringent thresholds (e.g., log<sub>10</sub> BF > 4 for primary signals and θ = log<sub>10</sub> 2 for exclusivity). We explored different θ parameters (Supplementary Table S2) and chose the one that minimized the inclusion of false positives, ensuring that our core gene sets (e.g., G<sub>1</sub>,G<sub>0</sub>,G<sub>2</sub>) are high-confidence discoveries.

      Rigorous Analytical Pipeline: As we note in the revised text, our conclusions are supported by a robust analytical pipeline. This includes trio-based phasing validated by simulation (Supplementary Table S1), the use of linear mixed models to control for relatedness and population structure, and the application of Bayes factors which inherently penalize variants with low minor allele frequencies, thereby reducing spurious associations.

      We believe these internal consistency checks and methodological rigor provide strong confidence in our findings. To further facilitate external replication, we will make the full list of POE eQTLs and eGenes available as a comprehensive resource (as noted in the Discussion and Supplementary Materials), enabling other researchers to validate these findings as appropriate datasets become available.

      Second, while Bayes factors are thoughtfully used to assess evidence of association, the paper does not fully explore how the chosen thresholds translate to the expected rate of false positives. For example, a minor allele frequency cutoff of 1% was applied, which seems somewhat arbitrary, and without reporting the allele frequency distribution of the identified eQTLs, it is unclear whether rare variants disproportionately contribute to the signals, potentially affecting the reliability of discoveries.

      We thank the reviewer for raising this important point regarding the calibration of our significance thresholds and the potential role of rare variants. We address this by clarifying the relationship between Bayes factors, prior odds, and false discovery rates, and by providing a more detailed characterization of the variants we identified.

      Bayes Factors and False Discovery: The reviewer is correct that the connection between a Bayes factor threshold and a false positive rate is not direct as it has to take into account of prior odds. As we briefly noted, for a given prior odds of association (e.g., 1 in 100 or 1 in 1000 for a cis-eQTL), a log<sub>10</sub> BF = 4 corresponds to a posterior probability of association (PPA) of 0.99 or 0.90 respectively. Consequently, 1 − PPA can be interpreted as the local false discovery rate (lfdr), as we have now explicitly stated in Section 2.2 (citing Soloff et al., 2024). Our choice of log<sub>10</sub> BF = 4 was therefore chosen to ensure a very low or modest lfdr (depending on the prior odds) for our primary findings.

      Minor Allele Frequency Threshold: The 1% MAF cutoff was indeed a pre-analysis filtering step. It was chosen based on the power afforded by our sample size of 1,477 trios. For variants rarer than 1%, our study is underpowered to detect associations, and any signals would be highly unstable. Importantly, the reviewer’s concern about rare variants disproportionately contributing to signals is further mitigated by our use of Bayes factors. As we note in Section 2.2, the prior used in our Bayes factor computation (with σ = 0.5 in the prior for effect sizes, as described in Section 4.4) inherently penalizes variants with small minor allele frequencies. This is because for a given effect size, the evidence for association is weaker for a rare variant than a common one. Thus, the combination of a pre-analysis MAF filter and the Bayesian analysis itself guards against spurious findings driven by very rare alleles.

      Allele Frequency Distribution: To directly address the reviewer’s request for transparency, in the revised manuscript we include a supplementary figure (e.g., Supplementary Figure S4) showing the distribution of minor allele frequencies (1000 genomes European descents) for the SNPs identified in paternal eQTL set S<sub>P</sub> and maternal eQTL set S<sub>M</sub>. This empirically demonstrate that our findings are not disproportionately driven by low-frequency variants and provide a more complete picture of the genetic architecture underlying these POE signals. We also add a sentence to the Results section (Section 2.5) summarizing this distribution.

      Third, the ancestry background of the study samples is not reported, which could be a confounding factor in the genetic analyses.

      We thank the reviewer for highlighting this omission. In the revised manuscript, we explicitly report the ancestry background of the Framingham Heart Study participants analyzed. Consistent with previous reports on this cohort, the vast majority of samples are of European descent.

      Crucially, as the reviewer suggests, population stratification can be a confounder in genetic studies. To mitigate this, our analysis employed a linear mixed model (Section 4.4) that includes a random effect with a covariance structure defined by the genetic relatedness matrix (GRM). This approach is specifically designed to control for spurious associations due to both subtle population structure and known relatedness among individuals, ensuring that our findings are robust to these potential confounders.

      Reviewer 2 (Public review):

      Summary:

      The authors have used 1477 sequenced trios with available gene expression data in the offspring to discover eQTLs that act in a parent-of-origin specific manner. The classified associated SNPs are tested for enrichment for GWAS hits, drug target genes, etc.

      Strengths:

      The manuscript presents an impressive analysis of a very rich data set of parent-of-origin eQTLs. To my knowledge, it is one of the largest studies of its kind, most analyses are sound, and the results are of interest to many in the field and potentially beyond. The different ideas of follow-up analyses are useful and make sense.

      Weaknesses:

      While in general the analyses are well-conducted, I noticed a major issue with the POE eQTL classification, which puts into question most of the downstream analysis. In light of this problem, most of the analysis would need to be rerun, which represents a major revision of the paper, but is straightforward to repair.

      We appreciate the reviewer’s concern and take it seriously. However, we believe the issue stems from a misunderstanding of our classification framework. We clarify our reasoning below, and we are confident that no re-analysis is necessary. In fact, our Bayesian approach was specifically chosen to avoid the very problem the reviewer raises.

      The major problem with the classification of POEs is that simply having significant maternal, but insignificant paternal effect is not an indicator of POE, this happens widely for SNPs with no POE whatsoever (it can happen by chance even when both maternal and paternal effects are the same and non-zero - the authors can see it via simulations under the null [maternal=paternal effect]).

      The reviewer raises a valid statistical concern: under the null hypothesis of equal maternal and paternal effects (β<sub>0</sub> = β<sub>1</sub>≠ 0), sampling variation could occasionally produce a scenario where one effect appears significant and the other does not. This is indeed a form of Type II error (failing to detect a true non-zero effect for one of the alleles).

      However, this is precisely why we chose Bayes factors over p-values. A key advantage of Bayes factors is that they are not blind to power. P-values are calculated solely under the null hypothesis and do not incorporate any information about the alternative hypothesis or the study’s power to detect it. Consequently, when power is low (e.g., due to minor allele frequency differences between paternal and maternal alleles), p-values can be misleading.

      In contrast, Bayes factors are computed under both the null and alternative hypotheses. They inherently incorporate power through the prior specification. As we note in Section 2.2, “Bayes factors penalize genetic variants with small allele frequencies to reduce false positives.” This means that a SNP where, by chance, one allele appears significant and the other does not—but where power is low due to allele frequency imbalance—will not receive a high Bayes factor, because the evidence is appropriately discounted.

      In order to be able to talk about POE, first, a significant difference between maternal and paternal effects needs to be claimed. Therefore, none of the 4 sets of POE eQTLs are justified. To me, the only relevant criterion to pick POE SNPs is the P-value when comparing the maternal and paternal effects.

      We respectfully disagree with the reviewer’s assertion that our approach to POE eQTL classification are not justified. There are multiple biologically meaningful patterns of parent-of-origin effects, and our classification scheme was designed to capture this diversity:

      (1) Paternal-specific eQTL (β<sub>0</sub> = 0, β<sub>1</sub> ≠ 0)

      (2) Maternal-specific eQTL (β<sub>0</sub> ≠ 0, β<sub>1</sub> = 0)

      (3) Opposing eQTL (β<sub>0</sub> ≠ 0, β<sub>1</sub> ≠ 0,β<sub>0</sub> × β<sub>1</sub> < 0)

      (4) Genotype eQTL (β<sub>0</sub>= β<sub>1</sub> ≠ 0)

      The reviewer’s proposed test (H<sub>0</sub>: β<sub>0</sub> = β<sub>1</sub>) collapses these distinct biological scenarios into a single binary outcome. For example: A purely paternal-specific eQTL (β<sub>0</sub> = 0, β<sub>1</sub> ≠ 0) would indeed show a significant difference, and would be captured by the reviewer’s test. However, a gene like ZNF890P in Table 1, where both effects are significant and in the same direction but of different magnitudes, would also show a significant difference. In the reviewer’s framework, this would be classified as a POE eQTL, yet biologically it behaves more like a genotype eQTL with an allelic imbalance. Our framework correctly separates these cases.

      Moreover, the reviewer’s proposed test is a nested special case of our broader approach. As we note in our response, our paternal-specific test (H<sup>0</sup>: β<sub>0</sub> = β<sub>1</sub> = 0 vs H<sub>1</sub>: β<sub>0</sub> = 0,β<sub>1</sub> ≠ 0) is a more constrained hypothesis that yields a subset of the SNPs that would be identified by the reviewer’s difference test, were it to have sufficient power. Our approach is therefore more conservative for classifying paternal- or maternal-specific eQTLs, not less.

      The definitions of the 4 groups are based on somewhat ad hoc priors, BF thresholds, etc. Also, in Section 4.6, the value of theta is arbitrarily chosen (along with the threshold of 4 to declare POE). In my opinion, the clean treatment of the 4 groups would start with a significant P-value (beta-maternal vs beta-paternal). Within this set, you can then use the original criteria presented in the paper, but only among these associations where there is solid evidence of different parental effects.

      We take strong issue with the characterization of our prior specifications and thresholds as “ad hoc” or “arbitrary.” In Bayesian analysis, prior specification is a principled and transparent modeling choice, not an arbitrary one.

      (1) Choice of log<sub>10</sub> BF = 4 threshold: As stated in Section 2.2, this threshold was chosen based on explicit considerations of prior odds and posterior probability of association. For a prior odds of 1:1000 (a reasonable guess for cis-eQTLs), this BF corresponds to a posterior probability of association of 0.91. If one prefers a more optimistic prior odds of 1:100, the PPA becomes 0.99. The threshold is therefore grounded in decision theory, not whim.

      (2) Choice of θ in Section 4.6: We explicitly state that we explored multiple values of θ(0, log<sub>10</sub> 2, log<sub>10</sub> 3) and chose θ = log<sub>10</sub> 2 because it “produced minimum G<sub>1</sub> and G<sub>0</sub> that contain known imprinted genes.” This is a principled, data-driven calibration step using positive controls, not an arbitrary selection. The transparency of this process is a strength, not a weakness.

      (3) Comparison to p-value thresholds: The reviewer suggests that p-value thresholds are somehow less arbitrary. However, the conventional p-value threshold of 0.05 is itself a historical convention with no universal justification. Moreover, as we note, p-values do not account for power differences across SNPs. A p-value of 5 × 10<sup>−8</sup> from a SNP with 40% MAF is not comparable to the same p-value from a SNP with 1% MAF, because the power to detect the association differs dramatically. Bayes factors automatically adjust for this through the prior, making them more comparable across variants, not less.

      In revision, we added a section in supplementary to review relationships between p-values, Bayes factors, and FDR.

      Recommendations for the authors:

      Reviewer 1 (Recommendations for the authors):

      Here are some suggestions to improve the study:

      (1) Provide information about the ancestry background of participants and consider including ancestry principal components in the eQTL models, as is commonly done, to account for population structure.

      We thank the reviewer for this suggestion. In the revised manuscript, we explicitly state that the participants in the Framingham Heart Study are predominantly of European descent, consistent with previous publications from this cohort. Regarding population structure, we respectfully note that our analysis already employs a linear mixed model (Section 4.4) that includes a random effect with a covariance structure defined by the genetic relatedness matrix (GRM). This approach is widely regarded as more robust than including a limited number of principal components, as it accounts for both fine-scale population stratification and known relatedness simultaneously.

      (2) Conduct sensitivity analyses using different Bayes factor cutoffs to assess the robustness of the findings.

      We appreciate the reviewer’s concern about threshold robustness. In fact, we already conducted a form of sensitivity analysis during the classification step. As described in Section 4.6 and shown in Supplementary Table S2, we explored multiple values of θ (0, log<sub>10</sub> 2, and log<sub>10</sub> 3) and observed how they affected the composition of our gene sets. The choice of log<sub>10</sub> BF = 4 for significance was similarly grounded in posterior probability calculations (Section 2.2). To further address the reviewer’s point, we add a Supplementary Table S3 for counts of eQTL and eGenes under different Bayes factor threshold. This demonstrates that our most significant claim, the abundance of POE eQTL, are not overly sensitive to the specific cutoff.

      (3) In the GWAS examples for KCNQ1 and CDKN1C, the assessment of whether the SNPs act as eQTLs for the two genes is based on a single BF threshold, which may be influenced by differences in gene expression levels. The authors could compare the corresponding effect sizes of these SNPs on both genes to provide a more nuanced investigation. While the limitation of missing data from other tissues is discussed in the paper, it remains possible that KCNQ1 plays a role in tissues more relevant to T2D.

      This is an excellent suggestion for a more nuanced investigation. We re-examined the effect sizes for the SNP rs2237892 in our published results. For gene CDKN1C, the paternal log<sub>10</sub> BF<sub>1</sub> = −0.477 and maternal log<sub>10</sub> BF<sub>0</sub> = 4.94, the normalized maternal effect in joint analysis is −4.86 vs −0.74 for paternal. Unfortunately, the published results has no eQTL for KCNQ1, which according to our selection creteria means maximum log<sub>10</sub> BF < 3 for all tests (genotype, paternal , maternal, joint). The concern for different gene expression level may affect BF is valid. We preempt this pitfall by quantile normalization of gene expression levels after controlling for GC content (as documented in Method Section). We agree with the reviewer that the lack of data from pancreatic tissues is a limitation. We add a sentence in revelant section to acknowledging that while whole blood is a valuable and accessible tissue, replication in T2D-relevant tissues (e.g., pancreas, adipose) would be an important future direction, and our findings provide a hypothesis for such targeted investigations.

      Reviewer 2 (Recommendations for the authors):

      Major comments:

      There are some literature elements missing:

      (1) Hofmeister has a newer and larger study [https://pubmed.ncbi.nlm.nih.gov/40770099/].Please cite that too; it also has POE pQTLs, which is relevant.

      (2) POE in pigs has been explored [https://www.nature.com/articles/s41467-02562243-6], please cite it.

      (3) An insightful review covering the mechanisms of POE for gene expression (https://www.sciencedirect.com/science/article/pii/S2352154618300482) should be cited.

      (4) Further studies on POE in gene expression in social insects (https://royalsocietypublishing.org and in mice (https://www.biorxiv.org/content/10.1101/2023.08.24.554674v1.full) are also relevant.

      We thank the reviewer for bringing these important references to our attention. We incorporated the suggested citations in the revision to provide a more comprehensive context for our work, including the newer POE pQTL study by Hofmeister et al., the findings in pigs, and the mechanistic review.

      While it’s OK to report and rank SNPs by BF, it is necessary to show association P-values as well. It is not explained in the text around the Table how the P-value is obtained in the Table. And it is important to show how their priors translate to FWER control. What is the FWER when picking SNPs at a certain BF value? 1-PPA and local FDR depend on the choice of the prior, but we need a prior-independent measure of FDR/FWER.

      We appreciate the opportunity to clarify. The p-value presented in Table 1 (column “P”) is indeed the frequentist p-value testing the null hypothesis of equal maternal and paternal effects (H<sub>0</sub> : β<sub>0</sub> = β<sub>1</sub>), as described in Section 4.5. We included this to provide a familiar metric for readers, but our discovery framework relies on Bayes factors for the reasons outlined in Section 2.2.

      Regarding error control, the reviewer is correct that 1-PPA is a local FDR that depends on the prior. We chose to control the local rate of false discoveries rather than the Family-Wise Error Rate (FWER) because FWER control (e.g., via Bonferroni) is often excessively conservative for exploratory analyses like eQTL mapping, especially given the correlation among tests due to LD.

      Our Bayesian approach provides a more nuanced measure of evidence at the level of each individual test, which is precisely what is needed for prioritizing SNPs with parent-of-origin effects.

      The demand for a prior-independent measure of FDR is conceptually problematic. Any probabilistic statement about a specific hypothesis being true or false necessarily requires a prior—this is a fundamental consequence of probability theory. Frequentist FDR, while prior-independent in one sense, does not provide a probability that a particular finding is false; it is a long-run error rate over many tests. Methods like q-values, often described as “prior-free,” still depend on implicit assumptions (e.g., the estimate of π<sub>0</sub>, independence of tests, and a mixture of effect sizes).

      In our specific context of cis-eQTL analysis, these assumptions are particularly questionable. LD induces correlation among nearby SNPs, violating the independence required for stable π<sub>0</sub> estimation. Moreover, effect sizes in a region are not randomly mixed—SNPs in high LD tend to have similar effect directions and magnitudes, which can bias the mixture model underlying q-value approaches. Our Bayesian approach, by modeling each SNP individually, avoids these cross-SNP assumptions.

      Importantly, while posterior probabilities depend on the choice of prior (π<sub>0</sub>), we have verified that our conclusions are robust across a wide range of plausible π<sub>0</sub> values (0.9,0.99,0.999). Given our extremely stringent Bayes factor threshold (BF<sub>j</sub> > 10<sup>4</sup>), the posterior probability for a maternal effect exceeds 0.90 for any π<sub>0</sub> < 0.999. Thus, the prior dependence is practically irrelevant for the SNPs we report.

      In revision, we added a section in Supplementary to describe the connections between p-value, Bayes factor, and FDR. We hope this will clarify that a (seemingly) prior independent FDR has a hidden assumption that cis-eQTL analysis is likely to violate.

      The major problem with the classification of POEs is that simply having significant maternal, but insignificant paternal effect is not an indicator of POE, this happens widely for SNPs with no POE whatsoever (it can happen by chance even when both maternal and paternal effects are the same and non-zero - the authors can see it via simulations under the null [maternal=paternal effect]). In order to be able to talk about POE, first, a significant difference between maternal and paternal effects needs to be claimed. Therefore, none of the 4 sets of POE eQTLs are justified. To me, the only relevant criterion to pick POE SNPs is the P-value when comparing the maternal and paternal effects. The definitions of the 4 groups are based on somewhat ad hoc priors, BF thresholds, etc. Also, in Section 4.6, the value of theta is arbitrarily chosen (along with the threshold of 4 to declare POE). In my opinion, the clean treatment of the 4 groups would start with a significant P-value (beta-maternal vs beta-paternal). Within this set, you can then use the original criteria presented in the paper, but only among these associations where there is solid evidence of different parental effects.

      We respectfully disagree with the reviewer’s assertion that a significant difference between maternal and paternal effects is the only valid criterion for defining POE, and we maintain that our classification is statistically sound and biologically meaningful.

      The Problem with the “Difference-Only” Approach: The reviewer’s proposed filter (a significant p-value for β<sub>0</sub> ≠ β<sub>1</sub>) is a single hypothesis test. Our goal was to classify eQTLs into multiple, distinct biological categories (paternal-specific, maternal-specific, opposing, etc.). The “difference-only” test collapses these categories. For example, a purely paternal-specific eQTL (β<sub>0</sub> = 0,β<sub>1</sub> ≠ 0) and a gene like ZNF890P (β<sub>0</sub> ≠ 0, β<sub>1</sub> ≠ 0, β<sub>0</sub> > β<sub>1</sub>) would both show a significant difference. In the reviewer’s framework, they would be lumped together, obscuring the fact that one is an imprinted gene and the other is a standard eQTL with allelic imbalance. Our framework correctly separates them.

      Bayes Factors are Not “Ad Hoc”: The choice of prior (σ = 0.5) follows established literature for linear model Bayes factors (Servin and Stephens, 2007). The threshold of log<sub>10</sub> BF = 4 was chosen based on its relationship to posterior probability (0.91-0.99 given reasonable prior odds), which is a transparent and principled decision rule. The selection of θ in Section 4.6 was calibrated using a positive control set of known imprinted genes, ensuring our definitions were conservative and accurate. This is the opposite of arbitrary.

      The Suggested Procedure Has Low Power: One can run the following simple R code to verify. We simulate maternal alleles xx and maternal alleles yy, then simulate phenotype with β<sub>xx</sub> > 0 and β<sub>yy</sub> = 0 (maternal effect only). We fit the joint model and compute p-values for the null β<sub>xx</sub> = β<sub>yy</sub> as suggested by reviewer. From the joint fit, we also extract p-values based on the null β<sub>xx</sub> = 0 and β<sub>yy</sub> = 0 respectively. The simulation was repeated 1000 times and p-values were stored in a matrix.

      We call positives based on suggested procedure, and compare number of positives called using marginal p-values at two threshold of 1×10<sup>−5</sup> and 1×10<sup>−6</sup> to declare significance. We used threshold of 0.01 to declare insignificance.

      The result demonstrates that the suggested procedure has a much lower power compared to the procedure based on marginal statistics.

      For the above reasons, the follow-up enrichment analysis is somewhat questionable. Most enrichments are non-significant, and it is likely because the SP and SM groups are diluted with SG SNPs. The P1-P9 groups have nothing to do with POE, and although the observation of increased enrichment for GWAS SNPs with increased pleiotropy is interesting, it is irrelevant for POE.

      We will address the dilution concern below. We agree that P1-P9 groups are not directly related to POE. But this is an interesting observation non-theless. As we found such an observation is missing in the literature, we ask to keep it in the paper.

      In the same way, section 2.7 is not supported; the claimed maternal and paternal POEs are heavily diluted by simple marginal associations. The same holds for sections 2.82.10. A striking example is Table 3: for clinical trial targets, paternal/maternal eQTLs behave just like simple marginal eQTLs (G<sub>G</sub>). A similar pattern emerges for combined target enrichment.

      The reviewer’s concern that our S<sub>P</sub> and S<sub>M</sub> sets are “diluted with S<sub>G</sub> SNPs” is precisely the issue our Bayes factor thresholds were designed to prevent. By requiring one effect to be significant and the other to be below a low threshold (θ), we explicitly excluded SNPs where both effects are significant and in the same direction (which defines S<sub>G</sub>).

      Regarding Table 3, the reviewer’s interpretation differs from ours. The fact that paternal eQTLs (G</sub>P</sub>) show significant enrichment for drug targets, while genotype eQTLs (G<sub>G</sub>) also show enrichment, does not imply dilution. Rather, it suggests there is an overlap in the biological importance of these gene sets, which is expected. The key message of the finding is the asymmetry: G<sub>P</sub> is significantly more enriched than G<sub>G</sub> (p=0.035 for combined targets), a pattern that would be washed out if G<sub>P</sub> were merely a diluted version of G<sub>G</sub>. This asymmetry supports the interesting biological hypothesis (Moore and Haig, 1991) we discuss. The non-significance for G<sub>M</sub> further highlights this asymmetry.

      I’m not sure how MR would be biased by POE: MR is conducted only if there is a marginal association, i.e., the average maternal and paternal effects are significant. If the expression is causal for a trait, the POE effect is propagated to the outcome; hence, the SNP effect on the exposure will be equally biased as the SNP effect on the outcome, and these cancel out, and the causal effect remains unbiased. Can the authors propose a concrete example of maternal/paternal effects that demonstrates their claimed bias?

      We thank the reviewer for this insightful question, which allows us to clarify our point with a concrete example from our data.

      Consider a scenario where one wishes to use Mendelian Randomization (MR) to test whether the expression of gene NECAB3 causally influences a particular trait (e.g., obesity). The reviewer is correct that if the causal effect is homogeneous, the average effect might still be captured. However, the bias we caution against arises in stratified analyses or in the interpretation of the genetic instrument itself.

      Take the SNP rs4911348 and its effect on NECAB3 (Figure 2). The genotype model shows no marginal association. Therefore, if a researcher were conducting a standard MR study using this SNP as an instrument for NECAB3 expression, they would discard it as an invalid instrument due to the lack of a marginal association. They would miss the true underlying biology entirely. The causal effect of NECAB3 on the trait would be masked in the full population.

      More subtly, even if a SNP has a marginal association, using it as an instrument while ignoring POE can lead to incorrect effect estimates in population subgroups defined by parent of origin. This is analogous to ignoring effect modification. For instance, if a treatment (exposure) has a different effect depending on which parent it came from (which is impossible, but the genetic propensity for the exposure does), failing to account for this can bias the instrumental variable estimate if the instrument’s strength varies by an unmeasured factor (parental origin).

      Our advice to “check the list of POE SNPs” is a practical caution: if the instrument for an exposure exhibits strong POE, the standard MR assumptions about the homogeneity of the instrument’s effect may be violated, potentially leading to biased estimates or incorrect conclusions about causality.

      Minor comments:

      (1) In Table 1, the last column header should be -log10(P), not ”P”.

      The column labelling is an editorial choice to prevent table overflow. This particularly labelling was explained in the caption.

      (2) While BFg/0/1/j are explained in the text, these notations should be explained in the Table caption as well.

      Added explanation in caption.

      (3) It should also be mentioned in the Table 1 caption how these top 10 SNPs were chosen.

      These are sentinel eQTL for each gene. We think the first paragraph of Section 2.3 explains clearly.

      (4) “may ”acquires” a cis-eQTL through” → ”may ”acquire” a cis-eQTL through”.

      Corrected. Thank you.

      (5) “which retained 16, 969 genes out of total 58103”, I assume the 58103 are transcripts, not genes.

      You are absolutely correct. We added transcripts after 58103.

      (6) In Equation (1), Z is not defined. In this concrete setting, isn’t it simply the identity matrix?

      Yes. Z is the identitity (loading) matrix for human study. We added a sentence to clarify in revision.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (on non-trivial pattern transformations):

      (3) All modelling is confined to one spatial dimension, and the very definition of a "non-trivial" transformation is framed in terms of peak positions along a line, which clearly must be reformulated for higher dimensions. It's well-known that diffusions in 1, 2, and 3 dimensions are also dramatically different, so the relevance of the three-class taxonomy to real multicellular tissues remains unclear, or at least should be explained in more detail.

      Reviewer #2 (on non-trivial pattern transformations):

      (5) The definition of non-trivial pattern formation is provided only in the Supplementary Information, despite its central importance for interpreting the main results. It would significantly improve clarity if this definition were included and explained in the main text. Additionally, it remains unclear how the definition is consistently applied across the different initial conditions. In particular, the authors should clarify how slopebased measures are determined for both the random noise and sharp peak/step function initial states. Furthermore, the authors do not specify how the sign function is evaluated at zero. If the standard mathematical definition sgn(0)=0 is used, then even a simple widening of a peak could fulfill the criterion for non-trivial pattern transformation.

      There was indeed a problem on how we defined non-trivial pattern transformations in the original version. This definition was not clear enough beyond 1D. We now provide a simple clear definition in the main text that applies to all dimensions (“P1” and “P2” in the second page of the introduction).

      As we now explain through the main text, even if the solution of the heat/diffusion equation depends on the dimension of the system, our classification of gene networks (and the mathematical analyses we use) does not depend on the dimensionality of the system. However, some aspects of the specific pattern transformations possible from these networks depend on the dimensionality of the system. In the current version of the article, every time we explain something about the resulting patterns in 1D, we also explain it for the resulting patterns in 2D and 3D. We also have added figures for the 2D cases (in current Fig.1 and Fig.9). We now explicitly explain how the possible resulting patterns in space can depend on the boundaries and shapes of the system (i.e. the distribution of cells in space) (see specially the 5th paragraph of the discussion).

      The criticisms about “slope-based measures” mentioned by reviewer 2, is now addressed in a paragraph at the end of the introduction (here we added it):

      “It is worth noting that these three basic initial patterns correspond to spatially discontinuous functions: in homogeneous with noise initial patterns, white noise is discontinuous by definition; in spike and combined spike-homogeneous initial patterns, there is a concentration discontinuity between cells on the edge of the spike and nearby cells outside the spike. However, once extracellular signal diffusion begins, these sharp boundaries are smoothed into differentiable gradients, where critical points can be properly defined (e.g., at the center of the initial spike).”

      The main concern among these relates to the validity of our linearization of the model equations and the extension of the results obtained for the linear system to the fully nonlinear system. In this regard, the reviewers’ comments are:

      Reviewer #1 (on linearization):

      (2) A central step in the model formulation is the linearisation of the reaction term around a homogeneous steady state; higher-order kinetics, including ubiquitous bimolecular sinks such as A + B → AB, are simply collapsed into the Jacobian without any stated amplitude bound on the perturbations. Because the manuscript never analyses how far this assumption can be relaxed, the robustness of the three-class taxonomy under realistic nonlinear reactions or large spike amplitudes remains uncertain.

      Reviewer #2 (on linearization):

      (2) Most of the proofs presented in the Supplementary Information rely on linearized versions of the governing equations, and it remains unclear how these results extend to the fully nonlinear system. We are concerned that the generality of the conclusions drawn from the linear analysis may be overstated in the main text. For example, in Section S3, the authors introduce the concept of dynamic equivalence of transitive chains (Proposition S3.1) and intracellular transitive M-branching (Proposition S3.2), which pertains to the system's steady-state behavior. However, the proof is based solely on the linearized equations, without additional justification for why the result should hold in the presence of nonlinearities. Moreover, the linearized system is used to analyze the response to a "spike initial pattern of arbitrary height C" (SI Chapter S5.1), yet it is not clear how conclusions derived from the linear regime can be valid for large perturbations, where nonlinear effects are expected to play a significant role. We encourage the authors to clarify the assumptions under which the linearized analysis remains valid and to discuss the potential limitations of applying these results to the nonlinear regime.

      We used three linearizations in the original version of the manuscript. One was to analyze hierarchic networks (in the Hierarchic networks section). In the new version of the article we do not use any linearization to study the hierarchic networks, so this problem is solved.

      The second linearization was in section S3 on transitive chains. We realized that this section is not really necessary at all for the article so we deleted it.

      We keep the third linearization but we now explain why such linearization is useful and valid in a section called “Linear stability analysis”. Thus, through this section we justify this choice (explicitly in its two first paragraphs).

      Regarding Reviewer 2 concerns about large perturbations, we acknowledge that the phrasing using “arbitrary height” may have been confusing. As we now explain in the linear stability analysis section, linear stability analysis assumes perturbations to be small.

      For the homogeneous-with-noise initial pattern, as we explain, these perturbations are assumed to be small because they are actually molecular noise.

      For the spike initial pattern and hierarchic networks the perturbation is not necessarily small. However, by the definition of the spike and combined homogeneous-spike initial patterns, all cells outside the spike start with the same concentration of the extracellular signals that are secreted from the spike (e.g. zero). Thus, even in the case in which extracellular signals concentrations in the spike would be unrealistically high, the amount of extracellular signal diffusing from it can be considered small by simply considering it at a small enough time interval. Thus, right outside the spike the diffusion of extracellular signals from the spike can be treated as a continuous small perturbation for which one can study the stability, as we do in the “Linear stability analysis section”. This we now explain at the end of the introduction and in the “Linear stability analysis” section when we talk about the initial patterns again.

      In the following, we respond to the remaining concerns raised by the reviewers:

      Reviewer #1 (Public review):

      (1) The Results section is difficult to follow. Key logical steps and network configurations are described shortly in prose, which constantly require the reader to address either SI or other parts of the text (see numerous links on the requirements R1-R5 listed at the beginning of the paper) to gain minimal understanding. As a result, a scientifically literate but non-specialist reader may struggle to grasp the argument with a reasonable time invested.

      We acknowledge that the original version of the main text may not be as clear as we intended. Initially, we believed that placing the more technical mathematical passages in the Supplementary Information would make the main text more accessible to readers. We were wrong. We have now moved crucial parts of the supplementary to the main text and adapted the rest of the text accordingly. The most important of those is the new “Linear stability analysis” section and the associated dispersion relation (e.g. Fig.6).

      Reviewer #2 (Public review):

      (1) We have serious concerns regarding the validity of the simulation results presented in the manuscript. Rather than simulating the full nonlinear system described by Equation (1), the authors base their results on a truncated expansion (Equation S.8.2) that captures only the time evolution of small deviations around a spatially homogeneous steady state. However, it remains unclear how this reduced system is derived from the full equations -specifically, which terms are retained or neglected and why- and how the expansion of the nonlinear function can be steady-state independent, as claimed. Additionally, in simulations involving the spike plus homogeneous initial condition, it is not evident -or, where equations are provided, it is not correct- that the assumed global homogeneous background actually corresponds to a steady state of the full dynamics. We elaborate on these concerns in the following:

      We are actually simulating the full nonlinear system described by Equation (1). In the current version we are more explicit about this. As we describe in the introduction and, now, through all the text several times (e.g. in the last paragraph of the model section and in the paragraph before the linear stability section), the aim of the article is to describe necessary requirements for non-trivial pattern transformations. We did not intent to describe all necessary requirements nor sufficient requirements. These requirements are at the level of gene network topology not at the level of f or its parameters. In other words, we just claim that gene networks having specific topological features can lead to some specific types of non-trivial pattern transformations but not to others. We do not say for which specific fs (or its parameters) these pattern transformations are possible, we just say that this can happen for some f, as long as these fulfill our requirements. We do show, however, that without some specific topological requirements there are non-trivial pattern transformations that are not possible, no matter the f (this explicitly stated in the last paragraph of the model section and in the paragraph before the linear stability section). Thus, all the simulations shown in the figures are just examples, with specific fs, of the types of non-trivial pattern transformations possible from each type of gene network topology.

      In all simulations we used the f of the Maini-Miura model. We could have chosen other ones but we happen to chose that f. The presentation of the Maini-Miura model has been revised to improve clarity (equation S6.1 in SI). This model we are simulating fully, we are not doing any linearization for the simulations. That may not have been explained clearly enough in the previous version of the article. We just happen to make a change of variable that may have been confused as a linearization. In the current version, the existence of a homogeneous steady state is parameterized by a tunable g<sup>*</sup>, that can be chosen as for spike initial patterns or g for noise-homogeneous and spike-homogeneous initial patterns. We have also included a proof that the model equations satisfy our conditions R1-5. Indeed, the model is non-linear as long as σ<sub>i</sub>≠0 for some gene product (as we explicitly assume).

      It is assumed that the homogeneous steady states are given by g_i=0 and g_i=c_i, where 1/c_i = \mu_i or \hat{\mu}_i, independently of the specific network structure. However, the basis for this assumption is unclear, especially since some of the functions do not satisfy this condition -for example, f5 as defined below Eq. S8.10.5. Moreover, if g_i=c_i does not correspond to a true steady state, then the time evolution of deviations from this state is not correctly described by Eq. S8.2, as the zeroth-order terms do not vanish in that case.

      In the revised manuscript, homogeneous steady states are parameterized by a tunable g<sup>*</sup>, which can be chosen as for spike initial patterns or g for noise-homogeneous and spike-homogeneous initial pattern. Function f(g) in (S6.1), as well as the specific non-linear entries used in certain simulations, are constructed such that g<sup>*</sup> is indeed a steady state of the system and that conditions R1-R5 are satisfied. We have also corrected some typos in section S6 (previously section S8) of the Supplementary Information, that we believe may have induced the confusion indicated by this reviewer.

      Additionally, the equations used contain only linear terms and a cubic degradation term for each species g_i, while neglecting all quadratic terms and cubic terms involving cross-species interactions (i≠j). An explanation for this selective truncation is not provided, and without knowledge of the full equation (f), it is impossible to assess whether this expansion is mathematically justified. If, as suggested in the Supplementary Information, the linear and cubic terms are derived from f, then at the very least, the Jacobian matrix should depend on the background steady-state concentration. However, the equations for the small deviation around a steady state (including the Jacobian matrix) used in the simulations appear to be independent of the particular steady state concentration.

      As described above we just chose an example f to exemplify the non-trivial pattern transformations possible from each class of gene network topologies. There is no special reason to include, or exclude for that matter, cubic cross-species interactions since the point is just to exemplify the types of possible pattern transformations from each type of gene network topology.

      In addition, we believe that part of the reviewer’s concern may have arisen from a notational ambiguity in the previous version of the manuscript, which has now been corrected: the matrix appearing in f(g) has been renamed from J to W<sup>T</sup>. As stated in the main text, the jacobian of the regulation function f(g) evaluated at the homogeneous steady state must coincide with the transpose of the network weight matrix. With the current equations (S6.1), we have , from which we easily get . Also, it is clear that the Jacobian of f(g) is not independent of g.

      This is why we believe that the differences observed between the spike-only initial condition and the spike superimposed on a homogeneous background are not due to the initial conditions themselves, but rather result from a modified reaction scheme introduced through a questionable cutoff.

      "In simulations with spike initial patterns, the reference value g≡0 represents an actual concentration of 0 and therefore, we must add to (S8.2) a Heaviside function Φ acting of f (i.e., Φ(f(g))=f(g) if f(g)>0 , Φ(f(g))=0 if f(g){less than or equal to}0) to prevent the existence of negative concentrations for any gene product (i.e., g_i<0 for some i)." (SI chapter S8).

      This cutoff alters the dynamics (no inhibition) and introduces a different reaction scheme between the two simulations. The need for this correction may itself reflect either a problem in the original equations (which should fulfill the necessary conditions and prevent negative concentrations (R4 in main text)) or the inappropriateness of using an expanded approximation which assumes independence on the steady state concentration. It is already questionable if the linearized equations with a cubic degradation term are valid for the spike initial conditions (with different background concentration values), as the amplitude of this perturbation seems rather large.

      The Heaviside function does not preclude inhibition, it precludes gene product concentration to be negative. In the current version of the article we do not use the Heaviside function but another similar, but continuous, function. Having this function can indeed affect the dynamics but: 1) does not violate our requirements on f 2) Does not affect which non-trivial pattern transformations are possible from which gene network topology. Without this function non-trivial pattern transformations are still possible from the spike initial pattern through hierarchical networks, in the way we describe in the article. The Heaviside function (and the one we now use) simply allows that to happen more easily, i.e. for a larger range of parameter values. With this function large inhibitions do not lead to negative gene products concentrations while without it, this can happen for some parameter combinations. None of the arguments nor proves in our article requires the Heaviside, or any similar function. Again this is simply because our aim is to identify topological requirements that are necessary, but not sufficient, for non-trivial pattern transformation. So an f that leads to negative gene products concentrations for some parameter combinations but to non-trivial pattern transformations for others, is still valid example of our points (although not the most interesting or realistic example f).

      We distinguish between the spike and combined spike-homogeneous initial patterns simply because they are biologically quite different, i.e. in the former the gene product in the spike is only expressed in the spike and nowhere else. As we describe in the current version the pattern transformations possible from these two different initial patterns are very similar. In the same way, which gene network topologies can lead to which types of non-trivial pattern transformations is not affected by using the Heaviside functions or not (although this can affect the range of parameter values in which this happens).

      Lastly, we note that under the current simulation scheme, it is not possible to meaningfully assess criteria RH2a and RH2b, as they rely on nonlinear interactions that are absent from the implemented dynamics.

      The implementation of nonlinear entries in f(g) whenever they are needed is now made explicit in the corresponding subsection in the main text and in section S6 in the Supplementary Information. This entries also satisfy conditions R1-R5 around the steady state given by g<sup>*</sup>. Again we should insist that the simulated fs are nonlinear (as now explicitly explained in the SI).

      (3) Several statements in the main text are presented without accompanying proof or sufficient explanation, which makes it difficult to assess their validity. In some cases, the lack of justification raises serious doubts about whether the claims are generally true. Examples are:

      "For the purpose of clarity we will explain our results as if these cells have a simple arrangement in space (e.g., a 1D line or a 2D square lattice) but, as we will discuss, our results shall apply with the same logic to any distribution of cells in space." (Main text l.145-l.148).

      The result of which gene network topologies can lead to pattern transformations are based on a linear stability analysis and some logical arguments. As we now explain through the text none of them depends on the number of dimensions nor on the shape of the arrangement of cells. The geometry of the domain can influence the specific form of the resulting patterns, but it does not alter the broader type of resulting patterns (e.g., periodic patterns, peaks emerging around a spike, etc.) that a given gene network topology can produce. We now explicitly discuss these dependencies in the 5th paragraph of the discussion.

      "For any non-trivial pattern transformation (as long as it is symmetric around the initial spike), there exists an H gene network capable of producing it from a spike initial pattern." (Main text l.366f).

      We now provide a more detailed justification of this statement and the limits of its applicability. This is now in section: “The ensemble of possible pattern transformations from spike initial patterns in H networks“. To make this section easier to understand, however, we have also done changes through all the hierarchic networks sections.

      "In 2D there are no peaks but concentric rings of high gene product concentration centered around the spike, while in 3D there are concentric spherical shells." (Main text l. 447ff).

      This result pertains specifically to pattern transformations arising from spike initial patterns. As defined in the text, spike initial patterns are radially symmetric (at least far away from the boundary). Since diffusion preserves radial symmetry, pattern transformations from spike initial patterns in two or three dimensions reduce to effectively one-dimensional transformations along each radial direction. In this framework, each pair of concentration peaks symmetric with respect to the spike in one dimension corresponds to a ridge surrounding the spike in two dimensions, and each ridge in two dimensions becomes a spherical ridge shell around the spike in three dimensions. In the current version we explain what happens in 1D but also, in the same places, what happens in 2D and 3D (and we have added figures to visualize this in 2D, e.g. Fig.1 and Fig.9)).

      (4) The study identifies one-signal networks and examines how combinations of these structures can give rise to minimal pattern-forming subnetworks. However, the analysis of the combinations of these minimal pattern-forming subnetworks remains relatively brief, and the manuscript does not explore how the results might change if the subnetworks were combined in upstream and downstream configurations. In our view, it is not evident that all possible gene regulatory networks can be fully characterized by these categories, nor that the resulting patterns can be reliably predicted. Rather, the approach appears more suited to identifying which known subnetworks are present within a larger network, without necessarily capturing the full dynamics of more complex configurations.

      We acknowledge that our explanation regarding the combination of sub-networks may have been too brief. We now provide a more detailed description in the section “Gene networks combining different classes of subnetworks” and in its sub-sections. There we explore the different ways in which signal subnetworks can be combined (upstream, downstream, in series, in parallel, etc.). However, this section cannot be understood (and that may have been the problem in the original version of the manuscript) without the linear stability analysis section that is now in the main text, and the associated discussion on the dispersion relation and results related to it. These are important because they apply to all gene networks and, thus, constrain the possible gene network topologies and the types of possible pattern transformations. In other words, whichever ways gene networks are combined, they will always be RD-stable (i.e. no pattern transformation) or RD-unstable of the first (periodic resulting patterns) or second kind (other patterns we discuss). In the current version, we combine this fact with other arguments to describe the types of pattern transformations possible by gene networks combining the different classes of subnetworks.

      (6) The manuscript lacks a clear and detailed explanation of the underlying model and its assumptions. In particular, it is not well-defined what constitutes a "cell" in the context of the model, nor is it justified why spatial features of cells -such as their size or boundaries- can be neglected. Furthermore, the concept of the extracellular space in the one-dimensional model remains ambiguous, making it unclear which gene products are assumed to diffuse.

      We now clarify all these points in the first three paragraphs of the “Methods: the Model” section. We have also included a figure for that clarification (Fig.3).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I suggest the following changes for each weakness I mentioned in the Public Review:

      (1) Presentation

      (R1.1) (a) Add a one-page "Key Requirements" table (e.g., immediately after the Model section) that lists every requirement code (R1-R5, I1-I2, RH1-RH2, etc.), its one-line statement, and the SI section where it is proved.

      In the new version of the article each requirement has its own paragraph starting with the requirement label, e.g. R1 (in bold): ….. We introduce each requirement there where they are justified or proven, otherwise the reader may not know where do they come from. We have also hyperlinked all requirements and most equations so that the reader can easily go back to the explanation of each requirement and equation.

      (R.1.2) Provide more figures illustrating the general structure of networks when you describe them; the network sketches could be folded into a single summary figure, so the reader sees all motifs at once. For example, in lines 304-311, it took me a while to understand if the requirement means just A -> k - ... ⊣ j, or it additionally requires A->...->j (through another pathway). It seems that the full requirement is A → k ⊣ j together with an independent positive route A → j. A figure describing the network structure, or at least a schematic "inline" plot in the spirit of what I just wrote, could help. This is just one example, but the text consists of a constant flow of such "diagrams encrypted in prose".

      We have followed the reviewer’s suggestions. Not all fit in a single figure so we have constructed new figures 4 and 5 for that purpose.

      (R.1.3) (b) Also consider supporting the main text with some key formulas and arguments from SI. My overall suggestion here is that it would be great to make the main text less prosaic and more self-consistent, if the journal requirements allow it.

      After the suggestions by both reviewers, and for the sake of clarity, we have actually moved (and clarified) several key parts of the SI into the main text. These include the whole “Linear stability analysis” and “Positive regulatory loops determine the kind of RD-instability” sections. These parts, although quite mathematical, facilitate the understanding of our results.

      (2) Linearisation

      (R.1.5) It's clear that keeping non-linearity is complicated and maybe redundant, but please, discuss the assumption of linearity explicitly, especially in the scope of relevance for the real systems, and explain why it's not important, if so. I guess that relaxing this assumption may affect the argumentation in many places, for example, equation (3) of the main text could break (i.e., if the signaling molecule can be consumed in some reaction of A+B->AB kind).

      We agree that the original version was not explicit enough about the reasons for the linear approximation. The first and last paragraphs of the section “Linear stability analysis” are explicitly devoted to justify this linearization. Moreover, the hierarchical network section is now written without using the linearization.

      We are not sure we understand which is the problem with the A+B→AB reaction. We are not assuming any specific f function, just the ensemble of functions that fulfill our requirements (R1 to R5). It is only for the simulations that we have to use a specific f. The reactions suggested by the reviewer could represent an f of the form d[AB]/dt=fAB([A]*[B])-m*[AB]**n for AB and d[A]/dt=-fAB([AB]) and d[B]/dt=-fAB([AB]), where fA and fB are functions that decrease with their arguments. We see no reason why there cannot be a fAB that fulfills our requirements. For example fAB=[A]*[B]/(K+[A]*[B])-m*[AB]. See also related comments in the public comments file.

      (R.1.6) Please, provide a separate section where you reformulate the definition of "non-trivial pattern transformation" for two- and three-dimensional domains, and summarize in this section why the analysis provided for 1D is relevant for higher-dimensional systems. By now, I'm not convinced.

      There was indeed a problem with the way we described non-triviality beyond 1D in the original version of the article. We have now refined the definition of pattern transformations so that it is understandable in 2D and 3D. This definition is presented in the introduction already (in P1 and P2). We have modified figure 1 accordingly.

      Reviewer #2 (Recommendations for the authors):

      Major Issues

      (1) Mathematical Proofs

      (R2.1) We strongly recommend that the authors revisit the mathematical derivations or provide a clear and rigorous justification for the assumptions made therein. These assumptions currently appear unjustified or overly simplistic, especially in light of the nonlinear dynamics the authors aim to describe. The authors should comment on why they expect their results to generalize to all complex network structures, as claimed, and not only apply to the simplified examples analyzed in the paper.

      The article has now been restructured to that end. Concerning the assumptions, they are now all explicitly described in the “Methods: the model” section. Concerning the derivations they are through all the results section. A major change in this line has been the moving of part of the supplementary into specific sections in the main text (and the consequent adaptation of the rest of the text). There are important points of the derivation that may have been buried into the old supplementary and that are crucial to understand the whole argument in the article. In fact, a large part of the results section is just a long argument to show that there are essentially only three classes of gene network topologies that can lead to non-trivial pattern transformations. These arguments are summed up in the last paragraph of the new section “Positive regulatory loops determine the kind of RD-instability” and in the first paragraph of the discussion. In brief:

      (1) Pattern transformation requires gene networks with extracellular signals

      (2) Applying previous mathematical results we show (given the broad requirements on f we have) that pattern transformation is only possible in gene networks that contain positive regulatory loops.

      (3) Applying previous mathematical results we show that in the gene networks in which these loops are extracellular, the only possible non-trivial pattern transformations lead to periodic resulting patterns.

      (4) Applying previous mathematical results we show that in the gene networks in which these loops are INTRAcellular, the only possible non-trivial pattern transformations do not necessarily lead to periodic resulting patterns.

      (5) Using simple logical arguments we also show that no non-trivial pattern transformations are possible in gene networks without negative interactions.

      (6) All the above points combined shows that there are only three classes of gene networks capable of nontrivial pattern transformations. 1) Those with intracellular positive loops, extracellular signals that do not affect themselves and some negative regulation by those (that we call hierarchic networks) 2) Those with intracellular positive loops and extracellular signals that affect themselves negatively (that we now call over-Turing networks) 3) Those with extracellular positive loops and an extracellular negative loops (that following previous work by others are called Turing networks).

      (7) Following previous research and different developmental arguments we explore the types of patterns transformations each of these three classes of gene networks can lead to. These types are characterized only in broad and potential terms. We say nothing about the parameters values for which any gene network leads to any specific pattern transformation. What we say is which types of pattern transformation may be possible (for some possible parameter combination) and which ones are not possible from gene network topology alone (based on the types of loops and so on).

      (R.2.3) Additional to the examples provided in the Public Review, claims such as "despite the large amount of theoretically possible gene network topologies, all gene network topologies necessary for pattern formation fall into just three fundamental classes and their combinations" (l. 34ff)

      This statement was originally intended as an introduction of the text following after it but it seems now clear that this was not apparent enough. This statement has been deleted but we convey a similar message letter in the text, now once its justification is provided. In fact, the justification for this statement is the summary we just described in the previous point (R.2.2) and it is discussed over the main text and summarized in the last paragraph of section “Positive regulatory loops determine the kind of RD instability”.

      (R.2.4) and "The same applies to the topologies we found not to be able to lead to non-trivial pattern transformation" (S7) are not or inadequately justified and should be either substantiated or significantly toned down.

      The same comments that above apply.

      (R.2.5) (a) We advise the authors to argue why it is enough to prove key results by considering linear dynamics (see S2-S7). While linearization is a common technique, the authors themselves emphasize the importance of nonlinearities in pattern formation throughout the paper.

      In the current version we provide an explicit justification for this in the section “Linear stability analysis”, especially in its first paragraph. Moreover, for the analysis of the hierarchical networks we do longer use any linearization.

      (R.2.6) (b) To make linear analysis meaningful, we suggest restricting the initial conditions to small fluctuations (e.g., small spikes or noise), which would justify using linearization to investigate the onset of non-trivial pattern formation. Alternatively, the authors should attempt to generalize the results to fully nonlinear dynamics, ideally for a broader class of functions f.

      As we now explain, the homogeneous-with-noise initial pattern already correspond to small perturbations around the homogeneous steady state (due to molecular noise). In addition, for the spike and spike–homogeneous initial pattern we now explicitly consider spikes of small amplitude. We acknowledge that the use of larger spikes in the previous version could lead to misunderstandings regarding the validity of the linear approximation, even though it does not contradict the assumptions underlying the analysis. In these initial patterns, pattern formation arises because the signal secreted from the spike diffuses into the surrounding domain, so that cells outside the spike experience only small deviations from the equilibrium concentration.

      Larger spikes may induce stronger deviations in cells located very close to the spike; however, because the spike occupies a region that is very small relative to the total domain size, these local effects do not influence pattern formation in the bulk of the domain. A similar situation occurs with boundary effects in cells located near the domain limits, which likewise do not affect the pattern formation process away from the boundaries. We have clarified this point in the revised manuscript, both in the final sentences of the Introduction and in the description of the initial conditions in the fourth paragraph of the “Linear stability analysis” section, where we explicitly state that each initial pattern can be interpreted as a perturbation of an otherwise homogeneous pattern.

      (R.2.7) (c) The assumptions required for the proofs should be explicitly stated and justified. At present, the logic behind the chosen constraints on f is unclear, and the flow of the argument suffers as a result.

      The actual justification for the requirements (i.e. constraints) on f are biological (and we now explain them more explicitly when we introduce these requirements). Most of the mathematical proofs do not require these requirements except when we explicitly say so.

      (R.2.8) (d) The illustrative functions provided in some of the proofs in the SI (e.g. S5.2.1 "To see this, let us consider, for example, that they are both quadratic monomials of the form f_k(g_A)=B_k g_A^2 and f_j(g_A)=B_j g_A^2") do not satisfy the authors' own stated conditions (e.g., this function violates requirement R4 (l.197 f)). More suitable examples should be selected to ensure consistency between assumptions and illustrations.

      We have changed the whole section (based on the comment R.2.9 from the same reviewer). We now provide arguments in the main text that generally do not rely on specific fs.

      (R.2.9) (e) Currently, all mathematical results are confined to the appendix. We recommend including key insights from the proofs in the main text to improve readability and to allow the main claims to stand on their own. For example, the section on the requirements RH2a and RH2b (l. 320 - l. 335)) would benefit strongly from the insights from S5.2.1

      We agree. We have moved the linear stability analysis and the dispersion relation section to the main text. We have also moved what used to be S5.2.1.

      (2) Simulations

      The simulations raise, as mentioned in the Public Review, several concerns regarding their generality and validity.

      (R.2.10) (a) We recommend validating the simulation results by comparing them with simulations of the full nonlinear equations. The authors should at least provide the equations for the full dynamics and explain how the expansion is performed and why it is valid. This also includes verifying the assumed steady states (g_i=0 and g_i=c_i, where 1/c_i = \mu_i or \hat{\mu}_i).

      We are simulating the whole non-linear equations. Here it is important to stress, as we do now in the main text, that our results apply to any f, as long as it fulfills our R1-R5 requirements. However, for the simulations in the figures we have to use a specific f (since there is an infinite amount of fs that fulfill our requirements). Again the figures are just examples to visualize the types of resulting patterns and gene networks we talk about.

      In the original version we may not have been clear enough about the equations used for the simulations. The presentation of the Maini-Miura model has been revised to improve clarity (equation S6.1 in SI). In particular, the existence of a homogeneous steady state is now parameterized by a tunable g<sup>*</sup>, that can be chosen as for spike initial patterns or for homogeneous-with-noise and spikehomogeneous initial patterns). We have also included a proof that the model equations satisfies our conditions R1-5. Indeed, the model is non-linear as long as σ<sup>i</sup>≠0 for some gene product (as we explicitly assume).

      The derivation of this cubic model from a separate expansion of general reaction-diffusion dynamics can be found in the original paper (Miura & Maini, 2004), with further applications to pattern formation that supporting its validity in subsequent works (Marcon et al., 2016; Diego et al., 2018). Importantly, this expansion is independent of the linearization performed in the main text of our article to derive the dispersion relation. The reference to this separate expansion in the previous version was included solely for contextual purposes; however, we have removed it in the revised manuscript to avoid potential confusion.

      (R.2.11) (b) The use of a Jacobian that is independent of the steady-state contradicts the assumption of nonlinearity (requirement R2 (l. 192f)) of f. We ask the authors to clarify this.

      We believe this concern arises from a notational ambiguity in the previous version of the manuscript, which has now been corrected: the matrix appearing in the regulatory term has been renamed from J to W<sup>T</sup>. As stated in the main text, the jacobian of the regulation function f(g) evaluated at the homogeneous steady state must coincide with the transpose of the network weight matrix. With the current equations (S6.1), we have , from which we easily get . Also, it is clear that the Jacobian of f(g) is not independent of g.

      (R.2.12) (c) In Figure S3 and similar simulations, the implementation of the nonlinear terms is ambiguous. The function f shown does not correspond to the Jacobian, and it remains unclear how these components are ultimately implemented in the simulation code. Additionally, as mentioned, it does not fulfill the necessary conditions for the global steady state.

      The implementation of nonlinear entries in f(g) whenever they are needed is now made explicit in the corresponding subsection of section S6 in the SI. With the new notation it becomes clearer that the fs used can fulfill the necessary conditions for the global steady state.

      (R.2.13) (d) The given function f_8 in S8.10.2 cannot correspond to the mentioned network since the number of gene products does not match the Jacobian and the network.

      This was a typo that has now been corrected.

      (R.2.14) (e) The given parameters for the figures in the SI do not match the figures. Please check and ensure that the correct figure is referenced (e.g., S8.2 Figure 3)

      This was a typo in the numeration of the subsections in the SI that has now been corrected.

      (R.2.15) (f) It is unclear which units are used, and the units used for the non-dimensionalization should be provided so one can relate them to biological systems.

      It is now explicitly stated in the revised version that the model equations are formulated in arbitrary units. This implies that the model dynamics are consistent with the characteristic units of any particular biological system under consideration. No non-dimensionalization of the model equations has been considered.

      (3) Conceptual and Structural Clarity

      The manuscript suffers from a lack of structural clarity, which affects both readability and scientific coherence.

      (R.2.16) (a) In one of the central figures (Figure 4) supporting their main claim, the naming of the network is not consistent with the main text. The network category referred to as "Over-Turing" is never mentioned in the main text. We suspect this should actually be labeled as the "noise-amplifying network."

      Indeed. This has now been corrected. We now use only the term “Over-Turing” in the article.

      (R.2.17) (b) The Supplementary Information includes an analysis of dispersion relations to classify patternforming networks, but this approach is not mentioned or referenced in the main text.

      This part of the SI has been moved to the main text and the dispersion relation has been fully and explicitly integrated in the overall argument of the article.

      (R.2.18) (c) In relation to Figure 6, we found that the concept of "diversity of possible final patterns" would benefit from a clearer definition and explanation. It is not immediately evident how this diversity is measured or what criteria are used to compare different networks. For instance, it is unclear why the Over-Turing network - which generates both periodic and noisy patterns - is considered to exhibit low diversity, whereas the Turing networks, which produce only periodic patterns, are described as having high diversity.

      This was just a large typo. The figure has been corrected. The reasons for this differences are now described in the last three paragraphs of the section “The ensemble of possible pattern transformations from H gene networks and spike initial conditions” for the hierarchical networks and in the last paragraph of the section “Pattern transformations in L- subnetworks from spike-homogeneous initial patterns ”, for the noise amplifying networks and in the seventh paragraph of the section “Pattern transformations in the combination of L+ and L- subnetworks” for the Turing networks.

      (R.2.19) (d) Additionally, the dependence of final patterns on initial conditions is not clearly described. It seems that this relationship is only analyzed for non-trivial pattern formations, but this is not explicitly stated. Clarifying these points in the caption of Figure 6 would greatly help readers understand the interpretation and significance of the results presented in this figure.

      Indeed, we have done nothing for the trivial pattern transformations. We are now more explicit about this already from the introduction. This article is only concerned with non-trivial pattern transformations. For each type of gene network we now provide a more detailed description of how the resulting pattern depends on the initial pattern (in the sections for each gene network).

      (R.2.20) (e) The significance statement is simply a verbatim repetition of parts of the abstract. This defeats its purpose, which is to articulate the broader implications of the work. We urge the authors to rewrite this section with a focus on significance rather than summary.

      We have now corrected this.

      (R.2.21) (f) We suggest including a dedicated figure to illustrate the biological model, depicting cells, intracellular and extracellular compartments, and the presence or absence of boundaries between adjacent cells. Such a figure would significantly enhance readers' understanding of the system being discussed.

      We have now done that. See new figure 3.

      (R.2.22) (g) We encourage the authors to strengthen the 2D and 3D results presented in the paper by adding supporting citations, sharing implementation details, or providing a more in-depth analysis of these systems. If such additions are not feasible, it may be best to remove references to the 2D and 3D systems to maintain clarity and focus.

      In the new version of the article we explain why our results on which gene networks can lead to pattern transformation do not depend on the dimensionality of the system. In fact, none of our proofs or arguments assumes or requires a specific number of dimensions. The networks are the same no matter the number of dimensions. The types of possible patterns can be seen as manifesting themselves differently depending on the number of dimensions. In the current version of the manuscript we explain now, every time we explain a resulting pattern, how the pattern is in 1, 2 and 3 dimensions and why. We have added Figures 1 and 9 for that purpose. As we explain in the text, the resulting patterns that are noisy would be noisy no matter the number of dimensions and the ones that are based on a spike in the initial pattern have necessarily radial symmetry (in any number of dimensions). Similarly the periodic patterns will be periodic no matter the number of dimensions (although some aspects of it will change). Similarly, in the 5th paragraph of the discussion we discuss the effects of the shape of the system and the boundary. There was a problem with the definition of pattern transformation we used, but this has now been corrected, in P1 and P2 in the introduction.

      (R.2.23) (h) The results section lacks a consistent structure. Section titles do not clearly indicate which phenomena or initial conditions are being analyzed, making it hard for readers to track the logical progression of the study.

      Now the results start with some introductory results with the subsections:

      “Basic requirements on gene networks capable of pattern transformation”

      The rest of the results are split into four clearly differentiated sections:

      “Gene network classification”

      “Linear stability Analysis”

      “Positive regulatory loops determine the kind of RD-instability”

      “Hierarchical Networks”

      “Emergent networks”.

      “Gene networks combining different classes of subnetworks”

      The last three sections have several sub-sections inside.

      We think that the titles of the sections are self-explanatory since hierarchical networks contain only H subnetworks while the emergent networks contain L+ or L- subnetworks and the last major sections is about how all these can be combined.

      Minor Issues

      (1) Notation and Terminology

      (R.2.24) (a) Variable naming is inconsistent throughout the paper. Terms like g_A(x) and A(x) (S5.2.1) are used for gene network concentrations without consistent usage. The naming of genes in networks also varies between the main text, SI, and figures. I.e., sometimes genes are labelled with small, sometimes with large letters, and sometimes with numbers.

      This has now been corrected.

      (R.2.25) (b) It would improve clarity to use distinct notations for intracellular vs. extracellular concentrations and gene expressions. Ensure networks and examples are consistent across all figures, captions, and supplementary materials. For example, RH2a and RH2b have different networks in the main text compared to the SI.

      As we now explain in the third paragraph of the “Methods: the model” section we consider, for simplicity, that gene products are either intracellular or extracellular. In that sense there is no possible ambiguity. As explained in that section, again for simplicity, we do not consider the receptor nor the signal transduction pathways of signals. This means that an extracellular gene product can “directly” regulate intracellular gene products. Because of that, we think that using different notations for extracellular and intracellular gene products would make things more confusing. We have corrected the misnaming between main text and figures.

      (R.2.26) (c) We suggest using distinct notation for the gene product itself and for its small deviation from a homogeneous steady state in the SI. This would help clarify whether specific statements apply only within the linearized regime or can be generalized to the full nonlinear dynamics.

      We do that in the new version of the article.

      (R.2.27) (d) Line 327 contains a mistake: g_k = g_j should be expressed as a proportional relationship. The division by g_A also seems unnecessary - please revise.

      This is now explained in a different way so this mistake does not apply.

      (2) Model Description

      (R.2.28) (a) Justify why boundary effects and spatial separation between cells can be neglected in the model.

      This is now discussed in the 5th paragraph of the model section. We do not claim that boundary effects are negligible. We claim, instead, that which are the gene networks that can lead to pattern transformations do not depend on the boundaries. The same occurs for the types of resulting patterns, in the coarse way we use, possible from each gene network and initial pattern.

      As stated in the first two paragraphs of the model section, the spatial separation between cells can be ignored because we assume there are many cells in the system and these are evenly spaced and sized (at least roughly). That is usually the case in animal development, although not always (there are exceptions in the very early stages of many marine invertebrates), and we do not claim to know exactly what happens in those cases: as we stated in the first paragraph of the introduction we assume systems made of many small cells.

      (R.2.29) (b) State explicitly that only extracellular gene products are assumed to diffuse - this is currently only mentioned in the SI.

      This is now explicitly stated early on in the first three paragraphs of the model section and also after the introduction of the model equations (1)-(3).

      (R.2.30) (c) In the Supplementary Information, the authors state that both extracellular and intracellular gene products can exhibit non-zero diffusion, which appears inconsistent with the conceptual framework and probably is a typographical error.

      This was indeed a typographical error. It is now corrected.

      (3) Assumptions and Requirements on f

      (R.2.31) (a) The equation for requirement R5 is incorrect as written in the main text and should be reformulated more rigorously. The condition should be stated for all constant values of g_i (and g_j) to avoid misinterpretation; otherwise, one might assume all matrix elements must have the same sign.

      This has now been corrected.

      (R.2.31) (b) Clarify what restrictions on f prevent pathological nonlinearities like 1/(g_k + \epsilon), which would contradict the assumed behavior at high concentrations.

      We do not understand this criticism. 1/(g_+\epsilon) fulfills our requirements on f and we do not see how is that pathological. We are unsure of what the reviewer means by the assumed behavior at high concentrations.

      (4) Figures and Captions

      (R.2.32) In Figure S3b, the diagram shows gene 5 being activated by gene 4, yet the caption states this is a negative regulation - please correct.

      This has now been corrected.

      (5) Readability and Formatting

      (R.2.33) (a) Improve navigation by hyperlinking references to equations, figures, and requirements throughout the document.

      In the new version we have inserted these hyperlinks.

      (R.2.34) (b) Adding hyperlinks to the requirements would additionally help the reader to keep track of them

      In the new version we have inserted these hyperlinks.

      (We.2.35) (c) Correct inconsistent or mismatched equation numbers and references. E.g. SI S5.1 is not referring to the correct equation (the equation it should be referring to would be Equation 3), and the reference to Figure 7 in part of the dispersion relation is wrong (as far as we see, this should be Figure 5).

      This has all been corrected now.

      (R.2.36) (d) Clarify ambiguous language in the introduction. For instance, the description of spike patterns (lines 136f) as a single cell spike contradicts the stated width (SI) and the visual representation involving 500 cells from the figures.

      This has now been corrected.

      (R.2.36) (e) The discussion of 2D and 3D simulations appears limited to the "noise amplifying" network. It's unclear whether a similar analysis was done for other network types.

      In Figures 1 and 9 and through the text we discuss all types of patterns in 2D and 3D.

      (6) Typos

      (R.2.37) Typos in the text (The following is just a small selection of the typos we came across. Since there are quite a few throughout the manuscript, we may not have caught all of them. We kindly recommend that the authors carefully proofread the full text to ensure consistency and clarity):

      We have corrected all the indicated typos and proofread the whole manuscript and SI.

      Reviewer #3 (Recommendations for the authors):

      Major concern:

      (R.3.1) Pattern formation can be induced by the positional information, and reaction-diffusion/Turing mechanisms is a foundational idea in the field. As in the references the manuscript cited, these paradigms were already clearly articulated and synthesized (e.g., Green & Sharpe's work (2015)). Moreover, the search for minimal network topologies that can generate Turing patterns has been extensively explored in Zheng et al. (2016). The novelty of the present work is unclear. It might offer a fresh perspective on an established problem, but it does not seem to present fundamentally new biological or mathematical advances.

      If the authors wish to strengthen the novelty and impact of the manuscript, they should consider explicitly acknowledging prior work and positioning their contribution as a formal extension or generalization, not discovery. To enhance the practical relevance of their work, the authors could demonstrate how their framework can be used to predict or classify gene network behaviors in pattern formation that are not easily identifiable through experimental approaches alone. For example, they could show how their classification helps distinguish between Turing, hierarchical, and noise-amplifying dynamics in complex or ambiguous biological systems, thereby offering a guiding tool for experimental design or interpretation.

      Indeed, the gene networks we identify have been identified before. We were and we are quite explicit about it, in the discussion, and we do cite the relevant work on that (including the one suggested by the reviewer). The novelty of the work is not identifying these gene networks, nor minimal ones, but showing that these are all the possible ones for pattern transformation (that there is no new type of network), this has not been done before (not even intended) and we are very explicit about that being our results (first paragraphs of the discussion).

      Minor concern:

      The writing style and language usage can be improved for clarity. Some explanations in the results and discussion can benefit from tight editing to eliminate redundancy and improve readability.

      We have corrected all the indicated typos and proofread the whole manuscript and SI.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer 3 (Public review):

      Comments on revised version:

      The current version of the manuscript is clear and complete. Kudos to the authors for their thorough revisions. My only remaining point concerns the definition of "report": "We define a report as any explicit behavioral response (whether verbal, manual, or otherwise) that communicates a participant's subjective state." It would be helpful to clarify whether this definition is intended to exclude purely internal, explicit self-reports that are not externally expressed. As currently formulated, the definition appears to require overt behavioral communication. However, this raises a conceptual issue in relation to the no-report paradigm literature, where the distinction between report, metacognitive access, and overt motor/verbal expression is precisely at stake.

      Could the authors specify whether "report" is meant to (i) be restricted to externally observable, behaviorally expressed reports, or (ii) extend to internally generated, explicit metacognitive judgments even when they are not communicated? Clarifying this point would help situate the manuscript more precisely within ongoing debates on the role of report in identifying neural correlates of consciousness.

      We thank the reviewer for prompting us to make this subtle but important distinction explicit. We agree that the two senses of "report", i.e., (i) externally observable, behaviorally expressed reports and (ii) internally generated, explicit metacognitive judgments that are not communicated, are conceptually distinct and that this distinction is precisely at stake in the no-report paradigm literature. We fully agree that sense (ii) (disentangling NCCs from covert metacognitive access) would be a valuable direction for future research. However, because the intracranial studies reviewed in the manuscript focus exclusively on distinguishing NCCs from overt behavioral reports, our definition is intentionally restricted to sense (i).

      To clarify this point in the manuscript, we added the following sentence at lines 111–114:

      "Note that the no-report intracranial studies described here attempt to distinguish NCCs from externally observable, behaviorally expressed reports, and not from internally generated metacognitive judgments that are not communicated."

    1. Author response:

      The following is the authors’ response to the original reviews.

      We have made several major changes in response to the comments and we feel that the manuscript is considerably stronger. In brief: 1. We have added substantial content about homeostasis and EI balance to the introduction. 2. We have addressed concerns about physiological relevance by performing calculations to show that the free calcium in our solutions is well within the physiological range, by citing previous studies showing that short-term plasticity is consistent across 33-38 ℃, and by doing simulations scaled to physiological temperatures to show that the key computational effects are retained. 3. We have addressed concerns about readability by extensive text rewrites, reformatting most of the figures, and by splitting figures into smaller, more focussed ones. 4. We have organized over 20 statistical evaluations and comparisons between our model and experiments into a table. 5. We have carried out additional calculations to examine how the optimal frequency for mismatch detection depends on parameters, and to show that mismatch detection remains even in the presence of stimulus jitter. 6. We have stated more clearly how our proposed mechanism for mismatch detection is based on transient plasticity-mediated skewing of EI-balance, and have added a schematic for the last figure to show this.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study uses optogenetics to activate CA3, while recording from CA1 neurons and characterizing the excitation/inhibition (E/I) balance. They observe use-dependent alterations in the E/I balance as a result of STP, and they develop a model to describe these observations. This is a very ambitious paper that deals with many issues using both experimental and modeling approaches.

      Strengths:

      This paper examines important principles regarding the manner in which synaptic circuitry and use-dependent synaptic plasticity can transform inputs and perform computations.

      Weaknesses:

      The use of selective ChR2 expression in CA3 cells is a good approach, but there are numerous issues that cause concern regarding the applicability of their slice recordings to physiological conditions and that make some aspects of their results difficult to interpret. Experiments are not performed under physiological conditions (high external calcium and low temperature), which makes the interpretation of their findings difficult.

      Calcium: We would like to reassure the reviewer that the free calcium levels in our solutions were at ~1.27 mM, well within the physiological range, since our aCSF solution used calcium buffers as well as CaCl2. We have added a section to the methods to show this calculation.

      Temperature: Klyachko and Stevens (J. Neurosci 2006) show that the facilitation, augmentation and filtering properties of the CA3-CA1 network were consistent between 33 and 38 degrees C, thus spanning our conditions of ~33 degrees C. Additionally, we have performed simulations to show that the mismatch detection computations remain pronounced (or are even strengthened) when simulation rates for kinetics and channels are scaled to physiological temperatures. Using a Q10 of 2, the scaling term for kinetics is ~37% faster. The outcomes are presented in Figures 7 and 9. We now state these points at the start of the results section:

      “Our bath solution had physiological levels of free ions including calcium (methods), and recordings were performed at 32-33 ℃ which has been shown in rats to yield similar short-term plasticity properties as at physiological temperatures (Klyachko and Stevens 2006b).”

      We have added a new section to the discussion “Relevance to in-vivo computation” in which we enumerate the caveats but also the points of convergence between our study and physiological conditions, to strengthen the interpretability of our results.

      In addition, the reliability of stimulating action potentials in CA3 pyramidal cells needs to be determined, particularly during high-frequency trains. If it is unreliable, there are alternative approaches that might prove to be superior, such as the use of somatically targeted ChR2.

      We acknowledge that somatically targeted ChR2 might have slightly improved the sparseness of stimuli, but even such localized expression could lead to unreliability if the position of the soma with respect to the illumination is such that the stimulus is near threshold. Instead, we have adopted a data-driven estimation of CA3 reliability. We reanalyzed our optically-triggered field potential readouts from CA3, to estimate their reliability individually and over trains (Figure 1).

      “Notably, the distribution of field amplitudes was very tight (Figure 1E), more so than the corresponding EPSPs (Figure 1H). Together with previous work using a similar optical stimulus system [6] we interpret this to say that the spiking responses from CA3 neurons to optical stimuli were consistent from trial to trial. The field response showed a slight decrease over the course of the pulse train of approximately 2% per pulse (regression fit slope=0.02, r2=0.05). We attribute this to ChR2 desensitization.”

      As a further bound to any functional outcomes of CA3 spiking (un)reliability, we point out that CA3-CA1 release probability is low (p~0.2). Any reduction in CA3 reliability is equivalent to reducing the probability of synaptic release, which is already treated as a stochastic process in our simulations. We were able to compare this to experiment as follows: We explicitly modeled the effect of different synaptic volumes as a surrogate for changing p_release in Figure 6-figure supplement 1, and mapped this to our data in Figure 6 D.

      “Then we compared the probability that each optical stimulus would elicit an EPSP (Figure 6 D). As expected, 15-square patterns (yellow dots) frequently gave an EPSP (77.5±11.7%), while 5-square patterns failed about half the time (51.4±16%). The simulated runs matched this (Table 1). The probability of failure reduced with increasing volume of the simulated presynaptic boutons, because larger volumes experienced smaller chemical noise (stochasticity) in synaptic release (Figure 6-figure supplement 1). We note that for the purposes of eliciting a postsynaptic response, any unreliability in optical stimulus-triggered firing of the CA3 neuron folds into the probability term for stochastic synaptic release. By matching this metric to experiment, we fine-tuned the volume scaling term for the presynaptic boutons to 0.2”

      In addition, a clearer, more detailed discussion of their model that distinguishes it from previous modeling studies would be helpful (and would make it seem less incremental).

      This is a good suggestion, as we regard our model as very substantially different from previous studies. We have incorporated this in the discussion as below:

      “Our current model is distinct in that it is truly multiscale, closely constrained by experiment, yet runs on modest hardware. It incorporates the network, a conductance based model of a CA1 pyramidal neuron, and chemical kinetic models of a population of stochastic synapses on its dendrite.

      Our network model is much reduced compared to models with exhaustive cellular and network-level detail44. Its simplicity enables extensive exploration of the network parameters and comparison with recorded activity under a series of well-controlled stimulus patterns (Figures 4-9).”

      We also point out that our proposed mechanism for mismatch detection is an advance over previous ones:

      “Leaving aside the obvious differences between auditory cortex and hippocampus, we frame our model as a transient differential tilt in EI balance (Figure 3, Figure 8A,B, Figure 10B), in distinction to the fresh-afferent model. This makes our model robust over a wide range of stimulus and network conditions (Figure 9), and has the functional implication that transient responses remain at about the same amplitude over a prolonged stimulus sequence (Figure 8B, Figure 10B), rather than declining.”

      Reviewer #2 (Public review):

      Summary:

      The authors investigate EI balance in the CA3-CA1 projections, emphasizing synaptic depletion and the implied rebalancing of excitatory and inhibitory projections onto a single CA1 Pyramidal cell. They present physiological results with optical stimulation in CA3 and measuring various response features in CA1, showing signatures consistent with the adjustment of EI balance. In particular, the authors emphasize a transient effect where the neuron escapes from EI balance, which can be used for mismatch detection. They partially replicate these results in a computational model that looks at detailed properties of synaptic plasticity in CA1.

      Strengths:

      The authors provide compelling evidence that non-specific modulation of synaptic plasticity, combined with their differential effects on excitatory and inhibitory neurons, can be used by CA1 excitatory neurons to detect changes in the population activity of CA3 neurons. Indeed, they provide insight into the potential computational role of transient EI imbalance.

      Weaknesses:

      The authors observe that "little is known about how EI balance itself evolves dynamically due to activity-driven plasticity in sparsely active networks." This is an overstatement, or better an understatement, given the extensive literature on EI balance (e.g. Wen W, Turrigiano GG. Keeping Your Brain in Balance: Homeostatic Regulation of Network Function. Ann Rev Neurosci. 2024. https://doi.org/10.1146/annurev-neuro-092523-110001 PMID:38382543). This way of framing the question does a disservice to the field and fails to contextualize the current research properly.

      We agree that we could have presented this better. Our focus was on short-term (<1 second) EI balance changes, but our statement did not set this context clearly. We rewritten and expanded the introduction to place our work in context of the substantial previous work on plasticity and homeostasis in EI balance.

      The evidence is incomplete because the authors do not show a specific relationship between synaptic change in CA1 and EI balance adjustment, i.e., the alternative could be that this is an unspecific effect unrelated to the specific regulation of EI balance and its functional role in the hippocampus and the cortex.

      We don’t quite follow this point. We have devoted Figures 2 and 3 to showing a specific relationship between short-term plasticity on CA3->CA1 synapses, and EI balance. In Figure 2 we show how E and I responses evolve over a pulse train. In Figure 3 we explicitly show the plasticity in E and I synapses, and then map it onto EI balance. In Panel 3E to G all these points come together and we show how gamma (the measure of nonlinearity of summation) evolves over a series of pulses in parallel with plasticity in E and I. We have added some new data in Figure 7A, B to show how E and I contribute to mismatch detection.

      Indeed, the paper drifts from addressing EI balance to elucidating the mismatch detection.

      We acknowledge that we did not sufficiently articulate the role of EI balance terms in our subsequent analysis of mismatch detection. We have added several figure panels (Figure 7A, B), added a summary schematic (Figure 10) and redone the text and discussion. With these changes we make the point that mismatch detection can be better framed as a transient shift in EI balance.

      “we frame our model as a transient differential tilt in EI balance (Figure 3, Figure 8A,B, Figure 10B), in distinction to the fresh-afferent model. This makes our model robust over a wide range of stimulus and network conditions (Figure 9), and has the functional implication that transient responses remain at about the same amplitude over a prolonged stimulus sequence (Figure 8B, Figure 10B), rather than declining.”

      The second shortcoming is that they do not show that the stimulation of the CA3 neurons occurs in a physiologically realistic regime.

      We have responded to the concerns about calcium concentration and temperature above in the response to the first reviewer. From the text:

      “Our bath solution had physiological levels of free ions including calcium (methods), and recordings were performed at 32-33 °C which has been shown in rats to yield similar shortterm plasticity properties as at physiological temperatures (Klyachko and Stevens 2006b).”

      In addition, there is a concern about the mapping between physiological activity and our stimuli. It is true that the patterned stimuli we delivered were artificial. We make the point that they are nevertheless a much closer map to sparse physiological patterns than conventionally obtained through Schaffer collateral volleys:

      “We use optical patterned stimuli to stimulate a cross-section of CA3 neurons with a variety of distributed patterns, theta, and other frequency rhythms. These stimuli are sparser and more dispersed than Schaffer collateral electrical stimuli which tend to stimulate adjacent fibres and in most cases are very strong.”

      We have added a section to the discussion “Relevance to in-vivo computation” to more completely address these points.

      Nor do they analyze what the impact will be of the excitatory transient in "mismatch detection", and CA1,

      We are unsure what the reviewer means by the excitatory transient. At the level of CA3, we observe a narrow optically triggered field response for each light pulse. At the level of CA1, we monitor the responses due to activation of E and I synapses, and are able to observe peaks for each of the light pulses. We have analyzed all these features in figures 1 through 3, and they are also explicitly included in the model. Based on the reviewer’s comment we have further characterized the field responses in CA3:

      “We observed a small amount of ‘ringing’ of the field response which we interpret as either CA3 spiking in a burst, or recurrent activation of the CA3 neurons (Figure 1 supplement 2). The ringing was down to ~5% within 8 ms, supporting our treatment of the optical input as a tightly time-delimited event, and setting a low bound to any contribution to patterns by recurrence.”

      When this would occur at the level of the whole population, i.e., the physiological impossibility of triggering uncontrolled chaotic excitatory responses.

      Again, we are unsure what population or chaotic responses the reviewer has in mind. As mentioned above we have further characterized the field readouts of population responses in CA3 and have established tight limits on recurrent activity (Figure 1-figure supplement 2). In case the reviewer is looking for the outcome at the entire CA1 network as a whole, our experiment figures 1GH,J,K,L show sharp, single peak CA1 neuronal responses.

      In particular, when we consider CA3 as an attractor memory system, the range of deviations (mismatches) that a CA1 neuron can be exposed to and detect, given the model presented in this paper, might be below those generated due to CA3 pattern-completion dynamics.

      While this is an interesting question for further work, our study focuses on a tighter question, that of mismatch detection downstream of the CA3. As indicated above and in Figure 1figure supplement 2, our field and patch recordings show that under our stimulus conditions, the internal dynamics of the CA3 produce minimal delayed or recurrent signals. Thus, by design, the CA3 layer in our system acts as an almost pure input layer with minimal internal dynamics. In the discussion we address some of the possibilities that may arise from pattern computations in CA3 and other upstream areas:

      “We speculate that upstream areas may encode higher order stimulus features such as gaps, duration, intensity, localization, and frequency steps into distinct input patterns. Our proposed EI-balance shift mechanism could be a common end-point for all of these. This would transform quite complex mismatch detection tasks into a uniform computation of pattern change, generalizing the mechanism to stimuli which were previously considered to require a more complex network-level implementation”

      In addition, the match between the model and the physiological results is not fully quantified, leaving it to the reader to make a leap of faith.

      While the original version had numerous points of comparison between physiology and model, we agree that the values were scattered. In this revision we have tabulated them and performed additional statistical comparisons between model and data for a total of over 20 comparisons for the cell electrophysiology and network readouts (Table 1). We have also organized the preceding chemical kinetic comparisons in the supplements to Figure 4. We regard our study as one of very few to undertake quantitative experimental comparisons over such a range of readouts, experiments, and scales.

      In addition, the manuscript suffers from poor analysis and presentation. The work could be improved by putting more effort into translating results into insightful metrics.

      We acknowledge that the presentation needed improvement. We have performed a major rewrite and reorganized many of the figures. As mentioned above, we have tabulated numerous metrics (Table 1) and have characterized EI balance and its evolution due to plasticity in a pulse train (Figures 2 and 3). For higher-level metrics, the new figures now extensively explore how mismatch sensitivity depends on parameters, stimulus patterns, and repeat frequency (Figures 7, 8, 9). We have added a discussion section “Relevance to invivo computation”

      Overall, the authors have not achieved their original aim to show that the observed phenomenon is relevant to computation in CA1 or the brain outside of a highly controlled in vitro setup and reductionist single cell model.

      We feel that with this revision we have more clearly shown that our measurements are relevant to in-vivo computation, both through improved clarity and additional analysis. We have added a section “Relevance to in-vivo computation” in the discussion which enumerates the steps we have taken to support the relevance of our study. In the revision we have also performed several modelling extrapolations which encompass in-vivo conditions, such as testing jitter and frequency range. In a broader sense, in vitro work by design, is meant to be highly controlled so as to be able to get at mechanisms, and in our study we have delivered a range of physiologically relevant stimulus combinations to bridge the gap.

      The authors combine several techniques for in vitro whole-cell patch-clamp recordings with patterned optical stimulation of the CA3 network in the mouse hippocampus, which is consistent with the state-of-the-art.

      They introduce a metric of similarity between expected and observed response patterns, called gamma. The name is confusing given the wide use of the label gamma for oscillation frequencies above 20 Hz. Gamma is calculated as (E*O)/(E-O). This means that gamma approximates infinity as the difference goes to 0, to mention one of the problems. This metric is not interpretable, and it is not clear why the authors did not follow a standard approach, e.g., likelihood, correlation, or percent error.

      We acknowledge the potential for confusion, however we felt it would be more confusing to change nomenclature. The metric gamma is derived from previous published work (Bhatia et al, eLife 2019) describing nonlinearities in summation, which is cited. In that study and the current one, there was no instance in which gamma became unreasonably large. It is true that the term gamma is used for many concepts, but we feel that the contexts are so different between summation nonlinearity and oscillation frequencies that confusion is unlikely. We have taken care with the wording in the text to further disambiguate the usage.

      The authors aim to replicate the physiological results with an "abstract model of the hippocampal FFEI network. In practice, this is a conductance-based model of a single CA1 neuron, including chemical kinetics-based multi-step neurotransmitter vesicle release. This is an abstraction from the FFEI network that the paper starts with.

      We stress that the full model was used for all simulations except synaptic chemistry parameter fitting. We have clarified this point in the text and discussion section. From the text following Figure 4:

      “We used this full model, with optical stimulus, CA3, Interneurons, CA1 neuron, probabilistic connectivity, and presynaptic signaling chemistry, for all subsequent calculations in this study.”

      The model has 256 integrate-and-fire CA3 neurons, 256 interneurons, plus 200 inhibitory and 100 excitatory synapses onto the CA1 neuron, in each of which we have distinct multistep transmitter release kinetics.

      It raises the question whether this is the right level at which to model the computational impacts of EI imbalance on CA1 neurons. Given the highly reduced model they have elaborated, the generalization to the complete CA3-CA1 network that the authors suggest can be achieved in the discussion is overoptimistic. Network models of CA3 and C1 must be considered, together with afferents from the entorhinal cortex to accomplish this generalization.

      We hope we have clarified that we do indeed base all our calculations on the full FFEI model converging onto the CA1 neuron whose connectivity influences circuit function, and we feel that this is necessary and sufficient for our goals in this study.

      While the role of the recurrent CA3 network and EC would be interesting topics for future work, the scope of our study is to model the computational impact of EI imbalance in the FFEI network of CA3-> CA1 on CA1 neurons.

      The authors reveal a potentially interesting physiological feature of CA1 excitatory neurons under very specific stimulus conditions.

      We thank the reviewer for considering the work as interesting. We would like to clarify, however, that our stimulus conditions are actually multidimensional. Specifically, we have varied frequency, pattern, and number of inputs for burst stimuli, and we have also examined Poisson train inputs. In the model we have examined spiking responses, and theta modulated stimuli. In the revision we have also included jittered synaptic input, and obtained frequency dependence of the mismatch detection. To our knowledge this is among the more multidimensional stimulus-response and modeling studies on this system.

      It could warrant follow-up studies to place EI imbalance in a physiologically realistic context.

      Reviewer #3 (Public review):

      Summary:

      This work shows experimentally and computationally that single CA1 neurons can perform mismatch detection on patterned CA3 inputs and that STP and EI balance underlie this detection.

      Strengths:

      It has been known that STP can enhance the EPSP when the corresponding presynaptic input exhibits abrupt changes in firing rate. This work provides experimental evidence and further computational support for the hypothesis that the basic computation through STP is useful for detecting abrupt changes in the spatial pattern of synaptic inputs at the Schaffer collaterals. Further, their results indicate the novel view that mismatch detection is most efficient when gamma-frequency bursting inputs exhibit mismatches between theta cycles.

      Weaknesses:

      Their model assumes that patterned activities in CA3 do not have overlaps. However, overlaps between memory engrams have been shown. Therefore, this assumption may not hold, and whether the proposed mechanism is valid for overlapping CA3 inputs needs further clarification.

      We see that our account of the methods needs clarification, since we explicitly incorporate overlap in our model. First, from the experiments themselves, we say that we expect overlap:

      “This was also consistent with the observation of a wide field of excitability around individual CA3 neurons [6] (Figure 1-figure supplement 1). From this we expect that there is some overlap in the sets of CA3 neurons activated by different patterns, and this overlap increases with more stimulus squares.”

      In the model, we systematically examine the effect of overlap and have added several figures to make the point (Figure 9 Bi, Figure 9Ci, Figure 4-figure supplement 6, Figure 7figure supplement 1, Figure 9-figure supplement 1).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The use of selective ChR2 expression in CA3 cells is a good approach, but there are numerous issues that cause concern regarding the applicability of the slice recordings to physiological conditions and that make some aspects of the results difficult to interpret.

      Weaknesses:

      (1) Some aspects of this study seem somewhat incremental. There is a rich literature on the study of excitation and inhibitory synapses and the issue of EI balance. There are a great many related studies that are not cited (off the top of my head: Pouille and Scanziani 2001, Mittmann, Chadderton and Hausser 2004, Atallah and Scanziani 2009, but there are many, many more). A great many of the ideas presented in this study have already been published previously (Klyachko and Stevens, 2006, and numerous other related manuscripts).

      We agree that the topic of EI balance has a very substantial literature. We have incorporated many of the mentioned articles and others in our introduction and discussion. Our study explicitly links several strands of work on EI balance with short-term plasticity and spatial patterning:

      “The current study integrates several research themes of EI balance, short-term plasticity, and network computation to systematically characterize and model the properties of a network with feedforward inhibition. We complete the experiment-model-prediction-testing loop and show that differential changes on E and I synapses may provide a mechanism for single neurons to extract interesting features of spatiotemporal inputs through STP (Asopa and Bhalla 2023), while keeping mean activity steady.”

      We find that our sparse optical stimulation protocol gives qualitatively distinct results, and is amenable to investigation of more complex spatial pattern dependent effects. We have explicitly discussed the mentioned paper by Klyachko and Stevens, and numerous others, to point out where our study differs. From the discussion:

      “For example, studies using field electrode stimulation of the Shaffer collaterals report a sustained shift to excitation during burst input (Klyachko and Stevens 2006a). In contrast, our sparse optical patterned stimuli results in a small window of escape from EI balance around pulse 2 or 3 in a burst (Figure 3), following which both E and I undergo depression to restore balance (Figure 3, 8). Thus, spatial patterning intersects with short-term plasticity to add another layer of timing control through gating of E-I balance.”

      (2) There are multiple technical issues that call into question the relevance of this study for physiological conditions and the study of STP.

      (a) Their experiments were performed in elevated external calcium (2 mM) compared to physiological calcium (1.1-1.5 mM). This will have a major influence on the probability of release and short-term plasticity.

      This concern does not take into account the composition of our solution, which incorporated calcium buffers to give free calcium levels of ~1.27 mM. We have provided detailed calculations in the methods section.

      “Our bath solution had physiological levels of free ions including calcium (methods), and recordings were performed at 32-33 °C which has been shown in rats to yield similar short-term plasticity properties as at physiological temperatures (Klyachko and Stevens 2006b).”

      (b) Their experiments were performed at reduced temperatures (32-33 {degree sign}C). This is alright for many studies, but this is an important deficiency for the particular issue of EI balance and STP, and the relevance of conclusions based on these conditions.

      Klyachko and Stevens (J. Neurosci 2006) show that the facilitation, augmentation and filtering properties of the CA3-CA1 network were consistent between 33 and 38 degrees C, thus spanning our conditions of ~33 degrees C. Additionally, we have performed simulations to show that the mismatch detection computations remain pronounced (or are even strengthened) when simulation rates for kinetics and channels are scaled to physiological temperatures. Using a Q10 of 2, the scaling term for kinetics is ~37% faster. The outcomes are presented in Figures 7 and 9.

      (c) I like the selective expression of ChR2 in CA3 pyramidal cells, but they have not provided any information on the effect of stimulation on the firing of CA3 cells (Extended Data Figure 1 is not enough). Is it reliable for single stimuli or stochastic?

      We have used field recordings in the CA3 to put tight bounds on the properties of CA3 firing (Figure 1, Figure 1-figure supplementar 2.) The field recordings show that on average the firing is highly reliable. We explicitly characterize the probability of eliciting EPSPs through Poisson patterned stimuli in Figure 6 D. As discussed in the text (excerpted below) any stochasticity in firing folds into the parameters for p_release, and synaptic firing is itself stochastic.

      We note that for the purposes of eliciting a postsynaptic response, any unreliability in optical stimulus-triggered firing of the CA3 neuron folds into the probability term for stochastic synaptic release.

      Do CA3 cells fire once or multiple times?

      This was a useful point, and we examined our field potential data more closely based on this.

      “We observed a small amount of ‘ringing’ of the field response which we interpret as either CA3 spiking in a burst, or recurrent activation of the CA3 neurons (Figure 1-figure supplement 2). The ringing was down to ~5% by the third peak which occurred within 8 ms, supporting our treatment of the optical input as a single brief event, and setting a low bound to any contribution to patterns by recurrence.”

      Are the spikes precisely timed, or do they vary?

      Based on the field potentials, the spikes are precisely timed (Figure 1-figure supplement 2D, E).

      “fEPSP Peak Width distribution centred around 1.2 ms, but no peak was wider than 1.6 ms, suggesting tight synchrony in case multiple CA3 neurons were spiking.”

      Are there use-dependent changes in the ability of optogenetic stimulation to evoke spiking?

      Yes, and this is characterized in figure 1 panel I. The decrement is about 2% per pulse.

      The CA3 regions are highly interconnected with recurrent collaterals. Does stimulation during trains alter the activity in the CA3 region as a result of these collaterals?

      Based on the CA3 field recordings, almost all CA3 activity is optically triggered (Figure 1figure supplement 2). Figure 1C shows narrow fEPSPs in a burst.

      This would be a particularly important issue during trains. Would they have gotten more readily interpretable results if they had used a somatically targeted ChR2 variant?

      We feel it is unlikely that a somatically targeted ChR2 would change outcomes. All our analysis assumes overlap of excitation of CA3 pyramidal neurons, that is, a given spot illuminates multiple cells to different degrees, and that there will be neurons which are activated by more than one spot. Somatic targeting does not eliminate activation due to scattering and out-of-focal-plane illumination.

      In extended Figure 5, they show stimulus patterns used to stimulate. I need some more explanation. Are they stimulating in the cell body region only, or are they stimulating in the vicinity of dendrites?

      Extended Figure 5 (now Figure 4-figure supplement 6) indicates the stimulus patterns in the model. The experimental illumination pattern was 336µm x 187.2µm oriented so that the long axis of the pattern lay along the CA3 cell body layer (methods). Sample stimulus patterns are illustrated in Figure 1 panels D, G and J. Given scatter and out-of-plane illumination we expect that dendrites will also be stimulated. This is corroborated in Figure 1figure supplement 1 where we find that in addition to a strong ‘receptive field’ at the soma, there is a dispersed region of weaker activation. We cannot say definitively whether this dispersed region is due to light scatter, out-of-plane illumination, or dendritic activation. However, even somatically targeted ChR2 would elicit multi-neuron activity due to scatter and out-of-plane soma activation.

      If that is the case, there are a great many complications that arise, and it seems to be an approach that could unreliably activate a great many CA3 cells.

      We have now put in a paragraph to discuss this, and to set bounds to the unreliability.

      “To monitor the strength and consistency of the total resultant optogenetic activation of the CA3 layer, we used an extracellular field electrode in the CA3 stratum radiatum (Figure 1A, methods). The field response correlated well with optically-driven CA1 PC depolarization (Figure 1E-G), and scaled with the size of the pattern (Figure 1F). This was also consistent with the observation of a wide field of excitability around individual CA3 neurons (Bhatia et al. 2019) (Figure 1-figure supplement 1). From this we expect that there is some overlap in the sets of CA3 neurons activated by different patterns, and this overlap increases with more stimulus squares. Notably, the distribution of field amplitudes was very tight (Figure 1E), more so than the corresponding EPSPs (Figure 1H). Together with previous work using a similar optical stimulus system (Bhatia et al. 2019) we interpret this to say that the spiking responses from CA3 neurons to optical stimuli were consistent from trial to trial.”

      We also note that any CA3 firing unreliability folds into the stochastic release terms, as discussed in an earlier point.

      (d) As far as I can tell, they did not examine the effects of blocking NMDA receptors in their slice experiments. This seems like a very important experiment to perform if they really want to understand EI balance.

      The reviewer is correct that we did not block NMDA receptors. While this would have teased apart contributions of NMDAR and AMPAR to the overall response, our analysis of EI balance required the intact synapse and hence this decomposition (which has been done in previous studies) was not needed for our analysis.

      Based on a-d it is not clear that their conclusions regarding EI balance and STP are relevant under physiological conditions, and their findings are difficult to interpret.

      We have addressed the concerns about physiological conditions when it comes to the Ca2+ levels and temperature. We do not feel that points c and d alter the interpretation of our findings.

      Minor:

      (3) Their model has only 1 type of interneuron, whereas there are many. CA3-interneuron synapse has very different plasticity for different types of interneurons, and different types of interneuron synapses onto different parts of the CA3 cell. They need to justify lumping all of these types of interneurons.

      We agree that our model had a coarse-grained representation of interneurons as a single class. We feel this is an appropriate level of detail because it fits well for our experiments, and keeps the model tractable.

      “We have, of course, simplified the network, most notably in the use of only one inhibitory interneuron class which maps to parvalbumin-positive fast-spiking interneurons with perisomatic connectivity. This level of detail was chosen as it was able to quantitatively fit a large number of observations with minimal circuit complexity.”

      (4) How many parameters can they adjust in their model? It seems that with so many parameters, their model is not very good at times (extended Figure 3B and E, for example).

      Our model has 6 free parameters for the network (Table 1), and another 7 parameters each for the E and I presynaptic plasticity models (Figure 4A and Supplementary Data). The presynapse plasticity parameters are directly assigned from the burst response recordings using the parameter fitting as described in the Methods. Normalized RMS differences between model and experiment for presynapse parameters are presented in Figure 4-figure supplements 1-3, panel F. Most traces lie below 0.3, which is a good fit. We have now tabulated numerous comparisons between model and experiment (Table 1). In all but 1 of 20 tests, the model value lies within the experimental range.

      “Overall, we were able to quantitatively replicate almost all features of the experimental dataset in our multiscale model incorporating presynaptic signalling, postsynaptic electrophysiology, and abstracted network connectivity and responses. Between the datasets in Figure 4-figure supplements 1 to 3, Figure 5, and Figure 6, we were able to substantially constrain the parameters in our model, from chemical to cellular physiology to network.”

      Additionally, we have included a new Figure 9 to systematically do parameter sweeps. From this we conclude:

      “...mismatch detection in our model is robustly present and can be tuned over a wide range of network parameters and model assumptions, with the notable exception that it is absolutely dependent on the presence of STP.”

      (5) They use the term short-term potentiation (STP), but plasticity is not just enhancement; there is also depression. That is why many others opt for the more inclusive "short-term plasticity".

      We agree that this was unclear. We meant to use “Short Term Plasticity” and have now clarified this in the text.

      Reviewer #2 (Recommendations for the authors):

      The paper is poorly written and would benefit from a more careful preparation of the manuscript. In the opinion of this reviewer, it does not meet the expected quality for a paper of this type. Reviewing the paper was somewhat frustrating, requiring puzzling through details that were not well described. Also, failing to put clear labels on figures and their low quality did not help.

      We have worked substantially on the readability in the revision. We have made numerous changes to the text and figure legends, and have reworked several figures, with the goal of addressing concerns about readability.

      The introduction lacks proper context for EI balance and the hippocampus.

      We have substantially rewritten the introduction to more clearly place our work in the context of the relevant literature. We touch upon short-term plasticity and computation, on homeostasis, on EI balance and on network correlates of plasticity such as mismatch detection.

      The data analysis is superficial, and insufficient effort is put into compressing complex data into insightful metrics.

      We have done substantial rewrites to address this concern. There are two kinds of metrics we have developed for this study: those that measure the goodness of fit between simulations and data (consolidated into Figure 4-figure supplements 1 to 3 and in Table 1), and those which capture high-level features such as sublinearity of summation due to EI balance (Figure 3), selectivity for mismatch detection (Figures 7 to 9), and peak frequency for mismatch selectivity (Figure 9). We have also performed additional simulations as per reviewer suggestions, which give metrics for dependence of transition detection on network parameters, and for sensitivity of mismatch detection to input spike jitter.

      The only attempt to do this was the gamma measure, which left one wanting (see above).

      We have responded to the points about the gamma measure above.

      Figures are low-quality, labels are missing,

      We have substantially reworked figures, their labels, and legends. The automated mapping from our high-resolution figures to PDF seems to have blurred many of the figures, however, links to the originals should be there in the revision.

      And the analysis stays too close to the data without presenting a clear quantitative synthesis and insight.

      Please see response above. We have tried to balance the process of characterizing numerous readouts and making a model that closely matches experiment, with the high-level insights by way of computational outcomes such as mismatch detection in a variety of more physiological contexts (pulse trains and theta patterned inputs, Figures 7 to 9).

      Key results and mapping between physiology and the model are kept subjective and not quantified.

      Please see response above. We have consolidated our comparisons between physiology and experiments into Figure 4-figure supplements 1 to 3 and Table 1.

      In addition, the similarity measure gamma, which is introduced to express the relationship or the modulation of the response, is mathematically naïve and not well-motivated. It will approach infinity when expected and actual values become more and more similar. While this might be the range where sensitivity is required.

      Please see response above. The metric gamma is derived from previous published work (Bhatia et al, eLife 2019) describing nonlinearities in summation, which is cited. In that study and the current one, there was no instance in which gamma became unreasonably large. It is true that the term gamma is used for many concepts, but we feel that the contexts are so different between summation nonlinearity and oscillation frequencies that confusion is unlikely. We have taken care with the wording in the text to further disambiguate the usage.

      Some detailed observations:

      P2: What is an "interesting" feature?

      We have replaced the word “interesting” with “salient”:

      “We complete the experiment-model-prediction-testing loop and show that differential changes on E and I synapses may provide a mechanism for single neurons to extract salient features of spatiotemporal inputs through STP (Asopa and Bhalla 2023), while keeping mean activity steady.”

      P6 L110: However, over the pulse train, E and I underwent distinct STP profiles (Figure 1 M).

      What makes them distinct?

      This panel is now removed. A clearer account is presented in Figure 2D,E and F:

      “The EPSC showed a trend of early potentiation followed by depression (Figure 2D, 2E), while the inhibition underwent depression from the start (Figure 2 D, Fi)”

      P6 L115: Why can recurrent excitation in the CA3 segment be excluded?

      We thank the reviewer for pointing us to a more detailed analysis, which is now presented in Figure 1-figure supplement 2. We have added the following text:

      “We observed a small amount of ‘ringing’ of the field response which we interpret as either CA3 spiking in a burst, or recurrent activation of the CA3 neurons (Figure 1-figure supplement 2). The ringing was down to ~5% by the third peak which occurred within 8 ms, supporting our treatment of the optical input as a single brief event, and setting a low bound to any contribution to patterns by recurrence.”

      P8 F2A: How are the responses normalized?

      In the text we state:

      “All the PSPs of an 8-pulse train were normalised to the probe pulse.”

      We have added this line into the legend.

      “Traces were normalised to a reference pulse 0, delivered 300ms before the burst.”

      Explain why, given this normalization, the 15 square stimulation is less effective than the 5 square one.

      We acknowledge this was unclear. In the revised text we explain:

      “For the EPSCs, the 15-square trials had a higher reference pulse and higher stimulus overlap (discussed below), hence their normalised peak values were smaller (Figure 2E).”

      F2D: Where do you show that the biphasic response is a statistically significant deviation?

      Thank you for pointing out this missing analysis. We have added it in Figure 3A.

      P9 149: E should be E&F.

      Corrected.

      P9 L150: Explain the "ii" indexing.

      Corrected.

      P10: It is a bit clumsy to call the measure gamma. For general observation on the equation, see the general remark above.

      Please see discussion on this. We are reusing a published term.

      P10 L175: How do your results and F3G show divisive inhibition?

      In the current study we’re not setting out to show divisive inhibition, as that work has been published (Bhatia et al, eLife, 2019). We’ve corrected the text accordingly.

      “Using responses from the reference pulse, we replicated earlier observations (Bhatia et al. 2019; Wehr and Zador 2003) showing divisive normalisation, and obtained a median gamma of 7.16 (95% CI = 4.76 - 10.2)(Figure 3G).”

      Becomes

      “By comparing observed vs. expected responses, we replicated earlier observations (Bhatia et al. 2019; Wehr and Zador 2003) showing sublinear summation, and obtained a median gamma of 7.16 (95% CI = 4.76 - 10.2) (Figure 3G).”

      P16: How does F5 demonstrate a good match between model and physiology?

      We acknowledge we left this out. In F5E we show the model and experiment distributions over different frequencies. Our previous analysis only reported frequency dependence, and now we have added the comparison of response amplitudes. We have inserted the analysis and consolidated the results into Table 1.

      P18 l281: 15-square patterns (yellow dots) almost always gave an EPSP, while 5-square patterns frequently failed.

      Where can I see this? It is mentioned in the caption, but legends are absent.

      In the original source file and in the original confirmation pdf from eLife, the figure legend is present, and has an entry for panel C and D.

      “C,D:probability of trigger to generate a peak in the EPSP trace”

      In the revised version we have quantified these values and put the comparisons into Table 1:

      “Then we compared the probability that each optical stimulus would elicit an EPSP (Figure 6 D). As expected, 15-square patterns (yellow dots) frequently gave an EPSP (77.5±11.7%), while 5-square patterns failed about half the time (51.4±16%). The simulated runs matched this (Table 1).”

      P19 l307: Overall, we were able to replicate numerous features...

      Please be specific. What exactly did you replicate? How is it statistically demonstrated?

      This is a good point, we have updated the text to more systematically work through comparisons and metrics. We have also added some further metrics for features of the responses in Figures 5 and 6. As a way to organize all our comparisons we have added Table 1.

      P22 l341: The transient responses must be proportional to the overlap. Please quantify this effect more precisely.

      In Fig 7 panels L and O we had previously quantified the amplitude of transient responses with respect to two parameters closely related to overlap: pattern sparseness and probability of connections from CA3 to CA1. In Figure 9Bi we show that there is a complex and frequency-dependent relationship between overlap and mismatch responses. In the revision in figures 8 and 9 we have recast the “pattern sparseness” term as the more intuitive “overlap”. These are related almost linearly with a negative slope (Figure 9-figure supplement 1).

      P22 l342: What does "in E" mean?

      Should be Figure 7E for the original version. In the revised paper we have removed this panel.

      l347: I cannot follow. How do these single traces (7C-E) show these effects?

      We acknowledge that the figure and legend did not clearly indicate the timings of the transitions. We have completely redone and reduced figure 7 to simplify the presentation. The timing of transitions between patterns is now indicated using red triangles.

      What does denser connectivity refer to?

      Denser connectivity refers to a higher value for probability of connection between CA3 and CA1. In the revised version we have changed the figure to refer to stimulus overlap:

      “None of the transitions in Figure 8D (dense stimuli, 34% overlap) were significant, but two transitions in Figure 8E were significant (sparse stimuli with 2.5% overlap, p = 1.53e-5 and 6.1e-5).”

      P26: It is unreasonable to expect a reader to put this puzzle together.

      We acknowledge that this is a large and complex figure. In response to the reviewer’s input we have split the figure between Figures 7 and 9, and removed some panels, so as to make it easier to navigate.

      Reviewer #3 (Recommendations for the authors):

      (1) Which parameters are crucial for determining the preferred frequency (i.e., gamma frequency) for mismatch detection? This point should be addressed further.

      This is an interesting suggestion and we have performed additional simulations to address it. It turns out that the frequency tuning is very broad, over almost the entire gamma range from 40 to 200 Hz, and is indeed tuned by simulation parameters. We have placed these findings in Figure 9 in the new version of the paper.

      (2) The meanings of horizontal and vertical color bars should be explained in the legend of Figure 2A. Do they show the average values over columns and rows? A similar question applies to Figure 3G.

      We have removed the marginal heatmaps from Figures 2 and 3 as they were not contributing to the interpretation.

      (3) I wonder whether the proposed mismatch detection is tolerant against timing jitters in repeated presynaptic spike patterns. This information allows us to infer the accuracy required for neural code using population spike patterns.

      This is a good suggestion. We have run additional simulations to quantify this. It turns out that jitter has a clear effect on mismatch detection, and affects 5-square (low-overlap) patterns differently from high overlap (15 square) patterns. The latter see a boost in selectivity with 6 ms jitter. This comparison is now in Figure 7E ii and 7 Eiii

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This study makes a valuable contribution to understanding how negative affect shapes food-choice decision making in bulimia nervosa by leveraging a mechanistic drift diffusion model to quantify the weighting of tastiness and healthiness attributes. The evidence is solid, supported by a randomized crossover design and generally appropriate statistical analyses. However, the interpretability of the findings is limited by ambiguities in the affect manipulation, particularly regarding whether neutral and negative inductions yielded reliably distinct affective states at the time of task performance in the bulimia nervosa group. Consequently, session-related differences in model parameters cannot be unequivocally attributed to negative affect rather than to uncontrolled state or contextual factors, and clearer separation of affective conditions alongside analyses aligned with the paired data structure would strengthen the conclusions.

      We thank the Editor and Reviewers for their careful summary of the study's strengths and for their constructive feedback.

      The eLife Assessment identified two specific limitations that qualified the strength of evidence:

      (1) ambiguity regarding whether the two affect inductions yielded reliably distinct affective states in the BN group at the time of task performance, and (2) analyses that were not fully aligned with the paired data structure. We have directly addressed both concerns in this revision. We provide explicit statistical evidence confirming that neutral and negative inductions yielded distinct affective states in the bulimia nervosa group; and we have re-analyzed all DDM parameters using updated mixed-effects regressions with an unstructured covariance matrix that appropriately accounts for the paired data structure. For completeness, we have also added the requested difference-in-difference analysis. Both approaches yielded conclusions consistent with those originally reported.

      In light of these revisions, we would be grateful if the Editorial Team would consider whether the strength of evidence rating might be updated from "solid" to "convincing." All changes in the revised manuscript are marked in blue.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Using a computational modeling approach based on the Drift and Diffusion Model (DDM) introduced by Ratcliff and McKoon in 2008, the article by Shevlin and colleagues investigates whether there are differences between neutral and negative emotional states in:

      (1) The timings of the integration in food choices of the perceived healthiness and tastiness of food options in individuals with bulimia nervosa (BN) and healthy participants (2) The weighting of the perceived healthiness and tastiness of these options.

      Strengths:

      By looking at the mechanistic part of the decision process, the approach has potential to improve the understanding of pathological food choices.

      Weaknesses:

      I thank the authors for revising their manuscript.

      I still notice that the authors did not go through their manuscript to look for wordings refering to a prediction interpretation of their results while I already highlighted the inappropriateness of this wording in my two first rounds of reviews: e.g. there is still "we used zero-inflated negative binomial models to predict the three-month frequency" and I can find other statements like this. The design of their study does not allow such claims.

      We thank the Reviewer for identifying cases where the term “predicted” may mislead readers about the causal nature of our claims. We have made the following edits (changes are italicized):

      Methods (lines 516-518): “For these exploratory analyses, we used negative binomials to test the association between parameter estimates and the three-month frequency of retrospectively reported Objective Binge Episodes (OBE) and Subjective Binge Episodes (SBE).”

      Figure 5 (lines 881-882): “Affect-induced changes in information onset were associated with more frequent subjective binge episodes.

      The authors answered my major concern regarding the experimental induction towards a negative or a neutral state before running the food decision task. My concern is: BN patients already seemed to be already in a high negative state before undergoing the neutral induction, while these patients are in a lower negative state before undergoing the negative induction. It is therefore not surprising that patients seem to report a similar level of negative state after the two inductions (according to the figure of the authors' previous article). Of note is that the additional analysis the authors ran within the BN group only provides a significant result: this result shows that there has been an induction but does not rule out that patients were in the exact same magnitude of negative state to perform the task as the figure in their previously published article suggests it. The major issue is to show that:

      (1) As compared to the neutral induction, there has been a higher variation in negative state after as compared to before the negative induction.

      (2) The magnitude of the negative state after the negative induction is higher than the magnitude of the negative state after the neutral induction.

      The first point shows that the induction worked. The second point shows that the participants are in two distinct states. Without showing the second point, it may be possible that one induction increases the negative state of participants to the same level as the one of the second induction that has not increased anything.

      Within this context, how is it possible to associate, in patients, a difference in the DDM between the two sessions to a negative state (which is one of the main focus of the article) rather than to another parameter that has not been captured? A similar situation would be in an experiment studying the consequence of stress, a stressfull induction over relaxed participants attending the lab has high chances to raise the level of stress of those participants to the same level as the one that the same participants would experience after a neutral induction when these participants attend the lab with an already high level of stress. In that case, would it be approrpiate to claim that a difference at a task performed after the induction would be related to stress while the participants would be at the same level of stress when performing the task despite the fact that the induction worked ?

      In the experiment performed by the authors, the additional analysis to perform would be a paired sample t-test (or the appropriate non-parametric test) to check whether the magnitude of negative state of BN patients was different between the negative and neutral conditions after the induction only. If not, associating the difference at the DDM with negative states in BN is highly misleading.

      We thank the Reviewer for pressing on this point, and we apologize that our previous response did not make this sufficiently explicit. We agree with the Reviewer that two things must be demonstrated: (1) that the negative induction produced a greater change in negative affect than the neutral induction, and (2) that the magnitude of post-induction negative affect was higher following the negative induction than the neutral induction. We had included the results of analyses addressing both points in the Supplementary Materials of our previous submission, but we appreciate that we had not made this clear in our response.

      Regarding point (1), the mixed-effects model in Supplementary Table S1 yielded a significant Affect Condition × Timing interaction (β = 20.43, SE = 6.35, t = 3.22, p = 0.002), confirming that negative affect increased significantly more from pre- to post-induction in the negative condition than in the neutral condition. This is further supported by within-BN-group analyses in the Supplementary Materials: the negative affect induction produced a large, significant increase in negative affect (mean difference = 20.36, SE = 4.21, t = 4.84, p < 0.0001, Cohen's d = 0.97), whereas the neutral induction was not associated with a significant change in negative affect (mean difference = 7.16, SE = 4.21, t = 1.70, p = 0.327, Cohen's d = 0.34).

      Regarding point (2), we directly compared post-induction negative affect between conditions within the BN group, as requested by the Reviewer. The magnitude of negative affect was significantly higher following the negative mood induction than after the neutral mood induction (mean difference = 17.40, SE = 4.21, t = 4.13, p = 0.0003, Cohen's d = 0.83). This large effect size confirms that participants with BN were in meaningfully distinct affective states when performing the food decision task under the two conditions.

      Together, these analyses establish (1) that the induction worked as intended, and (2) that the two post-induction states were both statistically and practically distinct. We have added explicit language to the manuscript to make both of these points clear (lines: 181-185):

      Critically, post-induction negative affect within the BN group was significantly higher following the negative affect induction than after the neutral affect induction (mean difference = 17.40, SE = 4.21, t = 4.13, p < 0.001, Cohen's d = 0.83; see Supplementary Materials for full details), confirming that BN participants completed the food decision task under meaningfully distinct affective states across the two sessions.

      I read carefully the authors' answer related to mixed models: they claim that mixed models take into account correlations within their repeated data. The specification of the structure of the covariance matrix allows to control only partly for that. I notice that the authors did not specify the structure of that matrix: the article they refer to justify the appropriateness of their analyses is not adapted. The specification of the structure of the covariance matrix needs to address, in a mixed model, the difference in handling 4 repeated data per participants that cannot be paired as compared to 4 repeated data that can be paired (two per session with one before and one after the neutral or negative priming sessions, if I count right). Of note is that a covariance structure that is left free of constraint for the fit of the model does not capture appropriately the pairing of the data: it has all chances to capture the covariance in a different way. And a covariance structure that has constraints has more chances to lead to a model that cannot be estimated because of an absence of convergence of the algorithms.

      By the way, a single two-sample t-test (or a Mann-Whitney test if appropriate), and not a set of multiple paired-sample t-test as the authors suggest, would answer the goal of the authors to test for what they call the three-way interaction in their comment. This test would be performed between the two groups of participants (BN/controls) with the computation for each participant separately: (assessment after neutral induction-assessment before neutral induction)-(assessment after negative induction-assessment before negative induction). This analysis answers points 1, 2 and 4 they raise together with my point of controlling for the paired data. I would have agreed with their choice of a mixed model if they had an unbalanced dataset within each participant.

      We thank the Reviewer for this clarification, and we apologize that our previous response did not adequately distinguish between two different sets of analyses: (1) analyses of DDM parameter estimates, which involved four observations per participant (2 affect conditions × 2 food types); (2) trial-level analyses of choice and response time behavior, where each participant contributed many trials per condition and the dataset is genuinely unbalanced across participants due to trial exclusions – precisely the situation where mixed-effects models with participant-level random slopes are appropriate. The concern about covariance structure applies specifically to the DDM parameter analyses, but does not apply to our trial-level analyses.

      We also want to clarify a point about the task design that may have caused confusion. The Food Choice Task was administered only once per session, after the mood induction (i.e., once after negative mood induction, and once after neutral mood induction). As detailed in Figure 1, the task was not completed pre-induction. The four observations per participant in the DDM parameter analyses therefore reflect 2 affect conditions × 2 food types assessed within each condition, not a pre/post structure. This does not change how we address the concern about covariance structure, as there is still a nested feature of interest (food type within condition), but we wanted to correct this misunderstanding explicitly.

      For the DDM parameter analyses, we agree with the Reviewer that the original random effects structure did not adequately account for the paired nature of the four within-person observations.

      We have addressed this in two ways.

      First, we re-estimated the mixed model specifying an unstructured covariance matrix using the nlme package, which places no constraints on the correlation pattern among the four withinperson observations. We acknowledge the Reviewer's point that an unconstrained covariance matrix is not guaranteed to recover the within-session pairing structure. We explored whether a more constrained specification would be preferable. Specifically, we tested a nested random effect of affect condition within subject, which would directly encode the pairing of Low-Fat and High-Fat observations within each session. However, this model failed to converge. This is not a numerical issue but a fundamental identification problem: with only two observations per session per subject, the session-level and residual variance components cannot be separately estimated. We therefore selected the unstructured model as a more conservative option. Importantly, even if the unstructured model does not explicitly encode the pairing, it is a more general mathematical formula which would not impose incorrect constraints on the correlation structure.

      Consistent with our original findings, the mixed model with an unstructured covariance matrix yielded a significant three-way interaction (Group × Condition × Food Type: β = 0.28, SE = 0.12, t = 2.36, p = 0.020). All simple effects analyses have been updated to reflect the models with this covariance structure, and these are reported in the updated Supplementary Tables.

      Second, following the Reviewer's suggestion (adapted to the actual design structure, in which the Food Choice Task was administered once per session after the mood induction rather than before and after), we computed a difference-in-difference score for each participant's relative attribute onset parameter (τ<sub>s</sub>) following the affect inductions: (negative condition, high-fat − negative condition, low-fat) − (neutral condition, high-fat − neutral condition, low-fat). This score directly encodes the paired structure by construction, bypassing the covariance specification problem entirely. Consistent with the Reviewer's recommendation to use a non-parametric test where appropriate, we used a Wilcoxon rank-sum test (equivalent to Mann-Whitney U) to compare these difference scores between groups. The results confirmed that BN participants showed significantly larger food-type-specific changes in τs following negative affect induction relative to HC (W = 156, p = 0.018). We then applied this approach to all other DDM parameters (i.e., ω<sub>taste</sub>, ω<sub>health</sub>, α, τ<sub>ND</sub>, and z), and report these results alongside updated mixed-effects model results in the Supplementary Materials. The conclusions drawn from the difference-in-difference analyses were consistent with those from the mixed-effects models across all parameters.

      Both approaches converge on the same conclusion and we report both sets of complementary results in the manuscript: the updated mixed-effects models address the full factorial design in a single framework, while the added difference-in-difference analyses explicitly resolve the covariance specification problem by encoding the paired structure directly into each participant’s score, as the Reviewer recommended.

      Reviewer #2 (Public review):

      Summary:

      Binge eating is often preceded by heightened negative affect, but the specific processes underlying this link are not well-understood. The purpose of this manuscript was to examine whether affect state (neutral or negative mood) impacts food choice decision-making processes that may increase likelihood of binge eating in individuals with bulimia nervosa (BN). The researchers used a randomized crossover design in women with BN (n=25) and controls (n=21), in which participants underwent a negative or neutral mood induction prior to completing a food-choice task. The researchers found that despite no differences in food choices in the negative and neutral conditions, women with BN demonstrated a stronger bias toward considering the 'tastiness' before the 'healthiness' of the food after the negative mood induction.

      Strengths:

      The topic is important and clinically relevant and methods are sound. The use of computational modeling to understand nuances in decision-making processes and how that might relate to eating disorder symptom severity is a strength of the study.

      Weaknesses:

      Sample size was relatively small, and participants were all women with BN, which limits generalizability of findings to the larger population of individuals who engage in binge eating. It is likely that the negative affect manipulation was weak and may not have been potent enough to change behavior. These limitations are adequately noted in the discussion.

      We thank the reviewer for their thorough description of the strengths and weaknesses of this study.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The research investigates the frequency-dependent effects of transcutaneous tibial nerve stimulation (TTNS) on bladder function in healthy humans and via a computational model. The authors report that low-frequency (1 Hz) TTNS accelerates the urge to void, while highfrequency (20 Hz) TTNS delays it, corroborated by a computational model suggesting brainstem-mediated mechanisms. The work bridges experimental and theoretical approaches to propose a novel framework for TTNS applications in urinary retention.

      Strengths:

      (1) The integration of human experiments and computational modeling is a major strength. The model successfully replicates bladder dynamics and provides mechanistic insights into frequency-dependent effects.

      (2) Identifies potential therapeutic applications for urinary retention, a condition with limited non-invasive treatments.

      (3) Figures are clear and illustrative, and supplementary materials provide essential methodological depth.

      (4) Controlled experimental design (eg., single-blinded, fluid/caffeine restrictions, etc), detailed computational model parameters and validation against animal data, transparency in data exclusion criteria and statistical adjustments.

      Weaknesses:

      (1) The study uses healthy participants; extrapolation to clinical populations (e.g., urinary retention patients) requires validation.

      The authors have included a statement noting this and explaining that future work will explore this.

      (2) The simulated bladder capacity (100-150 mL) is lower than physiological ranges (300400 mL). While the authors note this, the impact on model validity should be further addressed.

      The authors acknowledge that the simulated bladder capacity and voiding efficiency of the model are lower than human physiological ranges. They have added an additional explanatory paragraph detailing this limitation and proposing the animal training data as a possible cause. Despite these limitations we do not believe this prevents the model from being used to explore proof-of-concept hypotheses (e.g., presence of frequency dependence, potential mechanistic bases) as in the present paper.

      (3) The model omits nociceptive afferents, limiting its applicability to pathological conditions like overactive bladder.

      The authors acknowledge that this is a limitation of the model, and have included a paragraph in the paper’s discussion detailing the limited scope of our in silico approach and clarifying the extent to which the results may be interpreted.

      (4) The lack of significant differences in urge intensity between groups (despite timing differences) warrants deeper discussion. Is the primary effect on efferent activity (as suggested) rather than sensory perception?

      The authors acknowledge that this is a surprising result and as such have deepened the discussion of the pilot study results, including hypothesizing as to potential explanations and suggesting further research in the area.

      (5) One of the highlights of this study is the identification of the effect of low-frequency (1 Hz) tibial nerve stimulation (TNS) on facilitating bladder contraction. Although the authors have clarified this effect in healthy participants, it would strengthen the conclusion if a UAB animal model (e.g., PMCID: PMC7927909, PMC8163611, PMC7847056, PMC8799394) were used to evaluate the same effect.

      The use of animal models is out with the scope of this study which aimed to act as a proof of concept work using a primarily computational approach backed by preliminary human data. The authors acknowledge that this does limit the strength of the conclusions. However, several animal models have been utilized in previous work (as cited in the publication) that demonstrate an excitatory effect of low-frequency tibial nerve stimulation. This work builds upon these previous studies to strengthen the case for a frequency dependent effect of the intervention.

      Reviewer #2 (Public review):

      Summary:

      Tibial nerve (electrical) stimulation (TNS) has emerged over the past 15 years as a non-invasive method to treat bladder overactivity, but interestingly, new animal work has suggested that TNS could actually be used to excite the bladder when appropriately tuning the stimulation frequency, effectively inverting its effect, perhaps opening the door to treat different conditions (e.g., UAB). The present study tests how healthy people respond to low and high frequency TNS, with the authors showing that they can substantially delay people's first sensation of bladder fullness with high frequencies (20Hz, shown many times before) but also that they can slightly hasten people's first sensation with low frequencies (1Hz, new result in humans). Moreover, the authors develop a computational model of interconnected conductance-based simulated neurons arranged in a physiologically plausible circuit that reproduces some aspects of the frequency-dependent effects of TNS. Their simulations suggest that we might expect low-frequency TNS to also increase the duration of bladder contractions in humans. The study highlights a potential new research direction, optimizing TNS stimulation parameters to increase basal bladder excitability.

      Strengths:

      The main strength of the work is to call attention to a new possibility of inverting the effect of TNS in humans by manipulating stimulation frequency, opening new indications for the therapy. This is highly relevant because of the recent popularity of TNS and its non-invasiveness, which lends itself to rapid testing and evaluation for new conditions and a high willingness to adopt. The authors convincingly demonstrate a modest excitatory effect on bladder sensation with low-frequency TNS, which clearly warrants further investigation.

      The high-level design of the hypotheses, concepts, and experiments is clearly articulated in both the methods and in particularly clear diagrams, letting the reader focus their attention on the most important findings.

      It is rare to develop a new computational model of the lower urinary tract at a systems level, and even more so for it to incorporate circuits in the spinal cord and brainstem centers, and this work undoubtedly advances the field's ability to engineer such systems. Further, because the model is comprised of linked conductance-based point-neurons, it is an excellent tool to investigate how an arguably plausible wiring diagram for neural control of the LUT could result in stimulation frequency-dependent effects on pelvic efferents. It is a proof of concept demonstrating how their mechanistic hypothesis of TNS could be implemented neurophysiologically by the nervous system.

      Weaknesses:

      The main drawback of the work is the frequent over-interpretation of the results. The human study and computational model are both proof-of-principle studies because the experimental effect size and sample size are modest, and the computational model is poorly validated and does not generate physiologically typical cystometric responses in simulations that are designed to recapitulate nominal LUT behavior.

      Despite the stated caveats about the small effect in the human study, it should be emphasized throughout that this result is most reasonably interpreted as showing the possibility that TNS can have a low-frequency excitatory effect that merits follow-up, rather than a conclusive demonstration. The effect size is small (as the authors note) and should be placed in context with some minimally clinically important difference, if possible. The result is statistically significant, but even this may be subject to revision due to the small sample and the effect of post-hoc outlier removal and data analysis choices.

      Acknowledged, the authors have included caveats in the discussion making clear that the present results should be interpreted as a proof of concept rather than a definitive demonstration. We note that in combination with existing animal findings these results strengthen the case for the existence of an unexplored excitatory effect of TTNS in human beings that may have valuable clinical implications if generalised.

      Given the apparent mismatch between the model and the cystometric behavior at the systems level in the "normal" case (e.g., low capacity, low voiding efficiency, omitted pressure profiles, frequency, etc.) and the absence of quantitative model validation (e.g., it was not compared directly with any experimental data from human urodynamics or rodent cystometry, beyond the initial fit to the neural data, no sensitivity analyses were performed, no goodness of fit computed, etc.) the discussion should be much more circumspect about interpreting the results at a systems level and should probably contain a paragraph explicitly detailing the limitations of the model. The subsequent interpretation should focus narrowly on the neural circuitry, rather than things like contraction duration, where the model is at its strongest. As written, the authors over-interpret what the in silico study can reasonably be used to infer about LUT function.

      The authors have reworded the discussion section, including a limitations paragraph containing caveats about the interpretation of the results. We make clear that a systemslevel perspective should be maintained and that futher research is required to validate and generalise these results.

      More justification is needed for why the contraction duration of the model is the central focus of analysis, when it connects only tentatively to the human study results, which focus on urgency. While not necessarily incorrect, a clearer link or motivation should be offered for how this informs our understanding of frequency-dependent TNS afferent or efferent inhibition during filling (which was the focus of the human studies and the abstract). In other words, why doesn't the model reproduce the 1Hz excitation effect of expediting void onset (or urgency in the human study), and why is it justified to look at contraction duration as a surrogate measure?

      The authors acknowledge this issue, and have included an additional section to the discussion considering the disparity between afferent and efferent effects observed across the pilot study and computational experimentation. The need for further research within this area to disentangle the complex nature of the frequency dependence has been stressed.

      The authors claim that "voiding behavior occurred earlier [at 1Hz stim in the model]", pointing to Figure 6A as evidence, but this panel appears to show a single example model run where 1Hz voiding occurs only ~1s earlier (display makes this very hard to estimate). This is insufficient evidence to support the claim. Later, it is stated that "TNS did not ... void much earlier". The claims should be made compatible, and all such claims should have reasonable supporting evidence.

      The authors have included additional information in the supplementary materials to support the claim.

      This information includes the bladder volume profile of a number of simulations under 0Hz and 1Hz conditions as well as the average void-onset time (i.e., simulated time before first void).

      There are a number of reporting concerns that can be easily addressed:

      (1) Human Study:

      (a) To interpret the human study analysis, a fuller description of the "optional 10m inute extension" is necessary. How were participants presented with this option, how was blinding preserved, what fraction of participants accepted, and did phase 1 results influence their decisions to continue?

      The authors have included additional clarification detailing how blinding was maintained during the washout period. Additionally, we have included a section in the results which details participation rates for the washout period. Given that only one participant declined participation in the washout period we do not believe it is necessary to conduct an analysis on what factors influenced participation.

      (b) For reproducibility, details about the TNS parameters should be articulated, such as the method of determining "motor thresholds" (unless this is synonymous with "urge to urinate"), the shape of the stimulation pulses (e.g., biphasic, charge balanced), typical applied current, etc.

      The authors have included the requested information and added two figures to the supplementary materials detailing the parameters of the equipment and the exact electrode placement used during the pilot study.

      (2) The Computational Model

      (a) The code availability statement for this type of work is inadequate. The model used for simulations in this work, as well as the code used to initialize (and randomize synaptic connections), needs to be hosted publicly because i) a model this intricate is extremely hard to reproduce/verify without code, ii) simulations are an essential piece of the argument, iii) hosting code requires very little overhead. Although there is an appropriate level of detail in the model description, it would not be possible to reproduce the model in any reasonable amount of time (or at all) because of the implementation-level details that are, understandably, omitted from the methods (e.g., what is a "unit", what 'exactly' do the connections in the PMC and PAG diagrams relate to, what were the final parameters used for all conductances, which parameters were "matched" to the original papers and which were not, etc.).

      The authors have included a link to a public GitHub repository where any interested individuals may download and use the code on their own machines for their own purposes. The repository, which includes a readme file detailing the operation of the model, as well as the thoroughly documented code provide the necessary transparency as suggested by the reviewers. We hope that by making the code open-source in this manner further research efforts by any interested researchers will be stimulated.

      (b) Critical cystometric/urodynamic values that are typically analyzed to assess healthy LUT function are detrusor pressure (timeseries) and/or post-void residual or voiding efficiency (scalars). These should be included to verify that the model is representative of the "normal" case. This is especially important because the model's "normal" behavior appears to have extremely low voiding efficiency (Figure 6A).

      The authors acknowledge this limitation and as such have modified the simulation files to calculate and return: detrusor pressure, post-void residual, bladder capacity, and voiding efficiency (calculated post-hoc from these values). It should be noted however, that implementing this change required that the computational results be re-run using the new code. As such, the exact details of Figure 5 now differ slightly (though the high-level results and implications remain unchanged).

      While the high-level results surrounding the frequency-dependence of TTNS and the likely brainstem specific cause of this effect remain unchanged, there were minor changes in the results of the computational projection experiments that necessitated a re-write of a portion of the results section.

      Additionally, the authors have added a section exploring the low-voiding efficiency of the model at baseline and potential explanatory factors.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 6Cii, the high frequency is labeled as 10 Hz, but it should be 20 Hz. The authors should correct this in the figure legend.

      Acknowledged, the typo has been corrected.

      Reviewer #2 (Recommendations for the authors):

      (1) Data and Analysis:

      (a) Greater detail on analysis exclusion is warranted. What does it mean to have "greater than normal water intake"? Why was a large "urge duration" grounds for exclusion? Was its threshold set post-hoc, which group was that participant from, and does its inclusion (or not) affect the results of the analysis substantially?

      The authors acknowledge the issue of data removal. As such, to address this limitation an alternative analysis was conducted. Rather than frequentist methods, a Bayesian modelling approach and post-hoc ROPE analysis was conducted which included a greater proportion of the dataset (excluding only those who did not undergo neuromodulation, or who directly met the exclusion criteria for the study). This approach was taken as bayesian methods are better suited for smaller sample sizes such as the one utilised in the present work. The ROPE analysis provides additional evidence for a real-world relevance of the effect on bladder function. Though the authors acknowledge that these results are preliminary they hope they will provide initial evidence for the translation of a novel effect of TTNS into human participants.

      (b) It is my understanding that Figure 4C is a plot of G1Hz and G20Hz on the horizontal from 4A and G1Hz and G20Hz on the vertical from 4B-"before". Hopefully, this is correct, and perhaps there is some way to state more simply what data are being reported, as it took me some time to understand.

      The authors confirm that figure 4C is a representation of data from figure panels A, and B. Thee horizontal axis represents the temporal “"urge onset” and the vertical axis the subjective intensity experienced at this point. To clarify this, the authors adjusted the axis labels to make clear the data being reported. Additional clarification was also added to the figure legend.

      (c) The choice of units in Figure 6 makes interpretation harder than it needs to be. Although not SI units, the field commonly reports volume in ml and duration in seconds or minutes (certainly not ms). The horizontal on Figure 6A is especially confusing, since sim cycles are not clearly defined, nor is the reason for the 20ms of them, or if the 1000s of total simulation time means compute-time or simulated time. Is Figure 6A (20ms/cyc)(50000cyc)(1s/1000ms)*(1min/60s) = 16.67 min of simulated time? If so, does the model show >6 voiding events in that time under normal conditions (which probably requires some explanation, since that is unusual)? Later (L216), other terminology of "simulation run" is introduced and further complicates the interpretation of how much simulated time is passing.

      Acknowledged, the authors have updated the units used in figures througout the publication to match standard SI notation (Fig 4: M<sup>3</sup> -> ml, Fig. 5A:M<sup>3</sup> -> ml, 20ms cycles -> seconds, ms->seconds). Authors have also updated the language used in the figure and the paper to make clear that the figure is referring to 500 seconds (16.67 mins) of simulated time.

      (d) It appears that in Figure 6B that a contraction duration of 0ms means no contraction at all - unclear if that is also true for everything below the horizontal dashed line.

      (e) Using p-values for analyzing differences between average model outputs (Figure 6C) is not appropriate, since one can run the model as many times as needed, making any negligible effect size statistically significant.

      The authors acknowledge that the computational nature of the second analysis limits the statistical tests that may be reasonably applied. As such, they have rewritten the results and discussion section to instead compare mean differences/effect sizes without reliance on p-values specifically.

      (2) Clarity and Presentation:

      (a) Figure 3 should be removed since it describes an experiment not conducted in this study and whose data was used only for model fitting, not an integral component of the model concept, analysis, or results. A short description and a paper reference are sufficient.

      The authors acknowledge this feedback and have removed Figure 3 from the publication. We have instead provided a reference and brief description of the data used to fit the parameters of the model.

      (b) L46, based on my understanding, should read something like "...may be a frequency dependent of TTNS, where low frequencies up-regulate bladder activity while higher frequencies downregulate it."

      Acknowledged, this section has been reworded to improve clarity.

      (c) Generally speaking, there is nothing "paradoxical" about a frequency-dependent response to e-stim, which happens throughout the nervous system and even in the LUT with pudendal sensory stimulation. "Surprising", "useful", "underexplored", etc., are all closer to the authors' meaning.

      Acknowledged, the authors have avoided the use of the term paradoxical to better represent the original intent of the research findings.

      (d) I am used to "washout" rather than "runoff", but this is a journal style decision, and either is fine.

      Acknowledged, the authors have replaced the use of the term runoff with washout and adjusted figure 1 to reflect this change.

      (e) L51 "analytically" is a mathematical keyword reserved for closed-form solutions, which is not what the authors actually refer to. Something like "computationally" or "in silico" is closer to their meaning.

      Acknowledged

      (f) L172 "abnormality" should be "non-normality".

      Acknowledged

      (g) L148 "Like the original model", presumably referring to Gorski?

      Correct, wording has been changed to make this clear.

      (h) L208-220 Unclear precisely what is meant by "intensity of the voiding events" or "temporal nature of the cycle".

      Acknowledged, the authors have provided additional clarification to avoid confusion.

      (i) Figure 6C Is "baseline" the nominal model without stimulation, while the "all connected" is the nominal model with stimulation? And all the rest of the conditions indicate what was cut in silico?

      Acknowledged, authors have reworded the figure legend to improve clarity.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In this study, the authors set out to determine how two classes of kinase inhibitors, which stabilise a disease-relevant enzyme in either an active (Type I) or inactive state (Type II), influence its organisation and interactions with microtubule filaments in cells. Using the state-ofthe-art in-cell structural imaging approaches, they examine how these compounds affect the formation of protein filaments and their association with microtubules, and succeed in defining the underlying structural basis for these differences.

      A major strength of the work is the application of in-cell cryo-electron tomography combined with correlative imaging, which enables direct visualisation of protein organisation in a near-native cellular context. The data convincingly demonstrate that the Type I inhibitor compound stabilising the active state promotes extensive LRRK2 filament formation and microtubule bundling, whereas compounds stabilising the inactive state markedly reduce these interactions. The structural analysis further provides insight into how conformational states relate to filament organisation, including modelling of previously unresolved regions of the protein.

      These findings are internally consistent and align well with prior biochemical and structural studies, many of which were performed by the same team.

      There are, however, some limitations that should be noted. The experiments rely on overexpression of the I2020T mutant form of the LRRK2 protein, which is a rare variant, in a single cell type (293T cells), which may not fully reflect endogenous behaviour or wild-type LRRK2 in a physiological context. In addition, while the imaging data are compelling, the functional consequences of the observed filament formation and microtubule association remain unclear.

      The study therefore provides strong descriptive and structural insight, but more limited evidence linking these observations to cellular or disease-relevant outcomes.

      Overall, the authors largely achieve their aims, and the results support their central conclusion that different classes of kinase inhibitors have distinct effects on protein organisation in cells. The work represents an important advance in understanding how small molecules can reshape protein architecture in a cellular environment, with potential implications for therapeutic strategies. The methodological approach will also be of broad interest to the field, as it highlights the power of in-cell structural biology to study dynamic protein assemblies that are difficult to capture using traditional approaches.

      We thank the reviewer for their thoughtful and positive assessment of our work. We appreciate their recognition that in-cell cryo-electron tomography and correlative imaging provide a powerful approach for directly visualizing how small-molecule inhibitors reshape LRRK2 organization in a cellular environment.

      We agree that the use of overexpressed LRRK2I2020T in HEK293T cells represents an important limitation of the present study. This experimental system was selected because it enabled visualization and structural analysis of inhibitor-dependent LRRK2 assemblies in cells. However, the extent to which these observations apply to endogenous LRRK2, wild-type protein, other disease-associated variants, or physiologically relevant cell types remains to be established.

      We also agree that the functional consequences of inhibitor-dependent LRRK2 filament formation and microtubule association remain unresolved. The goal of the present study was to define how type I and type II kinase inhibitors alter the cellular organization and structural state of LRRK2. Our data demonstrate that these inhibitor classes have markedly different effects on LRRK2 filament formation and microtubule association in cells, and provide a structural framework for understanding these differences. Future studies will be required to determine how these assemblies influence LRRK2 signaling, microtubule-based processes, and diseaserelevant cellular phenotypes.

      We thank the reviewer for highlighting both the methodological significance of this work and its potential implications for understanding how therapeutic molecules remodel protein architecture in cells.

      Reviewer #2 (Public review):

      Summary:

      Mutations in Leucine-Rich Repeat Kinase 2 (LRRK2) are a major cause of Parkinson's disease. LRRK2 PD-related mutations all result in increased kinase activity. Therefore, LRRK2 has been the focus of the development of kinase inhibitors. So far, two classes of kinase inhibitors have been identified: type 1 LRRK2-specific inhibitors that stabilize LRRK2 in a closed active-like conformation and broad-range type 2 inhibitors that stabilize LRRK2 in an open inactive-like conformation. Basiashvili et al. used here in cell structural biology to study the effect of both type 1 and type 2 inhibitors on the localization and structural conformation of LRRK2-I2020T.

      Strengths:

      They showed that Type 1 and not Type 2 inhibitors induce LRRK2 filament/ on microtubules.

      Furthermore, they were able to build a structural map of full-length LRRK2 I2020T bound to a Type 1 inhibitor in a closed kinase confirmation. Together, this work thus confirms the data of previous studies that showed that LRRK2 Type 1 and 2 inhibitors differently affect filament formation.

      Weaknesses:

      All conclusions are fully supported by the provided data. However, as the authors indicated themselves, the physiological relevance of LRRK2 microtubule binding is questionable. Furthermore, although the authors used a full-length LRRK2 protein, like in previously published structures, the resolution of the N-terminal domains is rather poor. Therefore, it also remains unclear what we learn from this structure compared to the previously published structures.

      We thank the reviewer for their positive evaluation of our study and for recognizing that our conclusions are supported by the data.

      We agree that the physiological relevance of LRRK2 filament formation and microtubule association remains an important open question. Our study was designed to determine how type I and type II inhibitors affect the cellular organization and structural conformation of LRRK2. We explicitly acknowledge that future studies using endogenous LRRK2, disease-relevant cellular systems, and functional assays will be necessary to determine the biological significance of inhibitor-induced microtubule association.

      We also appreciate the reviewer’s comment regarding the resolution of the N-terminal domains. Although the N-terminal density does not support detailed atomic interpretation, its visualization provides information about the global organization of full-length LRRK2 within an inhibitorinduced, microtubule-associated assembly in cells. Importantly, our study does not claim highresolution structural determination of the N-terminal regions. Rather, the advance is the in-cell structural observation of full-length LRRK2<sup>I2020T</sup> in a type I inhibitor-stabilized, closed-kinase conformation, together with density indicating that the N-terminal repeat regions adopt an organization within the microtubule-associated lattice.

      We have revised the manuscript to clarify this point and to more carefully distinguish the structural information supported by the density from interpretations that would require higherresolution data.

      Reviewer #3 (Public review):

      Summary:

      This paper describes new insights into the effects of type-I and type-II LRRK2 inhibitors on HEK293T cells that over-express GFP-labeled LRRK2-I2020T. Using correlative light microscopy and cryo-electron tomography, a type-I inhibitor leads to the extensive decoration of microtubules with LRRK2, which is not seen for a type-II inhibitor. Subtomogram averaging reveals that LRRK2 binds to the microtubules in a closed-kinase conformation, with density for the N-terminal arms.

      Strengths:

      The paper is well written; the CLEM and cryo-ET appear to be done to a high standard. Consequently, I have only minor comments.

      Weaknesses:

      The resolution of the subtomogram averages is somewhat limited, but the authors have adequately limited the number of degrees of freedom in the fitting of their atomic models by only allowing rigid-body transformations of separate parts of LRRK2.

      The authors should include FSC curves between the rigid-body fitted atomic models and the various sub-tomogram average maps.

      We thank the reviewer for their positive assessment of the manuscript and for recognizing the quality of the correlative imaging and in-cell cryo-electron tomography analyses.

      We also appreciate the reviewer’s recognition that our interpretation of the maps was appropriately constrained by fitting domains as rigid bodies, rather than attempting unsupported high-resolution model refinement.

      We thank the reviewer for highlighting this and apologize for the oversight. We have added all the missing FSC curve plots of subtomogram maps presented in this study in Extended Data Figure 8.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I think the current study is OK as it is, and the authors have taken this as far as they can.

      In future work, for either the authors or others in the field, it will be important to determine whether endogenous LRRK2 can be recruited to microtubules in response to compounds that stabilise the active state, particularly in cell types that are more relevant to Parkinson's disease. Does this cause a roadblock that impacts microtubule-driven transport? Establishing whether such recruitment occurs under physiological expression levels will be critical for assessing the broader relevance of the findings.

      In addition, it would be valuable to evaluate whether these Type 1 compounds have detrimental cellular effects linked to altered endogenous LRRK2-driven microtubule association, and whether inhibitors that stabilise the inactive state offer a potential advantage by avoiding this phenotype.

      We thank the reviewer for insightful recommendations for future studies.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 5: What is map C, and how is it different from the other maps? The authors indicate that the resolution of the N-terminal domains is moderate. How certain are the authors of the fit of these domains? Since map C is not provided in the supplemental, it is not possible to check this.

      We apologize for this oversight. We have updated the text to reflect how the map C was calculated. Now the text reads:

      “Additionally, we performed subtomogram analysis in Dynamo on a larger LRRK2<sup>IT</sup>decorated lattice that contained three layers of LRRK2<sup>IT</sup> density around the microtubule; we refer to this average as map C. Refinement was focused on the central four LRRK2<sup>IT</sup> subunits to better resolve additional protein densities within this larger lattice. In map C (Fig. 5A; Ext. Fig. 7).”

      In addition, we updated the figure 5D-F to demonstrate clear fit of the N-terminal domains into the presented map. We also added an Extended Data Figure 7 to the supplemental materials to highlight the fit of the model in the map and highlight the areas that would correspond to the Nterminal domains of LRRK2. We hope these updates demonstrate a good fit and justify observations highlighted in the paper.

      (2) The authors convincingly confirm that LRRK2 Type 1 and 2 inhibitors differently affect filament formation and that type 1 LRRK2-specific inhibitors stabilize LRRK2 in a closed activelike conformation. However, from the way the paper is written, it is unclear what we learn from this new structural data. How similar is the current structure compared to the previous structures? What is the novelty?

      We thank the reviewer for noting that this is unclear and giving us the opportunity to highlight it in the manuscript. We have added the following sentence in the discussion:

      “However, how the N-terminal repeats of LRRK2 are organized when the protein is in its closedkinase conformation remained unresolved. Stabilization of LRRK2 in a closed-kinase conformation by MLi-2 treatment and microtubule association reduces conformational heterogeneity to permit structure determination of full-length LRRK2<sup>IT</sup> with the N-terminal repeats undocked from the catalytic core. Therefore, the key novelty of this structure is that it captures full-length LRRK2<sup>IT</sup> in a cellular, microtubule-associated closed-kinase state and shows that kinase closure is compatible with an undocked N-terminal architecture. This distinguishes the in situ closed-kinase state from previously described in vitro intermediate active states.”

      Minor comments:

      (1) "Its C-terminal catalytic region is composed of WD40, Roc GTPase, Kinase and COR (RCKW) domains."

      Suggest changing this to Roc GTPase, Cor, Kinase and WD40 (RCKW) domains for clarity/following of abbreviation.

      We have made this change.

      (2) "In the MLi-2 treated cells, LRRK2IT strands were organized around microtubules with a regularly spaced lattice, similar to the LRRK2IT strands in cells not treated without the inhibitor (Fig. 3A-E)"

      Phrasing, correct the underlined portion.

      We have made this change.

      (3) "While average pitch. rise, and handedness of the filaments of the rate GZD-824 treated LRRK2 filaments were similar..."

      Punctuation.

      We have made this change.

      (4) "Our results clarify the relationship between kinase conformation, repeat undocking, and microtubule association. Increased microtubule association observed for I2020T mutant favors repeat undocking, a prerequisite for kinase closure and filament assembly"

      Do the authors mean undocking by the N-terminal repeats or repeatedly undocking of these domains?

      We meant undocking of the domains, and have corrected the sentence to clarify this.

      (5) "Together, these findings provide a structural view of full-length LRRK2 in a closed kinaseconformation and capture a resolved snapshot along its conformational continuum"

      Needs a space.

      We have made this change, and thank the reviewer for pointing it out.

      (6) "Microtubule decoration by LRRK2IT has not been studied in cell types that endogenously express high levels of LRRK2, such as lung epithelial cells and brain-resident immune cells including microglia and macrophages44. Thus, it remains possible that aberrant LRRK2microtubule interactions occur under physiological expression conditions, potentially disrupting homeostatic intracellular transport and being further exacerbated by type I LRRK2 inhibitors, as suggested by in vitro studies23,45."

      Many studies have studied the localization of endogenous LRRK2, however were not able to detect filament localization on microtubules. Moreover, to my knowledge, there is also no clear evidence that type 1 inhibitors disrupt microtubule transport in cells expressing endogenous levels of LRRK2.

      Therefore, I suggest to rephrase or remove this paragraph.

      We agree that the current evidence does not establish that this occurs broadly in cells. However, to our knowledge, cells or tissues with high endogenous LRRK2 expression have not yet been systematically examined in this context. We therefore present sparse decoration of hyperactive LRRK2 on microtubules as a possibility rather than a strong conclusion. We have also previously shown that type I inhibitors disrupt microtubule transport in vitro, but determining whether a similar effect occurs in cells is ongoing work and beyond the scope of the present manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) P4: The first section of the Results refers to LRRK2 localising to microtubules in the presence of the type-I compounds, and to the cytosol with the type-II inhibitor. Aren't microtubules in the cytosol also?

      We meant cytosolic LRRK2, we have revised the text to reflect this. It now reads:

      In cells treated with MLi-2, we observed LRRK2<sup>IT</sup> in extended filaments, puncta, and diffuse in the cytosol (Fig. 1D-E; Ext. Fig 1A-D). In contrast, when cells were treated with GZD-824, LRRK2<sup>IT</sup> was mostly localized to puncta and distributed throughout the cytosol, with reduced filament formation (Fig. 1F-G; Ext. Fig 1E-H), in agreement with our previous work [23,24,40].

      (2) P4: second column, halfway down. I don't understand how the 16 and 8 neighbours are derived from Figure 3J-K. Perhaps indicate this in the figure?

      Thank you for bringing this to our attention. We have added an Extended Data Figure 5 to clarify this point. The Extended data figure 5 highlights and annotates the immediate neighboring LRRK2 densities in the MLi-2- and GZD-824-treated lattices, making clear how the 16 and 8 nearest-neighbor values were assigned from the observed lattice organization.

      (3) P6: first column, halfway down: perhaps make it explicit that only rigid-body fitting was performed because of the limited resolution?

      We have incorporated this useful suggestion. The text now reads:

      “We split this model in three parts: the WD40 and C-lobe of the kinase, the N-lobe of the kinase with ROC and COR domains, and the LRR and ANK domains, aligned and fitted each of these three to our map A (Fig. 4D-F). Given the limited resolution of the map A, we fit the model as three rigid bodies without atomic refinement.”

      (4) P6: same column near the bottom: what is map C? and how was it calculated? Also, it is not clear to me from Figures 5D-F whether the statement "clearly correspond to the LRR-ANK-ARM domains" is justified by the map. From Figure 5D-F, I see a rather poor fit in a low-resolution map. This needs to be toned down or better illustrated.

      We apologize for the oversight. We have updated the text to clarify how the map C was calculated. Now the text reads:

      “Additionally, we performed subtomogram analysis in Dynamo on a larger LRRK2<sup>IT</sup>decorated lattice that contained three layers of LRRK2<sup>IT</sup> density around the microtubule; we refer to this average as map C. Refinement was focused on the central four LRRK2<sup>IT</sup> subunits to better resolve additional protein densities within this larger lattice. In map C (Fig. 5A; Ext. Fig. 7).”

      In addition, we updated the figure 5D-F to better demonstrate the fit of the N-terminal domains into the presented map. We also added an Extended Data Figure 7 to the supplemental materials to further highlight the fit within the map and indicate the areas that correspond to the N-terminal domains of LRRK2. We hope these updates clarify how map C was calculated and better illustrate our interpretation of the additional densities.

    1. Author response:

      We would like to thank the reviewers for their careful analysis of our manuscript. We appreciate their insightful suggestions for improvement. We intend to address each of their comments in our revision, with the major points outlined below.

      (1) Reviewers 1 and 2 both highlighted the importance of the specificity of our genetic and optogenetic manipulations in the interpretation of our results. We agree that this point is essential. We will expand our discussion to incorporate more references demonstrating the specificity of our genetic approach, the networks engaged, and potential caveats.

      (2) We acknowledge the importance of validating the dystonic nature of our model as noted by Reviewer 2 and the value of more objective quantification of dystonic crisis as requested by Reviewer 1. We will discuss the potential as well as the difficulty of developing this kind of classification due to the non-stereotypic nature of dystonic movements and the lack of objective, measurable definitions even in clinical settings.

      (3) Reviewers 1 and 2 also requested additional discussion of the role of the iCNN to CL thalamus projection in driving dystonic crisis. We will clarify our claims on this point to more accurately reflect what we can confidently interpret from our current experiments and discuss the value of further experiments in the future.

      (4) We agree with Reviewers 1 and 2 that the effects of repeated stimulation are intriguing and deserve further investigation in the future. We will expand our discussion of this point to provide additional context and describe potential mechanisms that could explain our observed results, which may be tested in further studies.

      (5) Reviewer 1 noted that the clinical dataset could be discussed in more detail to support the translational relevance of our findings. We will provide additional information on the characteristics of our patient sample and potential confounding variables.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Yang et al. investigates the relationship between multi-unit activity in the locus coeruleus, putatively noradrenergic locus coeruleus, hippocampus (HP) sharp-wave ripples (SWR) and spindles using multi-site electrophysiology in freely behaving male rats. The study focuses on SWR during quiet wake and non-REM sleep, and their relation to cortical states (identified using EEG recordings in frontal areas) and LC units.

      The manuscript highlights differential modulation of LC units as a function of HP-cortical communication during wake and sleep. They establish that ripples and LC units are inversely correlated to levels of arousal: wake, i.e. higher arousal correlates with higher LC unit activity and lower ripple rates. The authors show that LC neuron activity is strongly inhibited just before SWR detected during wake. During non-REM sleep, they distinguish "isolated" ripples from SWR coupled to spindles and show that inhibition of LC neuron activity is absent before spindle-coupled ripples but not before isolated ripples, suggesting a mechanism where noradrenaline (NA) tone is modulated by HP-cortical coupling. This result has interesting implications for the roles of noradrenaline in the modulation of sleep-dependent memory consolidation, as ripple-spindle coupling is a mechanism favoring consolidation. The authors further show that NA neuronal activity is downregulated before spindles.

      Strengths:

      In continuity with previous work from the laboratory, this work expands our understanding of the activity of neuromodulatory systems in relation to vigilance states and brain oscillations, an area of research that is timely and impactful. The manuscript presents strong results suggesting that NA tone varies differentially depending on coupling of HP SWR with cortical spindles. The authors place their findings back in the context of identified roles of HP ripples and coupling to cortical oscillations for memory formation in a very interesting discussion. The distinction of LC neuron activity between awake, ripple-spindle coupled events and isolated ripples is an exciting result and its relation to arousal and memory opens fascinating lines of research.

      Weaknesses:

      I regretted that the paper fell short of trying to push this line of idea a bit further, for example by contrasting in the same rats the LC unit-HP ripple coupling during exploration of a highly familiar context (as seemingly was the case in their study) versus a novel context, which would increase arousal and trigger memory-related mechanisms. Any kind of manipulation of arousal levels and investigation of the impact on awake vs non-REM sleep LC-HP ripple coordination would considerably strengthen the scope of the study.

      Comments on revised version:

      The authors have added methodological details to the results section after the first round of reviews, improving the manuscript readability. Some points might still be improved, for example, the authors use a delta/gamma ratio to track cortical states for example, but there is no methods section corresponding to this metric. Authors write that higher SI corresponds to a lower arousal state that is associated with "more synchronized cortical population activity, higher ripple rate and reduced LC neurons firing" but there are no references or analysis to support this statement, only examples showing changes in SI over a few minutes.

      We thank Reviewer #1 for the positive evaluation of our study and for highlighting its strengths and potential avenues for future investigation.

      We have specified in the Methods the calculation of SI as a delta/gamma ratio and provided the frequency ranges used for each band: “Artefact-free EEG signals were band-pass filtered using a Butterworth filter implemented in Matlab 2024a (MathWorks, Natick, MA). Subsequently, deltaband power (δ, 1–4 Hz), theta-band power (θ, 6–10 Hz), and the θ/δ power ratio were computed within contiguous 4-second epochs.”

      We agree with the reviewer and have acknowledged in the Discussion that incorporating behavioral assays will be essential for achieving a mechanistic understanding of the observed network dynamics and their functional role in memory consolidation. Such experiments are beyond the scope of the present study but represent an important direction for future research. We have also revised the Discussion to avoid overstated claims and to ensure that our interpretation remains appropriately supported by the current data. Discussion (last paragraph): “Conducting behavioral assays before electrophysiological recordings, along with spatially and temporally precise modulation of LC activity during recording sessions, will be essential for achieving a mechanistic understanding of network dynamics and its functional role for memory consolidation in future investigations.”

      Reviewer #2 (Public review):

      Summary:

      In this study, authors studied the synchrony between ripple events in Hippocampus, cortical spindles and Locus Coeruleus spiking. The results in this study together with the established literature on the relationship of hippocampal ripples with widespread thalamic and cortical waves, guided authors to propose a role for Locus Coeruleus spiking patterns in memory consolidation. The findings provided here, i.e. correlations between LC spiking activity and Hippocampal ripples, could provide basis for future studies probing the directional flow or the necessity of these correlations in the memory consolidation process. Hence, the paper provides enough scientific advance to highlight the elusive yet important role of Norepinephrine circuitry in the memory processes.

      Strengths:

      Authors were able to demonstrate correlations of Locus Coeruleus spikes with hippocampal ripples as well as with cortical spindles. Specific strength of the paper is in the demonstration that the spindles that activate with the ripples are comparatively different in their correlations with Locus Coeruleus than those which do not.

      Weaknesses:

      The claims regarding the roles of these specific interactions were mostly derived from the literature that these processes individually contribute to the memory process, without any evidence of these specific interactions being necessary for memory processes. There are also issues with the description of methods, validation of shuffling procedures and unclear presentation and the interpretation of the findings, which are described in points that follow. I believe addressing these weaknesses might improve and add to the strength of the findings.

      Comments on revised version:

      The authors addressed all of my major concerns during the revision. As a result, the study now provides convincing evidence as well as improved presentation of results, that makes this manuscript important to the broader field of neuroscience, beyond the specific sub-field.

      We thank Reviewer #2 for the positive assessment of our work and for recognizing both its strengths and its potential to stimulate future research in this area. We agree that assessing memory function is essential for understanding how noradrenergic signalling influences the network mechanisms underlying memory consolidation. While such experiments are beyond the scope of the present study, we acknowledge this important limitation in the Discussion and identify it as a key direction for future research. Discussion (last paragraph): “Conducting behavioral assays before electrophysiological recordings, along with spatially and temporally precise modulation of LC activity during recording sessions, will be essential for achieving a mechanistic understanding of network dynamics and its functional role for memory consolidation in future investigations.”

      We added more details in the Methods and expanded the Figure 4 legend to improve the results presentation.

      Reviewer #3 (Public review):

      This manuscript examines how locus coeruleus (LC) activity relates to hippocampal ripple events across behavioral states in freely moving rats. Using multi-site electrophysiological recordings, the authors report that LC activity is suppressed prior to ripple events, with the magnitude of suppression depending on ripple subtype. Suppression is stronger during wakefulness than during NREM sleep and least pronounced for ripples coupled to spindles.

      The study is technically sound and addresses a timely and important question regarding how LC activity interacts with hippocampal and thalamocortical network events across vigilance states. While the findings are interesting, they remain observational in nature. Following revision, the manuscript has substantially improved in both presentation and interpretation of the results, and most concerns have been addressed satisfactorily. I therefore only have a few minor considerations that the authors may wish to explore further in the current study or in future work, as these directions could provide additional mechanistic insight and would likely be of considerable interest to the field.

      The authors demonstrate clearly that tonic LC firing rates preceding ripples differ significantly between wake-associated ripples (highest LC firing), isolated ripples during NREM sleep (lower LC firing), and spindle-coupled ripples (lowest LC firing). They also appropriately note that baseline firing differences will naturally influence the magnitude of LC suppression, which they also observe (highest LC reduction for wake ripples, then isolated ripples and last spindle-coupled ripples).

      However, this aspect could be explored further, as it may provide additional insight into the regulation of spindle-associated ripple events. Since LC activity appears to decline gradually prior to ripple occurrence (Suppl. Figure 2), it would be interesting to test whether this gradual reduction helps organize the emergence of isolated versus spindle-coupled ripples. For example, isolated ripples may occur during the initial phase of LC decline, whereas spindle-coupled ripples may preferentially emerge when LC activity reaches its lowest levels. Such a relationship could also be consistent with the stronger synchronization observed for spindle-ripple coupling.

      Related to this point, it would also be informative to examine whether isolated spindles occur more randomly in time, whereas spindle-associated ripple events appear more temporally clustered. If a single isolated spindle occurs, the associated LC suppression might be more pronounced. In contrast, when multiple spindle-associated ripple events occur in succession, LC activity may already be reduced following the first event, resulting in smaller additional suppression preceding subsequent events. Exploring this possibility could help clarify how LC dynamics shape the temporal emergence of ripple-subtypes

      We are grateful to Reviewer #3 for the positive evaluation of our manuscript and for the constructive comments highlighting the significance of our findings and their implications for future studies. We agree that a more comprehensive investigation of cross-regional coupling and its modulation by the LC–NE system represents an important and still insufficiently explored area of research. Further elucidating the complexity of these interactions will be essential for understanding how noradrenergic signalling shapes large-scale brain network dynamics across behavioral states. We acknowledge it in the Discussion: “A more comprehensive investigation of cross-regional coupling and its modulation by the LC–NE system represents an important and still insufficiently explored area of research. Further elucidating the complexity of these interactions will be essential for understanding how noradrenergic signaling shapes large-scale brain network dynamics across behavioral states.”

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Figure 4: It would be helpful to show the unshuffled data at the front (it is hidden partly behind the unshuffled data). Also, the unshuffled data are not introduced in the text for this figure. Would be helpful. Please also add color bars to improve interpretability.

      To improve readability and facilitate interpretation, we revised Figure 4. Specifically, we 1) reordered the plots to present the unshuffled (ripple) data at the front; 2) expanded the figure legend to provide a more detailed description of the shuffling procedure; and 3) removed the unnecessary color fill from the box plots in panels B and C, while retaining the labels.

      Figure 7: The color coding appears wrong in panel F (mean curves in F do not correspond to time traces in G). This should be checked and corrected if necessary.

      We have corrected the colour coding in Figure 7.

    1. Author response:

      We thank the reviewers for their thoughtful and constructive comments, and we plan to implement many of their suggestions to improve the paper. We agree that the manuscript would benefit from a clearer and more evidence-based presentation of how feedback responses relate to subsequent learning responses. To address this point, we will perform additional analyses and modeling, including model-free analyses of the phasic and tonic components. These analyses will allow us to test whether the tonic component remains the dominant predictor of the learning response without relying on the specific assumptions of the tonic/phasic decomposition model.

      We also agree that the manuscript would benefit from a more detailed discussion of the mechanisms that may shape the temporal evolution of feedback responses and their relationship to subsequent learning. We will therefore expand the discussion of this issue and relate our findings to adaptive feedback control and continuous-time models of motor adaptation, which may provide useful frameworks for interpreting the relationship between feedback responses and learning responses.

      Finally, we agree that the scope and limitations of the current experimental paradigm should be discussed more explicitly when considering the generality of our findings. We will therefore discuss whether and how the present results may generalize to broader forms of sensorimotor learning and adaptation. We will also

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study analyses correlations between traits of Chinese frog species and their Red List status, finding differences between adults and larvae and thus pointing to the importance of considering different life-cycle stages in this and possibly other animal groups when assessing species extinction risks. The current study is, however, incomplete because of unclear threat categories for tadpoles, the omission of other key species traits, and insufficient statistical analysis.

      Thank you very much. We have revised the manuscript according to the reviewers' comments. The parts highlighted in red in the manuscript are the revised portions.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript shows that different traits of adults and larvae correlate with Red List status. The authors argue that this shows a big gap in the conservation of amphibians and that the traits of all life stages should be taken into account in amphibian conservation. Specifically, amphibian conservation should do more for the habitats where the larvae live.

      The manuscript is well written and easy to understand. The methods are sound.

      While the study will make an interesting contribution to conservation science, there are many things that I disagree with.

      (1) I don't think that amphibian larvae and their requirements are a "blind spot" as the title suggests. When reading the manuscript, I didn't learn how conservation practice should change in response to the results.

      Thank you very much for your suggestions. The description of the 'blind spot' was inappropriate, and we have revised it. Investigating the relationship between life history traits and threat status can help us understand which species are more vulnerable to extinction. Furthermore, we can predict the potential threat severity of species that have not yet been assessed. Because we still lack knowledge about the biodiversity of many taxonomic groups. For example, as of early 2024, over 34% of Chinese anuran species have been described in the last ten years, and 100 - 200 new species are still being discovered globally each year. Under these circumstances, given the current investment in biodiversity conservation, it is nearly impossible to assess the threat status of every species and develop conservation strategies. Therefore, predicting the threat status of species is very important for biodiversity conservation, as it will provide support for the subsequent formulation of specific conservation policies. Among the already described animals species, most have complex life history cycles. Moreover, species face threats not only at the adult stage; those with certain traits at other life stages may also be vulnerable to threats. For example, our study takes amphibians as an example and shows that groups with larger body sizes at the tadpole stage may face more serious threats.

      (2) I wonder whether the relationship between species traits and extinction risk is of great importance for conservation. If a species is Data Deficient on the IUCN Red List, then species traits could be used to predict its Red List category. However, for other conservation projects, I don't see how this would work. How would traits be linked to captive breeding, conservation translocation, pond construction or habitat management in general? In some cases, I can envision a link between species traits and pond hydroperiod.

      Thank you very much for your suggestions. Understanding the relationship between traits and threat status is of great importance for the conservation policies and the allocation of conservation resources, especially when conservation resources are insufficient. As mentioned earlier, the current conservation resources are insufficient to support us in surveying and assessing every Data Deficient (DD) species, not to mention the large number of new species being discovered each year. By predicting threat status, we can identify which groups or species should be prioritized for research, such as population size and distribution range surveys, so that specific conservation strategies can subsequently be developed.

      (3) Species traits are body size and morphological traits. That makes sense. However, one of the species traits was microhabitat. I find it far-fetched to call habitat a species trait. This is standard habitat ecology. It is well known that habitats matter and that different habitat types face different threats, and consequently, the species that live in those habitats. Furthermore, habitat and morphology may be confounded. For example, tadpoles in lentic and lotic habitats have very different morphologies. So is it habitat or morphology?

      Thank you very much for your suggestions. The type of habitat in which a species lives affects the threats it faces. In many studies on the relationship between extinction risk and traits, microhabitat or habitat type is widely used as a predictive variable. For example, in studies on Squamata, whether a species is distributed on islands or peninsulas has also been included as a trait. Following your suggestion, we have revised the sentences to refer to 'morphological traits and microhabitat information'. Many morphological traits of species are related to habitat selection, but not all traits associated with habitat selection have been measured or have sufficient data. Therefore, it is necessary to include microhabitat type as an independent variable. Additionally, we calculated the Variance Inflation Factor (VIF) prior to the regression analysis to ensure that the analysis was not affected by multicollinearity.

      (4) I don't know how the threat status of Chinese amphibians is determined. IUCN has multiple reasons why a species can be Red Listed. One reason is range size, and another reason is population decline. Personally, I don't think they should be pooled in an analysis because they are fundamentally different reasons why a species has a high extinction risk. A reduction in population size of greater than 30% in 10 years or 3 generations is not the same thing as a small distribution range. Another issue is that IUCN developed the Green Status of species. The Green Status shows that even a species which is LC on the Red List may be significantly depleted.

      Thank you very much for your valuable suggestions. The assessment method of the China Biodiversity Red List is the same as that of the IUCN Red List, both of which are based on population size and area of distribution. We fully agree with your point that analyses should be conducted according to specific threat types. Unfortunately, the full report of the latest version of the China Biodiversity Red List, released in 2023, has still not been published. Therefore, we were unable to perform the relevant analyses.

      (5) The species traits in Table 1 are mostly functional/morphological and body size related (and microhabitat). While there may be correlations between traits and Red List status, it is unknown whether this is correlation or causation. In addition, it is difficult to know the conservation interventions that may be necessary now that we know that relative head with and Red List status are correlated.

      Thank you for pointing out the important distinction between correlation and causation. Your comment is very insightful, and we have revised our manuscript to further clarify the scope and limitations of our study. The aim of our study is to identify which traits show statistical associations with extinction risk, thereby providing testable hypotheses for future research. We acknowledge that the mechanisms underlying the associations between certain morphological traits (e.g., head length, tympanum diameter) and extinction risk remain unclear, and these findings cannot yet be directly translated into well-established management measures. Nevertheless, the value of our study lies precisely in generating hypotheses about traits that warrant prioritized investigation of their causal mechanisms, as well as offering clues for the initial allocation of conservation resources. Following your suggestion, we have discussed the limitations of the study in the Discussion section of the manuscript.

      (6) In the discussion, the authors explain why body size and other traits may affect extinction risk and whether there is a causal relationship. I agree that body size may have a direct effect because larger species are harvested more frequently (it was interesting to learn that tadpoles are harvested as well). However, as macroecological studies show, smaller species often have larger populations than larger species. Abundance may matter.

      Thank you very much for your suggestion. Following your advice, we have revised the discussion section regarding body size.

      (7) I found it much harder to understand why relative head length and tympanum size correlated with Red List status. I wasn't convinced by the arguments in the discussion. Typanum size may be related to hearing and anthropogenic noise. Several studies are cited which show that frogs alter their calling behaviour in response to noise. Crucially, however, they describe changes in behaviour or properties of the advertisement call, yet none show that noise has effects on population viability. If some anthropogenic stressor affects individuals, then this does not mean that it will cause a population decline. When IUCN published the second global amphibian assessment, did they list noise as a major threat to amphibians?

      We appreciate your insightful comments and fully agree with your assessment. Indeed, the hypothesis that noise threatened anuran amphibians lacks direct evidence. While relevant studies indicate that anthropogenic noise causes auditory masking in anurans and reduces individual reproductive success, the IUCN has not listed noise as a primary threat to amphibians. Although acoustic communication is vital for amphibian reproduction and is susceptible to noise interference, there is currently no definitive evidence proving that noise extensively impacts amphibian survival. Therefore, in the revised manuscript, we retained it as a hypothesis to be tested and explicitly clarified that current evidence is limited to behavioral changes. Regarding the correlation with relative head length, we acknowledge that the underlying mechanism remains unclear; it may stem from phylogenetic signal residuals or unidentified ecological factors (such as diet or locomotor ability). In the Discussion, we revised this part as a correlation requiring further investigation.

      (8) There are statements that the tadpole stage is the most important stage: "a critical period for amphibian survival" (line 78-79). While there is high mortality in the tadpole stage, tadpole survival is rather unlikely to affect population survival. Many population models show this. See, for example, Biek et al. 2002 in Conservation Biology. Other papers have argued that the postmetamorphic juvenile stage is most important (Petrovan and Schmidt 2009 Biological Conservation).

      We greatly appreciate your comment. We agree that the original statement was overly absolute. The most critical life stage for population persistence can differ across species, and many studies have shown that other stages may be more important. Accordingly, we have revised this sentence as you suggested.

      (9) The authors repeatedly make the statement that amphibian conservation should focus more on the tadpole stage. I don't understand why this statement is made. For example, a major activity in amphibian conservation is the restoration and de novo construction of ponds (see Calhoun et al. 2014 PNAS, Moor et al. 2022 PNAS). Ponds are habitats for tadpoles. Others removed fish from amphibian breeding sites because fish prey on tadpoles (and adults; see Vredenburg 2004 PNAS). Semlitsch (2002 in Conservation Biology) argued that the management of pond hydroperiod is a critical element of amphibian recovery plans. Ponds should be temporary because this effectively removes predators that consume tadpoles. Clearly, the tadpole stage is not a neglected stage in amphibian conservation.

      Thank you for pointing this out. The literature you cited (Calhoun et al., 2014; Moor et al., 2022; Vredenburg, 2004; Semlitsch, 2002) convincingly demonstrates that the tadpole stage has received a certain degree of attention in amphibian conservation practice. Our original statement was indeed problematic. What we intended to convey is that information on the tadpole stage needs to be integrated into conservation assessment frameworks and conservation planning. For example, many studies on the relationship between functional traits and threat extent have not included tadpole-related information. Compared with our knowledge of adult amphibians, we know far less about tadpoles, and for many species, information on the tadpole stage is entirely lacking. Therefore, we call for tadpoles to receive greater attention in future research relative to the current situation.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Conceptual problems:

      (1) Many conservation measures for amphibians target larvae; thus, globally, this is not a blind spot. If this is different in China, it would be important to point this out.

      We thank the reviewer for the thoughtful comment. We recognize that the tadpole stage has indeed received attention in amphibian conservation practice, and our original statement was therefore imprecise. Our intended argument was that tadpole-stage information should be integrated into conservation assessment frameworks and conservation planning. For instance, many studies examining the relationships between functional traits and threat extent have failed to include data on tadpoles. Our understanding of tadpoles remains far more limited than that of adult amphibians, and for a large number of species, no information on the tadpole stage is available. Consequently, we advocate for substantially greater research attention to tadpoles than they currently receive. We have revised the text accordingly.

      (2) While traits may be used to predict Red-List status, it is not clear how they could inform conservation measures. This should be discussed.

      Thank you for your comment. The aim of our study is to identify which traits show statistical associations with extinction risk, thereby providing testable hypotheses for future research. We acknowledge that the mechanisms underlying the associations between certain morphological traits (e.g., head length, tympanum diameter) and extinction risk remain unclear, and these findings cannot yet be directly translated into well-established management measures. Nevertheless, the value of our study lies precisely in generating hypotheses about traits that warrant prioritized investigation of their causal mechanisms, as well as offering clues for the initial allocation of conservation resources. Following your suggestion, we have discussed the limitations of the study in the conclusion section of the manuscript.

      (3) The Red-List categories may not be appropriate to link traits to extinction risk. It would be important to explain how these are defined for China and how this may affect the analysis (e.g. linking larval traits to larval extinction risks would be difficult if Red-List criteria do not consider larvae).

      Thank you very much for your suggestions. The assessment method of the China Biodiversity Red List is the same as that of the IUCN Red List, both of which are based on population size and area of distribution. The assessment process is independent of species' morphological traits. Consequently, analyzing correlations between traits and Red List categories does not constitute circular reasoning or contain any inherent logical contradiction. On the contrary, it is precisely because the two are independent that statistically significant associations between traits and extinction risk can have predictive value and inform conservation actions. In the revised manuscript, we clarified the independence of Red List assessments and rephrase any potentially misleading wording (e.g., changing "threat category of tadpoles" to "threat category of the species (assessed based on adults)").

      Methodological problems:

      (4) Choice of traits. Are morphological traits sufficient (add e.g. fecundity)? Justify the use of habitat traits (also, if additional ones would be included: geographic and altitudinal ranges, habitat specificity).

      Thank you for your suggestion. We fully agree that traits such as geographic range, elevational range, fecundity, and habitat specificity have important effects on extinction risk. The core objective of this study is to compare the stage-specific differences in the associations between extinction risk and morphological and microhabitat traits of adults versus tadpoles. Moreover, spatial traits such as geographic range are inherently highly correlated with the threat status of species, and including them might mask life-stage-specific signals. We will acknowledge this limitation in the discussion and identify the above-mentioned traits as important directions for future research.

      (5) Model choice: models have high uncertainty, thus better use model averaging and AICc instead of AIC. Overall, the statistical analysis and model selection procedure are poorly described; only summary results are presented.

      We greatly appreciate the reviewer's suggestion. Accordingly, we re-analyzed the data following your advice. In addition, the description of the methods has been supplemented.

      (6) Caveats: the data only allow for correlational analysis; causation cannot be derived from observational data. Furthermore, with a limited number of species, the number of predictors should not be too large.

      Thank you for your suggestion. Studying the relationship between traits and species threat status is important in conservation biology. Although such studies can only reveal statistical associations between traits and extinction risk rather than infer causality, they can generate hypotheses to facilitate future research. Additionally, this type of study can help predict the threat severity of unevaluated species, which is highly valuable for developing biodiversity conservation plans. In this study, 299 species were included in the analysis, and nine predictor variables (eight morphological traits plus one microhabitat type) were used. The ratio of sample size to number of variables was approximately 33:1, and variance inflation factor (VIF) tests indicated that multicollinearity was within an acceptable range (VIF < 5). Therefore, the risk of model overfitting is low. We will add this clarification in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) My first major concern is the species threat categories for tadpoles. The authors obtained the extinction risk data from the China Biodiversity Red List or IUCN. However, the assessment of threat categories, whether by the China Biodiversity Red List or IUCN, is based solely on adults. That means that the threat categories for both adults and tadpoles are the same, which can be seen in Figure 1. Since there is no specific assessment of threat categories for tadpoles, I have concerns about whether it is reasonable to relate species traits of tadpoles to the extinction risk for adults. I think it is one of the reasons why there is no study examining the association between functional traits and extinction risk in tadpole stages.

      We thank the reviewer for raising this important point, as it addresses a key prerequisite issue. The Red List assessment evaluates species, not individual life stages. The threat categories of both the IUCN and China Biodiversity Red Lists are determined based on criteria such as population size and geographic range of the species. The assessment process is independent of species' morphological traits. Consequently, analyzing correlations between traits and Red List categories does not constitute circular reasoning or contain any inherent logical contradiction. On the contrary, statistically significant associations between traits and extinction risk can have predictive value and inform conservation actions. In the revised manuscript, we will explicitly clarify the independence of Red List assessments and rephrase any potentially misleading wording (e.g., changing "threat category of tadpoles" to "threat category of the species (assessed based on adults)").

      (2) My second major concern is about the Data Analysis. The authors built and compared three types of models, i.e., PGLS_BM, PGLS_OU, and GLS_no_phylogeny. They claim that the OU-based PGLS model provided the best fit for both adult and tadpole datasets. Although the result seems reasonable, it is not clear how the OU-based PGLS model was obtained and what it exactly means. It seems to be a full model including all the predictor variables. However, since eight morphological traits and one microhabitat data of both adults and tadpoles were collected, there should be 29-1=511 candidate models. Unless the best model has an Akaike weight (wi) > 0.90 in all the OU-based PGLS models, it has substantial model selection uncertainty. If this is the case, the model average should be used, and weighted estimates of regression coefficients and unconditional standard errors that incorporate model selection uncertainty are better statistical methods (Burnham & Anderson, 2002).

      Thank you very much for your suggestion. Species' traits are related to evolutionary relationships, with more closely related species tending to be more similar. In the original manuscript, the three models we compared (PGLS_BM, PGLS_OU, GLS_no_phylogeny) were intended to select the optimal evolutionary covariance structure. Since we were more interested in the differences between adults and tadpoles, after selecting the OU structure, we actually used a single full model that included all traits to estimate the regression coefficients for each factor. Following your advice, we have added a model averaging analysis and revised the manuscript accordingly.

      (3) In addition, the Second-Order Information Criterion AICc, but not AIC, should be used for model selection. You have at least 9 variables (eight morphological traits and one microhabitat data) or 11/13 variables for the parameter estimates (Table 1). However, you have only 299 species included in the analysis (n = 299), which is relatively small compared to the number of variables (n/k << 40). Therefore, the AIC corrected for small sample size (AICc) should be used.

      We greatly appreciate the reviewer's suggestion. Accordingly, we re-analyzed the data following your advice.

      (4) Previous studies found that amphibian species with large body size, restricted geographic and elevational ranges, low fecundity or high habitat specificity are frequently predicted to have higher extinction risk (Cooper et al., 2008; Sodhi et al., 2008; Botts et al., 2013; Lips et al., 2003; Murray & Hose, 2005). The authors only included morphological traits and one microhabitat data point in the analyses. I wonder whether they can collect more trait data associated with extinction risk, such as geographic and elevational ranges, fecundity traits, or diet/habitat specificity, so as to gain more insight into the study.

      Thank you for your suggestion. We fully agree that traits such as geographic range, elevational range, fecundity, and habitat specificity have important effects on extinction risk. The object of this study is to compare the stage-specific differences in the associations between extinction risk and morphological and microhabitat traits of adults versus tadpoles. Moreover, spatial traits such as geographic range are inherently highly correlated with the threat status of species, and including them might mask life-stage-specific signals. In the Methods, we acknowledge this limitation and identify the above-mentioned traits as important directions for future research.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study provides an important assessment of how body size influences the occurrence of macro-organisms in urban areas across the globe. Size in most plants, but only some animal families, was positively associated with urban tolerance. The data set is impressive, but the evidence for broad-scale conclusions is incomplete due to methodological issues that need to be resolved.

      We have substantially revised the manuscript to resolve the methodological issues raised, including clarifying the definition, calculation, and interpretation of urban affinity (formerly named urban tolerance), and tightening the scope of our conclusions to align directly with the evidence presented.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors integrate multiple large databases to test whether body sizes were positively associated with which species tolerate urban areas. In general, many plant families showed a positive association between body size and urban tolerance, whereas a smaller, though still non-trivial, percentage of animal families showed the same pattern. Notably, the authors are careful in the interpretation of their findings and provide helpful context for the ways that this analysis can be generative in shaping new hypotheses and theory around how urbanization influences biodiversity at large. They are careful to discuss how body size is an important trait, but the absence of a relationship between body size and urban tolerance in many families suggests a variety of other traits undergird urban success.

      We appreciate this thoughtful and balanced assessment of our work and fully agree with the reviewer’s interpretation. In particular, we share the view that the heterogeneous and often weak association between body size and urban affinity across many families is an important result in its own right, underscoring that no single trait is likely to explain urban success across the tree of life. As the reviewer notes, our intention was not to present body size as a universal predictor, but rather as a widely available, integrative trait that can help reveal where general patterns do and do not emerge. We view the lack of a consistent relationship in many families as strong motivation for future work that explicitly integrates additional functional traits and ecological contexts, and we have clarified this perspective in the revised manuscript.

      Strengths:

      The authors aggregated a large dataset, but they also applied robust filters to ensure they had an adequate and representative number of detections for a given species, family, geography, etc. The authors also applied their analysis at multiple taxonomic scales (family and order), which allowed for a better interpretation of the patterns in the data and at what taxonomic scale body size might be important.

      We thank the reviewer for highlighting these strengths of the study. Considerable effort went into assembling, harmonizing, and filtering these data across taxa, regions, and taxonomic resolutions, and we were deliberate in applying conservative thresholds to ensure that species-level urban affinity estimates were based on adequate and comparable sampling. We hope that, beyond the specific results presented here, the compiled dataset and analytical framework will serve as a valuable resource for future studies aiming to explore additional traits, taxa, or mechanisms underlying species’ responses to urbanization.

      Weaknesses:

      My main concern is that it is not fully clear how the measure of body size might influence the result. The authors were unable to obtain consistent measures of body size (mean, median, maximum, or sex variation). This, of course, could be very consequential as means and medians can differ quite a bit, and they certainly will differ substantially from a maximum. And of course, sex differences can be marked in multiple directions or absent altogether. The authors do note that they selected the measure that was most common in a family, but it was not clear whether species in that family that did not have that measure were removed or not. This could potentially shape the variability in the dataset and obscure true patterns. This may require additional clarity from the authors and is also a real constraint in compiling large data from disparate sources.

      We appreciate this important point and agree that heterogeneity in how body size is measured (e.g., mean vs. maximum values, sex-specific measures) is a real but unavoidable challenge when compiling organismal trait data across such a broad taxonomic scope. We would like to clarify that our analytical approach was explicitly designed to minimize the influence of this heterogeneity rather than ignore it. Specifically, for each family we retained all species for which at least one body size estimate was available, rather than removing species that lacked a particular measurement type. When multiple body size measures existed for a species, we selected the measurement type that was most commonly available within that family in order to maximize comparability among species while retaining sample size. Importantly, differences among body size measurement types (including units, measurement detail, and whether values reflected means, maxima, or sex-specific estimates) were further accounted for by (i) log-transforming all body size values and (ii) centering and scaling body size values within each measurement type, which was included as a random effect in the hierarchical models. This approach reduces the influence of systematic differences among measurement types on estimated relationships with urban affinity. We have added a sentence to the methods clarifying that species with a single measurement type were not removed from analyses:

      “Importantly, this procedure did not result in the exclusion of species lacking a particular body size measurement type; rather, all species with at least one available body size estimate were retained, with measurement heterogeneity explicitly accounted for through hierarchical modeling.”

      We agree that variation in body size definitions may still contribute residual noise and potentially obscure weak relationships, and we now emphasize this more clearly as a limitation of large-scale trait syntheses. However, because our primary inference focuses on the presence, absence, and direction of size–urban affinity relationships across families, rather than precise effect sizes, we believe our approach provides a robust and conservative test of whether body size consistently predicts urban affinity across taxa. We highlight this point in the limitations section of our manuscript:

      “One important limitation of our synthesis is the heterogeneity in how body size is measured across taxa, including differences among mean, maximum, and sex-specific estimates. While our analytical framework explicitly accounts for this variation through transformation, scaling, and hierarchical modeling with random intercepts (see Methods), residual measurement noise may still obscure weak size–urban affinity relationships. This challenge is inherent to large-scale trait syntheses that integrate data from disparate sources, and highlights the need for continued efforts to standardize trait databases and expand the availability of harmonized organismal trait data across the tree of life.”

      Reviewer #2 (Public review):

      I have completed a thorough review of this paper, which seeks to use the large datasets of species occurrences available through GBIF to estimate variation in how large numbers of plant and animal species are associated with urbanization throughout the world, describing what they call the "species urbanness distribution" or SUD. They explore how these SUDs differ between regions and different taxonomic levels. They then calculate a measure of urban tolerance and seek to explore whether organism size predicts variation in tolerance among species and across regions.

      The study is impressive in many respects. Over the course of several papers, Callaghan and coauthors have been leaders in using "big [biodiversity] data" to create metrics of how species' occurrence data are associated with urban environments, and in describing variation in urban tolerance among taxa and regions. This work has been creative, novel, and it has pushed the boundaries of understanding how urbanization affects a wide diversity of taxa. The current paper takes this to a new level by performing analyses on over 94000 observations from >30,000 species of plants and animals, across more than 370 plant and animal taxonomic families. All of these analyses were focused on answering two main questions:

      (1) What is the shape of species' urban tolerance distributions within regional communities?

      (2) Does body size consistently correlate with species' urban tolerance across taxonomic groups and biogeographic contexts?

      We thank the reviewer for their careful reading of the manuscript and for this generous and accurate summary of the study’s aims, scope, and contributions. We appreciate the recognition of our group’s broader body of work using large biodiversity databases to quantify species’ associations with urban environments, and we are grateful for the reviewer’s acknowledgement that this study extends those efforts to an unprecedented taxonomic and geographic scale. We agree with the reviewer’s articulation of the two core questions motivating the paper, and we have revised the manuscript to ensure that these questions are stated clearly and addressed consistently throughout.

      Overall, I think the questions are interesting and important, the size and scope of the data and analyses are impressive, and this paper has a potentially large contribution to make in pushing forward urban macroecology specifically and urban ecology and evolution more generally.

      Thanks! We see this work as an effort to move beyond species-by-species descriptions of urban responses toward a community- and distribution-level perspective, where the shape of species’ urban associations themselves becomes an object of study. By framing species’ distributions along an urbanization gradient as a collective property of regional species pools, our approach opens a complementary way of thinking about how urbanization filters biodiversity.

      Despite my enthusiasm for this paper and its potential impact, there are aspects that could be improved, and I believe the paper requires major revision.

      Some of these revisions ideally involve being clearer about the methodology or arguments being made. In other cases, I think their metrics of urban tolerance are flawed and need to be rethought and recalculated, and some of the conclusions are inaccurate. I hope the authors will address these comments carefully and thoroughly. I recognize that there is no obligation for authors to make revisions. However, revising the paper along the lines of the comments made below would increase the impact of the paper and its clarity to a broad readership.

      We appreciate the detailed comments provided and have addressed each point in turn - see detailed responses below. We took these concerns seriously and undertook a substantial revision of the manuscript. In summary, we clarified the conceptual framing of “urban tolerance” (now referred to as “urban affinity”), explicitly defined the metric and its interpretation, added equations and a step-by-step methodological roadmap, and expanded justification for our regional stratification. Where appropriate, we refined language in the Results and Discussion to ensure conclusions are tightly aligned with what the metric can and cannot support. We agree that these revisions materially improve the clarity, rigor, and interpretability of the study, and we appreciate the reviewer’s perspective on how doing so strengthens the paper’s contribution and accessibility to a broad readership.

      Major Comments:

      (1) Subrealms

      Where does the concept of "subrealms" come from? No citation is given, and it could be said that this sounds like an idea straight out of Middle Earth. How do subrealms relate to known bioclimatic designations like Koppen Climate classifications, which would arguably be more appropriate? Or are subrealms more socio-ecologically oriented? From what I can tell, each subrealm lumps together climatically diverse areas. It might be better and more tractable to break things in terms of continents, as the rationale for subrealms is unclear, and it makes the analyses and results more confusing. The authors rationalized the use of subrealms to account for potential intraspecific differences in species' response to urbanization, but that is never a core part of the questions or interpretation in the paper, and averaging across subrealms also accounts for intraspecific variation. Another issue with using the subrealm approach is that the authors only included a species if it had 100 observations in a given subrealm, leading to a focus on only the most common species, which may be biased in their SUD distribution. How many more species would be included if they did their analysis at the continental or global scale, and would this change the shape of SUDs?

      We thank the reviewer for raising this point and agree that the rationale for using subrealms required clearer explanation. Next to allowing potential intraspecific differences in urban affinity across regions, our subrealm-based approach also provides a practical way to partition global biodiversity into ecologically meaningful regional assemblages while maintaining sufficient sample sizes for analysis. Urban affinity is likely to vary geographically within species due to differences in climate, habitat availability, urban form, and evolutionary history. By calculating urban affinity within subrealms rather than globally, our approach allows species to exhibit region-specific urban affinities while ensuring that comparisons are made among species co-occurring within the same regional ecological context. We have substantially revised the Methods to explicitly define subrealms, cite their origin, and clarify why this spatial stratification is appropriate for our study:

      “Accounting for geographic context through subrealm stratification

      To account for geographic heterogeneity in both species’ distributions and the baseline levels of urbanization, we stratified our analyses by global biogeographic subrealms (N=52; Fig. S1). Subrealms represent an intermediate hierarchical level within the One Earth [82] (https://www.oneearth.org/bioregions/) bioregionalization framework, grouping the 185 terrestrial bioregions into broader units that reflect shared species pools and ecological contexts while maintaining meaningful regional structure. This scale represents a practical compromise between analyzing data at the finer bioregion level (which would result in many regions with insufficient observations for robust analysis) and broader classifications such as continents or the 14 biogeographic realms, which aggregate ecologically distinct regions and species pools. This regionalization has been widely used in macroecological and biogeographic research to contextualize species–environment relationships because subrealms capture meaningful gradients in biotic assemblages that are not accounted for by climatic classifications alone [83,84].

      This stratification allows species’ associations with urban environments to be interpreted relative to the environments available within the regions they occupy. This is important, as previous work has shown that species’ responses to urbanization are constrained by biogeographic context, because regional species pools reflect shared evolutionary, ecological, and historical filters [23]. Previous work has also shown that urban associations among species are context-dependent, and interpreting species’ responses without accounting for regional baselines conflates availability of urban environments with species’ affinity to them. This distinction is critical because identical levels of urbanization (e.g., VIIRS radiance) can have different ecological meanings across regions with different species pools and land-use histories. It avoids conflating species’ urban affinity with global differences in urban availability.”

      We chose subrealms rather than Köppen climate classifications or continental units because our objective was not to partition species by climatic similarity per se, but to evaluate species’ associations with urban environments relative to the ecological and biogeographic contexts in which they occur. Climatic classifications such as Köppen are highly effective for addressing climate–species relationships, but they do not explicitly capture differences in species pools, evolutionary history, or land-use legacies that strongly shape how species interact with urbanization. Likewise, continents often aggregate ecologically disparate regions and species pools, potentially obscuring meaningful variation in baseline urbanization and species’ realized distributions.

      Importantly, urban affinity in our framework is a relative, context-dependent metric, explicitly interpreted within regions. Identical levels of urbanization (e.g., VIIRS radiance values) can have different ecological meanings across regions with distinct species pools, land-use histories, and settlement patterns. Stratifying analyses by subrealm therefore avoids conflating species’ affinity to urban environments with global or continental differences in the availability and intensity of urban land cover. We have clarified this distinction and motivation in the revised Methods (see responses below).

      Regarding the concern that requiring ≥100 observations per species per subrealm biases analyses toward common species: we agree that this threshold focuses the analysis on well-sampled species. This choice was intentional and follows previous work showing that such cutoffs are necessary to robustly characterize species’ responses to urbanization using occurrence data. While a global or continental analysis would indeed include additional, rarer species, it would also substantially increase uncertainty and conflate species’ responses across ecologically distinct contexts. Our study is therefore best interpreted as a macroecological synthesis of common species, which are also the taxa that disproportionately structure urban communities and drive the shape of Species Urbanness Distributions (SUDs). We now clarify this scope and limitation more explicitly in the introduction:

      “Our aim is to identify broad, cross-taxonomic patterns in species’ urban affinity at a global scale, rather than to resolve the specific causal mechanisms driving urban success or failure within individual taxa or cities.”.

      As well as in the discussion:

      “Our synthesis complements taxon-specific, presence–absence trait studies by identifying broad, cross-taxonomic patterns that can motivate and contextualize more mechanistic analyses [17,23].”

      Finally, while alternative spatial stratifications are possible, the central patterns we report particularly the skewed shape of SUDs—are robust to the use of regional context rather than absolute global metrics. Exploring how SUDs change under different spatial frameworks (e.g., continents, climate zones) is an interesting avenue for future work, but we feel is beyond the scope of the present study.

      (2) Methods - urban score

      The authors describe their "urban score" as being calculated as "the mean of the distribution of VIIRS values as a relative species specific measure of a response to urban land cover."

      I don't understand how this is a "relative species-specific measure". What is it relative to? Figures S4 and S5 show the mean distribution of VIIRS for various taxa, and this mean looks to be an absolute measure. Mean VIIRS for a given species would be fine and appropriate as an "urban score", but the authors then state in the next sentence: "this urban score represents the relative ranking of that species to other species in response to urban land cover".

      We agree that the wording in the original manuscript was unclear and conflated two distinct steps in the workflow. We have now revised the Methods to clearly distinguish between (i) the urban score, which is an absolute, descriptive summary of the mean VIIRS radiance associated with a species’ occurrence locations, and (ii) urban affinity, which is the relative, region-specific metric derived from the urban score. Specifically, we rewrote the methods to have distinct steps as subheadings, as follows: (1) urban score; (2) subrealms and why; (3) urban affinity. In the revised Methods, we explicitly define the urban score:

      “an absolute descriptive summary of the urbanization levels associated with a species’ occurrence locations within a given subrealm”.

      We no longer describe the urban score itself as “relative” or as a ranking among species. Relative comparisons among species arise only in the subsequent step, where species-specific urban scores are expressed relative to the regional background level of urbanization within each subrealm to derive urban affinity.

      We refer the Reviewer to the revised version which we feel is much clearer (lines 428-479)!

      That doesn't follow from the description of how this is calculated. Something is missing here. Please clarify and add an explicit equation for how the urban score is calculated because the text is unclear and confusing.

      The previous response, where we discuss the description, hopefully clarifies this. Further, we have revised the Methods to clearly define the urban score and to include an explicit equation. In the revised manuscript, the urban score for species s is calculated as the mean VIIRS radiance across all occurrence locations of that species:

      where n<sub>s</sub>is the number of GBIF occurrence records for species s, and L<sub>i</sub> is the VIIRS nighttime lights radiance value extracted at the location of occurrence i. We also clarify in the Methods that this urban score is an absolute summary statistic of observed urbanization at species occurrence locations

      (3) Methods - urban tolerance

      How the authors are defining and calculating tolerance is unclear, confusing, and flawed in my opinion.

      Tolerance is a common concept in ecology, evolution, and physiology, typically defined as the ability for an organism to maintain some measure of performance (e.g., fitness, growth, physiological homeostasis) in the presence versus absence of some stressor. As one example, in the herbivory literature, tolerance is often measured as the absolute or relative difference in fitness of plants that are damaged versus undamaged

      (e.g., https://academic.oup.com/evolut/article/62/9/2429/6853425?login=true).

      On line 309, after describing the calculation of urban scores across subrealms, they write: "Therefore, a species could be represented across multiple subrealms with differing measures of urban tolerance (Fig. S4). Importantly, this continuous metric of urban tolerance is a relative measure of a species' preference, or affinity, to urban areas: it should be interpreted only within each subrealm". This is problematic on several fronts. First, the authors never define what they mean by the term "tolerance". Second, they refer to urban tolerance throughout the paper, but don't describe the calculation until, where they write (text in [ ] is from the reviewer): "Within each subrealm, we further accounted for the potential of different levels of urbanization by scaling each species' urban score by subtracting the mean VIIRS of all observations in the subrealm (this value is hereafter referred to as urban tolerance). This 'urban tolerance' (Fig. S5) value can be negative - when species under-occupy urban areas [relative to the average across all species] suggesting they actively avoid them-or positive-when species over-occupy urban areas [relative to the average across all species] suggesting they prefer them (i.e., ranging from urban avoiders to urban exploiters, respectively). They are taking a relativized urban score and then subtracting the mean VIIRS of all observations across species in a subrealm. How exactly one interprets the magnitude isn't clear and they admit this metric is "not interpretative across subrealms".

      This is not a true measure of tolerance, at least not in the conventional sense of how tolerance is typically defined. The problem is that a species distribution isn't being compared to some metric of urbanness, but instead it is relative to other species' urban scores, where species may, on average, be highly urban or highly nonurban in their distribution, and this may vary from subrealm to subrealm. A measure of urban tolerance should be independent of how other species are responding, and should be interpretable across subrealms, continents, and the globe.

      We thank the reviewer for this careful and important critique. We agree that the term “tolerance” is commonly used to describe the ability of an organism to maintain performance (e.g., fitness, growth, physiological homeostasis) in the presence of a stressor, and that our metric does not measure tolerance in this mechanistic or fitness-based sense. To address this directly and unambiguously, we have revised the manuscript to explicitly define the term “urban affinity” as opposed to urban tolerance. 

      In the revised Methods, we also reorganized and clarified the calculation of urban affinity, introduced explicit notation, and provided a formal equation. Specifically, we now define urban affinity for species s in subrealm r as:

      where U<sub>s,r</sub>is the mean VIIRS radiance across all occurrence locations of species s within subrealm r, and Ū<sub>r</sub>is the mean VIIRS radiance across all occurrence records of all species in that subrealm. This transformation centers species’ urban scores on the regional background level of urbanization, yielding a relative measure of spatial association with urban environments.

      We agree with the reviewer that this metric is not interpretable as an absolute measure of affinity, and we now state this explicitly. Urban affinity values are, by construction, relative measures, interpretable only within subrealms, and they quantify whether a species tends to occur in more or less urbanized environments than is typical for that region. The magnitude of the metric therefore reflects deviation from the regional baseline, not a universal or global scale of urbanization, and is not intended to be compared directly across subrealms.

      We respectfully disagree, however, that this makes the metric flawed. Rather, it reflects a deliberate analytical choice aligned with our research questions. Our goal was not to estimate absolute urban exposure or physiological performance, but to compare species’ realized spatial associations with urban environments within shared biogeographic contexts. Because baseline urbanization levels, settlement history, and species pools vary strongly across regions, a globally absolute metric would conflate species’ affinities with regional availability of urban environments. By contrast, a relative, region-centered metric allows meaningful comparisons among species that coexist within the same ecological and biogeographic setting. This approach follows a growing body of macroecological work that infers species’ environmental affinities from spatial distributions rather than direct performance measures (e.g., Callaghan et al. 2020; 2021; 2023), and we now cite these studies explicitly.

      I propose the authors use one of two metrics of urban tolerance:

      (i) Absolute Urban Tolerance = Mean VIIRS of species_i - Mean VIIRS of city centers Here, the mean VIIRS of city centers could be taken from the center of multiple cities throughout a subrealm, across a continent, or across the world. Here, the units are in the original VIIRS units where 0 would correspond to species being centered on the most extreme urban habitats, and the most extreme negative values would correspond to species that occupy the most non-urban habitats (i.e., no artificial light at night). In essence, this measure of tolerance would quantify how far a species' distribution is shifted relative to the most highly urbanized habitat available.

      (ii) % Urban Tolerance = (Mean VIIRS of species_i - Mean VIIRS of city centers)/MeanVIIRS of city centers * 100%

      This metric provides a % change in species mean VIIRS distribution relative to the most urban habitats. This value could theoretically be negative or positive, but will typically be negative, with -100% being completely non-urban, and 0% being completely urban tolerant.

      Both of these metrics can be compared across the world, as it would provide either absolute (equation 1) or relative (equation 2) metrics of urban tolerance that are comparable and easily interpretable in any region.

      In summary, the definition of tolerance should be clear, the metric should be a true measure of tolerance that is comparable across regions, and an equation should be given.

      We thank the reviewer for this thoughtful and constructive suggestion, which raises an important conceptual issue regarding how “urban tolerance” should be defined and quantified. We agree that any such metric must be clearly defined, interpretable, and accompanied by an explicit equation, and we have revised the manuscript accordingly to clarify both our definition and its intended interpretation.

      The alternative metrics proposed by the reviewer anchoring species’ distributions to city centers or to the most highly urbanized habitats represent a valid and intuitive absolute framing of urban tolerance. Indeed, a closely related approach was explored and evaluated in Callaghan et al. (2020; https://doi.org/10.1016/j.ecolind.2020.106905), where species’ occurrence-based urbanness scores derived from VIIRS night-time lights were compared against abundance-based estimates of urban tolerance using explicit urban–non-urban contrasts. That study further demonstrated that urbanness scores depend on the choice of spatial baseline (e.g., regional buffers around cities versus continental extents), and showed that different baselines capture complementary, but not identical, aspects of species–urban associations.

      In the present study, we deliberately adopt a relative, regionally contextualized metric (now referred to as urban affinity), expressing each species’ mean VIIRS association relative to the background urbanization of the biogeographic subrealm in which it occurs. This choice reflects our goal of comparing species’ relative affinities to urban environments within shared ecological and biogeographic contexts. Importantly, identical VIIRS values can correspond to very different ecological conditions across regions, and anchoring all species to city centers or global urban maxima risks conflating species’ affinities with regional differences in urban availability and infrastructure.

      We now make this distinction explicit throughout the manuscript, including by (i) defining urban affinity as a relative, occurrence-based measure of urban affinity (rather than physiological or fitness-based tolerance), (ii) providing an explicit equation for its calculation, and (iii) clarifying that these values are interpretable within, but not across, biogeographic subrealms. We view absolute, city-center–anchored metrics and relative, regionally normalized metrics as complementary approaches, each suited to different questions; the latter is most appropriate for the macroecological, comparative analyses pursued here.

      (4) Figure 1: The figure does not stand alone. For example, what is the hypothesis for thermophily or the temperature-size rule? The authors should expand the legend slightly to make the hypotheses being illustrated clearer.

      We now expanded the legend so that the figure and hypotheses presented can be understood based on just the figure and its legend; we did so by explaining the illustrated hypotheses as requested by the Reviewer. The figure legend now reads as follows:

      “Fig. 1: Conceptual framework illustrating hypothesized mechanisms linking urban affinity to interspecific body-size shifts. These include dispersal and mobility constraints under habitat fragmentation [44,45], thermophily and the temperature–size rule driven by the urban heat island effect [15,30], size-biased competition and survival [94,95], and size-biased human preferences [64]. Urban fragmentation of habitat resources can select for increased mobility (e.g., larger butterflies) or reduced mobility (e.g., larger seeds) depending on isolation severity. Elevated urban temperatures favor thermophily, which often negatively correlates with size as it affects the heat balance via thermal inertia. Similarly, these higher temperatures generally favor smaller-bodied adult ectotherms because they accelerate development and reduce time available for growth (i.e., temperature-size rule). In plants, the increased CO<sub>₂</sub> and nutrient availability associated with anthropogenic environments due to heating- and traffic-related CO2 emissions and eutrophication provides a competitive advantage to larger plant species, and human preferences too may favor larger species (e.g., tree-lined streets), whereas smaller species may be advantaged in colonizing built infrastructure.”

      (5) SUDs: I don't agree with the conclusion given on line 83 ("pattern was consistent across subrealms and several taxonomic levels") or in the legend of Figure 2 ("there were consistent patterns for kingdoms, classes, and orders, as shown by generally similar density histograms shapes for each of these").

      The shapes of the curves are quite different, especially for the two Kingdoms and the different classes. I agree they are relatively consistent for the different taxonomic Orders of insects.

      We agree that our original wording overstated the similarity of distributions across taxa and regions. We have revised the text to clarify that the consistency we refer to pertains primarily to central tendencies rather than identical distributional shapes. To address this directly, we conducted additional analyses comparing urban affinity distributions across subrealms for taxonomic groups with the largest sample sizes. These results, now presented in new Supplementary Figures (Fig. S2-S4), show that while distributional shapes vary among higher taxonomic groups, median values and overall spread are broadly similar within comparable taxonomic levels. We have updated the Results text and the Figure 2 legend accordingly to reflect this more precise interpretation. 

      “These patterns in central tendency were broadly consistent across subrealms and taxonomic levels, although distributional shapes varied among higher taxonomic groups (Fig. 2).”

      “To evaluate this more formally, we compared distributions across subrealms for groups with the largest sample sizes and found that while distributional shapes varied among higher taxa, median values and overall spread were broadly similar within comparable taxonomic levels (Fig. S2–S4).”

      Figure 2 caption: “There were consistent patterns for kingdoms, classes, and orders (B) as shown by similar central tendencies despite variation in distributional shape.”

      We refer the Reviewer to the revised manuscript and supplementary material, but show the kindom level in Fig S2.

      More broadly, our goal in introducing Species Urbanness Distributions (SUDs) is not to argue that their exact shapes are invariant, but rather to provide a generalizable framework for describing how assemblages are structured along an urbanization gradient. In this respect, SUDs are conceptually analogous to Species Abundance Distributions (SADs), where the precise functional form has long been debated, yet the framework itself has proven extremely valuable for ecology. We therefore emphasize the utility of SUDs as a descriptive and comparative tool for quantifying community-level responses to urbanization, rather than as a claim about strict uniformity in distributional shape across taxa or regions.

      Reviewer #3 (Public review):

      Summary:

      This paper reports on an association between body size and the occurrence of species in cities, which is quantified using an 'urban score' that can be visualized as a 'Species Urbanness

      Distribution' for particular taxa. The authors use species records from the Global Biodiversity Information Facility (GBIF) and link the occurrence data to nighttime lighting quantified using satellite data (Visible Infrared Imaging Radiometer Suite-VIIRS). They link the urban score to body size data to find 'heterogeneous relationship between body size and urban tolerance across the tree'. The results are then discussed with reference to potential mechanisms that could possibly produce the observed effects (cf. Figure 1).

      We thank the reviewer for this clear and accurate summary of the study. We agree that the primary contribution of this work lies in the scale and taxonomic breadth of the analysis, and in introducing a framework (Species Urbanness Distributions) for quantifying species’ relative affinities to urban environments using globally available data. We have revised the manuscript to further clarify the scope of inference and the distinction between descriptive macroecological patterns and mechanistic explanations.

      Strengths:

      The novelty of this study lies in the huge number of species analyzed and the comparison of results among animal taxa, rather than in a thorough analysis of what traits allow species to persist under urban conditions. Such analyses have been done using a much more thorough approach that employs presence-absence data as well as a suite of traits by other studies, for example, in (Hahs et al. 2023, Neate-Clegg et al. 2023). The dataset that the authors produced would also be very valuable if these raw data were published, both the cleaned species records as well as the body sizes. The paper could strongly add to our understanding of what species occur in cities when the open questions are addressed.

      We appreciate highlighting the novelty of the taxonomic breadth and scale of our analysis. We agree that our approach is complementary to more detailed, taxon-specific trait studies based on presence–absence data. In response, we have further emphasized this distinction in the Discussion:

      “Our synthesis complements taxon-specific, presence–absence trait studies by identifying broad, cross-taxonomic patterns that can motivate and contextualize more mechanistic analyses17,23.”

      We also agree that the cleaned occurrence data and body size information represent a valuable resource, and all data will be made available, with the exception of some body size datasets which we are not able to make available.

      Weaknesses:

      I value the approach of the authors, but I think the paper needs to be revised.

      In my view, the authors could more carefully validate their approach. Currently, any weakness or biases in the approach are quickly explained away rather than carefully explored. This concerns particularly the use of presence-only data, but also the calculation of the urban score.

      The vast majority of data in GBIF is presence-only data. This produces a strong bias in the analysis presented in the paper. For some taxa, it is likely that occurrences within the city are overrepresented, and for other taxa, the opposite is true (cf. Sweet et al. 2022). I think the authors should try to address this problem.

      We thank the reviewer for raising this important point. We fully agree that GBIF occurrence data are subject to well-known sampling biases, including uneven geographic coverage, observer effort, and taxonomic focus. These limitations are now more explicitly acknowledged in the revised manuscript. At the same time, GBIF currently represents the only global biodiversity database that allows the scope of analysis undertaken here, spanning thousands of species across multiple taxonomic groups and regions. Systematic monitoring datasets that provide presence–absence data are typically restricted to particular taxa (often vertebrates or plants) and are geographically concentrated in the Global North, which would substantially limit the taxonomic and geographic breadth of our analysis.

      Importantly, our objective was not to estimate absolute species-specific responses to urbanization, but rather to examine relative patterns of urban affinity across species and families within comparable regional contexts. To address this, we structured our analyses at the subrealm level, which aggregates observations across large spatial extents and reduces sensitivity to fine-scale sampling biases associated with individual cities or urban–rural gradients. In addition, we restricted analyses to species with ≥100 observations per subrealm to focus on well-sampled taxa and reduce the influence of extremely sparse occurrence records. While these steps cannot fully eliminate sampling biases inherent to occurrence data, they substantially mitigate their influence when examining broad comparative patterns.

      Recent work has also evaluated the performance of GBIF data in urban biodiversity contexts. For example, Sweet et al. (2022) compared GBIF-derived species richness patterns with independent state-level biodiversity databases across cities and surrounding regions, finding that GBIF provided comparable or broader coverage across taxa and spatial extents. Their analysis showed that species richness was consistently higher in the surrounding region than in the city itself, suggesting that GBIF data capture broad urban–regional biodiversity gradients rather than systematically overrepresenting urban occurrences. Although our analysis differs in design, these results support the use of GBIF as a valuable resource for examining large-scale biodiversity patterns.

      More broadly, occurrence databases such as GBIF have become widely used for analyzing species–environment relationships at macroecological scales. While they may be insufficient for estimating precise species-specific environmental tolerances, they are informative for identifying broad patterns across taxa and regions. Our goal here is therefore to identify large-scale comparative patterns in urban affinity and generate hypotheses about trait– urbanization relationships, which can subsequently be tested with more structured monitoring datasets where available.

      Another important consideration is that our analyses focus on comparative differences among species within shared taxonomic and geographic contexts, rather than absolute estimates of urban affinity. Sampling biases in occurrence databases are often structured by observer behaviour (e.g., detectability, accessibility, or taxonomic interest), meaning that species recorded by similar observer communities are likely subject to similar sampling biases. Under these conditions, relative differences among species are expected to be preserved even when absolute occurrence frequencies are biased. This logic is consistent with the widely used target-group background approach in presence-only species distribution modelling, where species recorded by similar observer groups (often within the same taxonomic group) are used to control for shared sampling bias. Previous work by Callaghan et al. (2021; https://doi.org/10.1111/gcb.15670) performed additional validation analysis comparing our distribution-based urban affinity metric with estimates derived from occupancy modelling using well-sampled European butterflies (see Fig. S5 from the Callaghan et al. 2021 paper). The strong positive relationship between these approaches suggests that the broad patterns identified here are unlikely to arise solely from sampling artifacts.

      Finally, in the revised manuscript we now include additional comparisons among well-sampled taxonomic groups (see responses to other comments throughout our response document for details), which show substantial variation in urban affinity even among taxa with extensive sampling. These results suggest that the patterns reported here are unlikely to arise solely from sampling artifacts, but instead reflect meaningful ecological variation in how species interact with urban environments.

      The authors should compare their results to studies focusing on particular taxa where extensive trait-based analyses have already been performed, i.e., plants and birds. In fact, I strongly suggest that the authors should compare their results to previous studies on the relationship between traits, including body size and occurrences along a gradient of urbanisation, to draw conclusions about the validity of the approach used in the current study, which has a number of weaknesses.

      We agree that explicitly situating our findings within the existing trait-based urban ecology literature strengthens both interpretation and validation of our approach. We had already referenced several relevant studies (e.g., Hahs et al. 2023 and others) in the Introduction and Discussion, but we recognize that these comparisons were not sufficiently explicit. We have now added text to the Discussion directly comparing our results with previous trait-based studies across taxa:

      “Our results are broadly consistent with prior taxon-specific trait-based studies (eg., Hahs et al.[17]), but also highlight that relationships between body size and urbanization vary across taxa and analytical frameworks. For example, global syntheses and regional studies have reported positive, negative, or null size–urbanization relationships depending on clade and spatial scale. A recent global analysis that compiled empirical occurrence data for multiple terrestrial faunal taxa across cities worldwide reported broadly similar body-size responses to urbanization [17]. For four of the five groups that overlap with our analysis—amphibians, bats, bees, and birds—the direction of the body-size relationship with urbanization was consistent between studies. The only exception was carabid beetles, which tended to be smaller-bodied in highly urbanized environments in that analysis, whereas we detected no significant size effect for this family. Studies on birds, for example, have found mixed results, including positive associations to urbanization in some regional assemblages [45], no global relationship in others [46] or an overall negative relationship globally [23], and negative relationships in particular clades such as raptors [40]. Such discrepancies likely arise because different studies quantify urbanization differently, focus on different spatial grains, or analyze different components of species responses (e.g., presence– absence, abundance, or occurrence distributions). Additionally, a study on multiple taxa including butterflies and moths found a positive relationship in butterfly and moth community-weighed mean body size with increases in urbanization level, similar to our findings [31]. Researchers have also found that smaller-bodied dung-associated beetles potentially benefit from urban environments, which is similar to the negative association we found between urbanization and body size in beetles [47]. Our approach complements these studies by estimating occurrence-based urban associations across thousands of taxa simultaneously, allowing comparison of how consistently body size predicts urban affinity across taxonomic groupings rather than within a single lineage. In this sense, variation among published results does not contradict our findings but instead reinforces the conclusion that body size is a context-dependent filter whose direction and strength depend on ecological setting, taxonomic scope, and the urbanization metric used.”

      These additions highlight that published relationships between body size and urbanization vary widely across taxa, spatial scales, and analytical approaches. For example, prior studies have reported positive, negative, or null size–urbanization relationships depending on clade, geographic extent, and how urbanization or occurrence is quantified. Even within birds alone, the literature spans positive regional relationships, null global relationships, and negative relationships in particular clades such as raptors. We now explicitly discuss these contrasts and clarify that such discrepancies are expected because different studies measure different components of species’ responses (e.g., presence–absence vs. abundance vs. occurrence distributions), use different spatial grains, or focus on different taxonomic subsets.

      We emphasize that our analysis is not intended to replace taxon-specific trait studies, but rather to complement them by providing a macroecological synthesis across thousands of species simultaneously. Importantly, the heterogeneity we observe among families is itself a key biological result, indicating that body size is not a universal predictor of urban affinity but instead a context-dependent filter whose direction and strength vary across ecological and phylogenetic settings. We now state this interpretation more clearly in the revised manuscript.

      They should be be more careful in coming up with post-hoc explanations of why the pattern found in this study makes sense or suggests a particular mechanism. This reviewer considers that there is no way in which the current study can disentangle the different possible mechanisms without further analyses and data, so I would suggest pointing out carefully how the mechanisms could be studied.

      We agree that our study cannot disentangle the causal mechanisms underlying species’ responses to urbanization. Our intent in discussing potential mechanisms was not to claim definitive explanations, but rather to situate our findings within existing ecological theory and to highlight plausible, non-exclusive pathways that may generate the observed patterns. To make this clearer, we have revised the Discussion to explicitly frame these interpretations as hypotheses rather than conclusions, and to emphasize that testing the underlying mechanisms will require additional data and approaches, such as targeted trait datasets, experimental manipulations, and longitudinal or within-city studies:

      “Because our synthesis is correlative and macroecological in nature, the mechanisms discussed above are best viewed as hypotheses that can be evaluated through future work combining experimental, trait-based, and longitudinal data.”.

      Additionally, we modified our overall goal to make it clear that this is not inherently a mechanistic study per se:

      “Our aim is to identify broad, cross-taxonomic patterns in species’ urban affinity at a global scale, rather than to resolve the specific causal mechanisms driving urban success or failure within individual taxa or cities.”.

      More details should be given about the methodology. The readers should be able to understand the methods without having to read a number of other papers.

      We have substantially revised and expanded the Methods section to ensure that all analytical steps can be understood directly from the manuscript without requiring consultation of prior publications. In particular, we now (i) provide a clear conceptual roadmap of the workflow at the start of the Methods, (ii) define all key metrics explicitly, including equations for both the urban score and urban affinity, and (iii) clarify the interpretation, assumptions, and limitations of each step. We also added text explaining the rationale for subrealm stratification and the intended interpretation of relative values. Together, these revisions make the methodological framework fully transparent and self-contained (see revised Methods and related responses above and below).

      References:

      Hahs, A. K., B. Fournier, M. F. Aronson, C. H. Nilon, A. Herrera-Montes, A. B. Salisbury, C. G. Threlfall, C. C. Rega-Brodsky, C. A. Lepczyk, and F. A. La Sorte. 2023. Urbanisation generates multiple trait syndromes for terrestrial animal taxa worldwide. Nature Communications 14:4751.

      Neate-Clegg, M. H. C., B. A. Tonelli, C. Youngflesh, J. X. Wu, G. A. Montgomery, Ç. H. Şekercioğlu, and M. W. Tingley. 2023. Traits shaping urban tolerance in birds differ around the world. Current Biology 33:1677-1688.

      Sweet, F. S. T., B. Apfelbeck, M. Hanusch, C. Garland Monteagudo, and W. W. Weisser. 2022. Data from public and governmental databases show that a large proportion of the regional animal species pool occur in cities in Germany. Journal of Urban Ecology 8:juac002.

      We have incorporated these (and additional new references) into our revised manuscript.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you see from the general comments above and the specific recommendations below, the reviewers are impressed by your comprehensive data set and the analytic approach. However, they ask you to clarify your measures of organism size, occurrence data (vs. presence/absence and corresponding sample-bias caveats), urbanness (lighting differences between cities and regions?), urban tolerance (measure should not be relative to other species and particular regions), and region ("subrealm" vs. more commonly used defintions of world regions such as continents). They also encourage you to compare your general results with more detailed local studies to better justify using size as the only, easily available trait.

      We thank the Editor for this clear synthesis of the key priorities for revision. We have carefully addressed each point and substantially revised the manuscript to improve clarity, methodological transparency, and interpretability. In particular:

      We clarified how body size data were compiled, harmonized, and modeled, including explicit description of how different measurement types (mean, maximum, sex-specific) were retained and statistically accounted for through scaling and hierarchical modeling. We now state these procedures explicitly in the Methods.

      We expanded the Methods and Discussion to clarify that our analyses rely on occurrence data rather than presence–absence or abundance data, and we now explicitly discuss the implications and limitations of presence-only datasets, including potential sampling biases and how these may influence inference.

      We strengthened justification for using VIIRS night-time lights as a continuous proxy for urbanization, added supporting citations, and clarified that spatial heterogeneity in lighting primarily introduces additional variance rather than systematic bias. We also explicitly describe how urbanization values were calculated and interpreted.

      We substantially revised the manuscript to clearly define urban affinity at the outset (including in the Abstract), distinguish it from physiological definitions of tolerance, and provide explicit equations and step-by-step descriptions of how both urban score and urban affinity are calculated and interpreted. We now emphasize that the metric is a relative, region-contextualized measure of occurrence-based urban affinity.

      We added full justification, citations, and methodological explanation for the use of biogeographic subrealms, clarified how they differ from continents or climate zones, and explained why this stratification is appropriate for the ecological questions addressed. We also clarified the scope of inference and limitations of this approach.

      We expanded the Discussion to explicitly compare our results with prior trait-based urban ecology studies across taxa (including birds and other groups), highlighting where results converge, diverge, and why such variation is expected across spatial scales, taxa, and analytical frameworks.

      Reviewer #1 (Recommendations for authors):

      (1) Abstract

      (a) Please define how tolerance is being used here

      We now use affinity throughout and it is defined in various places (see responses to other comments here).

      (b) The abstract should clarify at what taxonomic scale body size is assessed. It is unclear in the abstract as to whether the reader expects intraspecific measures and interspecific, and at what resolution.

      We have revised the abstract by adding one sentence explicitly stating the scale body size was assessed:

      “We then assessed whether body size, an integrative ecological trait fundamental to space use, mobility, metabolism, and environmental sensitivity, showed consistent associations with urban affinity among species and across 371 taxonomic families. Analyses were conducted at the interspecific level and focused primarily on variation among taxonomic families (provided with this paper is an accompanying application to view results).”

      (2) Results/Discussion

      (a) The species urbanness distribution and comparison with the species abundance distribution is an interesting and conceptually useful contribution to urban ecology and underscores how urbanization functions on biodiversity at scale.

      We thank the reviewer for this positive assessment and are encouraged that they view the Species Urbanness Distribution (SUD) as a conceptually useful contribution to urban ecology. We see SUDs as a flexible framework that can be extended in several important directions, including comparisons across additional traits, cities of differing size and configuration, and temporal analyses that track how urbanness distributions shift with ongoing urban expansion or restoration. More broadly, we hope that SUDs can provide a framework to think about a macroecological understanding of how urbanization filters biodiversity.

      (b) In our Lambert et al. (2023) study that you reference, we suggest that 'exaptation' may be valuable to explore in urban areas. Although body size wasn't the trait we were considering at that time, it may be worth putting your discussion around pre-adaptation in this context.

      We agree that exaptation provides a valuable conceptual lens for interpreting species’ responses to urban environments. We have revised the Discussion to explicitly frame species’ urban success in this context:

      “Such traits “pre-adapted” to urban conditions allow for some species to not only persist but thrive in urban environments where most species cannot. Framing these patterns through the lens of exaptation may be particularly useful, as traits that evolved under non-urban selective pressures may incidentally confer advantages in urban environments without having arisen in response to urbanization per se (sensu Lambert et al.[4]). We therefore speculate that the skewed shape of SUDs may reflect the uneven distribution of exaptive traits across species pools, rather than widespread adaptive evolution to urban conditions. 

      Consistent with this interpretation, if exaptive traits that facilitate urban persistence are unevenly distributed across species pools, most species would be expected to exhibit avoidance rather than affinity of urban environments. Indeed, we found that the median urban affinity is most often below one, indicating widespread avoidance among species.”.

      (c) Given the family-scale effect, it would be helpful to discuss how often species within a family co-occur in a given geographic region, how much other traits covary with size, etc. Do we have an a priori reason to expect family to be the taxonomic resolution at which body size seems to be most varied?

      Our exploratory and preliminary analyses revealed that variation in the body size– urban affinity relationship was strongest at the family level, which prompted us to focus our main analyses at this taxonomic resolution. (But we also present results on order as well). Families represent a biologically meaningful intermediate scale in taxonomy: species within families typically share broad morphological, ecological, and life-history characteristics, yet still exhibit substantial variation in body size and ecological strategies. Indeed, body size is well known to covary with multiple traits—including dispersal ability, metabolism, and space use—making it an integrative trait that captures several ecological dimensions simultaneously within and among families. These correlated traits likely contribute to the heterogeneous responses to urbanization observed among families.

      Using the family level also provides a practical balance between biological relevance and statistical robustness. Many families contain sufficient numbers of species to allow independent model estimation while avoiding the strong data imbalance that would arise at higher taxonomic levels. In addition, family is a commonly used unit in macroecological trait analyses (e.g., Roy et al. 2009; Smith et al. 2004), and it often reflects major morphological and ecological similarities among species, as reflected in taxonomic identification frameworks.

      Regarding co-occurrence, our analytical framework already accounts for geographic context by estimating urban affinity within subrealms. This ensures that species are compared within the same regional species pools and environmental contexts, rather than across globally disparate assemblages. Consequently, family-level effects emerge from comparisons among species that co-occur within shared biogeographic settings rather than from global taxonomic aggregation.

      We have added a short clarification in the manuscript to emphasize that body size functions as an integrative trait that covaries with multiple ecological attributes, and that family-level analyses represent a balance between ecological interpretability and data availability:

      “Because body size covaries with multiple ecological traits (e.g., dispersal ability and metabolic rate), we focused on family-level analyses to capture shared ecological strategies while still allowing sufficient variation among species to detect trait– environment relationships [39]”.

      (d) The result that body size shows a stronger effect in plants perhaps could suggest that plant records in GBIF are more sensitive to potential collection bias, perhaps due to detectability differences or preferences for where botanists and citizen scientists collect plant data? You mention ornamental plants late, but it may be worth discussing this here, too.

      We agree that this is a possible mechanism, which likely conflates detectability and ecological signal. We have expanded this point in the discusssion to better address this:

      “These human-driven preferences may also influence detectability and recording effort, as larger and more conspicuous plant species are more likely to be planted, maintained, and documented in urban environments, and thus be available in GBIF for our analyses. However, we suggest that this is not purely a sampling artifact, but such processes likely interact with ecological filtering to shape the realized size structure of urban plant communities.”.

      (e) I appreciate the additional taxonomic layering to the discussion. Seeing patterns at the family and order levels is helpful for generating new theory and predictions about how urbanization structures biodiversity at different taxonomic scales.

      We agree that examining patterns across multiple taxonomic scales is particularly valuable for generating testable hypotheses about how urbanization structures biodiversity, as different mechanisms may emerge or break down depending on the resolution of analysis. We hope this multi-scale perspective helps stimulate new theory and predictions about the ecological processes shaping urban biodiversity across the tree of life.

      (3) Methods

      (a) The methodology provides a scalable, consistent, and reasonable measure of both urbanness and species-level urban tolerance. The urban tolerance measure will, of course, not be useful for certain types of research (e.g., animal behavior), but it is appropriate for the resolution of this study.

      We agree that the urban affinity metric presented here is intended for broad-scale, comparative analyses and is not designed to capture fine-scale processes such as individual behavior or short-term demographic responses. Our goal was to develop a scalable and consistent measure that enables cross-taxon and cross-region comparisons at a global extent, which we believe is appropriate for addressing the questions posed in this study. We have sought to be explicit about this scope throughout the manuscript (e.g., to better alleviate Reviewer #1 concerns) and emphasize that the framework is complementary to, rather than a replacement for, more mechanistic or organism-focused approaches.

      (b) I'm concerned that the authors were not able to constrain their dataset to mean, median, or maximum, not potentially sex variability in sizes. Later in the methods, the authors state that they selected the measure of size that was most common within a family. Does this mean that species within a given family that didn't have that measure of body size were removed from the analysis?

      We appreciate this important point and agree that heterogeneity in how body size is measured (e.g., mean, maximum, or sex-specific estimates) is a real and unavoidable challenge in large-scale trait syntheses. Our analytical approach was explicitly designed to minimize the influence of this heterogeneity while retaining as many species as possible, rather than excluding species based on inconsistent trait metadata.

      Specifically, species within a family were not removed based on the availability of a particular body size definition. All species with at least one body size estimate were retained. When multiple measures existed for a species, we selected the measurement type that was most commonly available within each family to maximize comparability while preserving sample size. Remaining heterogeneity among measurement types (including units, measurement detail, and whether values reflected means, maxima, or sex-specific estimates) was explicitly accounted for through log-transformation and metadata-aware centering and scaling, with measurement metadata included as random intercepts in the hierarchical models. We have clarified this point in the Methods:

      “Importantly, this procedure did not result in the exclusion of species lacking a particular body size definition; rather, all species with at least one available body size estimate were retained, with measurement heterogeneity explicitly accounted for through metadata-aware scaling and hierarchical modeling.”

      In addition, our taxonomic modeling strategy was intentionally hierarchical. Species belonging to families that did not meet the minimum threshold for family-level modeling (≥10 species) were not discarded; rather, they were included in higher-level taxonomic analyses (e.g., order- or class-level models), ensuring that available information was retained wherever statistically appropriate. This approach reflects our broader goal of maximizing data inclusion while matching inference to the resolution supported by the data.

      Reviewer #2 (Recommendations for the authors):

      (1) Overlap between VIIRS and GBIF data: While it would have been nice for the GBIF records and VIIRS timescales to match, the degree of mismatch isn't overly large (2010-2021 vs 2015-2021), and any bias or inaccuracies should be minimal. I am mainly making this comment as a potential counterpoint to a possible criticism from other reviewers.

      We thank the reviewer for this helpful observation and agree with their assessment. While the temporal coverage of GBIF occurrence records (2010–2021) and VIIRS night-time lights data (2015–2021) does not perfectly overlap, the mismatch is relatively small and unlikely to introduce substantial bias, particularly given our focus on broad, global patterns of urban affinity rather than fine-scale temporal dynamics. We appreciate the reviewer highlighting this point as a potential counterargument to concerns about temporal alignment.

      (2) Line 87: "only a select few species seem to possess traits that enable them to thrive in urban...".

      This seems like an odd statement, given how many of these species have positive urban tolerance measures.

      Agreed that this was oddly worded. We have revised for clarity, focusing on the magnitude of urban affinity:

      “Similarly, much like the skewed distributions observed in SADs [24,26], the skewed shape of SUDs indicates that while many species exhibit some degree of urban affinity, a relatively small subset of species attain high levels of urban affinity and dominate urban environments.”

      (3) Line 81: "skewed shape of SUDs suggests that traits enabling species to tolerate urban environments are both rare and specific".

      Again, based on the shape of some of these curves, I'm not convinced that it is rare, and there is nothing about these curves that suggests it is something "specific". Indeed, urban tolerance could be very multivariate, and the authors' own results suggest this is indeed the case.

      We have revised the sentence to retain a focus on traits while avoiding overinterpretation of adaptation from the distributional patterns alone. The revised wording emphasizes the uneven expression of high urban affinity across species without implying rarity or trait specificity:

      “The skewed shape of SUDs suggests that traits enabling species to tolerate urban environments are unevenly expressed, given that only a handful of species show extreme urban affinity values, but our results suggest this is geographically widespread across taxa.”.

      We also agree with the likelihood that it is multivariate, and return to this in the conclusion in a stronger sense:

      “Although body size emerged as a predictor of urban affinity, we found not only substantial heterogeneity across families and orders, but also that body size filtering alone is unlikely to explain the consistently skewed SUD shape. Taken together, these patterns suggest that urban affinity likely emerges from multiple trait combinations rather than a single, universally advantageous trait, and that strong affinity to urban environments is not uniformly expressed across taxa, despite occurring broadly across regions.”.

      (4) Line 100: "UHI", avoid abbreviations unless absolutely necessary.

      We have removed this abbreviation throughout.

      (5) Body size: focusing on one trait seems like a shot in the dark, and so it isn't too surprising that this didn't reveal a strong or consistent pattern. However, I also recognize that collecting consistent trait data across so many taxa is challenging, and size is a low-hanging fruit that correlates with multiple traits. Perhaps discuss more the range of traits you think are most likely to predict urban tolerance.

      Body size is indeed the ‘easiest’ to collect, but we acknowledge that there are other traits which could be important, and body size correlates with multiple traits. We revised our discussion to be more comprehensive to discuss some of the additional traits, and be explicit about the shortfalls of body size:

      “Ultimately, the heterogeneous and sometimes weak relationships between body size and urban affinity suggests that body size alone cannot explain the emergence of extreme urban exploiters and the skewed shape of SUDs. Focusing on body size as a focal trait necessarily represents a simplification of the multidimensional processes underlying species’ responses to urbanization, driven in part by data availability when conducting a taxonomically-broad synthesis. Instead, urban affinity likely depends on multivariate trait combinations [17,58] that vary among taxa [59] and ecological contexts [60]. Traits that are likely to correlate with urban affinity include dispersal capacity, behavioral flexibility, diet breadth, reproductive strategy, thermoregulatory ability, and, in plants, life history traits such as growth form, clonality, phenology, and seed size. The diversity of trait pathways through which species may persist or thrive in urban environments is consistent with the pronounced taxonomic heterogeneity we observe and helps explain why body size alone does not yield a universal pattern.”

      (6) Figure S2: This figure and analysis appear to 'come out of nowhere'. I think this is distracting and tangential, and it should be removed. I have the same thoughts about Figure S3. While I do think a discussion of other traits to measure is well warranted and needed, the inclusion of "preliminary' results that aren't motivated by clear questions, appropriate context, and rigorous analysis should be discouraged.

      We have removed Figure S2 and Figure S3 in response to this comment.

      I hope the authors find my constructive comments useful in their revision process.

      This was a very thorough and thoughtful review. We are greatly appreciative of the opportunity and guidance to improve our work!

      Reviewer #3 (Recommendations for the authors):

      Here is a list of a number of further points that the authors may want to address:

      (1) Figure 1 somehow misses the fact that humans simply do not want very large animals in the city. We kill large predators if they come too close to cities, and the same for large herbivores such as wild boar or deer.

      We agree that direct human persecution and management of large-bodied species can influence which species occur in urban environments, particularly for large predators and herbivores. Such processes represent important mechanisms shaping urban species assemblages and represent an entire field of socio-ecological dynamics. We have now clarified this point in the Discussion by noting that human–wildlife conflict, management, and persecution could contribute to observed size–urbanization relationships for some taxa, and that disentangling these mechanisms represents an important direction for future research. We added some text to highlight this point):

      “Similarly, human–wildlife conflict and active management of large-bodied animals in cities may influence which species persist in urban environments, potentially constraining the upper end of the body size distribution. Taken together, these examples illustrate the importance of considering the socio-ecological context of urban species assemblages [65]”.

      (2) Line 270. So you removed all data from the grid-based survey?

      We did not remove all data originating from grid-based surveys or gridded products. Rather, we retained GBIF point-occurrence records and applied a standard spatial filtering step, removing only those individual observations with reported coordinate uncertainty greater than 1 km. This was done to ensure reliable alignment between species occurrence points and remotely sensed environmental layers. We have clarified this distinction in the Methods to avoid confusion:

      “Due to uncertainty in matching observations with remotely-sensed products, any GBIF observation with a coordinate uncertainty > 1 km was removed. This filtering step removed individual observations with high spatial uncertainty, rather than excluding entire datasets or survey types.”.

      (3) Line 278. Human population density?

      Yes, we have added ‘human’ here (and elsewhere in this section) to make this clearer to the reader.

      (4) Line 284. What is a pixel?

      We have modified the text to make this clearer:

      “VIIRS Stray Light Corrected Nighttime Day/Night Band Composites product, representing monthly composites, (i.e., this dataset in Google Earth Engine: NOAA/VIIRS/DNB/MONTHLY_V1/VCMSLCFG) with a native resolution of ~500 m<sup>2</sup>. We took the median of all monthly composites for each pixel (i.e., a single grid cell of the night-time lights raster representing a fixed ground area) to calculate a pixel-level urbanization value, measured in average radiance, and used imagery from January 2015 to January 2021 to calculate this median”.

      (5) Line 292. It seems to me that lighting is different in different types of cities with the same level of impervious surface, depending on local customs of how many lights are installed, left switched on, etc. I guess that petrol stations and strongly lit industrial areas both produce high levels of light, while for the industrial areas, there could be lawn or other vegetation?

      We thank the reviewer for this thoughtful observation and agree that night-time lighting can vary across cities with similar levels of impervious surface due to differences in land use, infrastructure, and cultural lighting practices. We do not interpret VIIRS night-time lights as a direct measure of any single urban feature, but rather as a continuous, integrative proxy for urbanization that captures the combined footprint of human activity, infrastructure intensity, and energy use. VIIRS radiance has been repeatedly shown to correlate strongly with human population density, built infrastructure, and urban extent, while being negatively correlated with vegetation cover (e.g., EVI). It is repeatedly used in remote sensing and urban sustainability literature. This approach is widely supported in the literature, for example:

      Panić et al. used night-time lights were to map spatial and temporal patterns of artificial lighting as a proxy for human population distribution and activity, distinguishing areas of urban and rural occupancy.

      (https://www.ceeol.com/search/article-detail?id=1035395)

      Zhou et al. used night-time light observations were to develop a globally consistent time series of annual urban extent, delineating urban clusters and quantifying global urban growth over decades. (https://doi.org/10.1016/j.rse.2018.10.015)

      Chakraborty & Stokes used night-time light time series with machine learning to detect and quantify urban change processes—identifying deviations from expected radiance trends to monitor diverse urban transitions.

      (https://doi.org/10.1016/j.rse.2023.113818)

      Zhao et al. reviewed night-time light remote sensing was for its broad capacity to quantify human activities and socioeconomic dynamics—such as urbanization, economic change, and environmental impacts—across scales.

      (https://doi.org/10.3390/rs11171971)

      Zheng et al. used VIIRS nightime lights across 30 global megacities to produce a classification scheme to disentangle urban land changes into five categories, and assess global urbanization processes. (https://doi.org/10.1016/j.isprsjprs.2021.01.002)

      Zhao et al. argue that nighttime lights provide a consistent dataset to model and interpret urbanization dynamics and use this to track urban dynamics in Southeast Asia. (https://doi.org/10.1016/j.rse.2020.111980)

      While localized mismatches may occur (e.g., brightly lit industrial areas with surrounding vegetation), such heterogeneity is expected to introduce additional variance rather than systematic bias in the measure of urbanization, making our inference conservative. We have clarified this interpretation and added additional supporting references in the Methods:

      “Previous work has shown that VIIRS night-time lights is negatively correlated with greenness measured through the Enhanced Vegetation Index (EVI) and positively correlated with human population density [69,71]. Although night-time light intensity can vary among cities with similar impervious surface due to differences in land use, infrastructure, and cultural lighting practices, at broad spatial scales it functions as an integrative proxy of urbanization [75,76,77,78,79,80], with localized heterogeneity contributing primarily to additional variance rather than systematic bias.”

      (6) Line 295. How did you reconcile the spatial uncertainty of >1km with an urbanization pixel of 150m2? For how many species did you have a higher uncertainty than pixel size? In my experience, your ca. 39m accuracy is a strong assumption for GBIF data.

      We would like to clarify that we do not assume species occurrence accuracy at the scale of the geohash blocks (i.e., tens of meters), and we do not interpret GBIF records as having ca. 39 m positional accuracy. The use of geohash7 (~150 m blocks) reflects a computational indexing choice, not an assumption about biological or observational precision. All GBIF observations with reported coordinate uncertainty greater than 1 km were removed prior to analysis, ensuring that retained occurrences were compatible with the effective spatial resolution of the remotely sensed urbanization data. Importantly, the effective spatial resolution of our urbanization metric remains that of the VIIRS night-time lights product (~500 m). Geohash encoding at a finer resolution was used solely to efficiently associate point occurrences with the appropriate VIIRS pixel while avoiding redundant extraction or averaging across adjacent pixels. This approach does not increase the effective spatial precision of the analysis, nor does it imply sub-pixel inference. We have clarified this in the Methods:

      “The VIIRS night-time lights data, with a native resolution of ~500 m<sup>2</sup>, was then matched to these blocks by assigning each geohash7 block the average VIIRS radiance value that intersects it. We do not assume positional accuracy at the scale of the geohash blocks, but geohash encoding was used solely for computational indexing, while the effective spatial resolution of the urbanization metric is that of the VIIRS data (~500 m). This approach allows us to avoid unnecessary redundancy in the data while maintaining the original VIIRS resolution”.

      (7) Line 296. Why this high resolution in the species data when your light data is 500m2?

      The apparent mismatch in resolution reflects a distinction between data handling resolution and analytical resolution. Species occurrence records were retained at their native point-level precision to avoid premature spatial aggregation and to ensure that each observation could be accurately matched to the appropriate VIIRS night-time lights pixel. The finer-resolution geohash encoding does not imply that species data were analyzed at that scale, nor does it increase the effective spatial resolution of the analysis. We note, however, that the reported spatial uncertainty of some GBIF records may approach or exceed the resolution of the VIIRS data. Retaining such records represents a deliberate trade-off between spatial precision and data coverage, and is necessary to maximize taxonomic and geographic representation in a global analysis of this scope. Importantly, any residual spatial uncertainty is expected to introduce additional noise rather than systematic bias, making our estimates of species–urban affinity relationships conservative.

      (8) If you could show how your results match the results of Hahs et al and others with respect to occurrence and traits, this would strengthen your approach.

      We agree that explicitly comparing our findings with prior trait-based studies strengthens the interpretability of our approach. We have now added text to the Discussion that directly compares our results with published analyses, including Hahs et al. (2023) and other taxon-specific studies. In particular, we highlight where our occurrencebased estimates recover similar body size–urbanization relationships (four of five taxa in Hahs et al.) and where they differ (e.g., carabids), and we discuss how such differences likely arise from variation in spatial grain, response variables, and definitions of urbanization. These additions clarify how our framework aligns with, complements, and extends existing trait-based work rather than replacing it.

      (9) I wonder whether you could run your analysis with simplified data. In the end, you do not talk much about how high the urban score is, so you may also aggregate values to "highly lighted", "lighted", "some light" and "dark" and re-do the analysis, after checking how these scores correlate with e.g. impervious surface in a slightly larger area than what you used (maybe 50x50m).

      Our analytical framework—and the concept of Species Urbanness Distributions (SUDs) in particular—relies on retaining the continuous nature of the underlying urbanization metric. Discretizing night-time light values would necessarily introduce arbitrary thresholds, reduce information content, and obscure subtle but ecologically meaningful variation in species’ relative affinities to urban environments. Because we focus on relative affinity patterns rather than absolute urbanization classes, maintaining a continuous metric is central to both our methodological approach and conceptual contribution. That said, we agree that exploring how continuous urban affinity scores relate to categorical urban classes or alternative urbanization proxies (e.g., impervious surface at different spatial grains) represents a valuable direction for future work. Such analyses could be particularly informative for translating continuous affinity metrics into applied conservation or urban planning contexts.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study investigates how the brain categorizes written words from different writing systems (e.g., alphabetic vs. non-alphabetic), shedding potential light on the neural basis of language's social‑categorization function. Overall, the evidence supporting the authors' claims is solid, though some analyses and key interpretations would benefit from fuller justification.

      Thank you for handling our manuscript! We’ve modified the manuscript according to the reviewers’ comments and suggestions.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study demonstrates, through a series of EEG and MEG experiments, that the human brain automatically categorizes words from alphabetic and non-alphabetic languages, and it unpacks the neural mechanisms of this process from multiple angles. The work examines not only univariate repetition-suppression (RS) effects, but also how repeating or alternating languages influences the representational similarity of words within and across language categories.

      Strengths:

      The univariate RS effects across multiple experiments lend support to some of the main conclusions

      Weaknesses:

      I have reservations about the logic underlying the multivariate analyses, and I believe the implications of the control experiments merit fuller discussion.

      (1) Question 1: Logic of the multivariate analyses

      The original text states:

      "The processing of intra-language similarity was quantified as correlation distances between neural responses to two words of the same language, which occurred more frequently and would be inhibited in the Rep-Cond (vs. Alt-Cond) due to habituation (Fig. 1c)...".

      I argue that this passage conflates two levels. Building a representational dissimilarity matrix (RDM) is a data-analysis step; it cannot be equated with a cognitive computation. Hence, there is no sense in which this computation occurs "more frequently" in one condition. RDM construction rests on the pairwise similarity of activity patterns, so even if a task engaged no cognitive computation of representational similarity, we could still compute an RDM. Conversely, if a task factor alters the RDM, we must explain how that factor changes the underlying neural patterns, not claim that it triggers specific cognitive processing. Therefore, I neither understand what "more frequent processing" the authors refer to, nor accept their account of the multivariate results.

      The multivariate result pattern, briefly, is that distances between words, both within and across languages, are larger under the repetition condition. One plausible interpretation is that a word representation comprises two parts: language-type (alphabetic vs. non-alphabetic) and fine-grained identity features (visual shape, orthography, semantics, phonology, etc.). Repetition of language type may, via RS, reduce the weight of the first component, thereby increasing the relative contribution of fine-grained features and amplifying inter-word differences. This could explain the multivariate findings.

      Thank you for these insightful comments regarding the logic of the multivariate analyses. In the revision, we’ve elaborated the rationale underlying our experimental design. Specifically, we’ve explained why the processing of intra-language similarity is expected to occur more frequently in the repetition condition (Rep-Cond) than in the alternation condition (Alt-Cond) whereas the reverse is true for the processing of inter-language difference. Importantly, we’ve clarified that the processing of intra-language similarity was assessed rather than defined by conducting the multivariate analyses. The multivariate analyses were conducted to assess correlation distances between neural responses to pairs of words, either within the same language or across different languages. We explained what smaller intra-language correlation distances and larger inter-language correlation distances mean for language-base categorization of words (see Page 7-8).

      We appreciate the alternative account of the observed neural repetition suppression (RS) effects in terms of language-type versus fine-grained identity (visual shape, orthography, semantics, phonology, etc.) feature processing. We included a paragraph in the revised Discussion to discuss how possible the early neural RS effect can be attributed to the processing of the fine-grained identity features of visual words. This discussion allowed us to clarify that the early neural RS effects related to visual words of familiar and unfamiliar languages highlight the early spontaneous language-based categorization as a unique process of visual words of alphabetic and non-alphabetic languages. However, our results do not exclude the possibility that the processing of the linguistic properties of visual words may contribute to the long-latency RS effect (see Page 37-38).

      Page 7-8

      “The processing of intra-language similarity occurs when two words of the same language are perceived repeatedly with short interstimulus intervals. Because words of the same language were repeatedly presented in the Rep-Cond and words of two different languages were displayed in the Alt-Cond, the processing of intra-language similarity occurred more frequently and would be inhibited in the Rep-Cond (vs. Alt-Cond) due to habituation (Fig. 1c). By contrast, the processing of inter-language difference takes place when two words of different languages are perceived with short interstimulus intervals. Since words of different languages appeared more frequently in the Alt-Cond (vs. Rep-Cond), we would expect RS of the processing of inter-language difference in the Alt-Cond (vs. Rep-Cond). The neural processing of intra-language similarity was quantified as correlation distances between neural responses to two words of the same language whereas the neural processing of inter-language difference was assessed as correlation distances between neural responses to two words of two different languages. The correlation distances from the multivariate analyses were further employed to assess how words of one language are clustered and how far words of two languages are separated in a two-dimensional (2D) space during language-based word categorization. Enhanced language-based word categorization is associated with smaller intra-language correlation distances, which reflect more densely clustered words of the same language, and larger inter-language correlation distances, which manifest further separated words of two different languages.”

      Page 37-38

      “How possible are the early neural RS effects within 200 ms after word onset observed in our study related to the processing of low-level perceptual features or high-level linguistic (e.g., orthography, semantics, phonology) properties of visual words? Our analyses of the ERPs to scrambled Chinese and English words in Experiment 2 did not show significant RS effect. Because only low-level visual features were preserved in the scrambled words, the ERP results provided no evidence that the early RS effects on the neural response to words can be attributed to habituation of perception of the low-level perceptual features. Furthermore, we found that the RS effects on the neural response to radicals and letters in Experiment 3 took place in a delayed time window and exhibited different scalp distributions (i.e., over the central region for radicals and occipital regions for letters) compared with the neural RS effects related to words. Thus the early RS effects on the neural response to words cannot be interpreted as habituation of perception of the middle-level units of Chinese and English words (i.e., radicals and letters) either. In addition, the early neural RS effects were similarly observed for both familiar (i.e., Chinese and English) and unfamiliar (i.e., Korean and Italian) languages and occurred earlier than the time window in which the processing of the linguistic properties of visual words takes place (Marinkovic et al., 2003; Hodgson et al., 2021; Zhu et al., 2022). Therefore, the early neural RS effects identified in our work were unlikely to be associated with the processing of the linguistic (e.g., orthography, semantics, phonology) properties of visual words since these properties of unfamiliar languages were unknown to the participants. Taken together, our findings of the early neural RS effects highlight an early word-level representation of alphabetic vs. non-alphabetic languages which distinguishes words from letters/radicals but is similar for familiar or unfamiliar languages. Our results, however, do not exclude the possibility that the processing of the linguistic properties of visual words may contribute to the long-latency RS effect around 300 ms after word onset. Further processing of the linguistic properties of visual words of familiar languages may follow the early language-based categorization of visual words, though this should be tested in future research.”

      (2) Question 2:

      For unlearned languages, people cannot distinguish lexical from sub-lexical levels. What, then, determines (i) the RS-effect difference between letters and radicals in familiar languages and words in unlearned ones, and (ii) the similarity of repetition effects between words in unlearned and familiar languages? An explicit account is needed.

      Thank you for this suggestion. In the revised manuscript, we’ve included a dedicated paragraph addressing these two issues. Specifically, we’ve provided a more precise account of the differences in repetition suppression (RS) effects between words and letters/radicals in familiar languages, as well as the similar RS effects observed for unlearned and familiar languages. We believe that our findings of the early neural RS effects highlight an early word-level representation of alphabetic vs. non-alphabetic languages which distinguishes words from letters/radicals but is similar for familiar or unfamiliar languages (see Page 37-38).

      Page 37-38

      “How possible are the early neural RS effects within 200 ms after word onset observed in our study related to the processing of low-level perceptual features or high-level linguistic (e.g., orthography, semantics, phonology) properties of visual words? Our analyses of the ERPs to scrambled Chinese and English words in Experiment 2 did not show significant RS effect. Because only low-level visual features were preserved in the scrambled words, the ERP results provided no evidence that the early RS effects on the neural response to words can be attributed to habituation of perception of the low-level perceptual features. Furthermore, we found that the RS effects on the neural response to radicals and letters in Experiment 3 took place in a delayed time window and exhibited different scalp distributions (i.e., over the central region for radicals and occipital regions for letters) compared with the neural RS effects related to words. Thus the early RS effects on the neural response to words cannot be interpreted as habituation of perception of the middle-level units of Chinese and English words (i.e., radicals and letters) either. In addition, the early neural RS effects were similarly observed for both familiar (i.e., Chinese and English) and unfamiliar (i.e., Korean and Italian) languages and occurred earlier than the time window in which the processing of the linguistic properties of visual words takes place (Marinkovic et al., 2003; Hodgson et al., 2021; Zhu et al., 2022). Therefore, the early neural RS effects identified in our work were unlikely to be associated with the processing of the linguistic (e.g., orthography, semantics, phonology) properties of visual words since these properties of unfamiliar languages were unknown to the participants. Taken together, our findings of the early neural RS effects highlight an early word-level representation of alphabetic vs. non-alphabetic languages which distinguishes words from letters/radicals but is similar for familiar or unfamiliar languages. Our results, however, do not exclude the possibility that the processing of the linguistic properties of visual words may contribute to the long-latency RS effect around 300 ms after word onset. Further processing of the linguistic properties of visual words of familiar languages may follow the early language-based categorization of visual words, though this should be tested in future research.”

      Reviewer #2 (Public review):

      Summary:

      This study investigates how the human brain categorizes visual words from distinct writing systems (alphabetic vs. non-alphabetic) as a neural basis for the social-categorization function of language. Using a repetition suppression paradigm combined with electroencephalography and magnetoencephalography, the authors conducted nine experiments with independent participants to identify the neural network underlying language-based categorization, characterize its temporal dynamics, and test whether this process operates independently of linguistic properties such as semantic meaning and pronunciation.

      Strengths:

      (1) The study employs a well-validated design with clear control conditions and systematically manipulates key variables, including writing system, language familiarity, and native language background. The use of nine experiments with independent participant samples strengthens the reliability and replicability of the results.

      (2) The work combines EEG and MEG, cross-validating findings across imaging modalities to support the reported neural effects. A combination of univariate, multivariate, and connectivity analyses is used to characterize neural responses and network interactions.

      (3) Results are consistent across multiple language groups and for both familiar and unfamiliar languages, supporting the generalizability of the identified neural mechanism beyond specific languages or prior experience.

      Weaknesses:

      The authors provide compelling evidence that the identified neural network supports the categorization of words by language, including computations of intra-language similarity and inter-language difference. However, the conceptual framing of this finding as directly reflecting the social-categorization function of language may be premature. While the task captures spontaneous language categorization, it does not involve social evaluation or intergroup processes. The connection to social categorization is inferred from prior literature rather than demonstrated within the current experimental design. Clarifying this distinction would strengthen the conceptual precision of the manuscript.

      Thank you for this important comment. In the revised Introduction and Discussion, we’ve clarified several related issues. First, prior research suggests that language can serve as a socially relevant category cue. Second, these findings imply that rapid categorization of words by language may occur in the human brain. Third, although our results identify a neural network supporting such rapid language-based categorization of visual words, they do not directly test how this process relates to social categorization of people (see Page 3-4; Page 39). Highlighting these points help delineate the scope of our findings and point to important directions for future research.

      Page 3-4

      “The social-categorization function of language revealed in these behavioral studies implicates that rapid categorization of words of different languages may occur in the human brain. Furthermore, the findings of infant studies (e. g., Liberman et al., 2017b) suggest that the neural process involved in categorization of words of different languages may develop even prior to the processing of linguistic properties (e.g. semantic meanings) of words. Nevertheless, up to date, there has been little neuroimaging research examining the neural mechanisms underlying automatic and fast categorization of words of different languages.”

      Page 39

      “Finally, it should be noted that the current work was initiated by the previous behavioral findings which suggest that language can serve as a socially relevant category cue but focused on the neural mechanisms underlying rapid language-based categorization of visual words. Although the previous findings suggest that the language-based categorization of visual words provides a cognitive basis of social categorization of people, our work did not directly test whether and how the neural processes involved in the language-based categorization of visual words are linked to social evaluation or intergroup processes which are critical for social categorization of people. To clarify this issue should promote deep comprehension of the neural mechanisms underlying the social-categorization function of language but is beyond the scope of the current study. Future research should investigate the connection between language-based categorization of words and social categorization based on other social cues (e.g., faces), which is pivotal to understanding of social interactions in real-world situations.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Revise the conceptual framing to clarify the relationship between the experimental results and the proposed social-categorization function of language. If the authors wish to retain the emphasis on social categorization in the title or discussion, they should explicitly explain how the observed neural mechanisms of language-based word categorization link to social evaluation, intergroup processes, or real-world social categorization. This clarification would strengthen the conceptual coherence and justify the use of social categorization within the current study's scope.

      Thank you for this and the following suggestions. In the revised Introduction and Discussion, we’ve clarified the following point: First, the findings of prior behavioral studies suggest a social-categorization function of language. Second, based on these behavioral findings, we predicted automatic and fast categorization of words by language. Our study tested this prediction using neuroimaging and investigated the neural mechanisms of language-type-based categorization of visual words. This is the main goal of our work. Third, to examine how the observed neural mechanisms of language-based word categorization link to social evaluation, intergroup processes, or real-world social categorization is important but beyond the scope of the current work. However, this is a very important question. Future research should test the connection between the neurocognitive processes involved in social categorization of people and the neural categorization of visual words by language revealed in our study. Consistently, the title of our paper “Neural categorization of visual words of alphabetic and non-alphabetic languages” and Discussion focus on contributions of our findings to understanding of the neural categorization of visual words by language rather than its connection to social categorization of people. Above all, we’ve clarified in the revision that our study was initiated by the findings of social function of language but was limited to the neural processing of visual words (see Page 3-4; Page 39). Thanks again for this comment.

      Page 3-4

      “The social-categorization function of language revealed in these behavioral studies implicates that rapid categorization of words of different languages may occur in the human brain. Furthermore, the findings of infant studies (e. g., Liberman et al., 2017b) suggest that the neural process involved in categorization of words of different languages may develop even prior to the processing of linguistic properties (e.g. semantic meanings) of words. Nevertheless, up to date, there has been little neuroimaging research examining the neural mechanisms underlying automatic and fast categorization of words of different languages.”

      Page 39

      “Finally, it should be noted that the current work was initiated by the previous behavioral findings which suggest that language can serve as a socially relevant category cue but focused on the neural mechanisms underlying rapid language-based categorization of visual words. Although the previous findings suggest that the language-based categorization of visual words provides a cognitive basis of social categorization of people, our work did not directly test whether and how the neural processes involved in the language-based categorization of visual words are linked to social evaluation or intergroup processes which are critical for social categorization of people. To clarify this issue should promote deep comprehension of the neural mechanisms underlying the social-categorization function of language but is beyond the scope of the current study. Future research should investigate the connection between language-based categorization of words and social categorization based on other social cues (e.g., faces), which is pivotal to understanding of social interactions in real-world situations.”

      (2) Clarify the consistency between the reported model order (5 ms lag) and the sampling rate after downsampling (250 Hz, corresponding to 4 ms per time point). If a discrepancy exists, clearly explain how the time-series data were processed.

      We clarified in the revision (see Page 53) that “because down-sampling was not applied to the GCA analyses, a 5-ms lag was used for prediction of the neural activity in one brain region using the neural activity in another brain region”.

      (3) For the representational similarity analysis (RSA), report reliability measures for the representational dissimilarity matrices (e.g., split-half reliability) to verify that the observed effects are stable given the number of trials per condition.

      Following this suggestion, we’ve conducted split-half reliability analyses and reported the results in the revised supplementary materials. The reliability analyses are also mentioned in the revised Discussion (see Page 40).

      Page 40

      “In conclusion, our EEG and MEG results revealed robust RS effects in the early neural responses to visual words of the same language. The reliability of these RS effects was confirmed across words of different familiar and unfamiliar languages, in samples of speakers with different native languages, and through split-half reliability analyses (see Supplementary Materials, Fig. S19). These effects were supported by the bilateral neural networks whose activity reflected computations of correlation distances between word pairs, capturing both intra-language similarity and inter-language differences during the categorization of visual words in alphabetic and non-alphabetic languages. Together, these findings advance our understanding of spontaneous, language-based neural categorization of visual words as a key basis of the social-categorization function of language.”

      (4) Provide complete statistical information for all significant results reported in the supplementary materials, including relevant test statistics (e.g., t-values, cluster p-values) in figure legends or a supplementary results table to improve transparency.

      Complete statistical information has been provided in the revised supplementary materials (see Tables S4 and S5).

      (5) Streamline the presentation of the nine experiments in the main text to emphasize the core conceptual and methodological logic, potentially using a schematic overview or flowchart to improve readability.

      As suggested, we’ve included an overview of the nine experiments in the revised Introduction. This overview helps understanding of the core conceptual and methodological issues in our work (see Page 6).

      Page 6

      “In nine experiments we recorded EEG/MEG signals from Chinese, English, and German speakers when viewing words of an alphabetic language and a non-alphabetic language (English and Chinese words, or Italian and Korean words) or of two alphabetic languages (English and German) in the Rep-Cond and Alt-Cond. We recorded EEG signals from Chinese participants to examine temporal neural dynamics of spontaneous language-based word categorization in Experiment 1. The similar paradigm was employed in Experiments 2 and 3 to investigate whether perceptual features or radical/letters of words are sufficient to generate spontaneous language-based categorization of visual words. The results in Experiment 1 were replicated in native English and German speakers in Experiments 4 and 5, respectively. Neural dynamics of categorization of words of two unlearned languages were further investigated in Chinese participants in Experiment 6. Finally, the neural networks supporting the spontaneous categorization of words of two learned or unlearned languages were localized using MEG in Chinese and English speakers in Experiments 7-9, respectively.”

      (6) Strengthen the transition between the discussion of the social-categorization function of language and the neural mechanisms of visual word categorization in the introduction.

      Following this suggestion, we’ve modified the Introduction to strengthen the transition between the discussion of the social-categorization function of language and research on neural mechanisms of visual word categorization (see Page 3-4).

      Page 3-4

      “The social-categorization function of language revealed in these behavioral studies implicates that rapid categorization of words of different languages may occur in the human brain. Furthermore, the findings of infant studies (e. g., Liberman et al., 2017b) suggest that the neural process involved in categorization of words of different languages may develop even prior to the processing of linguistic properties (e.g. semantic meanings) of words. Nevertheless, up to date, there has been little neuroimaging research examining the neural mechanisms underlying automatic and fast categorization of words of different languages.”

      (7) Briefly define the repetition suppression (RS) paradigm when first mentioned (i.e., reduced neural response to repeated stimuli from the same category, reflecting categorical processing) to improve accessibility for non-specialist readers.

      The RS paradigm is now defined in Introduction when being mentioned for the first time in the manuscript (see Page 5-6).

      Page 5-6

      “The present study investigated neural dynamics of categorization of visual words of two different (an alphabetic versus a non-alphabetic, or two different alphabetic) languages by combining EEG/MEG with a repetition suppression (RS) paradigm adopted from previous studies of social categorization of faces (Zhang et al., 2023b; Zhou et al., 2020). RS refers to the attenuation in neural responses to a repeated occurrence of stimuli that engage common neuronal populations or processes due to habituation (Grill-Spector et al., 2006). The RS paradigm consisted of an alternating condition (Alt-Cond), in which visual words of two different languages were presented alternately, and a repetition condition (Rep-Cond), in which words of one language were presented repeatedly (Fig. 1a). Neural responses to stimuli of the same category were attenuated in the Rep-Cond compared to Alt-Cond due to habituation and this RS effect disentangles the neural activities underlying categorization of faces and body silhouettes of a specific social group.”

      (8) Report detailed participant demographic information, including exact age range/mean age and gender ratio for each experiment, to meet standard reporting practices in neuroscience.

      We’ve modified Table S1 to include the information about exact age range/mean age and gender ratio in each experiment.

      (9) Correct minor typographical and grammatical errors, including These finding (line 59) and Chinse (line 223).

      These and other grammatical errors have been corrected in the revision.

    1. Author response:

      The following is the authors’ response to the original reviews

      Summary of revision for all referees:

      We thank referees for their constructive comments. To address their concerns, we now performed additional statistical analyses integrating both paired and unpaired data, performed positive controls for comparisons between NH- and CI- evoked iEEG measurements, developed tools for measuring and collected new experimental data on forward masking ECAP measurements in CI implanted rats (N=3), and reworked both manuscript text and figures to improve clarity. These most significant changes are summarized here, and a complete list of responses to reviewers and corresponding changes will follow.

      Summary of major changes to revised manuscript:

      (1) Statistical treatment of paired vs unpaired recordings using mixed-effects models (updates to all manuscript figures that compare NH vs CI); this largely confirmed the results reported in our original submission.

      (2) New analysis, controlling for information-theoretic cross-modality comparison (i.e., training with tone- and testing with cochlear implant-evoked iEEG measures, Fig. 8).

      (3) Clarification of methods (Supplemental Fig. 2 & manuscript text)

      (4) Additional experiments testing peripheral tuning of our 8-channel CI rodent model via forward masking ECAP measures across 3 animals (N=3, Supplemental Fig. 1)

      (5) Detailed response addressing robustness of tonotopy in NH and CI animals

      Public Reviews:

      Reviewer #1 (Public Review):

      Strengths:

      The study poses a timely, clinically relevant question with clear implications for CI strategy. The analytical toolkit is appropriate: µECoG captures mesoscale patterns; TCA offers a transparent separation of spatial and temporal structure; and mutual-information decoding provides an interpretable measure of single-trial discriminability. Within-subject recordings in a subset of animals, in principle, help isolate modality effects from inter-animal variability. Where analyses are most direct, the acoustic condition yields higher single-trial decoding accuracy, which is a meaningful and clearly presented result.

      We appreciate the comments on the strengths of our analytic approaches.

      Weaknesses:

      Parts of the statistical treatment do not match the data structure: some comparisons mix paired and unpaired animals but are analysed as fully paired, raising concerns about misestimated uncertainty.

      Please see our response to specific comment #2 above. In short, we agree with this critique of our original analyses, and in our revised manuscript we re-analyzed all NH vs. CI comparisons using linear mixed effects models that incorporate both paired and unpaired observations within a single framework. This allows us to include all animals, account for within-animal dependence for paired experiments (normal hearing and cochlear implant data from the same animal when available), and to align the statistical tests with the data shown in the figures. In almost every case, the mixed effects models confirm our original conclusions. Two comparisons that were previously nonsignificant now reach criterion for statistical significance (Fig. 2E, p=0.048 and Fig. 6F, p=0.027). We updated the manuscript to report these values and to clarify the use of mixed effects modeling in the methods under the section titled, “Linear mixed effects modeling.”

      Methodological reporting is incomplete in places; essential parameters for both acoustic and electrical stimulation, as well as objective verification of implantation and deafening, are not described with sufficient detail to support confident interpretation or replication.

      Please see our response to comment #5 below. We have revised our manuscript to now include this information in the methods.

      Figure-level clarity also undermines the message. In Figure 2, non-significant slopes for CI, repeated identification of a single "best channel," mismatched axes, and unclear distinctions between example and averaged panels make the assertion of spatial organisation unconvincing; importantly, the normal-hearing panels also do not display tonotopy as clearly as expected, which weakens the key contrast the paper seeks to establish.

      This is an important point, thanks- please see responses to comment #1 above. We note that conventional tonotopic maps in auditory cortex are characteristic frequency maps, i.e., maps of topographic organization for responses to lowest-threshold stimuli (often presented around 20-50 dB SPL). Our maps were constructed from stimuli presented at 70 dB SPL, thus blunting crisp tonotopy to some degree. Furthermore, we quantified spatial organization using a previously published method from the Polley lab (Romero & Hight et al. 2020), in which local tonotopic gradient vectors (magnitude and direction) were computed from GCaMP responses at each pixel and projected onto a unit circle. Mean vector strength across all pixels was then compared to a shuffled distribution as a measure of tonotopic organization. We applied the same procedure to our iEEG best-frequency and best-channel maps. Both map types yielded mean vector strengths that were substantially larger than those derived from shuffled maps (p < 10<sup>-10</sup>), indicating that our maps have a consistent tonotopic (for BFs) or cochleotopic (for CI channels) organization that is highly unlikely to arise by chance. This is now included in our revised manuscript.

      Finally, the decoding claims would be strengthened by simple internal controls, such as within modality train/test splits and decoding on raw ERP/high-gamma features to demonstrate that poor cross-modal transfer reflects genuine differences in the underlying responses rather than limitations of the modelling pipeline.

      Please see our response to comment #12 below. In short, we have now included this analysis in revised Figure 8.

      Reviewer #2 (Public Review):

      Strengths:

      The study includes interesting analyses of the sound and cochlear implant representation structure based on decoders.

      We appreciate the comment on how interesting our analyses are, thanks!

      Weaknesses:

      The observation that responses to cochlear implant stimulation (stimulation) are spatially organized is not new (e.g., Adenis et al. 2024).

      We agree that it is not particularly novel to report that there is spatial organization to cochlear implant stimulation. However, we believe that our direct comparisons (when possible, within animal) between normal-hearing and cochlear implant modality maps is unusual in the literature, including asking how decoders based on one set of responses might apply to responses evoked from the other modality. Adenis et al. (2024) is a fantastic study of pulse shape and monopolar vs bipolar stimulation modes with a 6-channel implant in guinea pig, but as far as we can tell this study does also not compare normal hearing maps prior to deafening and implantation to the cochlear implant maps in the same animals.

      The claim that spatial and temporal dimensions contribute information about the sound is also not new; there is a large literature on this topic. Moreover, the results shown here are extremely weak. They show similar levels of information in the spatial and temporal dimensions, and no synergy between the two dimensions. This is however, likely the consequence of high measurement noise leading to poor accuracy in the information estimates, as the authors state.

      Good point, please see our response to comment #1 below.

      The main claim of the study - the mismatch between cochlear implant and sound representation - is not supported. The responses to each modality are measured in different animals. The authors do not show that they actually can compare representations across animals (e.g., for the same sounds). Without this positive control, there is no reason to think that it is possible to decode from one animal with a decoder trained on another, and the negative result shown by the authors is therefore not surprising.

      Good point, thanks- please see our response to comment #2 below, where we describe this new control we have added.

      Reviewer #3 (Public Review):

      Strengths:

      The model combining micro-eCoG and cochlear implantation and the methodology to extract both the Event Related Potentials (ERPs) and High-Gammas (HGs) is very well designed and appropriately analyzed. Likewise, the PCA-LDA and TCA-LDA are powerful tools that take full advantage of the information provided by the cortical ensembles. The overall structure of the paper, with a paced and exhaustive progress through each step and evolution of the decoder, is very appreciable and easy to follow. The exploration of single-trial encoding and stimulus identity through temporal and spatial domains is providing new avenues to characterize the cortical responses to CI stimulations and their central representation. The fact that single trials suffice to decode the stimulus identity regardless of their modality is of great interest and noteworthy. Although the authors confirm that iEEG remains difficult to transpose in the clinic, the insights provided by the study confirm the potential benefit of using central decoders to help in clinic settings… the reviewer wants to reiterate that the study proposed by Hight et al. is well constructed, relevant to the field, and that the overall proposal of improving patient performances and helping their adaptation in the first months of CI use by studying central responses should be pursued as it might help establish new guidelines or create new clinical tools.

      We thank the Reviewer for the positive comments about the thoroughness of our analyses and clear organization of our manuscript.

      Weaknesses:

      The conclusion of the paper, especially the concept of distinct cortical encoding for each modality, is unfortunately partially supported by the results, as the authors did not adequately consider fundamental limitations of CI-related stimulation. First, the reviewer assumed that the authors stimulated in a Monopolar mode, which, albeit being clinically relevant, notoriously generates a high current spread in rodent models.

      Thanks, this is an important potential concern. Please see our response to comment #5 of Referee 1 and responses to comment #3 below. We agree that monopolar stimulation would be expected to be less spatially specific than bipolar or multipolar modes. However, we chose monopolar stimulation because it is the main clinical configuration in human CI users and therefore most relevant for translational purposes. For our revised manuscript, we made new ECAP measurements of peripheral (spatial and temporal) tuning via a forward masking paradigm and demonstrate that monopolar is effectively tuned (Supplemental Fig. 2). Together with additional single-animal maps in Supplementary Figure 3, together with our vector-strength analysis (Response Fig. 2), demonstrate that even under acute monopolar stimulation we observe structured cochleotopic organization in cortex, rather than the extremely low-pass patterns one might expect if monopolar spread was a major contaminant.

      Second, comparing the averaged BF maps for iEEG (Figure 2A, C), BFs ranged from 4 to 16kHz with a predominance of 4kHz BFs. The lack of BFs at higher frequencies hints at a potential location mismatch between the frequency range sampled at the level of the cortex (low to medium frequencies) and the frequency range covered by the CI inserted mostly in the first turn-and-a-half of the cochlea (high to medium frequencies). Looking at Figure 2F (and to some extent 2A), most of the CI electrodes elicited responses around the 4kHz regions, and averaged maps show a predominance of CI-3-4 across the cortex (Figure 2C, H) from areas with 4kHz BF to areas with 16kHz BF. It is doubtful that CI-3-4 are located near the 4kHz region based on Müller's work (1991) on the frequency representation in the rat cochlea.

      Please see our responses to comment #3 below.

      Taken together with the Pearsons correlations being flat, the decoder examples showing a strong ability to identify CI-4 and 3 and the Fig-8D, E presenting a strong prediction of 4kHz and 8kHz for all the CI electrodes when using a pure tone trained decoder, it is possible that current spread ended stimulating indistinctly higher turns of the cochlea or even the modiolus in a non-specific manner, greatly reducing (or smearing) the place-coding/frequency resolution of each electrode, which in turn could explain the coarse topographic (or coarsely tonotopic according to the manuscript) organization of the cortical responses. Thus, the conclusion that there are distinct encodings for each modality is biased, as it might not account for monopolar smearing. To that end, and since it is the study's main message and title, it would have benefited from having a subgroup of animals using bipolar stimulations (or any focused strategy since they provide reduced current spread) to compare the spatial organization of iEEG responses and the performances of the different decoders to dismiss current spread and strengthen their conclusion.

      Please see our responses to comment #4 below as well as our responses related to monopolar vs bipolar stimulation. We agree that for future studies, it will be important to do a heads-on comparison of the differences between bipolar and monopolar stimulation depending on electrode location and stimulation intensity.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      We thank the reviewer for commenting on the strengths of our manuscript, including appreciating the power and timeliness of our approach.

      (1a) Figure 2 does not convincingly support the claim that "tone-evoked and CI-evoked iEEG measurements are spatially organized," particularly for CI data: Figure 2C repeatedly highlights the same "best channel," and the slopes in Figures 2B and 2G are non-significant; there are also discrepancies between panels (A vs. C, F vs. H) and mismatched frequency ranges (0-16 kHz vs. up to 32 kHz), which should be clarified as exemplar versus averaged displays and harmonized in scale.

      (First we note that Reviewer 3 also raised related concerns about the robustness of tonotopy in our iEEG data.) We address these by comparing our maps to previously published tonotopic maps, and using an established quantitative analysis of tonotopic strength from Romero & Hight et al. (2020).

      First, to place our tone-evoked iEEG maps in context, we overlaid them on the same spatial scale and orientation as both single-unit tonotopy in rat primary auditory cortex (A1) from Polley et al. (2006) and iEEG maps obtained with the same surface array in Insanally et al. (2016). The rostral–caudal and dorsal–ventral axes and cortical extents are matched across panels. Our best-frequency maps (Figure 2C) qualitatively recapitulate the high-to-low frequency gradient and spatial layout reported in both of these prior studies, supporting our claim that tone-evoked iEEG captures canonical mesoscale tonotopy. We have updated the manuscript results section to directly reference these two studies, “The area and orientations of tone-evoked maps qualitatively match those published from single unit recordings (Polley et al. 2006) and published using similar iEEG arrays (Insanally et al. 2016).”

      Second, to quantify tonotopy in a way that is directly comparable to previous work, we reproduced the analysis of Romero & Hight et al. (2020), who examined tone-evoked GCaMP signals (Romero & Hight et al. (2020)). In that paper, local tonotopic gradient vectors (magnitude and direction) were computed at each pixel and projected onto a unit circle; the mean vector strength across all pixels was then compared to a shuffled distribution as a measure of tonotopic organization. We applied the same procedure to our iEEG best-frequency and best-channel maps (Fig. 2C-E). Both map types yielded mean vector strengths that were substantially larger than those derived from shuffled maps (p < 10<sup>-10</sup>), indicating that our maps have a consistent tonotopic (for BFs) or cochleotopic (for CI channels) organization that is highly unlikely to arise by chance. We cite this paper for these analyses related to Figure 2.

      (1b) Figure 2C repeatedly highlights the same ‘best channel’

      We agree that many CI-evoked maps are dominated by a single channel, as seen in our exemplar and in the additional animals shown in new Supplemental Fig. 3. In Fig. 2C, channel 5 emerges as the dominant best channel, as CI-evoked activity in this animal is broad and is strongest for channel 5 (Fig. 2A). This reflects a feature of iEEG signals rather than a plotting artifact. Biophysically, iEEG reflects spatially summed local field potentials that low-pass filter underlying neural activity; these far-field signals aggregate excitatory and inhibitory processes and are not expected to show the sharp single-neuron tuning seen in spike recordings. As a result, broad peaks centered on the most strongly driven channels are expected. We have added text in the results section discussing these limitations, overall maps reduced from iEEG responses were similar in size and orientation compared to single unit maps, “albeit at coarser gradients likely due to aggregate recordings of excitatory and inhibitory activity and low-pass filtering due to potentials originating far from recording sites.” We also added in the results section the comparison of spatial correlations (Fig. 2B,G) at the extremes of stimulus separation “electrode separations (CI 1 vs ≥5 electrodes, ERP: p=0.01, HG: p=0.04)” as analyzed by linear mixed effects models.

      (1c) Mismatched frequency ranges

      We constricted the range of frequencies plotted in some panels (e.g., Fig. 2C from 1.4-32 kHz to 1.4-16 kHz) to emphasize the compressed range of tonotopic gradients and patterns.

      (1d) The slopes in Figures 2B and 2G are non-significant

      We agree that non-significant group-level slopes indicate that CI-evoked tonotopy is weaker than tone-evoked tonotopy, and we now emphasize this point. At the same time, the data exhibit systematic structure: for both ERP and HG, mean spatial correlations decline monotonically with increasing CI channel separation (Fig. 2B,G). We also directly compared spatial correlations at the extremes of stimulus separations (1 vs. ≥5-channel separation) and found a significant difference. This is updated in the manuscript as: “At the extremes, the spatial correlations were always higher for small vs. large tone separations (NH 0.5 vs ≥3.5 octaves, ERP: p<10<sup>-4</sup>, HG: p<10<sup>-4</sup> Student’s one-tailed t-test) and electrode separations (CI 1 vs ≥5 electrodes, ERP: p=0.01, HG: p=0.04).”. Together with the strong deviation from shuffled maps in the vector-strength analysis (Fig. 2E), we argue that analysis of spatial correlations indicates that CI-evoked maps are not random but reflect a coarse underlying gradient. In addition, as tone-evoked maps exhibit tonotopy, we asked if CI stimulation itself is at least spatially tuned in the periphery. Using ECAPs with a forward-masking paradigm (new Supplemental Fig. 1), we show that probe-evoked ECAPs are significantly more suppressed by adjacent than by distant maskers (N = 3), demonstrating functional spatial tuning of CI electrodes in the cochlea. We have also replotted these results in comparison with the same measurements from a human CI user (Author response image 1). This supports the interpretation that peripheral input is spatially specific and that the weaker cortical cochleotopy likely reflects the properties and resolution of iEEG and acute CI stimulation rather than a complete absence of spatial organization. Overall, the new comparative figures and analyses are intended to make transparent that (i) iEEG robustly captures tonotopy for acoustic tones, and (ii) CI-evoked CI-evoked responses exhibit coarser, but statistically non-random, cochleotopic organization.

      Author response image 1.

      Here, we compare data from the new Supplemental Figure 1C,D with human data (N=1) for spatial & temporal tuning in the periphery, as assessed by forward masking ECAP measurements. A) Spatial tuning functions were averaged across all probe electrodes and 3 animals (left) and 1 human subject (right) (black, mean; gray: s.e.m..; orange, average of individual subjects). B) Temporal tuning functions were averaged across all probe electrodes and 3 animals (left) and 1 human subject (right) (black, mean; gray, s.e.m.; orange, average of individual subjects). Note: human subject is the first-author, a long-term cochlear implant user (>10 years) with significant open set speech perception.

      (2) The statistical approach is inappropriate where pairing is incomplete: a Student's paired two-tailed t-test is used despite not all data being paired; a linear mixed-effects model would be more suitable, whereas an unpaired test risks reduced power.

      We agree with this suggestion. As the reviewer notes (also raised by Reviewer 3), our original analyses did not fully exploit the partially paired structure of the data. In the initial submission we used paired t-tests when animals contributed both normal-hearing (NH) and CI measurements, which meant that animals with only NH or only CI data were excluded from those tests.

      To address this, we have re-analyzed all NH vs. CI comparisons using linear mixed-effects models that incorporate both paired and unpaired observations within a single framework. This approach allows us to (i) include all available animals, (ii) appropriately account for within-animal dependence when both conditions are present, and (iii) align the statistical tests with the data shown in the figures. In nearly all cases, the mixed-effects models confirm our original conclusions. Two comparisons that were previously non-significant are now significant in the positive direction: Fig. 2E (p = 0.048) and Fig. 6F (p = 0.027, linear mixed-effects models). We have updated the manuscript to report these values and to clarify the use of mixed-effects modeling in the methods under the section titled, “Linear mixed effects modeling.”

      (3a) Given the surgical complexity, objective verification of implantation and deafening is needed (e.g., eABRs for implant function and post-deafening ABR thresholds)”

      We agree that objective verification of both implant placement and deafening is critical, particularly given the surgical complexity of multichannel CI implantation in rats. Note that we previously extensively documented deafness in our cochlear implant rats with eABRs, histology of hair cell counts, and behavior (turning the implant off and seeing performance drop to chance). As we argued in Glennon et al. Nature 2023, the primary outcome measure and definition of deafness is behavioral, as anatomical and physiological markers are correlates of functional deafness but ultimately deafness must be defined in terms of behavioral performance. This is described in more detail below.

      We agree that objective verification of both implant placement and deafening is critical, particularly given the surgical complexity of multichannel CI implantation in rats. Note that we previously extensively documented deafness in our cochlear implant rats with eABRs, histology of hair cell counts, and behavior (turning the implant off and seeing performance drop to chance). As we argued in Glennon et al. Nature 2023, the primary outcome measure and definition of deafness is behavioral, as anatomical and physiological markers are correlates of functional deafness but ultimately deafness must be defined in terms of behavioral performance. This is described in more detail below.

      Implant placement: Our primary concern during surgery is to ensure that the CI array is correctly positioned along the cochlear spiral toward the apex. As shown in Author response image 2, once the bulla is opened and the cochleostomy is made at the junction of the temporal bone and the stapedial artery, the orientation of the cochlear spiral is clearly visible under the surgical microscope. We advance the 8-channel array only in the apical direction, and we require that all 8 electrodes pass through the cochleostomy. A complete insertion of all 8 electrodes cannot be achieved with a basal-ward trajectory, so full insertion provides a strong anatomical confirmation that the array is directed apically. The white band on the array, visible just basal to the cochleostomy (Author response image 2), serves as a consistent visual marker of complete insertion. We have added text and this figure to the Methods to clarify these criteria, “We required that all eight electrodes pass through the cochleostomy, confirming that the array was inserted in the direction of the apex.”

      Verification of deafening: We also share the reviewer’s concern about confirming profound hearing loss, particularly because some CI animals were presented acoustic tones to drive individual channels. We used the same mechanical-only deafening procedure described and validated in our previous work (King et al., 2016; Glennon et al., 2023), which was chosen to minimize systemic side-effects and maximize post-surgical survival, validated in three ways:

      - Histology: In N=4 deafened animals, inner hair cell loss was ~50% and outer hair cell loss was near complete at almost 100% in all animals.

      - Physiology: For N=14 rats, acoustic ABRs were substantial before deafening but statistically similar to baseline noise after deafening.

      - Behavior: For N=16 deafened rats, behavioral performance with implant on was d′: 1.7±0.1, but when implant was turned off in a subset of sessions, performance dropped to chance (d′: −0.05±0.1, P < 0.0001).

      Author response image 2.

      Visual confirmation of a successful electrode insertion. The direction of an 8-channel array being implanted toward the apex is clear under microscope. Full insertion of all 8 channels is further confirmed by the white band’s (located after basal electrode) proximity to the cochleostomy.

      This combination of histological, physiological, and behavioral evidence indicates that the mechanical-only deafening protocol produces profound hearing loss, with no functionally relevant residual hearing at intensities equal to or greater than those used in our study (70 dB SPL). Given this prior validation under identical surgical and experimental conditions, we are confident that our CI animals were effectively deafened and that the iEEG responses we report are driven by the implant rather than by residual acoustic hearing. We now clarify this in the Methods and explicitly cite our validation: “(mechanical only, as described and validated in Glennon et al. 2023).

      (3b) One CI animal did not learn the task (Fig. 1C), potentially reflecting implantation efficacy.

      Good point, thanks. For both humans and rats, cochlear implant performance can be highly variable, reflecting a number of factors in terms of device performance, training efficacy and motivation, or other technical or biological sources of heterogeneity. We note however that not all animals included in this study were behaviorally trained, and wanted to show the full range of variable performance for the subset of animals that were trained (N=4 typical hearing and N=3 cochlear implant rats, one of the 4 trained animals lost the implant before it could be re-trained on the cochlear implant version of the task). We now highlight this range of performance variability in the results section and explain why N=4 normal-hearing and N=3 cochlear implant rats.

      (4) The behavioural paradigm and cohort accounting are unclear: Figure 1C shows four NH-trained rats, yet subsequent analyses include only two NH-trained animals, which is confusing.

      We have now clarified the relation between the behavioral cohort and the iEEG cohort in the revised manuscript. The key point is that the animals in Figure 1C are defined by their behavioral training history (NH vs CI training), whereas inclusion in the iEEG analyses is defined by the specific stimuli collected during acute recordings, and these two categorizations are not always the same. In total, four rats underwent both iEEG recordings and behavioral training. Of these four, three were subsequently deafened, implanted with chronic CIs, and trained on the CI-driven task (Fig. 1C). With respect to the acute iEEG experiments, we obtained tone-only iEEG in 1 animal, CI-only iEEG in 2 animals, and both tone- and CI-evoked iEEG in 1 animal.

      Thus, the “NH-trained” label in Figure 1C refers to behavioral training status, not to the stimulus conditions used during iEEG recordings. All iEEG measurements were acute and performed immediately after surgery (for CI animals) or in the normal-hearing condition, before any CI behavioral training. Consequently, the behavioral cohort in Figure 1C is larger than the subset of animals that contributed to specific iEEG contrasts in later figures, which explains why some panels include only two NH animals.

      To clarify this, we have added a new Supplementary Figure 2 that provides a timeline for each animal, indicating when behavioral training occurred, when deafening and implantation occurred, and which stimulus conditions (tones vs CI) were used for each iEEG recording. We kept this figure in the Supplementary section because the focus of the manuscript is on evoked iEEG measurements rather than behavior, but the revised text now explicitly refers to this schematic when describing the cohorts “The combinations of animals that underwent behavioral training and acute iEEG measurements are shown in Supplemental Fig. 2.”

      (5) Methods lack essential details: specify acoustic stimulus types and intensities, CI stimulation parameters (e.g., current/charge per phase, phase width, rate, loudness setting), and the recording state (awake vs. anaesthetised), which is only implied in the discussion.

      We agree that these details are essential, and Reviewer 3 raised similar concerns about methodological clarity. We have now expanded the Methods to specify the acoustic stimuli, CI stimulation parameters, and recording state.

      Acoustic stimuli: We now describe the acoustic stimulus set in the Methods, which references Insanally et al. (2016). Briefly, tones were pure sinusoids spanning frequencies from 1.4 to 32 kHz (half octave spaced), presented at 70 dB SPL with a duration of 50 ms with 2ms cosine-squared ramps and at a pseudorandom sequence of 1.25 Hz. These parameters are now updated in the methods under “Stimulus presentation for cortical sensory mapping in normal hearing rats.”

      CI stimulation parameters: CI stimulation used standard clinical-style monopolar mappings. We now specify in the Methods that pulses were biphasic, charge-balanced, with 8 µs interphase gaps and 25 µs /phase (total pulse width = 58 µs); stimulation rate was 900 pulses per second (pps); and current amplitude (and thus charge per phase) was set individually for each electrode based on its ECAP threshold. All stimulation levels were within normal and safe limits: charge densities remained below the Shannon limit and within the electrochemical “water window.”

      Loudness setting: In this study, CI stimuli were presented primarily at a single level—each electrode was stimulated at its ECAP threshold level for the tone-to-CI mapping experiments. We have added these details in the methods under the “Stimulus presentation for cortical sensory mapping in cochlear implanted rats” subsection.

      Recording state: All iEEG recordings reported in the manuscript were acute and performed under anesthesia. This is now stated explicitly at the start of the Methods section.

      (6) Plasticity and training effects warrant further consideration: although the manuscript reports no difference between naïve and trained rats, Figure 3 suggests greater across-trial variability for CI than NH that is not evident in the trained subset; examining relationships among behavioural performance, decoder performance, across-trial variability, and training duration would strengthen interpretation.

      We agree that plasticity and training effects are central questions for cochlear implant research and that iEEG is well suited to study how cortical representations evolve with CI use. However, the current dataset was collected mainly to compare cortical encoding of acoustic versus CI stimulation under matched, acute conditions (not necessarily after behavioral training with the implant, and we note that most studies of physiological responses to cochlear implant function in non-human species also do not incorporate aspects of training). All CI-evoked iEEG recordings were obtained immediately after implantation, before any CI-based behavioral training. As a result, any training effects reflected in the iEEG data can only arise from prior normal-hearing training, not from experience with CI stimuli themselves. Only a small subset of animals (N = 3 of 10) underwent behavioral training with cochlear implants, and their training histories (duration, performance levels, CI hardware status) are not uniform. This yields insufficient statistical power to meaningfully examine correlations among behavioral performance, decoder performance, across-trial variability, and training duration. While we note the reviewer’s observation that across-trial variability appears qualitatively different in the small, trained subset, we do not believe the current data justify strong conclusions about training-related plasticity.

      (7) Differentiating the CI rats stimulated directly or through the microphone of the speech processor -at least in the figures - would be useful to allow the reader to assess whether both stimulation strategies give rise to similar results.

      We agree that it is important to distinguish between rats stimulated directly via CI hardware and those stimulated acoustically through a speech processor. We now show in new Supplementary Figure 2, which animals received direct electrical stimulation and which were driven acoustically through the processor microphone. We also now plot tonotopic and cochleotopic maps for all CI animals in Supplementary Figure 3, with the stimulation mode indicated for each animal. As discussed in our response to comment #2 of Reviewer 3, we also provide validation that acoustic tones can be used to selectively drive individual electrodes via the speech processor. However, the sample sizes for the two stimulation strategies are small (N = 4 rats with direct CI stimulation, N = 3 rats with acoustic CI stimulation). For this reason, we have chosen not to draw strong statistical conclusions about differences between direct vs acoustic CI stimulation in the present manuscript.

      (8) Typographical error at the end of the introduction ("To this end we have designed and manufactured..."), and in the first paragraph of the Discussion ("...that both that...").”

      Thanks, we have updated the manuscript accordingly.

      (9) Inconsistent terminology: use a single form (e.g., "normal-hearing") throughout.

      Good suggestion, thanks. We have updated all main manuscript to only use normal-hearing. We found and changed two instances in which we used the acronym NH in lieu of normal-hearing, once early in the results section and once in the legend for Figure 3.

      (10) In Figure 3D (temporal), there appears to be an extra data point for the NH-trained group.

      Thank you for flagging this mis-labeling, which Reviewer 3 also pointed out. We have switched the appropriate data point in Figure 3D from ‘trained’ to ‘naïve’.

      (11) In Figure 4D, the yellow line is not defined; based on Figure 6D, it likely represents shuffled/chance performance and should be labeled accordingly (including beneath the chance line on the plots).

      We have updated Figure 6 to indicate that the yellow line does indeed reflect shuffled/chance.

      (12) Figure 8 would benefit from a control demonstrating that poor cross-modal decoding reflects train-test distribution differences rather than weak decoders (e.g., train on a subsample of NH and test on held-out NH), and from reporting decoding on raw ERP/HG features in addition to TCA-derived data.

      Good suggestion, thanks; we have now added this control. We agree that a positive control is necessary to show that poor tone→CI decoding reflects differences of underlying representations rather than a failure of the decoder or modeling approach. (Reviewer 2 raised the same point.)

      To validate our cross‑modal analysis pipeline, we re‑implemented the full procedure used in Figure 8, but instead of training on tone‑evoked responses and testing on CI‑evoked responses, we trained and tested on independent sets of tone‑evoked trials from the same animals (tone→tone). For each tone in each animal, we withheld 10 trials as a test set. Using the remaining trials, we fit the original TCA model to obtain spatial and temporal factors (Fig. 8A). We then fixed these factors and re‑optimized only the trial factors on the withheld tone‑evoked trials (Fig. 8B). The LDA decoder was trained on the trial factors from the original TCA fit and tested on the re‑optimized trial factors from the withheld trials, using the same classification pipeline as in the main analysis.

      As shown in the top panels of Figure 8C,D, this positive control yielded robust tone→tone generalization: predicted tone frequencies closely matched the actual tones, decoder performance was significantly above chance, and prediction errors were tightly clustered around the true stimulus, indicating that the decoder was tuned to tone frequency. In contrast, when we trained on tone‑evoked responses and tested on CI‑evoked responses, information transfer was markedly reduced (Fig. 8E-G).

      These results demonstrate that the TCA+decoder pipeline can reliably transfer information across independent tone‑evoked datasets, confirming that the method captures shared structure when it exists. The poor cross‑modal transfer between tone‑ and CI‑evoked activity therefore is unlikely to be due to a weak decoder or to a failure of the modeling pipeline, but instead reflects a genuine mismatch between CI and sound representations in auditory cortex. We have updated Figure 8 and the Results section to describe this positive control analysis and clarify the interpretation.

      (13) Perception and interpretation of signals are mentioned several times in the introduction, although perception is not explored in the manuscript (only neuronal processing). This might be confusing.

      We appreciate the need to distinguish between neuronal encoding and perception. We also feel we have been careful not to invoke relationships to perception when presenting analyses on iEEG measurements, but we did identify an opportunity to further clarify this distinction between neuronal processing and perception by adding text in the intro, as follows “for the auditory system to interpret patterns of evoked neural activity and inform downstream auditory areas.”

      (14) Figure 1C. Why is the performance of CI rats so much lower than what was previously published (Glennon et al., 2023)? Did the training duration change?

      The three animals that were behaviorally trained on the normal-hearing (pre-deafening) and cochlear implant task (post-deafening) are within the distribution of the full set of animals from Glennon et al. (2023). However, we note that for Glennon et al. (2023), as one of our behavioral criterion was days to d’ > 1, animals were trained daily until reaching that level and not included in the initial data set if they did not reach that level. However, as we were including animals in this study of iEEG responses that were not trained at all, we felt it appropriate to include this third animal as well, that was trained just for 3 days before recordings were made. The two other animals were trained for 9 and 13 days. We have now included this information in the methods.

      (15) The p-values = 0.5 should be given with an additional digit.

      We previously rounded to the nearest single decimal digit, for all p-values greater than 0.10. We have updated the figures and manuscript text to ensure precision at least to the second digit.

      Reviewer #2 (Recommendations for the authors):

      We thank the Reviewer for their thoughtful comments on our study.

      (1) Less noisy recording methods based on spike detection would provide stronger claims.

      We agree that spike recordings, particularly isolated single-unit activity, are powerful for testing hypotheses about sensory encoding in auditory cortex, and we plan to incorporate such approaches in future work. However, our decision to use iEEG arrays in the present study was deliberate and central to the scientific and translational goals of the project.

      First, iEEG and related population-level approaches such as scalp EEG (e.g., Lalor and Foxe, 2010; O’Sullivan et al., 2015) and fNIRS (e.g., Bortfeld et al., 2009; Peelle, 2017) are widely used in humans and have been highly successful in decoding sound- and speech-evoked responses, revealing fundamental principles of how sound and speech are encoded in the human brain. Because speech is uniquely human and cochlear implants are primarily designed to restore speech perception, aligning our recordings with clinically relevant, human-used modalities enhances the translational relevance of our work.

      Second, iEEG arrays provide distinct advantages over modern multi- and single-unit electrophysiology. Even with high-density probes, the spatial sampling of neuronal activity does not match the coverage of the 60-channel iEEG arrays used here, which span large extents of auditory cortex. One might instead consider optical methods such as calcium imaging to interrogate topographical encoding at single-neuron and mesoscale resolutions, as has been done in normal-hearing mice (Romero and Hight et al., 2019). However, calcium signals are intrinsically slow, limiting access to the temporal precision that is critical for CI encoding, and these tools are unlikely to be available in humans in the foreseeable future, substantially reducing their translational value.

      Using iEEG arrays, we show that CI-evoked responses are topographically organized, consistent with prior work (Klinke et al. 1999, Bierer and Middlebrooks 2002, Middlebrooks and Bierer 2002, including Adenis et al., 2024 now referenced in the manuscript). Our study extends these findings by exploiting simultaneous recordings across both spatial and temporal domains, which are essential for several key analyses (Figs. 3-8), including quantification of trial-by-trial variability, decoding of stimulus identity from single trials, and cross-modal comparisons between normal-hearing and CI-evoked iEEG responses.

      Thus, we believe that the strength of this study is due to, rather than in spite of, its use of iEEG arrays. This approach uniquely allows us to test hypotheses about CI encoding across cortical topography and time using a modality that is directly translatable to human research and clinical practice. In response to the reviewer’s concern, we have also (i) improved the statistical treatment of our data (by adopting linear mixed-effects models that incorporate both paired and unpaired observations), (ii) added additional positive controls (see response to comment #2), and (iii) collected new data that further validate our rodent CI model. Together, these additions strengthen the support for our conclusions while preserving the key advantages of the iEEG-based approach.

      (2) A positive control is necessary to claim the mismatch between CI and sound representations.

      We agree. We now have added a positive control specifically designed to validate our cross-modal analysis pipeline in our revised manuscript. As also suggested by Reviewer 1, the goal was to test whether our method can successfully transfer information when the training and test datasets are matched in modality (tone→tone), thereby ensuring that the observed failure of cross-modal transfer (tone→CI) is not an artifact of the analysis.

      To do this, we re-implemented the full pipeline used in Figure 8, but instead of training on tone-evoked responses and testing on CI-evoked responses, we trained and tested on independent sets of tone-evoked trials from the same animals. For each tone in each animal, we withheld 10 trials as a test set. Using the remaining trials, we fit the original TCA model to obtain spatial and temporal factors (Fig. 8A). We then fixed these factors and re-optimized only the trial factors on the withheld tone-evoked trials (Fig. 8B). The LDA decoder was trained on the trial factors from the original TCA fit and tested on the re-optimized trial factors from the withheld trials, using the same classification pipeline as elsewhere in the manuscript.

      As shown in the top panels of Figure 8C,D, this positive control yielded robust tone→tone generalization: predicted tone frequencies closely matched the actual tones, decoder performance was significantly above chance, and prediction errors were tightly clustered around the true stimulus, indicating that the decoder was tuned to tone frequency. In contrast, when we trained on tone-evoked responses and tested on CI-evoked responses, information transfer was markedly reduced and not different from shuffled controls (Fig. 8E-G).

      These results demonstrate that the TCA+decoder pipeline can reliably transfer information across independent tone-evoked datasets, confirming that the method captures shared structure when it exists. The poor cross-modal transfer between tone- and CI-evoked activity therefore cannot be attributed to a failure of the modeling pipeline but instead reflects a mismatch between CI and sound representations in auditory cortex. We have updated Figure 8, the methods, and the results section to include this new important analysis.

      Reviewer #3 (Recommendations for the authors):

      We thank reviewer 3’s appreciation for study design and the appropriateness of analyses taken. We also appreciate the recognition of noteworthiness, specifically that stimulus identity can be decoded on a single-trial basis and of the potential benefit of using central decoders in clinical settings.

      (1a) Animal heterogeneity: It is difficult to keep track of the animals used in this study, and some received a different protocol of stimulation (sounds through the speech processor vs. direct stimulation) and were also trained in a behavioral task using different target stimuli (4kHz vs. 22.6kHz, also no mention of the CI electrode used as a target).

      We have now clarified the animal cohorts and stimulation protocols in our revised manuscript. We added a new Supplementary Figure 2 that schematizes, for each animal if it underwent behavioral training with pure tones in the normal-hearing condition, if tone-evoked iEEG measurements were collected, if CI-evoked iEEG measurements were collected (and whether stimulation was direct or via the speech processor), and if it subsequently received CI-based behavioral training. Regarding the behavioral targets, we now specify in the Methods that for normal-hearing training, the target stimulus was a 22.6-kHz pure tone. For CI-trained animals, the target was either CI channel 3 (n = 2 rats) or CI channel 4 (n = 1 rat). Details about stimuli targets during behavior have been added to the methods section under “Behavioral training for tone and implant channel detection.”

      (1b) There is no comparison of the CI maps from rats tested with the speech processor and directly stimulated. How different were they? Was the frequency allocation of each electrode the same for each animal? Since data might already have intrinsic variability because of the grid placement, the mechanical deafening, and the cochlear implantation in each animal, such heterogeneity in the 'background' and stimulation protocol might blur the authors' results.

      Our study focuses on cortical encoding of single-channel CI stimulation, so it is indeed important to ensure that the stimuli are effectively delivered by a single electrode, regardless of whether they are driven acoustically via the speech processor or by direct electrical stimulation.

      Stimulation mode and frequency allocation: The project began with single-channel stimulation achieved by presenting pure tones to the speech processor (N=3 animals) and later transitioned to direct programmatic control of individual electrodes (N=4 animals) to simplify the experimental setup. In both cases, the goal was to activate only one CI channel at a time.

      For the programming speech-processor animals, the validation protocol described in Glennon et al. (2023) is as follows:

      - Set the number of active channels in the processor to 1 (the clinical default is 8) to avoid spectral spread across electrodes.

      - Disabled all additional signal-processing strategies (e.g., Scan, ASC, ADRO, SNR-NR, WNR).

      - Used customized frequency allocation tables that mapped narrow frequency bands to individual electrodes, as shown in Glennon et al., 2023, Extended Data Fig. 2.

      To confirm that a given tone drove only the intended electrode, we recorded tone-evoked electrodograms—measurements of the output at each electrode—and verified that only the targeted channel was active (Glennon et al., 2023, Extended Data Fig. 2). Thus, although the initial CI drive was acoustic, the effective stimulation at the array was restricted to a single electrode with a well-defined frequency allocation.

      For the direct-stimulation animals, we used the same underlying frequency allocations to choose which electrode to stimulate, but the pulses were delivered programmatically rather than via the speech processor. In both modes, the center frequency associated with each electrode was therefore defined consistently across animals, and stimulation was confined to one channel at a time.

      Comparison of maps across stimulation modes: We now explicitly indicate the stimulation mode (speech-processor vs direct) for each CI animal in Supplementary Figure 2 and plot the maps for all animals in Supplementary Figure 3. Qualitatively, the spatial organization of CI-evoked maps is similar across the two stimulation strategies; we do not observe systematic differences in map structure that would suggest large biases introduced by the stimulation mode. However, the sample sizes for each group are small (N = 3 speech-processor, N = 4 direct). For this reason, we have not performed formal between-mode statistics and instead treat stimulation mode as a source of minor heterogeneity, alongside inevitable variability from grid placement, mechanical deafening, and cochlear insertion. Given the electrodogram validation (Glennon et al., 2023, Extended Data Fig. 2) and consistent frequency allocation tables, we are confident that both approaches produce single-channel activation with comparable effective frequency assignments.

      (1c) The number of animals used is also confusing. The authors report 7 NH and 7 CI animals (14 total), 4 NH and 3 CI were trained before being implanted (so 3 naïve NH and 4 naïve CI remain). Figure 1C reports that only 3 trained NH performed with the CI (let us call them 3 NH->CI). But then Figure 1E reports only 1 trained NH->CI and only 1 trained NH and 3 naïve NH that got implanted later. On the other hand, Figure 1E reports only 1 true naïve CI animal, the 3 others being naïve NH that got implanted. For the sake of clarity, I would encourage the authors to provide a timeline of the procedures/stimulation protocols coupled with a schematic distribution of the animals.

      To address this, we have added a new Supplementary Figure 2 that provides, for each individual animal a chronological timeline (NH recordings, deafening, implantation, CI recordings); if it was behaviorally trained in the NH condition, the CI condition, or both; if CI stimulation was delivered via the speech processor or by direct electrical stimulation; and which stimulus conditions (tone-evoked iEEG, CI-evoked iEEG) were collected. This schematic makes it clear how the reported totals arise (7 NH and 7 CI for iEEG; 4 NH-trained and 3 CI-trained behaviorally) and shows which specific animals contribute to each panel in Figure 1 and to the later iEEG analyses. We now reference Supplementary Figure 2 in the Results when introducing the cohorts to guide readers through animal accounting.

      (2a) Methods and statistics: Deafening is only mechanical, with no direct or postmortem proof that deafening was complete. The authors cite previous studies, but that would have been a good control to have since mechanical deafening isn't as accepted as the chemical deafening, like Neomycin, especially when some of your animals were stimulated with pure tones through the speech processor.”

      We agree that rigorous verification of deafening is essential, particularly when some CI animals are driven acoustically through the speech processor. Ototoxic approaches (e.g., systemic or local neomycin) are one established method, but their effectiveness can be sensitive to dose and delivery, and they introduce systemic side-effects that can complicate long-term survival and recovery.

      Our laboratory has used the mechanical deafening procedure since it was first described in King et al. (2016) and more recently in Glennon et al. (2023). In King et al., mechanical and ototoxic methods were combined, and we found that ototoxic methods provided no more additional robustness in deafening compared to mechanical lesion. Instead, the additional time required for ototoxic drug application reduced survival times in what was already a very complex and long surgical procedure for bilateral deafening and unilateral cochlear implantation.

      In Glennon et al. (2023) we intentionally employed mechanical-only deafening to minimize side-effects while still achieving profound hearing loss in implanted animals. Glennon et al. (2023) provides an extensive validation of this mechanical-only protocol under the same surgical and experimental conditions as the present study. As we mentioned in our response to comment #3a of Referee 1, we assessed deafness through three measures:

      Histology: In N=4 deafened animals, inner hair cell loss was ~50% and outer hair cell loss was near complete at almost 100% in all animals.

      Physiology: For N=14 rats, acoustic ABRs were substantial before deafening but statistically similar to baseline noise after deafening.

      Behavior: For N=16 deafened rats, behavioral performance with implant on was d′: 1.7±0.1, but when implant was turned off in a subset of sessions, performance dropped to chance (d′: −0.05±0.1, P < 0.0001).

      This convergent anatomical, physiological, and behavioral evidence demonstrates that the mechanical procedure produces profound deafness, with no functionally relevant residual hearing at levels ≥90 dB SPL. Also as we mentioned in response to comment #3a of Referee 1, we believe that the behavioral criterion is most essential and also least common in the literature. Because the tones used to drive the speech processor in the current study were presented at 70 dB SPL, we have no reason to believe that residual acoustic hearing contributed to any of the CI-evoked responses we report.

      We now cite these validation data explicitly in the methods under the section “Bilateral sensorineural hearing loss” as follows “(mechanical only, as described and validated in Glennon et al. 2023)” to make clear why we consider the mechanical-only approach sufficient for ensuring deafness in the present experiments.

      (2b) What motivated the selection of 15 Principal Components for the PCA? That might need to be justified, maybe by scree plot or variance plot (Eigen Values or CEV), as if too many PCs are selected, you are at risk of losing information. Side comment for TCA: why is it important that the number of latent factors exceeds the number of tones or stimuli? Is there a way to justify this statement?

      We thank the reviewer for raising this point. Our choice of 15 components/latent factors was motivated by both theoretical and empirical considerations, which are now made explicit in the manuscript.

      For the PCA analyses, we selected 15 principal components for two reasons. First, because our decoder must discriminate between 10 tone conditions, we reasoned that providing at least as many dimensions as stimuli would be beneficial, while also allowing for the possibility that some components may carry little or no stimulus-selective information. We therefore chose a modest number of components that exceeded the number of tones (10) but avoided unnecessarily high dimensionality. Second, we empirically examined the variance explained as a function of the number of components. As shown in the new scree plots (Supplemental Fig. 4A), the cumulative variance explained enters a near-linear, low-slope regime beyond ~15 PCs, indicating diminishing returns for including additional components. Thus, 15 PCs capture a substantial fraction of the stimulus-related variance while minimizing the risk of overfitting and retaining a consistent dimensionality across animals.

      For the TCA analyses, we used 15 latent factors to match the dimensionality used in PCA and to ensure that the latent space was sufficiently flexible to represent the 10 tone conditions without being under-parameterized. In practice, increasing the number of TCA components reduces reconstruction error (Williams et al., 2018), but with diminishing improvement beyond a certain point. We therefore systematically evaluated model error as a function of the number of latent factors and found that error decreased rapidly up to ~15 components and then plateaued (Supplemental Fig. 4B). This pattern parallels the PCA scree plots and supports 15 as a reasonable trade-off between model flexibility and parsimony.

      We have updated the Results clarify these choices, as follows “The number of components (15) was chosen based on PCA scree plots (Supplemental Fig. 4A), which showed that explained variance entered a near‑linear, low‑slope regime beyond this point demonstrating a similar plateau in reconstruction error (Supplemental Fig. 4B).”

      (2c) Legend of Figure 2E, J states that a Student's paired t-test was used, meaning that only the 'linked' points of the graph were used (thus, comparing only animals that got tested NH then implanted). This is usually the same across the manuscript. Why not include all the points with an unpaired t-test? Otherwise, why are all the points plotted if they serve no purpose? This choice should be justified.

      We agree with this concern, which was also raised by Reviewer 1. We have revised our statistical approach accordingly in our revised manuscript. In the original submission, we used paired t-tests when animals contributed both normal-hearing (NH) and CI data, which meant that animals with only NH or only CI measurements were excluded from those comparisons even though they were shown in the plots.

      To address this, we have re-analyzed all normal-hearing vs. CI comparisons using linear mixed-effects models that include both paired and unpaired data within a single framework. This approach ensures that every plotted data point contributes to the statistical tests, properly accounts for within-animal dependence when both conditions are present, and avoids the loss of power that would arise from either paired-only or purely unpaired tests.

      The mixed-effects results are consistent with our original interpretations, with two comparisons becoming significant in the updated analysis: Fig. 2E (p = 0.048) and Fig. 6F (p = 0.027). We have updated the Results and figure legends to describe the use of mixed-effects models and to report these revised p-values. Together with the new tonotopy and cochleotopy analyses described above, these changes strengthen the statistical support for our conclusions without altering the overall interpretation of the data.

      (2d) Side comment: There are inconsistencies on the bar plots of Figure 6C (Missing a purple point) and Figure 3D (Temporal has 3 purple points).

      Thank you for flagging this mis-labeling (which Reviewer 1 also noticed). We have correctly updated the appropriate data point from trained to naive for Fig. 3D and from naive to trained for Fig. 6C.

      (3a) Pure tones and CI-evoked responses maps: It is the reviewer's understanding that Figure 2 is an averaged representation for all animals. Why is the tonotopic shift so dim for ERPs? The averaged maps aren't very convincing. How were the gradients on an animal-to-animal basis since Figure 2D is only an example animal? Also, everything has been evaluated at 70dB, where selectivity might not be best. It would have been easier to follow the tonotopic gradient at the CFs where contrasts are higher.

      We agree that the strength and interpretation of tonotopy/cochleotopy in our iEEG data needed to be presented more clearly. Reviewer 1 raised closely related concerns, and we have substantially expanded the analyses and explanations in response. Here we highlight the points that address your specific questions.

      Single-animal vs. averaged maps: We included both exemplar maps and population summaries in Figure 2. The panels analogous to Figure 2D show single-animal best-frequency (BF) or best-channel maps; these were chosen because they exhibit clear, interpretable gradients. In the exemplar shown, there is a local high-frequency (HF) region along the medial edge of the array that transitions to lower frequencies toward the rostral edge. For CI-evoked best-channel maps in the same animal, we observe a parallel pattern in which basal electrodes (e.g., electrode 8, representing higher frequencies) occupy the HF region and apical electrodes (e.g., electrode 1, lower frequencies) occupy the LF region.

      Averaged ERP maps, by contrast, necessarily blur some of this structure because iEEG is a summed field potential and animal-to-animal differences in array placement, cochlear insertion depth, and anatomy introduce variability. We have softened the language in the text to reflect that ERP-based tonotopy is coarse and weaker at the population level, while emphasizing that robust gradients are evident in single animals and in HG-based measures.

      Quantitative assessment across animals: To move beyond visual impressions, we added quantitative analyses that mirror those used in Romero and Hight et al. (2020) for calcium imaging data (Romero and Hight et al. 2020 and Fig. 2). For each map we computed local tonotopic gradient vectors at every pixel and summarized their magnitude/direction on a unit circle, then compared the mean vector strength to shuffled maps. Applied to our BF and best-channel maps, this analysis shows that both are significantly more ordered than shuffled controls (p < 10<sup>-10</sup>), indicating that the maps are tonotopic/cochleotopic rather than random, despite the apparent dimness of the gradients in some averaged ERP plots. These new results are described in the revised manuscript and shown in Romero and Hight et al. 2020 and Fig. 2.

      Effect of intensity (70 dB SPL) and “dim” gradients: We agree that stimulus level influences the apparent sharpness of tonotopy. Higher intensities tend to broaden tuning and compress the dynamic range of BF maps. As we now discuss in more detail (adapted from our response to Reviewer 1), tones were presented at 70 dB SPL, so we expect maps to emphasize mid-frequency regions (around 8 kHz) and to show somewhat broader tuning than maps derived at threshold. For CI stimulation, we used ECAP thresholds to set intensity, which is effective in our preparation because animals can robustly discriminate individual electrodes and these electrodes evoke clear cortical activity (King et al., 2015; Glennon et al., 2023).

      In summary, we clarified which panels in Figure 2 show single-animal exemplars vs population summaries, added quantitative analyses demonstrating spatial correlations are greater for adjacent stimuli compared to far-apart stimuli, and expanded the discussion of how recording modality and stimulus level influence the visibility of tonotopic gradients. These changes are intended to make the evidence for tonotopy/cochleotopy in our iEEG data (and its limitations) more transparent.

      (3b) Since new experiments might not be available, it is the reviewer's suggestion to add a supplementary figure showing a couple of animal examples following the format of Figures 2A and 2C that have more contrasted gradients to strengthen the group data. In the case of the CI-evoked responses map, this might also provide another argument to dismiss the potential monopolar smearing.

      Good suggestion, thanks. We now include a new Supplementary Figure 3 that shows additional single-animal examples for both tone-evoked and CI-evoked maps, following the same format as Figure 2C.

      Regarding monopolar stimulation, we agree that monopolar configurations are expected to be less spatially specific than bipolar or multipolar modes because current returns to an extracochlear reference electrode, potentially broadening the spread of excitation. We nevertheless chose monopolar stimulation because it is the predominant clinical configuration in human CI users and therefore most relevant for translational purposes. We acquired ECAP measurements of peripheral (spatial and temporal) tuning via a forward masking paradigm and demonstrate that monopolar is effectively tuned (Supplemental Fig. 2). Together with additional single-animal maps in Supplementary Figure 3, together with our vector-strength analysis (Romero and Hight et al. 2020 and Fig. 2), demonstrate that even under acute monopolar stimulation we observe structured cochleotopic organization in cortex, rather than the fully smeared patterns one might expect if monopolar spread completely dominated.

      We also note that all CI-evoked iEEG measurements were made acutely, immediately after implantation and before any CI-based behavioral experience. It is possible that with longer-term use and plasticity, cortical cochleotopy could become sharper than what we observe here under acute conditions. In this sense, our data provide a conservative baseline showing that even at the earliest stages of CI use, monopolar stimulation already engages tonotopically selective regions of auditory cortex. A longitudinal comparison of acute versus chronic maps would be an interesting direction for future work but is beyond the scope of the current study.

      (3c) Side comments: The legends of Figures 2D and 2I should mention that this is an animal example and not group data, as the rest of the figures are group data.

      Thank you for this suggestion to improve figure clarity. We have updated all of our figures, where appropriate, to indicate whether data are single or groups of animals.

      (3d) In general, some of the legends should be revised because they are sometimes too "strong". As an example, Figure 3B, D legend states: "Variability of iEEG measurements across trials (root mean square, rms) was consistently higher for cochlear implant-evoked compared to tone-evoked activity", despite three of the statistical tests being non-significant. The manuscript is correct, on the other hand.

      Good point. We revised the legend for Figure 3 to be consistent with the figure and the manuscript.

      (3e) The example spatial map given in Figure 3A for CI might not be the best choice since it is showing a pretty reliable trial-by-trial response, while your group data proves the opposite.

      We understand the reviewer’s concern and agree that the exemplar CI map in Figure 3A appears relatively reliable on a trial-by-trial basis. This example was chosen deliberately from an animal in which we had both NH- and CI-evoked iEEG recordings, so that the reader could visually compare the two conditions within the same preparation. In this animal, as in the group data, the differences between NH and CI trial-by-trial responses are subtle rather than dramatic.

      Our group-level analysis shows that the RMS error across trials is consistently higher for CI-evoked than for NH-evoked responses, but the absolute differences are small (< 0.1) and relatively uniform across animals. The spatial maps plotted in Figure 3A are representative of this pattern: both conditions show reasonably robust evoked responses, with CI responses nonetheless showing slightly greater variability. To avoid implying a stronger qualitative difference than is supported by the data, we have revised the text to emphasize that (i) CI-evoked responses remain clearly detectable on single trials, and (ii) the key effect is a small but consistent increase in variability across animals, as captured by the RMS error metrics, “We noted that the differences were qualitatively subtle (Fig. 3A, right panel), they were consistent across animals (Fig. 3B).”

      (4a) Decoders for CI stimulation Regarding CI stimulation, Pearson's correlations were truncated at a spacing of 5 electrodes. Likewise, none of the LDA classifiers show prediction for channels past CI-6. Again, that choice should be justified, or the missing channels should be presented.

      We truncated the correlation between electrodes at 5 because beyond that, the estimated means are significantly noisy. These estimated means are noisy because the number of data are significantly reduced, also significantly increasing the standard error. For example, for the maximum stimulus spacing, the number of pairwise correlations is at maximum the number of animals tested (i.e., N=7). We believe it’s important to be transparent, so we have included the non-truncated version of the figure here in this public review (Author response image 3). We leave the figures in the manuscript untouched but have updated the Figure 2 legend justify this selection of data.

      Author response image 3.

      Expanded figures for spatial correlations and LDA performance. A) The same data from manuscript Figure 2 are re-plotted but with expanded x-axes to include up to 4.5 octaves and 7 channels. Due to the smaller numbers of data at these points, the estimates for the mean spatial correlations are noisier. In all cases, the mean correlations are significantly higher for the first data point compared to the last 3 (NH, ERP p<0.001; NH, HG p<0.001; CI, ERP p=0.005; and CI, HG p=0.39, linear mixed effects models). B) The same data from manuscript figure 4 are re-plotted but with expanded x-axes to include up to ±3.5 octaves and ±6 channels.

      (4b) Finally, retrained PCA-LDA on spatial-only and temporal-only for CI are absent in Figure 3D. Since the authors were pretty consistent in showing both NH and CI alongside in the rest of the paper, it would be coherent to add the CI counterpart to Figure 3D, or maybe with a supplementary figure.

      We agree that consistency can be improved by including classifiers for CI-evoked measurements, though presumably for Fig. 6C and not Fig. 3D. Figure 6 has been updated accordingly.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates the impact of Pink1 loss on glial function and neuronal health in a Drosophila model, highlighting the role of mitochondria-organelle contacts and key genes such as Ccz1, Vps13, Mon1, and Rab7. The work provides insights into cellular processes underlying neurodegenerative diseases, with a focus on glia-neuron interactions. While the findings are promising, the study lacks critical controls, detailed mechanistic evidence, and explanatory figures to strengthen its claims.

      Strengths:

      (1) The study addresses an important topic in neuroscience, exploring the mechanisms of Pink1 loss, which has implications for Parkinson's disease and neurodegeneration.

      (2) The focus on mitochondria-organelle contacts and their regulation by Rab7-mediated pathways is novel and provides a potential mechanism for neuronal dysfunction.

      (3) The identification of key genes (Ccz1, Vps13, Mon1, Rab7) and their potential roles in Pink1-related pathways adds valuable knowledge to the field.

      (4) The manuscript uses a combination of genetic tools, Drosophila models, and functional assays to approach the problem from multiple angles.

      Weaknesses:

      (1) Specificity of Mz-Gal4: The study lacks validation of Mz-Gal4 specificity, as it may also drive expression in a few neurons or other types of glia. Additional control experiments using nls-GFP with Elav, Repo, or Draper antibody staining or alternative glial drivers would be helpful.

      We have addressed this issue of Gal4 driver specificity based on new experiments in the revised manuscript.

      (2) DLG staining is central to the story but is not well-supported by high-resolution Z-stack imaging, which should be included in the supplementary figures.

      We have included these in the supplement.

      (3) The manuscript does not confirm whether the candidate RNAi (Ccz1, Vps13, Mon1, Rab7) directly influence Rab7-mediated membrane trafficking or mitochondria-lysosome contacts in Pink1 mutants.

      This is indeed the case. These more mechanistic experiments were not yet performed.

      (4) Using ERG as a readout for EG effects in the antenna is not a direct or appropriate assay. Alternative functional assays relevant to antenna glia should be considered.

      We made the assumption that ensheating glial function is conserved across brain regions and now make this explicit in the reworded manuscript.

      (5) A graphical explanation of the interactions and functions of the candidate genes in Pink1 KO mutants is missing. This would greatly enhance the manuscript's clarity.

      We have included such a scheme in the new manuscript.

      (6) The study lacks details on sample sizes, effect sizes, and reproducibility, which are necessary for robust conclusions.

      We have included these essential data in the reworked document.

      (7) There are repeated words on page 3 ("olfactory Olfactory Receptor Neurons") and a lack of explanation in Figure 3C regarding the most up-regulated and down-regulated genes and the significance of large red dots.

      We have included the requested information.

      Reviewer #2 (Public review):

      Summary:

      This study proposes a novel role for ensheathing glia (EG) in a Pink1-model of Parkinson's disease and shows that this cell population exibits the highest number of DEG in a pre-symptomatic stage. In the olfactory system, there seems to be morphological changes in this cell-type that resembles an 'activated' state and the authors further show that the neuronal loss of Pink1 is responsible for this defect. The authors go on to show that manipulation of Pink1 in EG also leads to some defects in the visual system and in the dopaminergic neurons (DAN) that innervate the mushroom body (MB), and performed a screen based on the 'on-transient' defect of the ERG to identify potential genes that may modulate the function of EG in synaptic regulation. They focus on several genes related to Rab7/Vps13, and performed some additional experiments in the visual system and MB to propose the role of vesicle/lipid trafficking in EG as a important factor for PD pathogenesis.

      Strengths:

      The study proposes functional and mechanistic connections between several genes that have been linked to PD (PINK1, VPS13A/C). I feel that the data presented in Figure 1 and Fig3A-C are performed with rigor and are convincing/novel. The selection of Drosophila to study the questions is also a strength and the lab has extensive experiences in this field and model organism.

      Weaknesses:

      There is one fundamental concern I have with the genetic experiments performed in this paper (especially in Fig 3D and Fig4, see major issue #1), and I feel that there is a bit of a disconnect between the EG 'activation' phenotype the author show in the olfactory system and the other two neuronal systems (visual system, MB DAN) that the authors investigate see major issue #2). Also, there are quite a bit of information that is not provided in the manuscript (see major issues #3 and #4), which makes me difficult to judge the rigor and interpretation of several experiments.

      Major Concern #1: A number of lines used in this study are referred to as "RNAi" lines but when I look at the actual genotypes of reagents listed in the table in the METHODS section, many are actually NOT RNAi lines. Quite a few lines, including lines that the authors use as RNAi against Ccz1, Rab7 and Mon1, are gRNA lines for the TKO (TRiP-CRISPR knockout) system. While these reagents can theoretically knock-out these genes in somatic cells if used in combination with UAS-Cas9, there is no mention that UAS-Cas9 was used in this work throughout the manuscript. Hence, when these lines are just crossed to GAL4 with or without the Pink1 mutant, they shouldn't be having any effects. Similarly, the strongest hit from their screen was a TOE (TRiP-CRISPR Over Expression) gRNA against PIG-A, which could allow overexpression of PIG-A if there is a UAS-dCas9::VP64. However, I also do not see any mention that such activator was introduced into the crossing scheme. Considering that 3 of the 4 'hits' from their screen are not RNAi lines, I am quite skeptical of the study. Similarly, except for Vps13, all reagents used in Fig4 are TKO gRNA lines. Therefore, if this experiment was conducted without an UAS-Cas9, most of the data shown here are problematic. Also, note that several of the 'RNAi' lines listed in the Table in the METHODS section are actually MiMIC alleles. While some MiMIC lines could function as strong LOF alleles (if they are inserted in the exon or in an intron of the gene in the same orientation as the gene), some of the lines are not expected to affect gene function (e.g. FASN2 and CG17712, MiMICs are in introns and face the opposite orientation). Hence, the rationale of including these reagents in the screen doesn't make much sense. The description of the modifier screen should be much more detailed in the RESULTS and METHODS section and if the UAS-Cas9/dCas9::VP64 transgenes were not introduced when the TKO/TOE reagents were utilized, what can be concluded?

      In addition, for the 4 genes that the authors further study in Fig4, there are many other reagents that the authors can use, including mutant alleles, previously characterized RNAi lines (e.g. Vps13) and dominant negative/constitute active lines (e.g. especially for Rab7). The authors should validate their results with independent reagents to really convincingly show that the same conclusions can be drawn for the Vps13/Rab7 related genes since this is the key takeaway message of this paper.

      Also, they do not show whether the manipulation of these genes in a wild-type background (they only show what happens in Pink1 mutants) affect ERG and MB DAN synapse morphology. If these manipulations alone dramatically affect these phenotypes, it would be very difficult to interpret their data.

      We sincerely thank the reviewer for spotting this major oversight regarding the use of the TKO (TRiP-CRISPR knockout) and TOE (TRiP-CRISPR Over Expression) systems and the MiMIC alleles. As the reviewer pointed out, these lines were not used as intended, therefore our results and conclusions regarding the genetic interactions between Pink1 and several genes (PIG-A, Rab7, Ccz1, CG10646, Mon1, FASN2, CG17712), are incorrect and based on a technical mistake. These results were removed from the manuscript. While our mistake compromises the data regarding PIG-A, Rab7, Ccz1, CG10646, Mon1, FASN2, CG17712, it does not affect the results and conclusions for most of the genes of the screening and for Vps13 where we did use RNAi lines.

      Also, in the reworked manuscript, we provide additional evidence that modulation of vesicle trafficking proteins involved in mitochondria–endoplasmic reticulum (ER) membrane interactions, such as Vps13 and Vps35, influences neuronal function and rescues Pink1 mutant phenotypes when selectively downregulated in EG.

      Major Concern #2: In Figure 1, the authors show some morphological evidence that EG are 'activated' in Pink1 mutants, but whether the same phenomenon occurs in the visual system and in the MB is not shown. Since all of the studies in Fig3D and Fig4 are done in the visual system and MB, it is not clear whether the visual system and MB phenotypes are related to 'activation' of EG.

      Also, in the RNA-seq data in Fig1A and Fig3C, is there any molecular evidence that EG are indeed 'activated'? The only evidence that the authors show to state that EG are 'activated' in young Pink1 null animals is based on increased CD8::GFP staining in the olfactory system.

      The authors cannot draw a strong conclusion that indeed EG are 'activated' based on these data (e.g. perhaps the expression level of CD8::GFP is just increased). Additional evidence that the EG are 'activated' could be provided by looking at the increase in Draper intensity (as reported by Doherty et al. and MacDonald et al. that the authors cite), not only in the olfactory system, but also in the visual system and in the MB. It would also be informative if the authors can look at morphology of the EG in the visual system and MB to convincingly that the data shown in Fig4 is relevant to EG 'activation'.

      In line with the identification of DEG across the ensheating glia cluster in our single cell sequencing (where we did not distinguish between EG of different brain regions) we made the assumption that EG-(dys) function is consistent in the Pink1 mutant and conserved across brain regions. Nonetheless, to make clear that we did not consistently analyze EG morphology in the different brain regions that we probed in functional assays, we added a note in the manuscript. Furthermore, we also toned down our conclusion that the EG in Pink1 mutants are in an activated state: we note the similarity in phenotype in Pink1 mutants and situations of neuronal damage (where EG are activated) but added that the phenotype in Pink1 mutants may also be the result of the mere upregulation of GFP expression/fluorescence.

      Major Concern #3: In Fig3, there is no clear explanation why they focus on the ON transients and ignore the OFF transients, and also why the difference in the depolarization is not quantified in Fig4.

      We included this explanation in the reworked manuscript: In the Drosophila ERG, the sustained depolarization primarily reflects phototransduction in photoreceptors (and is defective when photoreceptors degenerate), whereas the ON and OFF transients arise from second-order lamina neurons and are widely used as readouts of signal transfer. We wanted to assess function and focused on the ON transient because in general it provides an onset-locked, more robust readout of function (Vilinsky & Johnson, 2012).

      Major Concern #4: While the authors claim that mz709-GAL4 is a EG specific driver, do the authors know that this is indeed true in the tissues and stages that are studied here? The Ito et al,. paper that is cited in the METHOD section has only looked at the expression of this reporter in embryonic and larval stages. The authors need to that the authors should validate their findings with an additional EG specific driver and/or provide additional data that mz709-GAL4 is indeed specific to EG in the adult fly brain and eye. If mz709-GAL4 is expressed in other cell-types, the interpretation of many of the data in this paper becomes quite questionable. I believe the data in Fig3B is suggesting that mz709-GAL4 is indeed specific to glia cells and not expressed in neurons, but whether this driver is truly specific to EG (and not in other glial types), especially in the visual system (including the lamina as well as in the eye), is not obvious.

      We labelled animals that express UAS-HisTag-eGFP (used also in our paper) under control of MZ709-Gal4 with anti-Elav (a neuronal marker) and find no significant overlap (see below “recommendation for authors”), consistent with MZ709-Gal4 not driving expression in neurons. This is consistent with previous published work: Indeed, MZ709-Gal4 has been amply used in adult flies and shown to be ensheating glia-specific (Doherty et al., 2009; Li et al., 2023; Sehgal et al.,2018). In the lamina neuropil of the Drosophila eye, MZ709-Gal4 is expressed in the marginal glia (Stenesen et al., 2019) which are neuropil-associated glia and are equivalent to generic ensheathing glia (Kremer et al., 2017). MZ709-Gal4 is also expressed also in satellite glia (Stenesen et al., 2019), but these glia enwrap the cell bodies of the lamina neurons and not the neuropil where synapses reside.

      Recommendations for the authors:

      Reviewing Editor Comments:

      We strongly encourage you to very carefully edit this manuscript. The reviewers made many probing comments that you should consider carefully.

      Reviewer #1 (Recommendations for the authors):

      (1) Validate the specificity of Mz-Gal4 by performing experiments with nls-GFP and Elav antibody staining to ensure there is no neuronal overlap. Additionally, consider using alternative glial-specific drivers, such as Repo-Gal4 or WG-Gal4, to confirm the findings.

      We expressed HisTag-eGFP (used also in our paper) under control of MZ709-Gal4 and labelled fly brains with anti-Elav (a neuronal marker). We do not observe significant overlap between the labels indicating MZ709-Gal4 does not express Gal4 in neurons (Supplementary figure 1).

      As indicated, these observations are consistent with previous published work. MZ709-Gal4 has been amply used in adult flies and shown to be ensheating glia-specific (Doherty et al., 2009; Li et al., 2023; Sehgal et al., 2018; Stahl et al., 2018). In the lamina neuropil of the Drosophila eye, MZ709-Gal4 is expressed in the marginal glia (Stenesen et al., 2019) which are neuropil-associated glia and are equivalent to generic ensheathing glia (Kremer et al., 2017). MZ709-Gal4 is also expressed also in satellite glia (Stenesen et al., 2019), but these glia enwrap the cell bodies of the lamina neurons and not the neuropil where synapses reside.

      (2) Include high-resolution Z-stack imaging of DLG staining to strengthen the assessment of synaptic integrity and ensure the robustness of the conclusions. These images should be added to either the main or supplementary figures.

      We included 2 supplementary figures (2 and 3) showing Z stacks that were used to delineate regions of interest at the MBs for the quantification of dopaminergic neuron afferents invasion. Our approach is identical to the one we used in Kaempf et al. 2026 (Kaempf et al., 2026).

      (3) Demonstrate whether the candidate RNAi (Ccz1, Vps13, Mon1, Rab7) directly influence Rab7-mediated membrane trafficking or mitochondria-lysosome contacts in Pink1 mutants. Use an appropriate method to confirm changes in organelle contacts in response to the RNAi treatments.

      Ccz1, Mon1 and Rab 7 were removed due to the technical mistake we made. We did confirm and maintain that Vps35 and Vps13 downregulation in EG rescues neuronal defects in Pink1 mutants. In the reworked manuscript we present a possible mechanism that involves the role of Vps35 and Vps13 in regulating ER-mitochondrial contacts, in line with our previous work (Valadas et al., 2018), while not ruling out possible other mechanisms.

      (4) Provide an alternative functional assay or evidence to support the use of ERG as a readout for EG effects in the antenna. Consider using a more direct assay relevant to antenna glia function.

      We agree that a more direct functional assay of antennal glia would be a nice addition (e.g., single-sensillum recordings or glial/ORN Ca<sup>2+</sup> imaging). However, implementing such assays would require new experimental pipelines and substantial additional data generation that is beyond our current ability and the scope of this revision.

      (5) Add a graphical illustration explaining the proposed mechanism of how Ccz1, Vps13, Mon1, and Rab7 function in Pink1 KO mutants, highlighting their interactions and roles within specific cell types.

      We included a schematic of our working model in Figure 5.

      (6) Clarify Figure 3C by explaining the most up-regulated and down-regulated genes and the significance of the large red dots. This will enhance the interpretability of the data.

      We expanded the legend to this figure: The large red dots represent the genes that rescue Pink1<sup>KO-WS</sup> phenotype when downregulated, the dark green dots are the 50 top most deregulated genes (magnitude of deregulation) in EG in Pink1<sup>KO-WS</sup> compared to controls, while the light green dots represent whole the genes detected in our cell-type specific transcriptomic experiment.

      (7) Correct repeated words on page 3 ("olfactory Olfactory Receptor Neurons") for clarity and consistency.

      Of course, sorry for this.

      (8) Ensure that sample sizes, effect sizes, and the number of replicates are explicitly stated for all experiments. This information is essential for evaluating the robustness and reproducibility of the findings.

      We made sure we consistently added all this information in the revised manuscript.

      (9) Verify and ensure that all data, reagents, and code used in the study are accessible and appropriately documented, in adherence with eLife's publishing policies.

      We made sure all data, reagents and code are available and/or properly described.

      By addressing these recommendations, the authors will significantly improve the clarity, rigor, and reproducibility of the manuscript.

      Reviewer #2 (Recommendations for the authors):

      Minor Points.

      (1) All figures seem to lack titles.

      We fixed this error.

      (2) In the abstract, the authors say that Rab7 and Vps13 are mutated in PD patients but I couldn't find the reference/information for Rab7 (the authors do refer to papers that linked VPS13A/C variants to PD but no mention about RAB7A/B being linked to PD). Please discuss this in the paper or modify the abstract accordingly.

      We removed this statement for rab7 from the paper.

      (3) When referring to the human gene, Pink1 should be written as PINK1 according to the HGNC nomenclature rules.

      We made this change.

      (4) The authors say Vps13 has two mammalian orthologs but actually it has four (VPS13A/B/C/D). I guess two of the four is linked to PD so the authors should modify there statement to reflect this.

      This is a misinterpretation of what we meant and we have clarified our intention: Drosophila possesses 3 paralogues of Vps13 - Vps13, Vps13B, and Vps13D - which we also detected in our screening (Neuman et al., 2025; Velayos-Baeza et al., 2004; Vonk et al., 2017). Among these Vps13 is most similar to human VPS13A and VPS13C (Hanna et al., 2023; McEwan & Ryan, 2022).

      (5) The abbreviation 'CNS' is used in the first page of the intro but I don't see it being spelled out as "central nervous system".

      We have spelled out central nervous system in the first page of the introduction.

      (6) On the top of page 5, the authors state that they confirmed that the 'synaptic area of DAN show a decrease in aged (25 days) animals' but data is not shown. If they want to make a statement like this, I believe such data should be included in supplemental data. Since the phenotype in the aged animal is not relevant to this study, one could remove this statement regarding the aged animals if they prefer not to show the data.

      The decreased synaptic area of DAN in 25-day old Pink1 mutants is shown in figure 2C-D of the manuscript and is consistent with data shown in (Kaempf et al., 2026).

    1. Author response:

      The following is the authors’ response to the original reviews.

      (1) We bioinformatically examined the repeat compositions of MLSs (Figure 3B), which clearly indicated that all MLSs are composed of repetitive sequences to a much greater extent than the rest of the genome.

      (2) We confirmed the blockage of chromosome breakage by the 4R-CBS mutations using a telomere-anchored PCR assay (Figure 5C-E).

      (3) We examined the effect of the 4R-CBS mutations on the expression of genes encoded in 4R-MDS by RNA-seq (Figure 9). This analysis unexpectedly revealed that gene expression from 4R-MDS is not significantly affected in the mutants, allowing us to extend our discussion.

      (4) We added two authors, Alix Lemoine and Tomoko Noto, who performed the experiments for these revisions.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, Nagao and Mochizuki examine the fate of germline chromosome ends during somatic genome differentiation in the ciliate Tetrahymena thermophila. During sexual reproduction, a new somatic genome is created from a zygotic, germline-derived genome by extensive programmed DNA elimination events. It has been known for some time that the termini of the germline chromosomes are eliminated, but the exact process and kinetics of the elimination events have not been thoroughly investigated. The authors first use germline-specific telomere probes to show that the loss of these chromosome ends occurs with similar timing as other DNA elimination events. By comparative analysis of the assembled germline and somatic genomes, the authors find that the ends of each of the germline chromosomes are composed of a few hundred kilobases of micronuclear limited sequences (MLS) that are removed starting around 14 hours after the start of conjugation, which initiates sexual development. They then develop an in situ hybridization assay to track the fate of one end of chromosome 4 while simultaneously following the adjacent macronuclear destined sequence (MDS) retained in the new somatic genome. This allows the authors to more clearly show that these adjacent chromosomal segments are initially amplified in the developing genome before the terminal MLS is eliminated. Finally, they mutate the chromosome breakage sequence (CBS) that normally separates the MLS terminus from the adjacent MDS region, to show that strains that develop with only one mutant chromosome can produce viable sexual progeny, but it appears that both the MLS and the MDS from the mutant chromosome are lost. If both chromosome copies have the CBS mutation, the cells arrest during development and do not eliminate many germline-limited sequences and fail to produce viable progeny. Overall, this study provides many new insights into the fate of germline chromosome ends during somatic genome remodeling and suggests extensive coordination of different DNA elimination events in Tetrahymena.

      Strengths:

      Overall, the experiments were well executed with appropriate controls. The findings are generally robust. Importantly, the study provides several novel findings. First, the authors provide a fairly comprehensive characterization of the size of the MLS at the end of each germline chromosome. I'm not sure whether this has been published elsewhere. Second, the authors develop a novel method to study the fate of chromosome termini during development and use it to conclusively track the elimination of these termini. Third, the authors show that the elimination of these termini appears to occur concurrently with most other DNA elimination events during somatic genome differentiation. And fourth, the authors show that failure to separate these eliminated sequences from the normally retained chromosome alters the fate of these adjacent MDS and the loss of the cells' ability to produce viable progeny.

      Weaknesses:

      It appears the authors did extensive analysis of the MLS chromosome ends, but did not provide too much information related to their composition. If this has not been published elsewhere, it would be useful to describe the proportion of unique and repetitive sequences and provide more information about the general composition of the chromosome ends. Such information would help the reader understand the nature of these MLS and how they may or may not differ from other eliminated sequences.

      We now calculated the proportions of unique and repetitive sequences for each MLS, and these data are included in Figure 3B and described in the main text of the revised manuscript. A more comprehensive analysis of chromosome-end composition, including detailed characterization in the context of the complete MIC genome assembly, is beyond the scope of the current study and will be presented in a future publication.

      Although the development of the novel FISH probes for large chromosome ends allowed for these novel discoveries, the signal in several images was visible, but often quite faint. I'm not sure there is anything the authors could do to improve the signal-to-noise ratio, but one needs to stare at the images carefully to understand the findings.

      We have submitted higher-resolution images for the revised manuscript, which we believe much improve the visibility of faint signals.

      One main weakness in the opinion of this reviewer is that the authors did very little to understand why, when a terminal MLS and the adjacent MDS fail to get separated because of failure in chromosome breakage, both segments are eliminated. The authors propose that possibly essential genes in the MDS get silenced, and the resulting lack of gene expression is the issue, but this and other possibilities were not tested. The study would provide more mechanistic insight if they had tried to assess whether the MDS on the CBS mutant chromosome becomes enriched in silencing modifications (e.g., H3K9me3). Alternatively, the authors could have examined changes in gene expression for some of the loci on the neighbouring MDS.

      The 4R-CBS mutation causes two distinct defects that should be considered separately: (1) co-elimination of 4R-MLS and the adjacent 4R-MDS during uniparental transmission of the 4R-CBS mutation; and (2) a global block of DNA elimination during biparental transmission of the 4R-CBS mutation.

      For the first defect, 4R-MLS and 4R-MDS may simply co-segregate into the nuclear compartment where DNA elimination occurs when the chromosome break that normally separates 4R-MLS from 4R-MDS is blocked. In this scenario, no additional process, such as spreading of scnRNA production, heterochromatin formation, or gene silencing, would be required to induce co-elimination. This point was not clearly stated in the previous manuscript, and we have now added a discussion of it to the revised manuscript.

      The possibility of gene silencing within 4R-MDS was raised as a potential explanation for the second defect. To test this possibility, we performed RNA-seq analysis of wild-type and 4R-CBS mutant cells to determine whether gene expression from 4R-MDS is affected by mutations at 4R-CBS. Contrary to our expectations, we found that genes in 4R-MDS are not significantly down-regulated in 4R-CBS mutant cells compared with other genes. This result suggests that the DNA elimination defect in these cells cannot be explained by silencing of genes located within 4R-MDS. We have added these RNA-seq data to Figure 9 and described them in the Results section. We have also revised the Discussion to propose alternative possibilities that may guide future investigations.

      The other main weakness is that since the authors only mutated the end of one germline chromosome, it is not clear whether the elimination of the MDS adjacent to the terminal MLS on chromosome 4 when the CBS is mutated is a general phenomenon, i.e., would happen at all chromosome ends, or is unique to the situation at Chromosome 4R. Knowing whether it is a general phenomenon or not would provide important insight into the authors' findings.

      As was described in the manuscript, the short (CBS = 15 nt) target within AT-rich and repetitive regions prevent designing gRNAs specifically targeting some of the chromosome end CBSs. We tried to mutate the CBS sequences of the left end of the chromosome 3 (3L) and the left end of the chromosome 5 (5L) by the strategy we used to mutate 4R-CBS but failed. Therefore, to systematically mutate other chromosome-end CBSs, we need to establish a different strategy, such as combining template-based repairing to CRISPR-induced DSB. We have explained this technical limitation and stated that “Our data support a critical role for 4R-CBS in separating 4R-MLS from 4R-MDS, but it remains unclear whether all MIC chromosome ends are strictly CBS-dependent for their elimination.” in Discussion (Page 12).

      Reviewer #2 (Public review):

      Summary:

      Nagao and Mochizuki investigated how the germline (MIC) telomere was removed during programmed genome rearrangement in the developing somatic nucleus (MAC). Using an optimized oligo-FISH procedure, the authors demonstrated that MIC telomeres were co-eliminated with a large region of MIC-limited sequences (MLS) demarcated on the opposite side by a sub-telomeric chromosome breakage site (CBS). This conclusion was corroborated by the latest assembly of the Tetrahymena MIC genome. They further employed CRISPR-Cas9 mutagenesis to disrupt a specific sub-telomeric CBS (4R-CBS). In uniparental progeny (mutant X WT), DNA elimination of the sub-telomeric MLS was not affected, but the adjacent MAC-destined sequence (MDS) may be co-eliminated. However, in biparental progeny (mutant X mutant), global DNA elimination was arrested, revealing previously unrecognized connections between chromosome breakage and DNA elimination. It also paves the way for future studies into the underlying molecular mechanisms. The work is rigorous, well-controlled, and offers important insights into how eukaryotic genomes demarcate genic regions (retained DNA) and regions derived from transposable elements (TE; eliminated DNA) during differentiation. The identification of chromosome breakage sequences as barriers preventing the spread of silencing (and ultimately, DNA elimination) from TE-derived regions into functional somatic genes is a key conceptual contribution.

      Strengths:

      New method development: Oligo-FISH in Tetrahymena. This allows high-resolution visualization of critical genome rearrangement events during MIC-to-MAC differentiation. This method will be a very powerful tool in this area of study.

      Integration of cytological and genomic data. The conclusion is strongly supported by both analyses.

      Rigorous genetic analysis of the role played by 4R-CBS in separating the fate of sub-telomeric MLS (elimination) and MDS (retention). DNA elimination in ciliates has long been regarded as an extreme form of gene silencing. Now, chromosome breakage sequences can be viewed as an extreme form of gene insulators.

      Weaknesses:

      The finding of global disruption of DNA elimination in 4R-CBS mutant progeny is highly intriguing, but it's mostly presented as a hypothesis in the Discussion. The authors propose that the failure to separate MLS from MDS allows aberrant heterochromatin spreading from the former into the latter, potentially silencing genes required for DNA elimination itself. While supported by prior literature on heterochromatin feedback loops, the specific targets silenced are not identified. While results from ChIP-seq and small RNA-seq can greatly strengthen the paper, the reviewer understands that direct molecular characterization may be beyond the scope of the current work.

      As mentioned in our reply to Reviewer #1’s comment above, we performed RNA-seq on wild-type and 4R-CBS mutant cells at 13.5 hpm and 15 hpm and found that genes in 4R-MDS are not significantly downregulated in 4R-CBS mutant cells (Figure 9), suggesting that the DNA elimination defect in these cells cannot be explained by aberrant heterochromatin spreading. Therefore, the link between the chromosome break at 4R-CBS and general DNA elimination remains elusive and will be a very interesting subject for our future research. We have added these results and revised the discussion in the manuscript.

      Reviewer #3 (Public review):

      Programmed DNA elimination (PDE) is a process that removes a substantial amount of genomic DNA during development. While it contradicts the genome constancy rule, an increasing number of organisms have been found to undergo PDE, indicating its potential biological function. Single-cell ciliates have been used as a prominent model system for studying PDE, providing important mechanistic insights into this process. Many of those studies have focused on the excision of internally eliminated sequences (IES) and the subsequent repair using non-homologous end joining (NHEJ). These studies have led to the identification of small RNAs that mark retained or eliminated regions and the transposons that generate double-strand breaks.

      In this manuscript, Nagao and Mochizuki examined the other type of breaks in ciliates that were healed with telomere addition. They specifically focused on the sequences at the ends of the germline (MIC) chromosomes, which have received relatively less attention due to the technical challenges associated with the highly repetitive nature of the sequences. The authors used the Tetrahymena model and developed a set of new tools. They used a novel FISH strategy that enables the distinction between germline and somatic telomeres, as well as the retained and eliminated DNA near the chromosome ends. This allows them to track these sequences at the cellular level throughout the development process, where PDE occurs. They also analyzed the more comprehensive germline and somatic genomes and determined at the sequence level the loss of subtelomeric and telomere sequences at all chromosome ends. Their result is reminiscent of the PDE observed in nematodes, where all germline chromosome ends are removed and remodeled. Thus, the finding connects two independent PDE systems, a protozoan and a metazoan, and suggests the convergent evolution of chromosome end removal and remodeling in PDE.

      The majority of sites (8/10) at the junctions of retained and eliminated DNA at the chromosome ends contain a chromosome breakage sequence (CBS). The authors created a set of mutants that modify the CBS at the ends of chromosome 4R. CBS regions are challenging for CRISPR due to their AT-rich sequences, making the creation of the 4R-CBS mutants a significant breakthrough. They used the FISH assay to determine if PDE still occurs in these mutant strains with compromised CBS. Surprisingly, they found that instead of blocking PDE, its adjacent retained DNA is now eliminated, suggesting a co-elimination event when the breakage is impaired. Furthermore, in biparental mutant crosses, no PDE occurred, and no viable progeny were produced, indicating that the removal of chromosome ends is crucial for proper PDE and sexual progeny development. Overall, the work demonstrates a critical role for 4R-CBS in separating retained and eliminated DNA.

      We appreciate Reviewer 3’s assessment.

      Recommendations for the authors:

      Reviewing Editor Comments:

      All reviewers agree that this study makes an important contribution to the field; however, they also offered several suggestions for how the manuscript could be improved. In particular, we draw your attention to the comments from Reviewer #1, who suggests that the manuscript could benefit from additional information on the general composition of germline chromosome ends, where available.

      As noted in our response to Reviewer #1 in the Public Reviews above, we have included an analysis of the fraction of repetitive sequences for each MLS as Figure 3B in the revised manuscript, highlighting the highly repetitive nature of MLSs compared with the rest of the genome.

      Reviewer #1 (Recommendations for the authors):

      As mentioned in the weaknesses section, the authors could provide more information regarding the nature of the sequences that make up the terminal MLS. There have been reports that these are highly repetitive; is that the case? Also, did the authors identify common repeats that are not internal to mic chromosomes that could be used to track all terminal segments of the five chromosomes? This would complement their mic-telomere probe.

      As noted in our response to Reviewer #1’s Public Review above, we have added an analysis of the fraction of repetitive sequences for each MLS as Figure 3B in the revised manuscript, which confirms that MLSs are highly repetitive.

      Apart from the moderately conserved Telomere Associated Sequence (TAS), described by Kirk and Blackburn (1995) and of unknown function, we were unable to identify any obvious shared repeats unique to MLSs that could support the development of pan-MLS-specific probes.

      One major weakness is that the authors did little to determine the cause of the elimination of the adjacent MDS along the 4R-MLS when the CBS was mutated. It would really improve the study if the authors could show that:

      (1) Gene expression of genes on the MDS is reduced in 4r-CBS mutant progeny.

      (2) Heterochromatin modifications are unexpectedly acquired on the MDS in mutants relative to wild-type chromosomes.

      (3) Do scnRNA specific to the MDS region appear in the mutant progeny during development, but not in wild-type crosses?

      Any data that would help support the authors' hypothesis regarding how the MDS region is eliminated when the CBS is mutant would definitely strengthen the conclusions of the study.

      As noted in our response to Reviewer #1’s Public Review above, we performed RNA-seq on wild-type and 4R-CBS mutant cells at 13.5 hpm and 15 hpm. Our analysis showed that genes within the 4R-MDS are not significantly downregulated in 4R-CBS mutant cells (Figure 9), suggesting that the DNA elimination defect in these cells cannot be attributed to aberrant heterochromatin spreading. Therefore, the connection between the chromosome break at 4R-CBS and general DNA elimination remains unclear and represents an important avenue for future investigation. We have incorporated these results and revised the discussion accordingly in the updated manuscript.

      The other main weakness is that by mutating the CBS of only one chromosome arm, one can't know whether the loss of the MDS with the MLS in the mutants is generalizable for all chromosome arms or is unique to 4R. The authors noted that they were unable to make any other mutated CBSs. Another way to try to get to this question is to try to rescue the mutant by inserting a new CBS into the 4R arm such that some MLS remains linked to the 4R-MDS and see whether removing the mic telomere is the issue, or would a block of MLS attached to the 4R-MDS be sufficient to cause its elimination. I'm not sure where to exactly put the new CBS, but worth thinking about.

      To introduce a new CBS into 4R-MLS, we would need to insert a CBS-containing construct into the MIC by homologous recombination during conjugation and then select engineered transformants using a drug resistance marker expressed from the derived MAC. However, because 4R-MLS is still eliminated in the progeny of 4R-CBS mutants, the introduced marker would be lost from the MAC even if homologous recombination were successful. Therefore, although the strategy suggested by this reviewer is very interesting, several technical innovations are required to make such experiments feasible, leaving this approach for a future project.

      It seems somewhat curious that the mutation of the CBS completely blocks nuclear development. In Paramecium, the failure to complete internal DNA elimination events can lead to alternative telomere addition. The caveat being that, in Paramecium, telomere addition appears more promiscuous than in Tetrahymena. It would be helpful to know how absolute the failure to produce progeny is in these mutants. Is it zero progeny in 10<sup>6</sup>, 10<sup>7</sup>, 10<sup>8</sup> ..... mated cells? Can the authors provide a possible lowest possible frequency?

      The viability tests were performed using bulk mating of 2.5 × 10<sup>4</sup> cells for each cross. Because ~70-80% of mating pairs complete the conjugation process and produce exconjugants under our standard culture conditions, and because we did not detect any 6-mp-resistant progeny from MUT x MUT crosses, we estimate that the probability of obtaining viable progeny in these crosses was less than 1 progeny per ~2 × 10<sup>4</sup> mating pairs. The number of cells used for the viability assay is described in the “Viability Test of Sexual Progeny” section of Materials and Methods and the estimated frequency of progeny production from the mutants has been mentioned in Results section in the revised manuscript.

      The one implication of the study is that chromosome breakage and DNA elimination, two different events, are coupled. In most mutants that block scnRNA-directed DNA elimination, both IES excision and chromosome breakage occur. In the study by McDaniel, SL. et al (2016). DRH1, a p68-related RNA helicase, is required for chromosome breakage in Tetrahymena. Biology Open pii: bio.021576. doi: 10.1242/bio.021576, germline knockouts of DRH1 could complete IES excision, but not chromosome breakage, indicating that the processes can be uncoupled. It may be useful for the authors to discuss this previous work in relation to their finding that failure in chromosome breakage can lead to DNA elimination of neighboring sequences.

      So far, DRH1 is the only gene reported to be required for chromosome breakage without affecting DNA elimination in Tetrahymena. However, McDaniel SL et al. (2016) examined chromosome breakage at only two CBSs (distinct from 4R-CBS), and thus it remains unclear how broadly chromosome breakage, including that at 4R-CBS, is affected in the absence of DRH1. In addition, McDaniel SL et al. (2016) assessed DNA elimination at three different IESs using PCR, whereas our study examined elimination of the repetitive Tlr1 transposon using FISH. Therefore, without further analysis of the similarities and differences in chromosome breakage and DNA elimination phenotypes between DRH1 knockout cells and 4R-CBS mutants, it is difficult to draw meaningful conclusions. Accordingly, we have limited ourselves to stating the following in the Discussion of the revised manuscript: “Moreover, chromosome breakage can be inhibited without disrupting DNA elimination, as shown in cells lacking zygotic expression of the p68-like RNA helicase Drh1 (McDaniel et al., 2016).”

      Minor corrections:

      Page 7, line 3: the text "......inducing chromosome break" should either be "......inducing chromosome breaks" or "......inducing a chromosome break".

      Corrected as “inducing a chromosome break”.

      Page 13, line 13: "......large block...." should be "......large blocks......".

      Corrected as suggested.

      Reviewer #2 (Recommendations for the authors):

      The authors can experimentally validate that chromosome breakage at 4R-CBS is indeed disrupted by the mutations. A PCR-based assay testing de novo telomere addition is a standard tool. In addition, MLS-linked telomere should only appear transiently during conjugation in WT cells.

      Because it was previously unknown whether de novo telomere addition occurs at the ends of MLSs upon chromosome breakage, we tested this using a PCR-based assay. We detected telomere-added chromosome ends of 4R-MLS and 3L-MLS, which were undetectable until 10.5 hpm, appeared at 12 hpm, and gradually decreased by 18 hpm in wild-type cells (WT × WT cross). Importantly, the appearance of the telomere-added 4R-MLS end, but not the 3L-MLS end, was blocked in 4R-CBS mutants (Mut x Mut crosses), strongly supporting that the 4R-CBS mutations specifically disrupt chromosome breakage at 4R-CBS. These new data are shown in Figure 5C–E and described in the Results section.

      The high FISH background during conjugation may be caused by the abundant presence of dsRNA, which is resistant to RNase A treatment but may be degraded by RNase III.

      The high FISH background was observed in the parental MAC at 9 and 12 hpm (Figure 2, 4, and S2) where dsRNA accumulation was not detected in the previous studies (Woo et al. 2016; Shehzada et al. 2024). In contrast, the MIC at 3 hpm and the new MAC at 9 and 12 hpm, where strong dsRNA accumulation was detected, showed much weaker background FISH signals (Figure 2, 4, and S2). Therefore, we believe that dsRNA is not the main cause of the high FISH background.

      It is likely that the long MIC telomere is treated as IES and targeted for DNA elimination. Indeed, telomere-specific scnRNA is abundantly produced during conjugation (http://www.ncbi.nlm.nih.gov/pubmed/19460867).

      We have cited the suggested literature and the following description has been added in Discussion to relate the reported telomere-derived scnRNAs to the abundant scnRNAs produced from MIC chromosomal ends: “In addition, telomere-complementary scnRNAs were reported to be produced specifically during conjugation (Cao et al. 2009).”

      Global disruption of DNA elimination may be a direct effect (DNA excision machinery affected) or indirect (unrepaired DSB and checkpoint activation).

      It has been reported that unrepaired DSBs caused by loss of Ku80 (Tku80) do not block DNA elimination in Tetrahymena (Lin et al. 2012). Therefore, checkpoint activation by unrepaired DSBs, if it occurs, is unlikely to explain the DNA elimination defect observed in the progeny of 4R-CBS mutants. Nonetheless, this direct-versus-indirect issue would be relevant when considering whether disruption of specific 4R-MDS-encoded genes in 4R-CBS mutants could cause the DNA elimination defect. Our new RNA-seq analysis, however, suggests that this possibility is unlikely. Therefore, we did not add further discussion of this direct-versus-indirect issue.

      Minor points:

      The zoom-in boxes in most images are barely visible.

      We have modified the zoom-in boxes to make them clearer.

      Page 13: scnRNA precursors (Cai et al., 2025) (Cai et al., in press). Is it one paper or two?

      They are two papers and the latter was published reacently. We have updated the citation.

      Reviewer #3 (Recommendations for the authors):

      The manuscript is well-written, with clear data, thoughtful discussion, and concise presentation. I have only a few minor comments below.

      For Figure 4 and others, the right panel shows the stats and percentages, with positive and negative labels. It's a bit confusing at first glance. I think it can be clarified what positive and negative mean in the legend.

      The legends of Figure 4, Figure 6 and Supplementary Figure S2, have been modified as “The presence (Positive) or absence (Negative) of the 4R-MLS FISH signal in new MAC (An) in 50 cells per time point was examined.”

      The quality of the FISH images is low at their current resolution. It is difficult to get a clear view.

      In the initial version, some images were in low resolution when we combined them into a single pdf file for review. In the revised manuscript, the images have been replaced with high-resolution images.

      The co-elimination of neighboring 4R-MDS when 4R-CBS is mutated, can this be viewed as a fail-safe mechanism to ensure the elimination of the chromosome ends? Regardless, the result begs the question of the significance of end removal and remodeling of PDE. Some speculations in the discussion might be helpful.

      Because the neighboring 4R-MDS contains approximately 100 predicted genes, its co-elimination would likely be too risky to evolve as a fail-safe mechanism for ensuring chromosome-end elimination in every generation. Instead, we interpret this as an erroneous process that can still be compensated for through endoreplication of the remaining, normally processed 4R-MDS from the non-mutated copy.

      We further speculate that the connection between chromosome breakage at 4R-CBS and the essential PDE process may serve as an evolutionary pressure to preserve the 4R-CBS locus in a chromosome breakage-competent state. We have added the following discussion to the revised manuscript (Page 15): “The observed link between chromosome breakage at 4R-CBS and the essential DNA elimination process may reflect the biological significance of MLSs and the importance of their removal from the MAC. Coupling these processes may have evolved as a mechanism to ensure that only functional chromosome-end CBS loci are preferentially transmitted to future generations.”

      Figure 1, legend, line 3, "the sexual reproduction process", do you mean "the sexual reproduction proceeds or initiates"?

      We meant “conjugation” = “the sexual reproduction process”. To make this clearer, we have revised the legend as “conjugation, which is the sexual reproduction process of Tetrahymena”.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors propose that HSV-1 infection degrades the class I histone deacetylases HDAC1 and HDAC2. The MDM2 E3 ubiquitin ligase from the DNA damage response pathway is responsible for ubiquitinating these HDACs that are subsequently degraded via proteasomes. The authors hypothesize that HDAC degradation will cause hyperacetylation of viral chromatin and enable viral gene transcription.

      Strengths:

      The ubiquitination of HDAC1 & HDAC2 by Mdm2 and the mapping studies are clear.

      Comments on revised version:

      The authors enhanced their manuscript by more supportive data and providing clarification and the necessary corrections. However, a few more issues pertain:

      (1) In Figure 4j at 2 h post-infection we typically see the input virus and not progeny virus production. The input seems to have about 1-log difference that is expected to impact the results.

      We sincerely appreciate the reviewer's valuable comments regarding the timing notation. It should be noted that the "2 h" indicated in Figure 4j does not refer to two hours after the start of viral infection, but rather to two hours following medium replacement—after the virus has completed adsorption and internalization at 37°C (typically taking 2 hours), with this moment defined as the new time zero point (t = 0 h). Thus, this corresponds to approximately 4 hours post-infection (4 hpi). All subsequent sampling time points (4, 6, 12, and 24 h) are consistently defined according to this same system. This temporal framework aligns with previous studies: Nobe et al. (mBio 2025; DOI: 10.1128/mbio.00280-25) have clearly demonstrated that newly generated viral particles can be detected as early as 4 hours after HSV-1 infection, supporting the possibility of early progeny virus production at this time point in our experiment. We have accordingly revised the figure legend for Figure 4j to explicitly state the time reference ("t = 0 h defined as time of medium replacement post-adsorption") and added detailed procedural descriptions in the Methods section regarding adsorption, medium change, and sample collection time points to ensure clarity and reproducibility of the timing protocol.

      (2) Figs 1A, 1E, 2H it seems unclear why ICP4 becomes detectable at 12 h post-infection in HeLa cells? How about other a-genes? How about other cells? ICP4 is typically detectable within 2-3 h post-infection.

      We sincerely appreciate the valuable comments provided by the reviewers. Regarding the observation that ICP4 was detected only after 12 hours post-infection in HeLa cells, we re-evaluated our experimental conditions and reviewed relevant literature. The results indicate that at a higher multiplicity of infection (MOI = 5), ICP4 can indeed be reliably detected in HeLa cells as early as 2 hours post-infection (Author response image 1). Notably, Fouad S. El-Mayet et al. reported that under MOI = 1, ICP4 could not be detected until 8 hours after HSV-1 infection of mouse neuroblastoma Neuro-2A cells Figure 5A (Fouad S. El-Mayet et al., Antiviral Research, 2024, DOI: 10.1016/j.antiviral.2024.105870), although their early protein VP16 showed positive expression as early as 4 hours post-infection. This time difference is closely related to cell type: Neuro-2A is a highly susceptible neuronal cell line for HSV-1, exhibiting significantly faster viral gene expression kinetics compared to epithelial-derived HeLa cells. In contrast, HeLa cells are human cervical cancer epithelial cells with relatively low efficiency in initial transcriptional activation of HSV-1 and higher baseline expression levels of endogenous antiviral factors (such as interferon-stimulated genes), which may lead to a marked delay in the expression of early immediate-early genes like ICP4.

      Author response image 1.

      (3) In responses 2-2, Fig 5K: An infection without transfection has not been included. This is important to understand kinetics of infection in transfected cells.

      We sincerely appreciate the reviewer's insightful identification of this critical oversight. In all relevant experiments, we have strictly included empty vector transfection controls—serving as a baseline reference for each transfection group to eliminate potential influences from the transfection procedure itself and the vector background on viral replication, gene expression, and signaling pathways. The failure to clearly label this control in previous figure legends and main figures was indeed an omission in our presentation; we have now fully addressed this in the revised manuscript: all figures involving transfections (including Figures 3L, 3M, 5K, etc.) now explicitly indicate the "empty vector" control, and we have added detailed explanations in the figure legends and methods section regarding its role as an internal transfection control and procedural comparator. Once again, we thank the reviewer for their high level of professionalism in helping us enhance the completeness and scientific rigor of our data presentation.

      (4) Why HDAC1 with deleted NES does not accumulate or looks like it is degraded? Why then ICP4 does not accumulate?

      We sincerely apologize for the lack of clear labeling of the FLAG-HDAC1 ΔNES protein band in Author response image 2. This omission may have led reviewers to misinterpret its expression level as abnormal. After re-evaluation and improved annotation, Author response image 2 now clearly indicates the FLAG-HDAC1 ΔNES band its migration position corresponds to the expected molecular weight (slightly smaller than wild-type FLAG-HDAC1), and the band intensity is comparable to that of the empty vector and wild-type groups, indicating stable intracellular expression of this mutant protein without significant degradation. Therefore, its inhibitory effect on HSV-1 replication is not due to protein instability, but rather results from subcellular localization defects caused by the loss of nuclear export signal (NES): the ΔNES mutation causes HDAC1 to abnormally retain within the nucleus, ultimately leading to significant downregulation of ICP4 transcription and impaired protein accumulation.

      Author response image 2.

      Reviewer #2 (Public review):

      Summary:

      The authors discovered that HDAC1/2 are degraded in HSV-1 and PRV infections. They attempted to establish a new mechanism by which HDAC1/2 are translocated to the cytoplasm to be degraded in HSV-1 infection, and the degradation causes changes in histone acetylation to affect the DDR pathway.

      Strengths:

      (1) Interesting findings of HDAC1/2 degradation during HSV-1 and PRV infection, and it may impact more than the virology field.

      (2) Significant work to identify the ubiquitin site in HDAC1/2 and K63 linkage.

      Comments on revised version:

      The authors added experiments to address the previous comments. The added knockdown and overexpression experiments provided sufficient support for the proposed mechanism. The conclusions are now strengthened. However, a few essential controls are still missing.

      (1) Figure 3K: How does the expression level of Flag-HDAC1 variants compare to the endogenous HDAC1 level? The stripe probed by Flag antibody should be reprobed by HDAC1 antibody. Also, how does the K74R mutant affect histone acetylation? Moreover, the numbers between the panels are hard to read and have not been explained.

      We sincerely thank the reviewers for their insightful and constructive feedback. In response to the comment on Figure 3K, we performed antibody re-probing of the Flag-immunoprecipitated or Flag-immunoblotted membranes with a validated HDAC1-specific antibody. Consistent with robust transfection and expression, both wild-type Flag-HDAC1 and its mutants including K74R exhibited markedly elevated total HDAC1 protein levels relative to vector control, confirming efficient exogenous expression and protein stability. To directly assess functional consequences, we evaluated global histone acetylation status in parallel samples and found that the K74R mutant induces significantly greater deacetylation than wild-type Flag-HDAC1, as demonstrated by pronounced reductions in H3K56ac and H4K8 acetylation levels. Finally, to improve clarity and readability, we have revised the lane annotations in Figure 3K—increasing font size, enhancing contrast, and ensuring consistent alignment—and fully documented these modifications in the updated figure legend.

      (2) Figure 3M and 3L: DNA transfection per se frequently stimulates cell reactions that inhibit HSV-1 replication. Is the HSV-1 only sample transfected by empty vector or untransfected?

      We sincerely appreciate the reviewer's insightful identification of this critical oversight. In all relevant experiments, we have strictly included empty vector transfection controls serving as a baseline reference for each transfection group to eliminate potential influences from the transfection procedure itself and the vector background on viral replication, gene expression, and signaling pathways. The failure to clearly label this control in previous figure legends and main figures was indeed an omission in our presentation; we have now fully addressed this in the revised manuscript: all figures involving transfections (including Figures 3L, 3M, 5K, etc.) now explicitly indicate the "empty vector" control, and we have added detailed explanations in the figure legends and methods section regarding its role as an internal transfection control and procedural comparator. Once again, we thank the reviewer for their high level of professionalism in helping us enhance the completeness and scientific rigor of our data presentation.

      (3) Figure 4G-4J: What is the MDM2 knockdown efficiency?

      During the construction of the MDM2 knockdown cell lines, we first systematically validated the knockdown efficiency by qRT-PCR. As shown in Figure 4A, compared to the control group (shCtrl), MDM2 mRNA levels were reduced by approximately 60% in shMDM2 cells, and protein expression also showed a corresponding significant decrease, confirming that the cell line had been successfully established and exhibited stable gene silencing effects.

      (4) Figure 5F and line 400-401: "thereby preventing HDAC1 degradation-markedly impaired HSV-1 replication (Fig. 5F)." However, viral replication is not demonstrated in Figure 5F.

      We sincerely appreciate the reviewer for pointing out the error in the figure legend numbering. Upon verification, the experimental data referred to in lines 400–401 of the original text and in Figure 5F actually correspond to the revised new Figure 5J. We apologize for failing to update the figure references in the main text during the revision process due to an oversight. We have now uniformly corrected all relevant descriptions in the text to "Figure 5J" and conducted a comprehensive review of all figure numbers, table numbers, and cross-references throughout the manuscript to confirm there are no other similar errors.

      (5) Figure 5K: also need a control of empty vector. Furthermore, how does the HDAC1 ΔNES expression affect histone acetylation and DDR responses?

      We sincerely thank the reviewers for their thoughtful and constructive feedback on Figure 5K. With regard to the empty vector control: all pertinent experiments in this study were performed with rigorous inclusion of an appropriate empty vector control (pCMV-Flag or its isogenic backbone), serving as the definitive negative control. The prior absence of this control in the figure representation was unintentional and reflects an oversight in data presentation—not in experimental design—and we offer our sincere apologies. We have now incorporated the empty vector control bands into Figure 5K and revised the figure legend to explicitly identify and describe this control. In addition, per the reviewers’ recommendation, we conducted a comprehensive assessment of HDAC1 ΔNES function, specifically examining its impact on global histone acetylation and canonical DNA damage response (DDR) activation. Quantitative immunoblotting and immunofluorescence analyses revealed that HDAC1 ΔNES expression leads to significantly greater reduction in H3K56ac and H4K8 acetylation compared with wild-type HDAC1. Moreover, upon induction of DNA damage, HDAC1 ΔNES-expressing cells exhibit attenuated DDR signaling, evidenced by diminished γH2AX focus formation, reduced CHK2 phosphorylation (p-CHK2), and blunted p53 stabilization and activation consistent with impaired DDR initiation or propagation (see Author response image 3). Collectively, these data indicate that nuclear retention of HDAC1 due to NES deletion not only potentiates its chromatin-targeted deacetylase activity but also contributes to suppression of DDR signaling, likely through epigenetic modulation of damage-sensing chromatin domains.

      Author response image 3.

      (6) Statements listed below are better moved to discussion after all data being presented. They are quite a stretch when looking at each figure by itself.

      (i) Line 268-270: "Together, these findings indicate that HSV-1 selectively degrades class I HDACs, resulting in widespread histone hyperacetylation that fosters a chromatin state conducive to viral replication". ----may be okay for a statement.

      (ii) Line 291-292: "providing initial evidence that HSV-1 infection promotes DDR activation through downregulation of HDAC1 expression"

      (iii) Line 331-333: "Together, these results indicate that HSV-1 infection promotes K63-linked polyubiquitination of HDAC1/2 at conserved lysine residues, ultimately leading to their proteasomal degradation."

      (iv) Line 334-336 is a repeated sentence.

      We sincerely thank the reviewers for their thoughtful and constructive feedback. As noted, statements of mechanistic interpretation are not appropriate in the Results section; accordingly, we have relocated all such statements to the Discussion section. Furthermore, we have conducted a comprehensive line-by-line review of the manuscript to ensure that (i) every mechanistic inference is directly supported by experimental data presented in the Results, and (ii) integrative interpretations particularly those linking molecular observations to broader biological implications are confined exclusively to the Discussion.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) The central concept of geomapping as a broadly applicable strategy is wonderfully supported by the 17 successes documented in the paper. While this is actually, of course, a strength, the study does not include a comparative analysis across multiple sites with varying sampling outcomes for different bacterial types, which would be necessary to validate this claim more generally.

      We thank the reviewer for the point, and it is well taken. We addressed this below, where we give a full discussion.

      (2) Some elements, such as beta diversity comparisons and the metagenomics analysis of viral dark matter, would benefit from additional statistical analysis and clearer context.

      The reviewer is quite correct as to the importance of bringing statistical analysis to our metagenomic analysis. To that end, we performed statistical analysis on our metagenomic datasets. We performed statistical analysis on our metagenomic datasets. We approached this using MetaPop to analyze viral metagenomic sequence data at the interpopulation (macrodiversity) level. MetaPop's macrodiversity analysis includes raw population abundance, normalized population abundance, and α-diversity calculations. With normalized population abundance tables, we were able to generate heatmaps to view feature-level distinction between samples and biomes. Furthermore, we were able to calculate β-diversity based on Bray-Curtis dissimilarity. PCoA was performed, and to assess robustness, 2,000 features were randomly subsampled and analysis repeated across 1,000 bootstrap iterations. Resulting ordinations were aligned to a reference with Procrustes alignment. Mean coordinates and standard deviations were calculated for each sample, and scatter plots were generated. Supplementary Tables 6 and 8 and Supplementary Figure 4 have been added.

      (3) Claims about therapeutic cocktails would be better framed as speculative and/or moved to the discussion section.

      We thank the reviewer for their point, and it is well taken. Please see our more detailed response to this earlier in this reply.

      (4) The manuscript could be strengthened by elaborating on the scope and composition of the phage and bacterial isolate collections, which are important for interpreting the broader significance of the findings.

      We thank the reviewer for their point. We have added further details on the bacterial and phage isolate collections so the readers may draw the proper conclusions.

      Reviewer #2 (Public review):

      Weaknesses:
>

      While the authors acknowledge several limitations, some aspects require clearer framing or additional clarification. The proposed workflow focuses exclusively on aquatic environments as sources of phages, which may limit the diversity of hosts and phage types recoverable using this approach. Some interpretations, particularly regarding taxonomic classification and sampling saturation, would benefit from more cautious wording given current limitations in viral taxonomy and the observed data.

      The reviewer makes an excellent point. To try and address this, we made several edits to the main text of the discussion section to reframe and add clarification to our limitations. We also mention the limitation of our strategy to aquatic environments. Lastly, we addressed the final sentence below.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) To really demonstrate geomapping success would require more comparisons: choosing a variety of locations with and without high host levels and then analyzing the yes or no outcomes in terms of whether phages were found. This manuscript demonstrates 17 substantial and significant successes, very much worth sharing, but I am not sure it answers the central question posed in the title and abstract. This could potentially be accomplished through analysis of existing data related to the attempts, or it may be important to state this limitation in the discussion.

      We thank the reviewer for their insightful comment on the generalizability of our geomapping strategy and their emphasis on adding more comparisons to support the claim. While we did not test the 17 bacterial isolates (XΦROs) across multiple sites with varied host levels, we did incorporate an initial, broad preliminary screening panel designed to assess the presence and diversity of phage against multiple pathogens and laboratory strains across different sampled environments after site comparisons with the geomap (Fig. 2G). In another instance, we did without the guidance of the geomap (Supp. Fig. 2I). In each case, highly polluted waters located in densely populated areas (wastewater, Brays Bayou, and Buffalo Bayou) had higher success rates in phage recovery compared to other less polluted sites (Clear Creek, Galveston seawater, Hamilton Pool Preserve, and Pedernales Falls). This trend is consistent with previous reports [16,51]. The screening step functioned as an initial comparison and allowed us to gauge phage availability across different genera of bacteria prior to a focused geomap-guided phage hunt. In the geomap-guided experiment, we compared a low host-availability control site (Clear Creek) with two high host-availability sites (wastewater and Brays Bayou). More comparisons, especially incorporating even more sites with varied host availability, would strengthen the claim but not be invalid without it.

      In an ideal situation, we would perform the exact additional experiments proposed by the reviewer. However, we ran into logistical issues. As it turns out, phage quantities in various sites vary with weather, specifically rainfall (a finding that we intend to touch on in a future manuscript). As such, this variable must be matched. Even with the long summers of Texas, we were unable to perform multiple, serial PhiHD runs at multiple sites with similar weather conditions. As such we rely on the strength of preliminary screening (16S and plating for phage) to provide a clear guide as to where to place PhiHD experiments.

      We would still contend that the consistency between preliminary screening results and subsequent successful geomapping-guided discoveries (finding phages against 17 different isolates) gives sufficient evidence that geomapping is an effective strategy for identifying productive sampling sites. However, we acknowledge the excellent point of the reviewer and we have edited the main text to include this in the limitations, so that the reader may make their own evaluation.

      (2) Line 44: 24% of infections >> perhaps better to describe as isolates, as many infections contain multiple isolates of different types. More background in terms of the number of infections in general, comprising the 24% on the bacterial side, and also a description of the 350 phages in terms of known hosts, would be very interesting to add.

      This is an excellent suggestion that would add valuable background to the introduction and provide context for phage laboratories interested in the volume of cases handled by TAILΦR and the isolates they receive. We changed the main text to include information regarding this suggestion. We also provided a three extra tables to show 1) TAILΦR’s number of distinct case counts with number of bacterial isolates received, 2) TAILΦR’s phage library, and 3) number of isolates with NO phage. Please see Supp. Tables 1-3.

      (3) Line 92: To add a statistical test to the beta diversity comparisons beyond visual inspection of the location of the points on an ordination, I suggest adding a PERMANOVA.

      The reviewer makes an excellent point, and we agree that visual inspection should be backed up by rigorous statistics where possible. As such, we performed an analysis of molecular variance (AMOVA, done in Mothur) to compare samples derived from brackish, sea, fresh, and sewage. AMOVA showed statistical significance between the samples and confirms our visual inspection of the β-analysis. The text has been altered to reflect AMOVA data.

      (4) Line 112 (related to point 2 and Fig 1A): How many isolates in the collection for which 24% did not have a phage, and 35% were Pseudomonas?

      Related to changes in point 2, we added a supplementary table to provide insight into the number of bacterial isolates without a phage. In total, 104 (24% of 435 isolates in the TAILOR library) have no phage, and 35 belong to Pseudomonas aeruginosa.

      (5) Line 141: How is it known that Pseudomonas phage concentration increased by 95x? There could be unknowns / difficult to cultivate or phages without the right host. Consider describing as yields rather than the absolute concentrations.

      We appreciate the reviewer’s point that not all Pseudomonas phages are accounted for due to host specificity and potential unknowns. We agree completely that all we have are surrogates. We tracked the concentration of phages that infected our indicator strain Pseudomonas aeruginosa PAO1. We selected PAO1 because of its broad susceptibility profile and its role as a permissive host for isolating a wide range of Pseudomonas phages. While we acknowledge that not all Pseudomonas phages will be detected, PAO1 captures a wide breadth of Pseudomonas phages, enabling consistent comparisons between samples. Our concentration changes reflect a within-sample comparison between unprocessed material and the processed material. Aligned with reviewer’s concern, the values should not be taken as an estimate of absolute phage concentration in the samples. Rather, the values are method-dependent estimates of enrichment efficiency for PAO1-infecting phages. In tandem with PAO1-infecting phages, the concentration of other viral-like particles infecting other organisms is most likely increased with each concentration step. We have revised the entire manuscript to mention “phage yield,” rather than associate an increase to a concentration.

      (6) Line 224: 39.1% viruses> I believe this refers to vOTUs rather than viruses.

      The reviewer is correct; we appreciate the catch! We corrected the text to “vOTUs.”

      (7) Lines 224-230: How does this relate to the expected ~70% Dark Matter?

      As observed in many viral metagenomic studies, our dataset is also dominated by viral dark matter. Between 60.9% (outlier/singles from vCONTact2) and 66.5% (unclassified from PhaGCN) of vOTUs are in this group. To proceed with caution, we edited the main text to draw attention to the large percentage of viral dark matter in our metagenomic dataset. Although a substantial fraction of vOTUs is unknown, the remaining identifiable sequences provide some biological context, enable validation of sampling strategies and comparative analyses between samples.

      (8) Line 241: there are many perspectives on whether phage treatment should involve cocktails. If a phage is immunogenic and leads to antibody production that can neutralize other phages, one phage could ruin the game for others. Consider presenting this as a perspective, rather than a ground truth, and consider moving to discussion

      This is a very insightful input on this perspective. Our intent was not to present this as a definitive conclusion, but rather to highlight a broader need for a more diverse phage library. This is not limited to phage cocktail generation. We changed the introductory sentence to this paragraph to encompass a broader need for phage diversity and succinctly lead into the next sentence.

      (9) Figure 2b contains an R2 value of 0.7, and 2c has R2=0.76. Where does this come from? Maybe a PERMANOVA? Please describe in legend and/or methods+results.

      Thank you for catching this! β-diversity calculations were based on Bray-Curtis dissimilarity. We have adjusted the methods and results to incorporate this information.

      (10) The Rphi library is mentioned in several places, would be wonderful to have a bit more description of this collection.

      We thank the reviewer for their sharp eye, we definitely wanted to ensure the reader understands the significance of this. We added some descriptor sentences to better highlight and introduce the RΦ-library.

      (11) Consider adding a central success to the abstract, the fact that phages were found for 17 recalcitrant strains of various ESKAPE pathogens, yielding 35 phages after standard phage hunting and experimental evolution approaches had failed.

      We appreciate the reviewer for their emphasis on highlighting the success of our manuscript and made the appropriate changes. We added altered the last sentence to the abstract, and added another sentence to summarize our success.

      Reviewer #2 (Recommendations for the authors):

      (1) Figures 1C and 1D require a more detailed description.

      We thank the reviewer for noticing this. We have altered the figure legend to be more descriptive.

      (2) Raw and assembled sequencing data should be submitted to a public repository, and the accession numbers should be provided.

      We have uploaded the raw and assembled sequencing data to a public repository, and the accession numbers are provided in Supp. Table 9 and 12. For raw metagenomic shotgun sequences, BioProject accession is PRJNA1308632 (Supp. Table 2).

      (3) Line 89: The text states that the rarefaction curves plateaued; however, by definition, a plateau implies that the curve no longer increases. In the presented data, all curves continue to rise at the final sampling point. This does not affect the conclusions but suggests that sampling saturation has not been fully reached.

      This is a great observation by our reviewer. We agree with this point as the curves do not reach a complete plateau. We have revised the text to use more accurate language and clarify sampling depth.

      (4) Figure 2D: The heatmap normalized by Z-score within the selected taxa may give a biased impression of enrichment of certain taxa in specific environments, when in fact it only indicates enrichment relative to the other pathogenic taxa included in the analysis.

      The reviewer raises a great point, and we should have pointed this out directly. To avoid potential misinterpretations, we have revised main text to explicitly state that the heatmap displays relative enrichment to the other pathogenic taxa.

      (5) Given the variable taxonomic resolution achieved by 16S rRNA sequencing (genus or family level), it would be important to highlight that some detected taxa include non-pathogenic members. For example, Vibrio is common in seawater, yet only a few species are pathogenic to humans.

      We agree with this! We added a sentence to emphasize this point.

      (6) Figure 2G: The color scale bar is uniform across all panels; please adjust for accurate comparison.

      For more accurate comparisons between different samples and phage concentrations, we added a second color to assist with visualization.

      (7) The PCoA figures should specify which distance metric was used.

      We want to thank the reviewer for the catch, we should have mentioned that. Our PCoA was calculated based on Bray-Curtis dissimilarity. We have adjusted the main text and methods section to mention it.

      (8) Figure 3: The meaning of the colors in panels A and B should be clarified.

      We thank the reviewer for their keen eye. We changed the figure and figure legend to clarify. The colors on the map and PCoA represent influents from various wastewater treatment plants around Texas.

      (9) The manuscript jumps from Supplementary Figure 2 to Figure 6. In general, the order and referencing of supplementary materials are confusing. Supplementary tables and figures should not be intercalated within the same file.

      We thank the reviewer for their patience and apologize for the confusion. This occurred as we had multiple revisions to the manuscript and we did not update the sequence of the figures. To address the reviewer’s comment, we separated the supplementary tables and figures into apart. We also ensured that each main and supplementary figures and table were mentioned sequentially in the main text.

      (10) It is unclear why some vOTUs were observed in the 5 L collection but not in the concentrated sample (10/24; Supplementary Figure 3E). One would expect that the most abundant vOTUs in the 5 L sample should also be easily detected in the concentrate.

      The reviewer brings up a fantastic point. One would certainly expect that the most abundant vOTUs in the 5L samples would also be detected in the concentrated sample.

      We have several suspicions as to why several vOTUS were not detected in our concentrated samples. Because we used shallow shotgun metagenomic sequencing, as compared to deep sequencing, we may have obscured our ability to detect and quantify low-abundance taxa. Consequently, dominant taxa occupying a large portion of sequencing reads may have masked the detection of rarer species/vOTUs. Lower sequencing depth results in fewer total reads per sample and reduced sensitivity for rare, infrequent species/vOTUs to be detected. When their abundance falls below detecting limits, they may appear absent from a data set.

      Furthermore, we reached out to Novogene, who we outsourced for library preparation and shotgun metagenomic sequencing. According to Novogene, not all genetic material in a sample is used during their library preparation. The maximum amount of DNA to build a PCR-free metagenomic library at each time is approximately 1.5 µg of DNA. Although we submitted 184.4 µg of DNA (from the 400L-concentrate) and 25.6 µg of DNA (from the 5L-sample), we suspect only a fraction of the material was used for library preparation and subsequently sequenced. This may have limited the representation of low abundance vOTUs.

      (11) Line 227: The statement that "but only 33.5% of viruses could be classified to the family-level" requires caution. Since the traditional Siphoviridae, Podoviridae, and Myoviridae families were abolished, many viruses currently lack family-level classification. Therefore, this taxonomic level may not be ideal for assessing novelty, as many viruses closely related to known types remain unassigned.

      We appreciate the reviewer for bringing up this topic. We recognize the current limitation in the viral metagenomic landscape. Many viruses lack-family level classification due to ICTV taxonomic restructuring and the lack of reference genomes present in a database. A large fraction of viral sequences constitute “viral dark matter.” From a single metagenomic dataset, viral dark matter ranges from 60-90% of vOTUs. Within our own dataset, it is also dominated by viral dark matter. Between 60.9% (outlier/singles from vCONTact2) and 66.5% (unclassified from PhaGCN) of vOTUs are in this group. Although a substantial fraction of vOTUs is unknown, the remaining identifiable sequences provide some biological context, enable validation of sampling strategies and comparative analyses between samples. To complement vCONTact2 results, we utilized PhaGCN to classify each vOTU as a means to compare taxa derived from each sampled biome from one another and not to assess novelty of the metagenomic dataset. We aimed to provide measurable and interpretable context to our metagenomes. However, due to the substantial variability and uncertainty in the field and our dataset, we revised the text to highlight the large fraction of unclassified sequences and their implications.

      (12) Supplementary Figure 5: The legend does not clearly explain the two inner rings. One may correspond to GC skew, but this should be explicitly stated.

      Well spotted! We have made the appropriate corrections.

      (13) Line 325: The reference to "50 mL samples" is unclear-please specify which samples this refers to.

    1. Author response:

      We thank the reviewers for such positive and constructive feedback, and for their enthusiasm about our use of controllability and dynamical systems perspectives to understand learning variability. We are glad to see that they believe this work will be “highly impactful” and “directly motivate new learning experiments”. We agree that these findings suggest new experimental tests of dynamical constraints on learning, in BCIs and motor control as well as other computations that depend on neural dynamics, such as decision-making tasks. Combined with new tools for data-driven identification of latent dynamics, we are excited to see how dynamical constraints can help understand learning outcomes across different tasks, brain areas, and individuals.

      Based on reviewer comments, we identified three sets of analyses that will improve the clarity and strength of evidence for our primary conclusions.

      (1) As the reviewers identified, a central contribution of this study is to show that continuous within-class variability becomes explainable by considering underlying dynamical structure. We realize this was insufficiently emphasized in Figure 6. All regression models included group-specific intercepts, so improvements from dynamical features reflect prediction beyond class-level differences. To quantify this directly, we compared against an intercept-only model and evaluated prediction of within-class residual variability (mean-subtracted). Geometric features did not improve performance beyond class means, whereas dynamical features significantly improved prediction (p<10<sup>-5</sup> for both behavioral measures). Moreover, only dynamical features predicted within-class residual variability (cross-validated R<sup>²</sup> = 0.19 and 0.30 for learning speed and hit-rate change, respectively; p < 10<sup-8</sup>). We will add these analyses and revise the text to clarify this point.

      Author response image 1.

      Cross-validated R<sup>2</sup> for (left) learning speed and (right) change in hit rate, for true behavioral outcomes (total variability, blue) and after subtracting class means for OMPs and WMPs (residual variability, orange).

      (2) We appreciate the reviewers’ comments to clarify what changes in neural structure are small, and to provide a quantitative comparison to changes observed in the primate BCI experiments.

      We referred to published analyses of within-manifold perturbations (WMPs) in the primate BCI experiments, which reported <10% reduction in fractional variance within the intrinsic manifold for most sessions (Golub et al., 2017). (No comparable analysis was reported for OMP sessions.) For adaptation to WMPs, changes in variance within the intrinsic manifold in RNN models with input plasticity closely matched experimental observations (75th percentile: 94% of pre-learning variance in the model versus 90% in data), whereas recurrent plasticity RNN models produced substantially larger departures (78%). In fact, the entire distribution with recurrent plasticity was shifted to larger changes than those observed in most primate WMP sessions. A second comparison based on covariance changes along BCI dimensions (Figure 5 in [1]) yielded a similar conclusion. The authors estimated ~5-20% changes in covariance along both the intuitive and perturbed decoder dimensions during WMP sessions. For our RNN models trained with input plasticity, we observed similar changes: changes along the perturbed decoder were <10% although changes along the intuitive decoder were ~40%. We borrowed the terminology of “small” from the experimental findings in [1], where comparisons were made to alternative learning hypotheses (with predicted changes as >10-fold higher). These analyses now provide more quantitative evidence that neural reorganization under input plasticity is largely consistent with primate neural data. We will add these comparisons as a supplementary figure in the revised manuscript.

      Author response image 2.

      Proportion of maps with normalized variance in intrinsic manifold (IM) above a certain minimum value. Results with training RNNs on WMPs, with either input plasticity (blue) or recurrent plasticity (orange), overlaid on primate data from Golub et al, 2017 (black). Dashed lines indicate the 75th percentile value.

      We agree with reviewers that under input plasticity, both statistical and dynamical changes are relatively modest, particularly when compared to the behavioral changes. Rather than focusing on the magnitude of these changes, our regression analyses in Figure 6 highlight that the dynamical changes are a better predictor of continuous variability of behavioral outcomes. Moreover, OMPs are misaligned with both the intrinsic manifold and the controllable subspace. Thus, mean OMP learning performance alone cannot disentangle the contribution of these different sources of misalignment. By showing that variability within each class is explained by considering dynamics (Figure 4, Figure 6), and using the dissociation between task manifold and controllable subspace by varying controller architecture (Figure 8), we provide evidence that dynamical constraints provide a more comprehensive picture of learning variability, beyond categorical differences.

      (3) Finally, we tested whether the same dynamical features explain learning variability across the alternative controller architectures in Figure 8. They remained predictive of learning speed (cross-validated R<sup>2</sup> of 0.35 and 0.33 for low-D and high-D controller networks respectively), supporting the generality of the proposed dynamical constraints. We will add this analysis to the revised manuscript.

      As per reviewer suggestions, we will also perform additional analyses to examine the relationship of learning outcomes to initial behavioral metrics for different decoders, assess flowfield changes during the preparatory phase, report the relevant statistics for stated comparisons, and clarify that learning with only one set of inputs (either feedforward or feedback) was poorer.  We will also clarify several points raised by the reviewers, including:

      (i) the compatibility of overlapping confidence intervals of WMP/OMP learning outcomes with prior experimental data in Sadtler et al, 2014;

      (ii) the distinction between flow-field changes in the full neural state space (Figure 5D) and along behavioral readout dimensions (Figure 5E);

      (iii) that autonomous dynamics contribute to controllability and how differences in pre-trained autonomous dynamics across controller architectures could indirectly vary feedforward controllability (Figure 8); and

      (iv) the relationship between controllability and reachable manifolds in position-decoder BCIs.

      References:

      (1) [Golub et al, 2017]   Golub, M.D., Sadtler, P.T., Oby, E.R., Quick, K.M., Ryu, S.I., Tyler-Kabara, E.C., Batista, A.P., Chase, S.M. and Yu, B.M., 2018. Learning by neural reassociation. Nature neuroscience, 21(4), pp.607-616.

      (2) [Sadtler et al, 2014]   Sadtler, P.T., Quick, K.M., Golub, M.D., Chase, S.M., Ryu, S.I., Tyler-Kabara, E.C., Yu, B.M. and Batista, A.P., 2014. Neural constraints on learning. Nature, 512(7515), pp.423-426.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors conclude that mPFC is not required for avoidance, based on the minimal behavioral effects of optogenetic inhibition. While this interpretation is supported by the data, the choice of viral constructs could lead to an underestimation of the mPFC's role for other reasons. First, the choice of viral constructs could lead to an underestimation of the mPFC's role for several reasons. Specifically, the efficacy of eArch3.0 inhibition was not verified beyond histology, and its non-cell-type-specific nature could lead to disinhibition or compensatory activity in downstream regions. Although the authors' use of visual cortex (VI) inhibition as a control suggests that broad cortical inhibition does not impair avoidance, subcortical compensation cannot be ruled out. Additionally, Vgat-ChR2 targets only GABAergic neurons, potentially missing glutamatergic contributions. Addressing these limitations in the Discussion section would strengthen the manuscript.

      We thank the reviewer for these points. First, although we did not perform direct electrophysiological verification of eArch3.0 efficacy in mPFC in the present study, this construct has been extensively validated in prior work and is widely used to produce robust neuronal inhibition. In our experiments, the lack of behavioral effect with eArch3.0 inhibition converged with the results obtained using the independent Vgat-ChR2 approach, which we directly validated, supporting the conclusion that mPFC inhibition does not impair avoidance under these conditions. Our results are also consistent with previous studies showing that mPFC lesions do not impair avoidance behavior.

      Second, we agree that manipulating mPFC activity will necessarily influence downstream circuits, including subcortical regions, given the interconnected nature of these networks. Our goal was to test whether inhibiting mPFC activity alters avoidance behavior, not to isolate it from its targets. In this context, the absence of behavioral effects indicates that avoidance behavior can be supported without mPFC activity. While compensation is always a possibility, this usually reveals some impairment while compensation occurs, but we did not observe those effects. Our results are consistent with the idea that subcortical circuits normally mediate these behaviors.

      Finally, regarding Vgat-ChR2, activating GABAergic neurons is a well-established approach to suppress cortical activity, as these interneurons provide strong inhibition onto local glutamatergic neurons. Thus, this manipulation is expected to broadly reduce excitatory output in cortex. Indeed, the robust suppression of cortical activity we observed with GABAergic activation makes it unlikely that major glutamatergic contributions were missed.

      These points are in the paper, including the Discussion.

      Reviewer #2 (Public review):

      (1) There are few details on the linear mixed models in the methods. This section could be improved by including a mathematical description. More importantly, the reader never learns how accurately the models capture the data. Given that most conclusions rely on the models, it seems central to address this point carefully. For example, what is the explained variance, marginal, and conditional? Were the nested models compared to non-nested ones (e.g., AIC), what are the specific outputs of the likelihood ratio tests briefly mentioned in the methods?

      Model structure was defined a priori by the experimental design and hypotheses rather than selected through model comparison, but we verified the contribution of key model components (e.g., covariates, interactions, and random effects) using likelihood ratio tests comparing models. Regarding model performance, we now report for each model the marginal and conditional R<sup>2</sup> values (Nakagawa), which quantify variance explained by fixed effects alone and by the full mixed model including random effects. In addition, likelihood ratio test results for all fixed effects and interactions (χ<sup>2</sup> statistics) were already reported in the manuscript.

      (2) For several figures, there is a disconnect with the main text, in the sense that it is difficult to understand how statements in the main text connect with specific figure panels or bars in their graphs. This is particularly the case for the most complex figures, e.g., Figures 3, 4, and their supplements. It would be beneficial to introduce subfigure labels (A1, etc) and state explicitly in the main text what figure panel is described (in parentheses). Alternatively, breakdown the figures into multiple ones, decreasing ambiguity. This is important because it will help the reader better assess the strength of the results.

      We have significantly revised the manuscript to reduce ambiguity and thank the reviewer for each of their (28) requests, which we have implemented in full. We also added additional figure references to the Results to assist with readability. This has significantly improved clarity and readability.

      (3) It does not appear that the code and data used to produce the figures are made available. That would be very beneficial, given the complexity of the analysis and dataset collection procedures. It would also help readers better understand the results and probe their validity.

      As usual, we will share the full dataset in the VOR at Dryad after the revision is completed.

      Reviewer #3 (Public review):

      The main weakness, in my view, lies in the Results section. In the figures, the authors do not present any raw data, and the plots are shown as mean {plus minus} SEM without displaying the distribution of individual data points.

      We thank the reviewer for the recommendations. Individual data points are shown where appropriate (e.g., Fig. 1). However, most of our analyses involve repeated-measures, hierarchical data with multiple levels (cells and sessions nested within animals), where simple point overlays can be misleading or difficult to interpret without explicit linking across levels. We therefore use mean ± SEM visualizations for clarity in these summary figures, while preserving the full hierarchical structure in the statistical analysis through mixed-effects models. All data will be made available in the VOR to allow full inspection of the underlying distributions.

      It is both a strength and a weakness that the authors do not attempt to guide the reader through the Results section and instead present the findings with very little emphasis on the key outcomes of the GLM. While this approach is arguably the most transparent way to report results, it also makes the section quite difficult to follow and may discourage readers.

      I would recommend rewriting the Results section to make it more accessible to a broader audience. A similar issue applies to the figures: presenting all plots reflects a commendable commitment to transparency, but it would greatly benefit from a clearer narrative. As it stands, it is difficult to grasp the message of each figure by simply browsing through them.

      The full description (complexity) of the models is entirely in the legends and supplemental figures. This was done to make the results easier to follow. We have made all the changes noted above to facilitate readability while assuring there is enough transparency to assess the data. We think readability has significantly improved.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Below are a few specific suggestions related to the main weaknesses mentioned above.

      (1) P4 L9: The sentence starting with "However, most ..." sounds more like a statement than a contrast with the previous sentence. Therefore, please delete "However" and please add references to justify the statement.

      Done.

      (2) P8: Definition of movement peaks. It would be great to have three videos illustrating the mouse behavior in the three different movement peaks. This would allow the reader to better understand the differences between no peaks 3 sec prior, more than 5 seconds, and one example that does not fit these two categories. In addition, what percentage of all peaks to the no peaks 3 sec prior and more than 5 sec represent?

      We added the percentages. The “3 sec prior” represent ~23% and the “5 sec” represent ~31%. However, we do not think adding a single video of one movement per these 3 cases would be useful as the dataset is composed of thousands of these movements.

      (3) P8: Last paragraph. When you state that you performed a linear fit between DF/F and movement, do you mean speed? In addition, the statement "integrating both signals over a 200 ms window" is incomplete. How is the window selected? Is the window 200 ms around movement onset or movement peak speed?

      Yes, the movement variable used in the linear fit corresponds to speed. Regarding the 200 ms window, this analysis does not focus on specific behavioral events such as movement onset or peak speed. Instead, both ΔF/F and speed signals were segmented into consecutive 200 ms windows across the entire recording session, and the linear relationship was computed across these paired segments. Thus, the analysis captures the overall relationship between neural activity and ongoing movement, rather than eventaligned dynamics. We have revised the text to clarify both the use of speed and the implementation of the 200 ms window.

      (4) P14: Discussion of AA19 and AA39 tasks: It would be helpful to clearly specify what percentage of actions you would expect given no learning, is it the 23% action dashed line indicated in the top panel of Figure 2B?

      The expected percentage of actions under no learning is not fixed, as it depends on the rate of spontaneous (non–cue-driven) crossings. In these tasks, we estimate this baseline using behavior during the noUS condition, where the action rate is ~23% (Fig. 2B). In the AA19 and especially AA39 tasks, this baseline decreases because spontaneous inter-trial crossings (ITCs) are progressively reduced, leading to lower expected action rates under no-learning conditions. Thus, the 23% baseline derived from noUS is lower in the AA19/39 tasks. In other studies, we explicitly included NoCS (no-cue) trials to estimate chance performance; however, in the present design we rely on the noUS baseline and the observed changes in ITC rate. We have clarified this point in the text.

      (5) P15 L2: "Considering tone intensity (Fig. 2B), CS1 avoids latencies increased at medium and high intensities but not a low intensity." This is confusing. Are you referring to the AA39 triangles under CS1 in the middle panel, left? They are all above the dashed reference line. So the plot seems to contradict the statement. If you are referring to AA19, the red dots also seem to show the opposite of the statement.

      The dashed reference line reflects latency during the noUS condition and is included for visual reference; however, these values are not directly comparable to those in the AA tasks, as noUS latencies are largely unconstrained and reflect baseline behavior rather than learned responding. The statement in the text refers specifically to changes across AA conditions, consistent with our analysis approach throughout the manuscript, where values are compared to the immediately preceding condition. In this case, we are referring to AA39 (triangles) relative to AA19 (circles). Under this comparison, CS1 avoidance latencies increase at medium and high intensities, but not at low intensity, consistent with the statistical contrasts. We have revised the text to clarify the points.

      (6) P17: "Movement and neural measures subtract the baseline from the other three windows at a trial level." Do you mean to say that for each measure, the baseline was subtracted? How is baseline defined (over which time window)?

      The baseline is defined in that same paragraph as the −0.5 to 0 s pre-CS window. To improve clarity, we have revised the text to explicitly restate this definition in the sentence describing baseline subtraction.

      (7) P17: "Fig. 2-Supplement 2A,B shows model-derived marginal means of movement averaged across tone intensities." Some explanation needs to be provided, since the previous figures show a dependence of behavior on tone intensity. Are you doing this based on Fig. 2-S1?

      Yes, these results are derived from the same model of the full data shown in Fig. 2–S1. In this particular analysis, tone intensity was included in the model but not retained when computing marginal means and contrasts, effectively averaging across intensity levels. The rationale for this approach is that tone intensity was primarily used to increase behavioral variability, particularly error rates, which are otherwise low in this task. Averaging across intensity therefore improves statistical power and allows us to more clearly isolate the effects of the primary factors of interest. We have clarified this point in the text.

      (8) P18: "Orienting magnitude was strongly dependent on tone intensity...". However, in Figure 2-S2, there is no information about tone intensity. So how is the reader supposed to see this? Same issue on P19 when discussing the action window. Generally, the description of Figure 2-S1 and S2 is difficult to follow and should be improved. It is not clear that all panels are referred to in the text.

      We have revised the start of the Movement section to clarify how tone intensity is treated across analyses and figures. Specifically, tone intensity is included as a factor in all statistical models; however, for clarity of presentation, it is sometimes collapsed in figures to reduce dimensionality and to emphasize other task-related factors. This manipulation was introduced primarily to increase behavioral variability (particularly error rates), thereby improving sensitivity for estimating the effects of the other task variables.

      We have also clarified when we reference Fig. 2–S2 legend that, although intensity is not displayed in the figure for visualization purposes, it is included in the underlying model and its effects are reported in the supplement.

      (9) P22, 23: Windows are mentioned, but not defined or indicated in figures.

      We have clarified in the text that the same time windows defined for movement analyses (baseline, orienting, action, and from-action) were also used for the neural analyses.

      (10) P22: "Covariates were standardized within each window so that estimated marginal means reflected ΔF/F at average covariate values." It is unclear what was done exactly. What do you mean by "standardized"? Maybe give an example here and elaborate in the methods.

      By “standardized within each window,” we mean that covariates were z-scored within each analysis window (i.e., each covariate was transformed to have a mean of 0 and a standard deviation of 1 within that window). This ensures that estimated marginal means correspond to ΔF/F evaluated at the average covariate values within each window. We have clarified this in the Methods and Results.

      (11) P24-25: Indicating spurious action on Figure 3-S2 (and in Figure 3) would help the reader follow the argument in the main text.

      We clarified this in the legends by indicating that actions not classified as AA, PA, Escape, or PA Error are spurious actions.

      (12) P25: "After controlling for ..., but this includes the effects of aversive stimulation." The second part of this sentence was not clear.

      We have clarified this sentence to indicate that avoidance errors are followed by aversive stimulation (i.e., errors are punished).

      (13) P34L3: "Classs" -> "Class".

      Fixed.

      (14) P42 top paragraph: There are two references to Figure 5-S1 panel D, but there is no panel D on the figure.

      Fixed.

      (15) P57: The sentence starting with "Random effects were specified ..." is very difficult to follow.

      We have revised this sentence to improve clarity by separating the description of the random-effects structure from the model syntax.

      (16) P57: The windows analyzed are finally defined at the bottom of this page. The information also needs to be included early in the results to improve comprehension.

      This is now included in the main text when windows are first used in the movement section.

      (17) P58: Several R packages are mentioned by name, but without specifying that they are R packages, which would facilitate reading.

      We added R.

      (18) P58 top paragraph: "Tuckey's correction", do you mean "Tukey's HSD test"?

      We thank the reviewer for noting this. We used Holm-adjusted p-values for multiple comparisons (as implemented in emmeans) and have revised the text.

      (19) P63: "features extracted from F/F" do you mean "DF/F"?

      Yes, fixed.

      (20) Figure 1B speed plots: it is not possible to visualize the lines at the movement peak because they overlap completely. You can either add an inset on the left of the peak (for each panel), magnifying that region, or play with the transparency of the traces to improve visibility. There is a similar issue in Figure 5A, B. (Alternatively, if it is not possible to solve the issue graphically, explicitly state that traces overlap.)

      We have fixed this by making some traces dashed in Figure1 and 1-S1, which reveals the underlying traces. We also stated that the peak speed completely overlaps. In Figure 5, we stated that traces overlap as expected; transparency or dashing does not work well with the colors used in Figure 5 and in fact the overlap emphasizes the similarity of the movements.

      (21) Legend 1A: abbreviation CCF not defined. Is it anterior to the left? Abbreviation WM not defined. The right panels are unclear. The legend states that they show a schematic of the location of the optical fibers, but that was not clear. Do the dots indicate the location of the fibers? Is the green region indicative of V1? Same for dark gray in the mPFC panel. What are the lighter grey regions and the blue region? Does 'lateral' mean 'lateral from midline'? Please clarify these points.

      CCF is defined in Methods, and the typesetting process will adjust abbreviations as needed per the journal. We have defined MW and clarified all the other points in the legend.

      (22) 1B: "peaks taken at a fixed interval > 5 s", this is a bit confusing. If the interval is fixed, the exact time interval should be given. If it is > 5 s, then this suggests that it is not fixed. Do you mean "at intervals > 5 s"?

      Yes, fixed.

      (23) Figure 1-S1C: is the area the integral of the z-scored DF/F above zero DF/F? If so, it should have units of seconds (integral over dt of a dimensionless variable). Similarly, the Peak is a z-score value? In addition, is the time to peak in seconds? What is zero? Peak time of movement?

      We thank the reviewer for raising these points. We have clarified the terminology in the text and figure. Specifically, “area” was inaccurately labeled and refers to the mean z-scored ΔF/F within each analysis window (not a time integral). Peak values correspond to the maximum z-scored ΔF/F within the window, and time to peak is reported in seconds relative to the alignment point. We have also clarified the definition of time zero and included these definitions in Methods.

      (24) Figure 2-S1: It is not clear if this figure is obtained by averaging across all animals. Please explain in the legend.

      We clarified that values represent averages across mice.

      (25) Figure 2-S2: Are the speeds in A and B in units of cm/s (vertical axis)? This needs to be indicated.

      We have clarified in the figure legend that movement speed is expressed in cm/s.

      (26) Figure 5A, scale bar: It looks like a Delta is missing in front of F because the label reads 0.5 F/F instead of 0.5 DF/F. I am unclear why there are three colored traces for the speed panels. If the colors denote neuron classes, does this mean they were recorded in different sessions, allowing the authors to distinguish activation speed for each class separately?

      We fixed the scale bar typo. The speed traces in the bottom panels are shown to illustrate that movement is highly similar across activation types within each avoidance mode, indicating that the observed large differences in neural activity cannot be attributed to differences in movement. Minor differences in the speed traces arise because activation types are composed of neurons that can be recorded in the same or different sessions, and each activation type may not be present in every session. We added several sentences to this section that should fully clarify the issue.

      (27) Figure 4-S1 legend B: Please indicate why the two panels are missing for the PA case (for the confused reader).

      We have clarified in the legend that panels are not shown for correct CS2 passive avoids because these trials do not involve an action, and therefore from-action alignment cannot be defined.

      (28) Figure 5-S A, B: Units missing for speed.

      Fixed.

      Reviewer #3 (Recommendations for the authors):

      I cannot assess the scientific validity of the study design as it is too far away from my direct field of expertise. But I found the authors' arguments convincing, and the results sound pretty consistent with the little I know of the field. The recording methods are good and the statistical analysis robust. So my only recommendation for the authors would be to work on the figures to improve clarity.

      Thank you. We have introduced various changes that we hope will facilitate readability for a wider audience while preserving the necessary details.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Morgan et al. studied how paternal dietary alteration influenced testicular phenotype, placental and fetal growth using a mouse model of paternal low protein diet (LPD) or Western Diet (WD) feeding, with or without supplementation of methyl-donors and carriers (MD). They found diet- and sex-specific effects of paternal diet alteration. All experimental diets decreased paternal body weight and the number of spermatogonial stem cells, while fertility was unaffected. WD males (irrespective of MD) showed signs of adiposity and metabolic dysfunction, abnormal seminiferous tubules and dysregulation of testicular genes related to chromatin homeostasis. Conversely, LPD induced abnormalities in the early placental cone, fetal growth restriction and placental insufficiency, which was partly ameliorated by MD. The paternal diets changed placental transcriptome in a sex-specific manner and led to a loss of sexual dimorphism in the placental transcriptome. These data provide a novel insight on how paternal health can affect the outcome of pregnancies, which is often overlooked in prenatal care.

      Strengths:

      The authors have performed a well-designed study using commonly used mouse models of paternal underfeeding (low protein) and overfeeding (Western diet). They performed comprehensive phenotyping at multiple timepoints including of the fathers, the early placenta and late gestation feto-placental unit. The inclusion of both testicular and placental morphological and transcriptomic analysis is a powerful non-biased tool for such exploratory observational studies. The authors describe changes in testicular gene expression revolving around histone (methylation) pathways that are linked to altered offspring development (H3.3 and H3K4), which is in line with hypothesised paternal contributions to offspring health. The authors report sex differences in control placentas that mimic those in humans, providing potential for translatability of the findings. The exploration of sexual dimorphism (often overlooked) and its absence in response to dietary modification is novel and contributes to the evidence-base for the inclusion of both sexes in developmental studies.

      Comments on revised version:

      The authors have done a great job addressing my concerns. The description of the data analysis and the figures are now much clearer. The inclusion of the potential links between the microbiome and male reproductive fitness is informative and improves the flow of the discussion.

      Reviewer #2 (Public review):

      Summary:

      The authors investigated the effects of a low-protein diet (LPD) and a high sugar- and fat-rich diet (Western diet, WD) on paternal metabolic and reproductive parameters and feto-placental development and gene expression. They did not observe significant effects on fertility; however, they reported gut microbiota dysbiosis, alterations in testicular morphology, and severe detrimental effects on spermatogenesis. In addition, they examined whether the adverse effects of these diets could be prevented by supplementation with methyl donors. Although LPD and WD showed limited negative effects on paternal reproductive health (with no impairment of reproductive success), the consequences on fetal and placental development were evident and, as reported in many previous studies, were sex-dependent.

      Strengths:

      This study is of high quality and addresses a research question of great global relevance, particularly in light of the growing concern regarding the exponential increase in metabolic disorders, such as obesity and diabetes, worldwide. The work highlights the importance of a balanced paternal diet in regulating the expression of metabolic genes in the offspring at both fetal and placental levels. The identification of genes involved in metabolic pathways that may influence offspring health after birth is highly valuable, strengthening the manuscript and emphasizing the need to further investigate long-term outcomes in adult offspring.

      The histological analyses performed on paternal testes clearly demonstrate diet-induced damage. Moreover, although placental morphometric analyses and detailed histological assessments of the different placental zones did not reveal significant differences between groups, their inclusion is important. These results indicate that even in the absence of overt placental phenotypic changes, placental function may still be altered, with potential consequences for fetal programming.

      Comments on revised version:

      The authors have adequately addressed all my previous comments.

      We would like to thank the Editor and Reviewers for their consideration and thoughtful comments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It was a little difficult seeing exactly what had changed in the manuscript without going back to the original version as not all changes were marked yellow in the revised version. In future, I would recommend clearly labelling all changes to aid the referee.

      We apologise to the reviewer for the difficulty in seeing where the changes had been made. We acknowledge their comments for subsequent manuscripts and thank them for their time, consideration and comments.

      Small comments:

      (1) I noted the description of the statistical analysis now includes the addition of paternal age/diet duration in the generalised mixed model for the late gestation cohort. Was this also done for the early gestation cohort? If not, why not?

      For the data presented in Figure 6, each data point was obtained from a separate male. As such, we were not able to factor in male effects, as no male sired more than one litter (Figure 6A). Additionally, only one conceptus per male was analysed for ECP area and development meaning paternal age effects could not be accounted for.

      (2) The legend of Figure 2 states that "Data were analysed using either a one-way ANOVA with Holm-Sidak post hoc tests for multiple comparison respectively". Is some text missing here?

      We thank the reviewer for spotting this typographical error. This has now been corrected and reads “Data were analysed using a one-way ANOVA with Holm-Sidak post hoc tests for multiple comparison”.

      (3) Figure 1 remains low resolution in the reviewer's copy. If possible, it would be good to upload a higher resolution figure during production of the article.

      We apologies that the resolution of this figure was still low for the Reviewer. We have checked the dpi and it is 300x300. However, we will ensure the quality is as high as possible during production.

      Reviewer #2 (Recommendations for the authors):

      One minor remaining issue: the caption of Figure 3 still contains the phrase "non-fasting metabolic status", which should be deleted from this sentence.

      We thank the reviewer for spotting this typographical mistake. This has now been corrected.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Joint Public Review:

      In this study, the authors suggest that DuoHexaBody-CD37, a biparatopic CD37-targeting antibody, can induce direct cytotoxicity in diffuse large B-cell lymphoma (DLBCL) cells through antibody clustering and SHP-1 activation, independent of complement. They further propose that DuoHexaBody-CD37 inhibits cytokinemediated pro-survival signalling, suggesting a broader role for CD37-directed therapy in disrupting tumour supportive signalling networks.

      A strength of the study is the systematic in vitro characterisation of signalling responses to DuoHexaBodyCD37 across both malignant and normal B-cells. The inclusion of phosphoproteomic profiling and mutant constructs provides mechanistic detail, and the findings may be of interest to researchers working on antibody therapeutics in lymphoma.

      However, the evidence supporting key mechanistic processes - particularly the role of SHP-1 in mediating cytotoxicity and the requirement for Fc receptor crosslinking - is incomplete and would benefit from further functional validation. While CD37 has been explored previously as a therapeutic target, this study does add mechanistic insight into direct cytotoxicity and cytokine modulation. Nevertheless, the exclusive reliance on in vitro systems makes the translational relevance unclear. Overall, the study provides valuable insight into CD37-mediated signalling in lymphoma cells, but the evidence remains incomplete to support broader conclusions about therapeutic impact.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      In the manuscript, Singh and colleagues reveal a new mechanism via which DuoHexaBody-CD37 induces DLBCL cytotoxicity, which is independent of external factors, such as the effector cells and the complement system. As cited by the authors, the induction of B cell death has previously been demonstrated for antibodies directed against B cells, including anti-CD37 (otlertuzumab). Furthermore, the majority of these observations are made using in vitro systems, and it is not clear if this phenomenon happens in vivo or not?

      Thank you for pointing this out. We would like to refer to previous report that have demonstrated potent anti-tumor activity of DuoHexaBody-CD37 in vivo in cell line- and patient-derived xenograft models from different B-cell malignancy subtypes [PMID: 32341336]. Moreover, DuoHexaBody-CD37 ex vivo activity has been shown in primary tumor cell samples from a large cohort of newly diagnosed (ND) and relapsed/refractory (RR) patients with a broad range of B-cell malignancies, including chronic lymphocytic leukemia (CLL) and B-cell non-Hodgkin lymphoma, including diffuse large B-cell lymphoma (DLBCL) [PMID: 33324950]. We refer to these data in the introduction.

      The presented data suggest that DuoHexaBody-CD37 relies on Fc crosslinking for its optimal cytotoxic activity. Investigating which FcγR is needed for this purpose would have been useful, as FcγRIIb, for instance, has been shown to be important in supporting the therapeutic function of mAbs like anti-CD40.

      We thank you for this suggestion. To further investigate the role of specific FcγRs in effector cell-mediated Fc cross-linking, PBMC-mediated direct cytotoxicity was compared across various immune cell subsets: B cells (FcγRIIb), NK cells (FcγRIIIa, IIc), monocytes (FcγRI, IIa/b, IIIb), and T cells (no confirmed FcγR expression). Notably, all immune cells subsets expressing FcγRs exhibited similar or enhanced cytotoxicity against DLBCL cells compared to the total PBMC pool. These results indicate that DuoHexaBody-CD37 induced killing is independent of specific FcγR subtypes. We have added these new data to new Figure 1C.

      Specific comments:

      (1) Line 92:93: The authors should also cite the following reference for rituximab: https://pubmed.ncbi.nlm.nih.gov/19620786/ .

      We have added this reference to the revised paper (ref. 31).

      (2) Figure 1 and 2: Since cell death was only observed in the presence of crosslinking in Figure 1, Figure 2 should also investigate the clustering and internalization of CD37 in the presence of the same secondary antibody. It is likely that DuoHexaBody-CD37 will induce receptor internalization upon crosslinking.

      To further investigate internalization, we compared the surface availability of CD37 with and without Fc-mediated crosslinking of DuoHexaBody-CD37 across cell lines. Little to no decrease in the surface availability of CD37 upon Fc-mediated crosslinking (new Supplementary Figure 2) was observed.

      In addition, we performed cluster analysis studies in lymphoma cells treated with DuoHexabody-CD37 in the absence and presence of Fc-crosslinking (and respective isotype controls). We observed that DuoHexabodyCD37 by itself was already sufficient to induce CD37 clustering, which was further enhanced by Fc-crosslinking (new Figure 2A, B).

      (3) Figure 3A: the Y-axes should be clearly labelled.

      Done.

      (4) Figure 6: What is the reason for the selective use of different cell lines in Figure 6? Additionally, only 1 donor has been used for the IL-6 analysis.

      The reviewer is indeed correct in noticing that only one cell line has been used for the IL-6 analysis. We observed that HBL-1 cells were the only cell line that were sensitive to IL-6 treatment, in contrast to IL-4 and IL-21. We have added this sentence to the discussion to explain this better: “p-STAT3 downregulation upon DuoHexaBody-CD37 treatment in presence of IL-6 requires further investigation in additional IL-6-responsive cell lines, as HBL1 was the only IL-6-responsive lymphoma cell line tested in this study.”

      The data shown in Figure 6 are results from at least three independent experiments (each dot is an independent experiment, not a donor).

      Reviewer #2 (Recommendations for the authors):

      Singh et al uncover a novel mechanism of action for the DuoHexaBody-CD37 against DLBCL, whereby it is shown to induce direct cytotoxicity independent of complement and to activate the phosphatase SHP-1. DuoHexaBody-CD37 is also shown to reduce cytokine induced JAK/STAT signalling in DLBCL cells.

      Strengths:

      The authors provide novel insight into CD37 targeting across normal B cells, DLBCL and Burkitt lymphoma cells, which have the potential to inform clinical translation.

      Weaknesses:

      The mechanisms behind differences in signalling and apoptosis between normal B cells, Burkitt lymphoma, and DLBCL cells with CD37 targeting require further clarification. In particular, the contribution of SHP-1 to this effect is not clear and indeed is increased in both normal b cells and DLBCL cells.

      Key points that require addressing are below:

      (1) Viability of Burkitt lines was less affected than DLBCL in Figure 1- this should be compared with surface CD37 expression in these same lines to determine whether this accounts for the effect. This difference is a key finding for clinical translation.  

      We thank the reviewer for this suggestion and we have now performed flow cytometry analysis across DLBCL and Burkitt cell lines upon staining with two different anti-CD37 antibodies (WR17, M-B371) to quantify membrane CD37 expression (new Supplementary Figure 1B). These data show that CD37 expression levels are not directly related to DuoHexaBody-CD37 mediated cytotoxicity in the studied B cell lines. 

      (2) pSHP1 is increased in both normal B cells (lines 169-171, Figure 3C) and DLBCL and yet the authors state specific upregulation of pSHP1 in DLBCL as a reason for induced cytotoxicity in DLBCL (lines 183-185). This requires clarification and experimental confirmation. The authors should investigate normal B cells in the cytotoxicity assays as in Figure 1 for comparison. The authors should also confirm the importance of SHP-1 in this apoptosis process using specific SHP pharmacological agents, which are commercially available.

      To analyze the role of SHP1 mediated signaling in induced cytotoxicity of DLBCL, SHP1 knock outs (KO) were generated in HBL1 and OciLy7 cell lines using CRISPR Cas9 technology (new Supplementary figure 5A). The wild type and SHP-1 KO cell lines were then compared for differences in cytotoxicity after treatment with DuoHexaBody-CD37 with and without Fc-crosslinker. No differences in cytotoxicity were observed between the wild type and knock out cell lines (new Supplementary figure 5B), indicating that DuoHexaBody-CD37induced SHP1 signaling does not play a direct role in the increased cytotoxicity. We have added these new data to the results and rephrased the role of SHP-1 in the revised manuscript. 

      (3) It would be informative to assess caspase activation and PARP cleavage across normal B cells, DLBCL and Burkitt under these conditions for clarity on apoptosis induction.

      We thank the reviewer and we agree it would be informative to confirm apoptosis induction in the cell lines upon DuoHexaBody-CD37 treatment. We addressed this question by flow cytometric analysis of different lymphoma cell lines stained with/without Annexin V (apoptosis marker) and 7AAD (late apoptotic/necrotic marker) in presence or absence of DuoHexaBody-CD37, with and without Fc-crosslinking. These experiments demonstrate that Fc-crosslinking DuoHexaBody-CD37 leads to the induction of apoptosis across DLBCL cell lines (new Supplementary Figure 1A).

      (4) The regulation of JAK/STAT signalling by SHP-1 should be mentioned in the introduction and discussion as this is a key finding of the manuscript.

      Based on the new data on the role of SHP-1 (Suppl. Fig. 5), we have rephrased the text on the SHP1 in the discussion of the revised paper: “DuoHexaBody-CD37 treatment also led to an increase in SHP1 mediated signaling, however we could not confirm a direct role of SHP1 signaling in DuoHexaBody-CD37-mediated cytotoxicity. DLBCL cells may undergo signal rewiring upon SHP1 knockdown by altered levels of p‑AKT, p‑STAT3, and p‑STAT6, or SHP2 may compensate for the loss of SHP1. It is currently unclear what the biological implications are of the increased SHP1 signaling observed upon treatment with DuoHexaBody-CD37 in DLBCL cells.”

      (5) The authors state that DuoHexabody-37 is particularly effective at downregulating STAT signalling in the presence of IL-6 (lines 302-303) however, this is not statistically significant in the results section. There is a trend for a reduction, however, further experimental repeats would clarify this.

      We agree with the reviewer, and rewrote the text on IL-6 in the discussion: “p-STAT3 downregulation upon DuoHexaBody-CD37 treatment in presence of IL-6 requires further investigation in additional IL-6-responsive cell lines, as HBL1 was the only IL-6-responsive lymphoma cell line tested in this study.”

    1. Author response:

      The following is the authors’ response to the current reviews.

      Reviewer #1:

      Yet I think that important aspects of my critique of the first statement of the manuscript about the flaws of [SR] model remain unanswered.

      I believe that I have fully addressed the points in the earlier review. The reviewer had doubted that my results were correct, attributing them to “a poor setup of the model” on my part. The reviewer stated that if I were correct about the factor of >10<sup>43</sup> change in cmax, this would “naturally break down all the estimates and conclusions made in Siljestam and Rueffler” (S&R).

      It appears that the reviewer is now convinced that my results represent a faithful analysis of the models on which S&R based their claims. The reviewer now contends that these results, including the factor of >10<sup>43</sup>, present no difficulties for the claims of S&R after all. In fact, this enormous factor of >10<sup>43</sup> is now claimed to support the conclusions of S&R by invalidating my conclusions. I respond to these new and very different arguments in what follows.

      As I stated in the first round of review, the issue is not the enormity of this factor per se, but the fact that the compensatory adjustment of cmax conceals the true effects of changes in other parameters. These effects are large; small changes to the parameter values mostly eliminate the diversity that the model is claimed to explain.

      The model in [SR] is not phenomenological as none of the parameters or functional forms were derived empirically. Instead, it is a proof of principle demonstration that inevitably grossly simplifies the actual immune response.

      The hidden sensitivity of the results of S&R to paramater values is sufficient to invalidate them as a proof of principle. The manuscript goes further and explains how the problem "is not specific to the details of the models of Siljestam and Rueffler, but is inherent in the phenomenon invoked to allow high diversity" because "any change that affects condition by as much as the difference between MHC heterozygotes and homozygotes will eliminate high equilibrium diversity". This general principle addresses all of the reviewer's points.

      In reality, a new pathogen cannot reduce the "survival" by such a factor as it would wipe out any resident population. So to compensate for such an artifact, the additional factor cmax was introduced to buffer such an excess. There is no reason to fix cmax once for an arbitrary number of pathogens, because varying cmax basically reflects the observation that a well-adapted individual must have a reasonable survival probability.

      This is not a legitimate reason for making compensatory, diversity-promoting adjustments to cmax when evaluating sensitivity to other parameters. If the number of pathogens or their virulence changes, cmax obviously does not automatically change along with it. If the population or species consequently goes extinct, then it goes extinct. If it persists, it does so with the same value of cmax.

      The possibility of extinction arguably puts a minimum value on cmax, but it does not restrict it to a range of values that conveniently leads to high MHC diversity. In the examples that I analyzed, slightly decreasing the number of pathogens or their virulence, which increases survivability, eliminates diversity. This phenomenon obviously cannot be dismissed on the grounds that survivability would be too low for the species to exist.

      S&R in effect assume that the condition of the most fit homozygote remains fixed, regardless of the number of pathogens, their virulence, and myriad other differences between species. It is this assumption that is without justification.

      At the same time, there are many ways in which the numerical simulation may break down when the survival rates become of the order of 10^(-43) instead of one

      I am not sure what is meant by “the numerical simulation may break down”. Numerical error is not a tenable explanation of the lack of diversity observed in that simulation. The outcome is exactly what is expected from purely theoretical considerations: conditions of all genotypes fall on the steep part of the curve, making the mechanism proposed by S&R largely inoperative, so a pair of alleles forming a fit heterozygote comes to predominate. The numerical simulation is actually superfluous.

      Low survival rates are completely irrelevant to the effect of decreasing the number of pathogens or their virulence, which does not lower survival rates, but does eliminate diversity.

      so it comes to no surprise that the diversification, predicted by the adaptive dynamics, does not readily occur in the scenario with an addition or removal of the 8th pathogen with a very high virulence \nu=20.

      Whether or not it surprising, the lack of diversity is a problem for the claims of S&R, as there is no reason to expect the number of pathogens to have just the right value to produce high diversity. Furthermore, for many combinations of values of the other parameters (e.g., my v=19.5 and 20.5 examples), no number of pathogens leads to high diversity.

      Again, the general principle mentioned above makes the details that the reviewer refers to irrelevant. Nonetheless, some additional remarks are in order:

      (1) This comment ignores the fact that removal of a pathogen, or a slight decrease in “virulence”, eliminates diversity without lowering survival rates.

      (2) Small increases or decreases in v (virulence) eliminate diversity without having such large effects on condition.

      (3) In the example emphasized by the reviewer, mean survival rates are nowhere near as low as 10<sup>-43</sup>. Only homozygotes have such low fitness.

      (4) The adaptive dynamics predict the low diversity seen in the simulations, contrary to what the reviewer seems to suggest. Elimination of diversity is not an artifact of the simulation.

      (5) v\=20 was chosen because it is most favorable to the model of S&R in that it yields the highest diversity. Indeed, S&R only observed realistically high diversity with the narrow gaussians that the reviewer objects to. With lower values of v, diversity is much lower, but even this meager diversity is eliminated by small changes in parameter values (see below). If narrow gaussians and large effects of pathogens somehow invalidate results, then they invalidate the high-diversity results of S&R.

      I have doubts that the reported breakdown of the [SR] model with fixed cmax remains observable with less extreme values of m and \nu (say, for \nu=7 and m=3 plus or minus 1 used in Fig. 3 in the manuscript).

      These doubts are unwarrented. With the suggested parameter values, for example, increasing or decreasing m by 1 reduces the effective number of alleles to around 1 or 2. This can easily be checked using the simulation code of S&R, as detailed in my initial response and now in a Supplementary Text. Even without this result, the general principle mentioned above tells us that considering other regions of parameter space cannot rescue the conclusions of S&R.

      So I still find the claim that " the phenomenon that leads to high diversity in the simulations of Siljestam and Rueffler depends on finely tuned parameter values" is not well substantiated.

      What is unsubstantiated is the claim of S&R that “For a large part of the parameter space, more than 100 and up to over 200 alleles can emerge and coexist”. As my manuscript illustrates, this is an illusion created by the adjustment of one parameter to compensate for changes in others.

      The reviewer even acknowledges that “the choice of constants and functions...works in a limited range of parameter values”. Furthermore, the manuscript explains why this problem is inherent to the general phenomenon, not specific to the details of the model or parameter values.


      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It appears obvious that with no or a little fitness penalty, it becomes beneficial to have MHC-coding genes specific to each pathogen. A more thorough study that takes into account a realistic (most probably non-linear in gene number) fitness penalty, various numbers of pathogens that could grossly exceed the self-consistent fitness limit on the number of MHC genes, etc, could be more informative.

      The reviewer seems to be referring to the cost of excessively high presentation breadth. Such a cost is irrelevant to the inferior fitness of a polymorphic population with heterozygote advantage compared to a monomorphic population with merely doubled gene copy number. It is relevant to the possibility of a fitness valley separating these two states, but this issue is addressed explicitly in the manuscript.

      An addition or removal of one of the pathogens is reported to affect "the maximum condition", a key ecological characteristic of the model, by an enormous factor 10^43, naturally breaking down all the estimates and conclusions made in [RS]. This observation is not substantiated by any formulas, recipes for how to compute this number numerically, or other details, and is presented just as a self-standing number in the text.

      It is encouraging that the reviewer agrees that this observation, if correct, would cast doubt on the conclusions of Siljestam and Rueffler. I would add that it is not the enormity of this factor per se that invalidates those conclusions, but the fact that the automatic compensatory adjustment of c</sub>max</sub> conceals the true effects of removing a pathogen, which are quite large.

      I am not sure why the reviewer doubts that this observation is correct. The factor of 2.7∙10<sup>43</sup> was determined in a straightforward manner in the course of simulating the symmetric Gaussian model of Siljestam and Rueffler with the specified parameter values. A simple way to determine this number is to have the simulation code print the value to which c</sub>max</sub> is set, or would be set, by the procedure of Siljestam and Rueffler for different parameter values. I have in this way confirmed this factor using the simulation code written and used by Siljestam and Rueffler. A procedure for doing so is described in the new Supplementary Text S1. In addition, I now give a theoretical derivation of this factor in Supplementary Text S2.

      This begs the conclusion that the branching remains robust to changes in cmax that span 4 decades as well.

      That shows at most that the results are not extremely sensitive to c</sub>max</sub> or K. They are, nonetheless, exquisitely sensitive to m and v. This difference in sensitivities is the reason that a relatively small change to m leads to such a large compensatory change in c</sub>max</sub>. It is evident from Fig. 4 of Siljestam and Rueffler that the level of diversity is not robust to these very large changes in c</sub>max</sub>, which include, as noted above, a change of over 43 orders of magnitude.

      As I wrote above, there is no explanation behind this number, so I can only guess that such a number is created by the removal or addition of a pathogen that is very far away from the other pathogens. Very far in this context means being separated in the x-space by a much greater distance than 1/\nu, the width of the pathogens' gaussians. Once again, I am not totally sure if this was the case, but if it were, some basic notions of how models are set up were broken. It appears very strange that nothing is said in the manuscript about the spatial distribution of the pathogens, which is crucial to their effects on the condition c.

      I did not explicitly describe the distribution of pathogens in antigenic space because it is exactly the same as in Siljestam and Rueffler, Fig. 4: the vertices of a regular simplex, centered at the origin, with unity edge length.

      The number in question (2.7∙10<sup>43</sup>) pertains to the Gaussian model with v\=20. As specified by Siljestam and Rueffler, each pathogen lies at a distance of 1 from every other pathogen, so the distance of any pathogen from the others is indeed much greater than 1/v. This condition holds, however, for most of the parameter space explored by Siljestam and Rueffler (their Fig. 4), and for all of the parameter space that seemingly supports their conclusions. Thus, if this condition indicates that “basic notions of how models are set up were broken”, they must have been broken by Siljestam and Rueffler.

      ...the branching condition appears to be pretty robust with respect to reasonable changes in parameters.

      It is clear from Fig. 4 of Siljestam and Rueffler that the branching condition is far from sufficient for high MHC diversity.

      Overall, I strongly suspect that an unfortunately poor setup of the model reported in the manuscript has led to the conclusions that dispute the much better-substantiated claims made in [SD].

      The reviewer seems to be suggesting that my simulations are somehow flawed and my conclusions unreliable. I have addressed the reasons for this suggestion above. Furthermore, I have confirmed the main conclusion—the extreme sensitivity of the results of Siljestam and Rueffler to parameter values--using the code that they used for their simulations, indicating that my conclusions are not consequences of my having done a “poor setup of the model”. I now describe, in Supplementary Text S1, how anybody can verify my conclusions in this way.

      Reviewer #2 (Public review):

      (1) The statement that the model outcome of Siljestam and Rueffler is very sensitive to parameter values is, in this form, not correct. The sensitivity is only visible once a strong assumption by Siljestam and Rueffler is removed. This assumption is questionable, and it is well explained in the manuscript by J. Cherry why it should not be used. This may be seen as a subtle difference, but I think it is important to pin done the exact nature of the problem (see, for example, the abstract, where this is presented in a misleading way).

      I appreciate the distinction, and the importance of clearly specifying the nature of the problem. However, as I understand it, Siljestam and Rueffler do not invoke the implausible assumption that changes to the number of pathogens or their virulence will be accompanied by compensatory changes to c</sub>max</sub>. Rather, they describe the adjustment of c</sub>max</sub> (Appendix 7) as a “helpful” standardization that applies “without loss of generality”. Indeed, my low-diversity results could be obtained, despite such adjustment, by combining the small change to m or v with a very large change to K (e.g., a factor of 2.7∙10<sup>43</sup>). In this sense there is no loss of generality, but the automatic adjustment of c</sub>max</sub> obscures the extreme sensitivity of the results to m and v.

      (2) The title of the study is very catchy, but it needs to be explained better in the text.

      I have expanded the end of the Discussion in the hope of clarifying the point expressed by the title.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I would like to suggest to the author that they provide essential details about their simulations that would justify their claims, and to communicate with Mattias Siljestam and Claus Rueffler whether claims of the lack of robustness could be confirmed.

      The models simulated were modified versions of those of Siljestam and Rueffler. Thus, only the modifications were described in my manuscript. I have added a more detailed description of how c</sub>max</sub> was set in the simulations concerned with sensitivity to parameter values. In addition, the new Supplementary Text S1, which describes confirmation of the lack of robustness using the code of Siljestam and Rueffler, should remove any doubt about this conclusion.

      Reviewer #2 (Recommendations for the authors):

      I have no further recommendations. The manuscript is well written and clear.

      Thank you.

      Reviewer #3 (Recommendations for the authors):

      (1) Since this is a full report and not just a letter to the editor, it would benefit from a bit more introduction of what the MHC actually is and what the current understanding of its evolution is. Currently, it assumes a lot of knowledge about these genes that might not be available to every reader of eLife.

      I have added some more information to the opening paragraph. I would also note that this report was submitted as a “Research Advance”, which may only need “minimal introductory material”.

      (2) Some more recent literature on MHC evolution should be added, e.g., the review by Radwan et al. 2020 TiG, a concrete case of MHC heterozygote advantage by Arora et al. 2020 MolBiolEvol, and a simulation of MHC CNV evolution by Bentkowski et al. 2019 PLOSCompBiol.

      I have cited some additional literature.

      (3) Since much of the criticism hinges on the cmax parameter, its biological meaning or role (or the lack thereof) could be discussed more.

      I am not sure what I can add to what is in the first paragraph of the Discussion.

      (4) I find it difficult to grasp how the v parameter, which is intended to define pathogen virulence, if I understand it correctly, can be used to amend the breadth of peptide presentation. Maybe this could be illustrated better.

      I have attempted to make this clearer. The parameter v actually controls the breadth of peptide detection conferred by an allele, which, if not identical to the breath of presentation, is certainly affected by it. The basis of the “virulence” interpretation seems to be that narrower detection breadth can, according to the model, only decrease peptide detection probability, which increases the damage done by pathogens.

      (5) Please check sentences in lines 279ff on peptide detection and cost of . There seem to be words missing.

      There was an extraneous word, which I have removed. Thank you for pointing this out.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors investigate ubiquitylation of RPS27A/eS31 by the E3 ligase RNF25 in response to translational stress. Previous studies have identified RPS27A/eS31 ubiquitylation at Lys113 under conditions where translation factors are trapped in the ribosomal A-site. Here, the authors extend this work by testing whether additional translational stress conditions, including amino acid deprivation, induce RPS27A/eS31 ubiquitylation. They further show that GCN1 is required and explore a possible competition between RNF25 and GCN2 for GCN1.

      Strengths:

      This study expands on the range of stress conditions leading to RPS27A/eS31 ubiquitylation, reporting that it occurs in a variety of conditions associated with ribosome stalling, including amino acid deprivation. These observations are useful because they suggest that the RNF25 pathway may not require translation factors trapped in the ribosomal A-site, but may instead respond more broadly to translational perturbations associated with ribosome collisions.

      We wish to point out that our study in fact suggests that the RNF25 pathway is activated by translation factors in the A-site, in agreement with what has been previously proposed, and in addition by stalling conditions that are assumed to not trap translation factors in the A-site. We do not exclude that these conditions might be sampled by A-site binding quality control factors before recognition by RNF25.

      Weaknesses:

      The evidence supporting several of the major claims is incomplete, and additional controls and orthogonal approaches would greatly strengthen the evidence presented.

      We appreciate adding more controls to further substantiate our novel findings. In the course of the revisions we will focus our work on those experiments that do not merely reproduce established facts in the field.

      In particular:

      (1) It is unclear whether the different conditions used to induce translational stress lead to ribosome stalling or collisions. The model presented by the authors seems to rely on ribosomal collisions, but this is not shown. In addition, further investigating amino acid deprivation beyond the removal of Arg or Lys would strengthen the paper.

      We thank the reviewer for the comment. It is correct that we don’t formally show collisions.

      However, the conditions we use have been previously established in the field to induce ribosome stalls and/or collisions, which we may not have pointed out clearly enough. In the revised version, we will include all relevant citations, i.e. for ternatin (Oltion et al., 2023): collisions, anisomycin (Juszkiewicz et al., 2018, Sinha et al., 2020): collisions, emetine (Sinha et al., 2020): collisions, didemnin B (Juszkiewicz et al., 2018, Stoneley et al., 2022): accumulation of ubi-eS10 and changes in polysome profiles indicative of collisions, MMS (Stoneley et al., 2022): changes in polysome profiles indicative of stalls or collisions, starvation -Arg/-Lys (Darnell et al., 2018, Stoneley et al., 2022): accumulation of collided ribosomes only upon GCN2 inhibition, indicative of collisions.

      Secondly, we do not claim to induce collisions when describing the inhibition data (Figure 1 and Figure S1) and were careful to say that we use ‘conditions that cause ribosome stalling’.

      Thirdly, we conclude on collisions when interpreting the data on amino acid starvation (and in our model (Figure 6)), based on our data demonstrating that RNF25 activity in RPS27A/eS31 ubiquitylation is dependent on GCN1 (Figure 3), an established sensor of collided disomes (Pochopien et al., 2021). This conclusion is thus based on the current knowledge in the field.

      We will carefully screen the text for potential points of overinterpretation or confusion between stalling and collisions.

      To address the request of further investigating amino acid deprivation beyond the removal of Arg or Lys, we will include an additional experiment in which we will deplete another amino acid.

      (2) Ubiquitylation of RPS27A/eS31 by RNF25 is used throughout the paper as a readout of RNF25 activity and is assumed to be on Lys113 based on previous work, but is not formally shown here.

      It is established that Lys113 is the main target of RNF25, not only by our work (Montellese et al., 2020), but also by recent work of other groups to which we had referred in our manuscript (Gurzeler et al., 2023, Oltion et al., 2023, Zhao et al., 2026).

      To experimentally address this point, we will add an experiment testing ubiquitylation of RPS27A/eS31 in cells carrying the K113R mutation.

      (3) Rescue experiments of the different mutants used in this study with wild-type and different domain deletions (i.e., ΔRWD for RNF25, ΔRWD-binding for GCN1) would help confirm specificity and strengthen the mechanistic claims.

      Minimally, we will include rescue experiments for RNF25 (using WT, DRWD and enzymatically dead mutant) and, if possible, also for GCN1, which might be more challenging due to its large size and anticipated problems with cloning, cell line generation and protein expression.

      (4) The conclusion that RPS27A/eS31 ubiquitylation supports translation (Figure 4) is based entirely on polysome/monosome ratios, which are difficult to interpret without additional assays of translation output, elongation, or collision.

      It is correct that we base our conclusion on polysome profiles and agree that these are an indirect measure of translation output. However, this assay is well established in the field to show dysregulation of polysome/monosome ratio upon ribosome stalling (Garzia et al., 2017), (Wu et al., 2020), (Chatterjee et al., 2024), (Gurzeler et al., 2023).

      Elongation defects would be expected to lead to stalls and/or collisions (which we conclude on). However, we cannot exclude that there is more initiation when RPS27A/eS31 carries the K113R mutation, although this is hard to rationalize mechanistically and experimentally challenging to exclude. Therefore, to address the point, we will add a sentence that we cannot exclude indirect effects on initiation but consider these unlikely.

      (5) The idea that RNF25 competes with GCN2 for GCN1 binding is interesting, and related models have recently been proposed in RNA damage. The effect of GCN2 KO on RNF25dependent ubiquitylation appears modest, and the data would be strengthened by rescue experiments with wild-type GCN2 and GCN2 mutants defective in GCN1 binding. The authors propose: "that the RNF25 pathway acts as a first line of defence to resolve ribosome collisions, outcompeted by GCN2 binding to GCN1 under acute stress." This model would suggest a further increase in RPS27A/eS31 ubiquitylation upon Arg/Lys deprivation in GCN2 KO cells, since this is the condition in which GCN2 is expected to be activated and engaged with GCN1 (i.e., when it would be competing with RNF25), but no further increase in RPS27A ubiquitylation is observed. It is therefore not clear that these data support the proposed model. Contributing to this may be the fact that many of these assays are performed in a USP16 KO background, which may make it difficult to assess changes in RPS27A/eS31 ubiquitylation.

      We thank the reviewer for the comment. We measure on average a 50% increase in the level of ubiquitinated RPS27A/eS31 in GCN2 KO cells. Considering the large number of ribosomes in a cell (~10<sup>7</sup> per HeLa cell), this 50% increase (from 12.5 to 25% ubiquitinated RPS27A/eS31) amounts to an estimated number of 1,25 x 10<sup>6</sup> of RPS27A/eS31 molecules that get additionally modified, which is clearly a substantial difference, especially compared to the naturally very low levels of RNF25 (in the range of 23’000 molecules (Itzhak et al., 2016)).

      We respectfully disagree that performing experiments in USP16 KO background makes it difficult to assess RPS27A/eS31 ubiquitination. On the contrary. The natural levels of RPS27A/eS31 ubiquitination in WT cells are very low, making quantification sensitive to background fluctuations (see Figure S1). Therefore, in our experience, the usage of USP16 KO makes the quantitative analysis of RPS27A/eS31 ubiquitination robust, allowing us to analyse both increase and decrease in the levels of ubiquitination. We agree that with increasing collisions, the level of ubiquitinated RPS27A/eS31 reaches a plateau in USP16 KO, which may limit the observable increase. Therefore, the substantial 50% increase might indeed underestimate the effect as compared to WT cells. Still, the measurable increase is substantial and robust.

      To experimentally address the point of the reviewer, we will try generating GCN2 KO cells in a WT background, i.e. in absence of USP16 KO, to strengthen our model.

      (6) Given that several RWD domain proteins can interact with GCN1, and that DRG2 KO appears to affect RPS27A/eS31 ubiquitylation (Figure S5), the data do not support the GCN2specific title. The results are more consistent with a broader, incompletely characterized network of GCN1-associated RWD domain-containing proteins that seems to affect RNF25-dependent ubiquitylation rather than with a demonstrated RNF25-GCN2 competition mechanism. Further characterization of GCN2-dependent ISR activation (p-eIF2a and ATF4 WB) in the absence of RNF25 in Arg/Lys starvation will help shed light on the RNF25-GCN2 competition. The authors use K113R, but this is not shown to prevent RNF25 engagement with GCN1, so a RNF25 KO should be used.

      While we fully agree that our data point at a broader network of competition on GCN1, we wished to avoid an overstatement on other pathways than GCN2, since our experimental evidence on DRG2 is limited at the moment. As it stands, changing the title of the manuscript to a more general message, would indeed fuel the view that our claims are incomplete. But we are glad to reconsider this suggestion if further supporting evidence can be obtained in the course of the revision work.

      The reviewer suggests experiments on competition of RNF25 with GCN2. In contrast to the expectation of the reviewer, we do not expect KO of RNF25 to manifest in defects in ISR activation due to the low expression levels of RNF25. In the revised manuscript, we will make clearer that our model refers to competition in the other direction, i.e., of GCN2 with RNF25, which our data supports. The reverse competition of RNF25 with GCN2 is expected to be inefficient to enable a robust activation of the ISR by GCN1 when needed. In addition, other pathways (such as DRG2) might also contribute to the resolution of collisions in the absence of RNF25, affecting the level of ISR activation.

      We feel that further working out these competitive relationships will be interesting to perform in future work. Currently, it is also not clear whether all involved RWD-containing factors bind GCN1 with the same affinity, which is important to consider for the effectiveness of a mutual competition model as suggested by the reviewer.

      Reviewer #2 (Public review):

      Summary:

      The authors show that deprivation of Arginine and Lysine induces a ~50% increase in the ratio of ubi-RPS27A to RPS27A, and this induction requires E3 ubiquitin ligase RNF25. The authors show ZAKalpha and EDF1 are not required for steady state or ribosome stalling-induced ubiRPS27A, while GCN1 is required. The ratio of polysomes to monosomes is increased in RNF25 knockdown cells or when translation is activated by ISRIB in a RPS27A K113R mutant cell line. GCN2 KO cells indicate elevated levels of ubi-RPS27A, and overexpression of the GCN2 RWD domain reduces levels of ubi-RPS27A.

      Strengths:

      (1) The authors identified a novel pathway to sense amino acid deprivation, indicated by ubiRPS27A, previously implicated in ribosome stalling.

      (2) The authors find antagonism between two proteins known to act downstream of GCN1, giving insight into how signaling occurs from an upstream sensor of ribosome stalling to multiple downstream pathways.

      Weaknesses:

      (1) The authors suggest that, based on increased Polysome/Monosome ratios, there is more disome stalling in RNF25 KD cells and RPS27A K113R cells treated with ISRIB, but this readout is very indirect and could be driven by other changes in the cell other than ribosome stalling.

      We thank the reviewer for this important comment. We intentionally used ISRIB in Figure 4F, G to avoid possible effects on initiation, and the results are consistent with our model. While we agree that ISRIB itself might have indirect consequences, these should be the same for the control (WT cells) and the assay condition (K113R cells). We also show the data without ISRIB, which show a similar trend but are less robust (Figure 4D, E). It is very hard to exclude other possible effects which would selectively affect K113R cells in presence of ISRIB.

      (2) While the authors propose that GCN2 and RNF25 compete for binding to GCN1, no evidence was shown that RNF25 binds to GCN1 in cells, nor that the interaction increases when GCN2 is absent.

      The idea of RNF25 binding to GCN1 is based on a previously published work (Oltion et al., 2023, Seidel et al., 2026, Zhao et al., 2026). We will design additional experiments to potentially confirm the interaction between RNF25 and GCN1.

      (3) The use of USP16 to enhance the detection of ubi-RPS27A in many experiments brings the question of whether USP16 KO may alter the protein levels of any known regulators of ribosome collisions? (i.e. ZNF598, GCN1, EDF1, ZAKalpha, etc.) If USP16 KO causes changes in other important regulators of collisions, the authors could be identifying genetic interactions with USP16 in their experiments throughout the paper.

      Indeed, we can’t exclude the effect of USP16 KO on the expression levels of other collision sensors. We will experimentally confirm the levels of other ribosome collision sensors in USP16 KO cells.

      (4) In Figure 5E, the expression level of the GCN2 3K RWD domain looks to be lower than the WT RWD domain; perhaps this could be what is driving the smaller decrease of ubi-RPS27A seen with GCN2 3K vs WT.

      We thank the reviewer for pointing at this issue, which we will experimentally address in the revised version.

      Reviewer #3 (Public review):

      Summary:

      This study examines the role of RNF25 in translational quality control. Previous work indicated that RNF25 is activated by ribosomes stalled with defective elongation or termination factors bound in the A-site. Here, the authors provide evidence that RNF25 is activated by other treatments that evoke ribosome stalling, including amino acid starvation, where the A-site may be empty, leading to ubiquitination of RPS27A in a manner requiring the ISR collision sensor Gcn1, but not EDF1 and ZAKα, involved in the RQC and RSR surveillance pathways. They present some evidence from polysome profiling that RNF25 and its ubiquitination of RPS7A help resolve ribosome collisions and support translation elongation in basal conditions. They further show that KO of Gcn2 increases RPS27A ubiquitination in basal conditions, but not in amino acid-starved cells, and that RPS27A ubiquitination was reduced on overexpressing the WT RWD domain of Gcn2 but not a variant harboring substitutions of residues predicted to bind Gcn1. Based on these findings, they propose a model that, in response to ribosome stalling induced by various stresses, Gcn1 recruits RNF25 via the latter's RWD domain to ubiquitinate RPS27A and thereby resolve ribosome stalling and promote continued elongation. If collisions increase even further, GCN1 recruits GCN2 instead of RNF25 to elicit the ISR.

      Strengths:

      The data is convincing that a variety of triggers leading to diverse stalled ribosomal states, including amino acid limitation, can activate RNF25, suggesting that activation of this pathway does not require the presence of trapped protein factors in the ribosomal A-site but is a more general response to ribosome collisions. It is also convincing that Gcn1 is required for RNF25 activation under all of these conditions, which is consistent with previous findings that Gcn1 is required for RNF25 function in the presence of trapped elongation or termination factors. The finding that EDF1 and ZAK are not needed for RNF25 activation in amino acid starvation conditions is of interest for EDF1, given the recent claim that it is required for full ISR activation.

      Weaknesses:

      (1) The evidence presented from polysome profiling that RNF25 helps resolve naturally occurring ribosome collisions in basal conditions is not compelling, as eliminating RNF25 could be increasing the rate of initiation rather than increasing stalled ribosomes as the means of increasing the P/M ratio. The Rps27A-K113R mutation could have the same effect of increasing initiation, which could have been obscured by inhibiting the ISR with ISRIB.

      Our results indicate that P/M ratio increases upon ISRIB treatment of K113R cells compared to WT cells, aligning with the idea that ISRIB enhances initiation, causing increased loading of ribosomes on mRNA and consequent increased frequency of collisions. As outlined above, we agree that this experiment is indirect and results might be affected by secondary effects. However, we cannot rationalize how inhibition of the ISR by ISRIB would specifically obscure the effect for the K113R mutation but not the WT.

      (2) The evidence that RNF25 competes with Gcn2 for Gcn1 binding is also not compelling. While it's convincing that Rps27A-Ubi is elevated in basal conditions on eliminating Gcn2, loss of GCN2 would be expected to increase ribosome loading on mRNAs, potentially elevating the frequency of collisions and thereby stimulating RNF25 activity indirectly.

      We have not made sufficiently clear that we did not intend to claim that RNF25 efficiently competes with GCN2 (see also response to reviewer 1), which we do not expect due to the low levels of RNF25. Our manuscript is focussed on competition in the reverse direction, i.e. of GCN2 with RNF25.

      We agree that loss of GCN2 may increase ribosome loading on mRNA similar to ISRIB treatment, which could lead to more collisions by enhanced translation and hence increased Rps27A-Ubi. At the same time, however, this does not exclude that loss of GCN2 contributes more directly at the level of RNF25 recruitment. Therefore, the experiment also supports the competition model, and both effects together may contribute to the observed increase in ubiquitylated RPS27A/eS31. Without other evidence, the experiment would remain inconclusive.

      Therefore, to directly test the competition model, we had overexpressed the GCN1-binding RWD domain of GCN2, which leads to decreased levels of ubiquitinated RPS27A/eS31, lending direct support to the competition model of GCN2 with RNF25, which is consistent with similar models recently proposed by two other manuscripts (Seidel et al., 2026, Zhao et al., 2026).

      (3) It's also quite puzzling and left unexplained why they observed no further increase in Rps27AUbi on -Arg/-Lys starvation in the cells lacking Gcn2. Why wouldn't -Arg/-Lys starvation lead to further stalling and RNF25 activation in the absence of Gcn2? (Since Gcn2 KO increases Rps27A-Ubi in the presence +Arg/+Lys conditions, it can't be that Gcn2 is required for RNF25 function.) The same puzzling and unresolved observation was made in the cells lacking DRG2. One possible explanation for this conundrum is that low-level RNF25 abundance limits further activation.

      Over all of our experiments, we have observed that RPS27A-Ubi reaches a plateau of about 30% to 35% of total RPS27A in the USP16 KO background (GCN2 deletion or amino acid starvation). This plateau indeed limits seeing further increases. We do not know the underlying reason but note that under these conditions about one third of 40S subunits carry ubiquitin on RPS27A/eS31. As the reviewer suggests, RNF25 is expressed at low levels (in the range of 23’000 molecules, (Itzhak et al., 2016); see point 5 of reviewer 1), likely rendering it the limiting factor for further ubiquitination events.

      To circumvent the plateau issue, we will attempt to generate GCN2 KO cell lines in the WT background for the starvation experiments (see also response to reviewer 1, point 5).

      (4) The quantitative effects of overexpressing the Gcn2 RWD domain on Rps27A-Ubi, constituting their other evidence presented to support the competition model, are quite small in magnitude.

      We respectfully disagree with the reviewers’ comment concerning the magnitude of the effect. There is a ~27% decrease in ubiquitination, which is substantial considering the number of 40S ribosomal subunits and possible consequences of such change. It should also be noted that this is a transient transfection experiment not hitting all cells of the population. We will repeat the experiment, optimizing the expression of the negative control construct.

      Cited literature:

      Chatterjee S, Naeli P, Onar O, Simms N, Garzia A, Hackett A, Coyle K, Harris Snell P, McGirr T, Sawant TN et al. (2024) Ribosome Quality Control mitigates the cytotoxicity of ribosome collisions induced by 5-Fluorouracil. Nucleic Acids Res 52: 12534-12548

      Darnell AM, Subramaniam AR, O'Shea EK (2018) Translational Control through Differential Ribosome Pausing during Amino Acid Limitation in Mammalian Cells. Mol Cell 71: 229-243 e11

      Garzia A, Jafarnejad SM, Meyer C, Chapat C, Gogakos T, Morozov P, Amiri M, Shapiro M, Molina H, Tuschl T et al. (2017) The E3 ubiquitin ligase and RNA-binding protein ZNF598 orchestrates ribosome quality control of premature polyadenylated mRNAs. Nat Commun 8: 16056

      Gurzeler LA, Link M, Ibig Y, Schmidt I, Galuba O, Schoenbett J, Gasser-Didierlaurant C, Parker CN, Mao X, Bitsch F et al. (2023) Drug-induced eRF1 degradation promotes readthrough and reveals a new branch of ribosome quality control. Cell Rep 42: 113056

      Itzhak DN, Tyanova S, Cox J, Borner GH (2016) Global, quantitative and dynamic mapping of protein subcellular localization. Elife 5

      Juszkiewicz S, Chandrasekaran V, Lin Z, Kraatz S, Ramakrishnan V, Hegde RS (2018) ZNF598 Is a Quality Control Sensor of Collided Ribosomes. Mol Cell 72: 469-481 e7

      Montellese C, van den Heuvel J, Ashiono C, Dorner K, Melnik A, Jonas S, Zemp I, Picotti P, Gillet LC, Kutay U (2020) USP16 counteracts mono-ubiquitination of RPS27a and promotes maturation of the 40S ribosomal subunit. Elife 9  

      Oltion K, Carelli JD, Yang T, See SK, Wang HY, Kampmann M, Taunton J (2023) An E3 ligase network engages GCN1 to promote the degradation of translation factors on stalled ribosomes. Cell 186: 346-362 e17

      Pochopien AA, Beckert B, Kasvandik S, Berninghausen O, Beckmann R, Tenson T, Wilson DN (2021) Structure of Gcn1 bound to stalled and colliding 80S ribosomes. Proc Natl Acad Sci U S A 118

      Seidel AS, Nemcekova L, Grønbæk-Thygesen M, Shi X, Ramalho S, Mordente KC, Bekker-Jensen S, Haahr P (2026) RNF25 restrains GCN2 hyperactivation to sustain protein synthesis and cell proliferation in response to RNA damage. bioRxiv

      Sinha NK, Ordureau A, Best K, Saba JA, Zinshteyn B, Sundaramoorthy E, Fulzele A, Garshott DM, Denk T, Thoms M et al. (2020) EDF1 coordinates cellular responses to ribosome collisions. Elife 9

      Stoneley M, Harvey RF, Mulroney TE, Mordue R, Jukes-Jones R, Cain K, Lilley KS, Sawarkar R, Willis AE (2022) Unresolved stalled ribosome complexes restrict cell-cycle progression after genotoxic stress. Mol Cell 82: 1557-1572 e7

      Wu CC, Peterson A, Zinshteyn B, Regot S, Green R (2020) Ribosome Collisions Trigger General Stress Responses to Regulate Cell Fate. Cell 182: 404-416 e14

      Zhao S, Palma-Chaundler CS, Engel CM, Cordes J, Nixdorf D, Luo MY, Kaya S, Suryo Rahmanto A, van den Heuvel D, Mackens-Kiani T et al. (2026) RNF25 confers mRNA damage tolerance by curbing activation of the integrated stress response. Mol Cell 86: 1275-1292 e12

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study identifies mutations in alpha-tubulin that suppress Tau-induced neurodegeneration using the C. elegans model of Tauopathy, suggesting a potentially interesting role for microtubule properties in modulating Tau toxicity. These missense mutations cluster in the C-terminal Tau-interacting helix 12 region of alpha-tubulin genes (tba-1, tba-2, and mec-12). Further analysis, particularly using the strongest suppressor tba-2, shows that it rescues Tau-induced behavioral deficits and neuronal loss without significantly altering bulk tau-phosphorylation, aggregation, or binding to soluble tubulin. The authors suggest that altered microtubule properties underlie the neuroprotective effects, and manipulating microtubule properties may have therapeutic potential.

      Strengths:

      The study is conceptually interesting as it shows that Tau-induced neurotoxicity can, in this model, be partially uncoupled from canonical pathological hallmarks such as Tau-hyperphosphorylation and aggregation. The identification of multiple independent mutations in the same structural region of three alpha-tubulin genes provides support for the functional relevance of helix 12 in modulating Tau-induced toxicity. The authors demonstrate significant rescue of behavioral deficits (using motility and manual thrashing assays) and neuronal loss in both WT-tau and FTLD-associated TauV337M in combination with mutant alpha-tubulins, suggesting a general mechanism for tubulin-regulated modulation of Tau-toxicity. Moreover, the correlation between mutant tubulin expression levels and the extent of rescue supports a causal relationship.

      Weaknesses:

      One of the major claims of this manuscript is that altered microtubule properties suppress Tau toxicity. The only supporting evidence in this context provided by the authors is reduced taxol-stabilized microtubule mass, which does not fully explain neuronal loss or the rescue of behavioral deficits. What remains unclear is whether these mutations alter microtubule dynamics, catastrophe, lattice stability, or axonal transport.

      We agree with Reviewer #1’s critique that the evidence presented does not fully explain neuronal loss and requires further investigation. This first manuscript characterized the mutations discovered through forward genetic screening techniques and provided data to support the positive correlation mutant expression and level of suppression. We believe the studies and data presented here help to formulated the next testable hypotheses, and guide the next lines of experimentation. We are encouraged by Reviewer #1’s assessment that exploration of microtubule dynamics, catastrophe, lattice stability and axonal transport will be critical to testing the hypothesis that mutant tubulin drives suppression of tau toxicity through changes to microtubule properties. These suggestions are highly relevant and align with our priorities as we recently submitted an application for a 5-year research award to support these key questions.

      To address this specifically, the reviewer recommended “The microtubule-dependent axonal transport should be examined in tubulin mutants and compared with mutant tubulin + Tau conditions. Imaging of mitochondrial or synaptic vesicle markers, along with appropriate quantifications (velocity or run length), may provide a functional readout linking microtubule changes to neuronal survival.”

      We agree with the reviewer that these experiments will be highly valuable to further understand the mechanisms underlying suppression, and we have planned to complete these experiments upon receipt of funding that would directly support the completion of these experiments.

      The authors show that mutant tba-2 reduces total tau levels by ~45%. This level of reduction is likely significant but underexplored in the manuscript. Why are the Tau levels reduced? How is Tau getting cleared- is there enhanced autophagy or ubiquitin-proteasome pathway getting upregulated in tba-2 + Tau animals? Or one or more of the Tau species not detectable by the antibodies used in this study? The observation that the mec-12 mutant rescues Tau-induced phenotypes without altering Tau levels suggests that suppression can occur through Tau-independent mechanisms. This raises an important unresolved question regarding the extent to which suppression is Tau-dependent vs Tau-independent across different mutant alpha-tubulin genes, complicating the interpretation of the rescue phenotypes.

      We think the reviewer has addressed an important point that there may be both tau-dependent and tau-independent mechanisms at work here, and we will add greater nuance to this in our discussion. Additionally, we agree these two potential mechanistic pathways merit further exploration. To address this, we have planned to conduct experiments using reporter C. elegans lines crossed with our mutant tubulin/tau-transgenic lines to detect potential upregulation of these pathways as mechanisms for tau clearance.

      Given that Tau primarily associates with the microtubule lattice in vivo, measuring interactions with soluble tubulin may not fully capture biologically relevant binding dynamics and therefore does not exclude the possibility that these mutations alter tau-microtubule interactions at the lattice level or may affect the binding of other MAPs/regulators, thereby altering stability or trafficking.

      In the discussion we acknowledge the limitation of only examining the binding affinity between soluble tubulin and tau and intend to complete further studies with polymerized microtubules containing mutant α-tubulin. We will expand discussion of this in the text. Similar to reviewer 1, we have also concluded that the next line of experimentation will focus on mutant alpha-tubulin effects on the microtubule polymer such as changes to MAP interactions, stability and trafficking. We have applied for and hope to receive funding to address these questions in the near future.

      To address this concern specifically, we plan to conduct these experiments using C. elegans extracts to polymerize microtubules and subsequently test the binding of recombinant human tau. These co-sedimentation experiments are expected to be included in the revised manuscript.

      A large body of conclusions is drawn from behavioral rescue and biochemical assays. This limits the understanding of how molecular changes in tubulin might affect cellular mechanisms of neuroprotection. Are there changes in the neuronal microtubule organization, Tau localization, or its redistribution in the mutant alpha-tubulin background? Are there differences in soluble vs oligomeric vs insoluble Tau in mutant tba-2 and mec-12 animals?

      The reviewer raises relevant questions regarding elucidation of the mechanisms underlying mutant tubulin-mediated suppression at the cellular level. To address this concern we will analyze the cellular distribution of tau in neurons from mutant and non-mutant C. elegans.

      Ultimately, our goals are to identify and connect the underlying biochemical mechanisms with the observed prevention of cell death as Reviewer 1 has identified. Their suggestion to explore cellular-level changes such as mutant tubulin effects on tau distribution is highly relevant. We therefore plan to test this directly by imaging neurons in C. elegans strains expressing fluorescently labeled tau and/or immunohistochemical techniques to stain for tau in C. elegans neurons.

      The suppression of behavior in the co-pathology model is interesting but mechanistically insufficient, mainly because the underlying basis of suppression is not examined in these models. Moreover, it remains unclear whether tubulin-Tau genetically interacts with Aβ or TDP-43, and what cellular mechanisms account for the partial rescue observed in these co-pathology models.

      In agreement with Reviewer #1’s assessment, we have concluded these data, while interesting, do not substantially expand our understanding apart from the existing data. Without additional information regarding the underlying mechanisms, they do not provide substantial novel insights and we have therefore chosen to remove the co-pathology data sets from the revised version of the manuscript to refine the scope of the data and hypotheses discussed in this work.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Benbow et al. identifies, through a genetic screen, key tubulin mutants that, with high confidence, rescue tau-mediated ND phenotypes. This manuscript is well written, and the experimental results strongly support the authors' claims that these tubulin mutants can rescue ND-linked phenotypes in C. elegans while having little to no direct effect on Tau aggregation.

      Strengths:

      Benbow et al. use a relatively unbiased forward genetic screen to identify mutations associated with phenotypes that suppress tauopathy-related defects. The authors then logically focus on the various α-tubulin missense mutations identified in H12, which are known to localize to the external face of microtubules. The authors also carefully compare their established tauopathy-associated phenotypes in the WT TauH model, with and without specific α-tubulin mutations, using appropriate controls throughout. Lastly, the authors provide partial mechanistic insight into the α-tubulin mutant-mediated rescue, showing that these effects are independent of tau aggregation and tau phosphorylation, and instead suggest that the α-tubulin mutations may confer altered microtubule assembly properties based on the sedimentation assays.

      Weaknesses:

      While the claims are largely supported by the experimental outcomes, the authors at times do not provide enough detail in the text for readers to interpret the data sets independently. In addition, some claims appear to be slightly overstated relative to the data or the degree of error associated with those data.

      We appreciate the feedback regarding the need for additional clarity for independent analysis of the datasets. We will revise the figures and text to increase clarity for the readers. We will review statements and edit language in accordance with their degrees of error as appropriate.

      The authors measure tau binding affinities using soluble tubulin but do not assess tau binding to assembled microtubules. This is an important limitation, as the physiologically relevant interaction involves α/β-tubulin heterodimers, either free or incorporated into the microtubule lattice. Furthermore, the binding analysis appears to focus only on the D429N α-tubulin mutant, which further limits physiological relevance, as β-tubulin, which is also required for normal tau binding, is not explicitly considered.

      We acknowledge that the limited conclusions may be drawn from soluble tubulin interactions with tau and additional analysis with polymerized microtubules will be useful in understanding tau-microtubule binding affinity. The analysis was completed with isolated pools of tubulin from C. elegans, not recombinant mutant tubulin, so this is a heterogenous mixture of tubulin composed of α/β heterodimer subunits, and a mixture of the mutant isotype within the larger pool of wild type isotypes. While this further complicating the analysis, and is the likely source of variability, it incorporates the normal heterodimer subunit biochemistry.

      Given that tau prominently binds the microtubule lattice we agree with the reviewers that the assessment that experiments with polymerized microtubules containing mutant tubulin would offer a greater understanding of the effects of mutant alpha-tubulin on microtubule properties and potential mechanisms of toxic tau suppression. To test this directly we intend to complete co-sedimentation experiments using C. elegans extracts from wild type and mutant tubulin expressing C. elegans incubated with recombinant human tau.

      In conclusion, the thoughtful commentary and suggestions from reviewers will help improve the manuscript. We plan to complete the following experiments to address their concerns.

      (1) Assess tau localization in mutant tba-2 and mec-12 C. elegans as compared to tau-transgenic C. elegans without tubulin mutations. We plan to use immunohistochemical techniques and/or imaging of Dendra2-labeled tau to assess the sub-compartmental distribution of tau in C. elegans neurons. This addresses Reviewer #1’s question of whether the mutant tubulin changes tau localization in neurons.

      (2) Assess changes mutant-tubulin driven changes to tau affinity for polymerized microtubules. To address both reviewers concerns regarding the limitations of biding experiments with tau and soluble tubulin, We plan to use C. elegans extracts to tests whether microtubule polymers containing mutant alpha-tubulin alter tau-microtubule co-sedimentation.

      (3) Using C. elegans reporter lines we plan to assess whether tau clearance occurs in tba-2 mutant tubulin C. elegans through the upregulation of autophagy or ubiquitin degradation pathways.

      (4) Evaluate the neuroprotective effects of mutant alpha-tubulin in cholinergic neurons using a C. elegans strain expressing a fluorescent label specifically in cholinergic neurons.

      We plan to make textual revisions to increase clarity, aid in independent analysis of the presented datasets, and better address the possibility of both tau-dependent and tau-independent mechanisms. We appreciate the Reviewers attentive reading and thoughtful feedback for the improvement of this manuscript.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this article, the authors couple a 3d vertex model to the extracellular matrix and include activity through contractile springs at the edge. They study, sequentially, the distribution of shear stresses in liquid and solid spheroids, the correlation between stress and cell shape, and the spatial distribution of stresses. The authors find that stresses are higher in solid spheroids (somewhat unsurprisingly), but that the stress distributions are wider in the fluid spheroids. Moreover, stress and shape are not correlated with each other in solids (that seems to be due to vertex model peculiarities), but they are for liquids. In contrast, for solids, the stresses are concentrated at the interface. The authors attribute a lot of the phenomenology to strain-stiffening properties of vertex models as being akin to a network model (correctly in my opinion). Then they strain individual cells and confirm this link, though I missed any explanation of how they did this. Would it have to be within a medium for computational consistency?

      We thank the reviewer for this helpful comment. The current manuscript already describes this procedure in Sec. II.C, “Cell strain-stiffening with volume-preserving deformations,” where we state that individual cells are taken from the final spheroid configuration and then strained by imposing a prescribed volume-preserving deformation along their principal elongation axis. Figure 4 then compares the original and strained cells and shows the resulting increase in maximum shear stress.

      We agree, however, that this point was not explained clearly enough. In the revised manuscript, we will make explicit that this is a single-cell deformation test designed to isolate the intrinsic strain-stiffening response of the vertex-model cell. The cell does not need to remain embedded in a surrounding medium for this specific test, since the goal is not to simulate the full coupled cell–ECM dynamics, but rather to measure how the stress of an individual vertex-model cell changes under imposed strain.

      Indeed, single cells can exhibit strain stiffening as presumably can a spheroid. However, given that we are studying strain stiffening in the context of single/few cell breakout, we also plan to measure the stress in the breakout cells in the extended vertex model to determine the extent of strain stiffening given the surrounding medium of fibers and cells.

      Finally, they generate an extended vertex model, where they replace the single face linking cells with a double face and mechanoresponsive springs. This allows for stronger coupling of individual cell motion to eventual movement out of the spheroid.

      Strengths:

      Coupling a three-dimensional vertex model to the extracellular matrix, modelled as a crosslinked fiber model, is a computational tour-de-force. Adding activity through fluctuations at the interface is also of the correct symmetry (stresses), instead of the self-propulsion which has been used by other authors, and which is not compatible with Newton's 3rd law. This also allows for accurate back-and-forth mechanical coupling between the cells and the ECM.

      I would like to highlight that deriving vertex model stress tensors in full three dimensions is an open problem due to the complex topology. Any progress is valuable, and decomposing things into tetrahedra like here will allow for connections with, in particular, finite element approaches. Therefore, adding some of these results (eq. 13) to the main text would strengthen the paper in my opinion.

      Adding the nonlinear springs to the VM in the 3rd act is a good idea, and a first step to mechanical feedback. One might argue that at this point, removing the vertex model part would even be an option.

      Weaknesses:

      The paper is written in a very qualitative manner, with all of the model equations and analysis hidden in the supplementary information. I do not understand this choice, as it makes things fuzzy and hard to read. The conclusion is also very long and simply reiterates the previous points.

      At the same time, this paper is rather thin on new results and reads more like a handful of new simulations carried out using the method established in [10] (from largely the same authors). Moving some of the actual results to the main text would help, in particular, the 3d stress formulation and the definitions of different measures.

      We thank the reviewer for this constructive criticism. We agree that the main text was too qualitative and that placing most of the equations and definitions in the Supplement made the manuscript harder to read. In the revised version, we will move the essential technical material into the main text, including the 3D cell stress formulation, the definitions of maximum shear stress and cell-shape anisotropy, and the stress–shape alignment measure. Longer derivations and implementation details will remain in the Supplement.

      We will also shorten and reorganize the Discussion/Conclusion to avoid reiterating previous points. Finally, we will revise the presentation to make the new contributions beyond Ref. [10] clearer: the 3D polyhedral-cell stress formulation, the stress-distribution and spatialpatterning analyses, the single-cell strain-stiffening test, and the extended adhesion-spring model used to distinguish single-cell from multi-cell breakout. These changes should make the paper less qualitative and make the main results more visible in the body of the manuscript.

      Vertex models also have a very clear limitation: They cannot model the transition from a confluent to a non-confluent tissue, and individual cells or groups of cells leaving the spheroid. Even having a surface and having significant deformations of the surface are numerically dicey, so the current model is at the edge of what is feasible. The model as written can only do "invasion" by a single cell moving outward, and then another following it a bit (or not).

      I strongly suspect that further progress on 3d cell models will need particle-based models or models where cells are fully meshed surfaces (some of which are in development currently).

      However, none of these problems is mentioned anywhere in the text. The authors also do not review the increasingly broad zoology of other models.

      We thank the reviewer for raising this important limitation of standard vertex models. We agree that a strictly confluent 3D vertex model is not designed to fully capture the transition from a confluent tissue to freely migrating detached cells, and we will make this limitation explicit in the revised Discussion. However, the standard 3D vertex model can still capture collective spheroid deformation, surface remodeling, and local protrusive deformations prior to complete breakout. Thus, it remains useful for studying the mechanical state of the spheroid and the onset of outward deformation before full cell detachment.

      At the same time, we clarify that this very limitation motivated the extended vertex model introduced in Sec. II.D and Supplement G. In this model, cells no longer share interfaces as in a standard confluent vertex model; instead, neighboring cells interact through explicit, tunable cell– cell adhesion springs. This allows us to represent, in a coarse-grained mechanical way, the separation of a boundary cell from the spheroid and the motion of a follower cell behind it. Thus, while the model does not describe full post-detachment migration, it partially addresses the confluent-to-nonconfluent transition at the level needed to study the mechanical onset of breakout.

      We will revise the manuscript to make this distinction clearer and state that our goal is to identify minimal mechanical ingredients for incipient breakout—strain stiffening, adhesion weakening, and adhesion anisotropy—rather than to provide a complete model of long-time invasion.

      We will also note that the current Introduction already discusses several existing modeling approaches, including cellular automaton simulations, a 2D Voronoi model, phenotypeswitching/ECM-remodeling models, and the prior 3D vertex–fiber framework. However, we agree that this discussion should be broadened, and we will add a more explicit comparison with particlebased, phase-field, cellular Potts, and fully meshed deformable-surface models, which may be better suited for later-stage non-confluent migration.

      Reviewer #2 (Public review):

      Summary:

      The manuscript concerns the mechanisms by which cells in a spheroid embedded in the extracellular matrix can escape, either as single or multiple cells.

      Strengths:

      Overall, the manuscript is well written and easy to follow. The claims are mostly justified by the data. Some data can be better analyzed and presented to strengthen the conclusion.

      Weaknesses:

      (1) The description around Figure 2c is not exactly well supported by their results. While values close to 0 for sigma3 dot g3 for solid-like spheroids indicate little correlation between the direction of maximum stress and maximum elongation, this analysis alone does not imply that highly stressed cells are necessarily less globular. The dot product combines the magnitudes of the two vectors and the angle between them. For the distribution graph, it would be useful to have the cumulative frequency equal 1.

      We thank the reviewer for pointing this out. We agree that the interpretation of Fig. 2c should be stated more carefully. In our calculation, the vectors used in the dot product are normalized eigenvectors of the stress tensor and the gyration tensor. Thus, the plotted quantity measures only directional alignment between the principal stress direction and the cell elongation axis, not the magnitudes of stress or shape anisotropy. We will revise the text to make this explicit.

      We also agree that Fig. 2c alone does not support statements about whether highly stressed cells are more or less globular. It only quantifies alignment between stress and shape directions. To address this, we will add or refer to an additional analysis, such as the correlation between maximum shear stress and cell-shape anisotropy, or the shape-anisotropy distribution conditioned on high-stress cells.

      Finally, we agree that the distribution in Fig. 2c should be normalized more clearly. In the revised figure, we will plot the distribution as a probability density or cumulative distribution with total probability equal to one, and we will update the caption accordingly.

      (2) One of the central claims of the paper is that morphology alone is not a reliable indicator of mechanical state. Since the authors compute cellular stresses and cellular shape in their simulation (i.e., Figure 3a and b), can the authors directly plot these two quantities for individual cells in solidlike and fluid-like spheroids?

      We thank the reviewer for this helpful suggestion. We agree that a direct cell-by-cell comparison of cellular stress and cellular shape would strengthen the central claim that morphology alone is not a reliable indicator of mechanical state. In the revised manuscript, we plan to add scatter plots of maximum shear stress versus cell-shape anisotropy for individual cells in both solid-like and fluid-like spheroids.

      (3) There is experimental evidence showing the solid stress inside a spheroid is higher than at the periphery (e.g., https://www.nature.com/articles/ncomms14056). How does this cellular stress relate to these experimental measurements, since they are opposite to what is simulated here (i.e., the authors find max shear stress is lowest in the center and increases towards the boundary, which is opposite to what is measured?

      We thank the reviewer for raising this important point. We agree that the comparison with experimental stress measurements in compressed spheroids should be clarified.

      The main distinction is that the cited experiments measure local pressure, or isotropic compressive stress, from the volume change of embedded elastic beads. In contrast, Fig. 3 in our manuscript shows the cellular maximum shear stress, which reflects the deviatoric part of the cell stress tensor. These quantities do not necessarily have the same spatial profile: a region can be under high isotropic compression while having low shear stress. The loading conditions are also different. The experiments apply external osmotic/mechanical compression to the whole spheroid, whereas our simulations consider active cell–ECM coupling through contractile linker springs at the spheroid boundary. Thus, the elevated boundary shear stress in our model reflects local cell– ECM force transmission, not internal hydrostatic pressure. We indeed will revise the manuscript to make this distinction explicit, cite this experimental work, and avoid implying that maximum shear stress is directly comparable to measured solid pressure. Where appropriate, we will also discuss the isotropic component of the simulated cell stress tensor as a more direct comparison to pressure-based measurements.

      (4) It's worth pointing out that stress fibers aren't really prominent in cells in 3D spheroids. Nonetheless, cells moving on collagen fibers would have stress fibers and utilize contractile actomyosin bundles to generate traction forces.

      We thank the reviewer for this clarification. We did not intend to imply that prominent stress fibers are generally present in cells within the interior of 3D spheroids. The relevant statements in the manuscript were meant to refer to strained boundary cells or cells engaging collagen fibers during mesenchymal-like motion. We will revise the wording in Secs. II.C and II.D to make this distinction explicit and avoid suggesting that bulk spheroid cells generally contain prominent stress fibers.

      (5) In section 2D, it talks about the result that as the kcc associated with the boundary cell is decreased 10-fold for every 5 percent strain decrease in the fiber target spring length, can this result be shown? I have a hard time seeing where this came from.

      We thank the reviewer for this comment. The 10-fold decrease in kcc for every 5% decrease in the fiber target spring length was meant as a phenomenological adhesion-weakening protocol, not as a directly measured law. We agree that this was not made clear enough. In the revised manuscript, we will explicitly state this.

      (6) The results of single-cell vs. two-cell breakouts shown in Figure 5 b and c are very qualitative and should be accompanied by some quantitative comparison.

      We thank the reviewer for this helpful suggestion. We agree that the current presentation of Fig. 5b,c is too qualitative. In the revised manuscript, we plan to add a quantitative comparison between the single-cell and two-cell breakout cases. Specifically, we plan to track the displacement of the pulled boundary cell, the separation between this leader cell and its neighboring/follower cell, and the distance between the follower cell and the remaining spheroid as the fiber target length is decreased.

      Reviewer #3 (Public review):

      Summary:

      The authors describe a mathematical and computational approach used to compute stresses and cellular deformations in a multicellular spheroid embedded in a fiber network. This approach is then used to predict stress and cellular anisotropy distributions in "solid-like" and "fluid-like" spheroids. Simulations show that shear stresses in solid-like spheroids are large and concentrated at the boundary of the spheroid, yet cells do not align with the direction of the largest shear. Conversely, shear stresses in fluid-like spheroids are smaller and uniformly distributed in the spheroid. In this case, cellular elongation is more likely to be aligned with the direction of the largest shear stress. The model and simulations also predict a nonlinear stress-strain relationship that is indicative of strain stiffening. This strain-stiffening is more pronounced in fluid-like spheroids. In an extension of the preliminary polyhedral vertex model, in which cellular interfaces are shared, the authors incorporate mechanical cell-cell interactions via adhesion springs between neighboring vertices. Using this extension, they show that cell breakout is more likely to occur in fluid-like spheroids, where cells are more likely to elongate and stiffen, allowing for larger forces to be exerted on the surrounding fiber network. Furthermore, the authors state that anisotropic cellcell adhesion is required for multicell streaming during breakout.

      Strengths:

      The modeling and computational approach used in this research is this work's biggest strength. Treating the embedded spheroid as a set of polyhedra, where each polyhedron represents a single cell, is a mechanically robust, yet still tractable way to model multicellular spheroids in three dimensions. Starting with expressions for constraining cell volume and surface area as well as a surface energy term, the authors derive an expression for an averaged stress tensor for each polyhedron. This allows the authors to approximate the stress in each polyhedral cell that is caused by cellular deformations during mechanical interactions with the extracellular fiber matrix. This is a clever and robust approach that is based on fundamental mechanical principles that allow one to make reasonable predications about the mechanical state of the spheroid under a variety of conditions.

      Weaknesses:

      The weakness of the manuscript is the exposition. There are significant pieces of critical information missing from the manuscript that would make the presented work significantly more understandable and better support the authors' claims. Most importantly, many necessary details of the model are missing. I was able to get a better understanding of some of these details by reading the authors' earlier work (ref [10] in the submitted manuscript), and for this reason, I do feel that this work has value. However, several descriptions must be added for the paper to be more readily understandable.

      These include

      (1) A better explanation of what drives motion, in particular in the case where no external fiber network is present.

      We thank the reviewer for pointing this out. We agree that the source of motion should be described more clearly. In the embedded simulations, motion arises from overdamped dynamics driven by the forces from the total mechanical energy, including spheroid mechanics, fibernetwork elasticity, and active contractile linker springs at the boundary. The shortening of the linker-spring target lengths provides the active cell–ECM pulling, while effective fluctuations promote cell-shape fluctuations and rearrangements.

      When no external fiber network is present, these linker-mediated cell–ECM forces are absent. The spheroid then evolves only through vertex-model mechanical relaxation, surface tension, cell rearrangements, and effective fluctuations. We will clarify that this no-network case is a control for the intrinsic spheroid stress state, not a simulation of ECM-driven invasion.

      (2) What physically distinguishes fluid-like spheroids from solid-like spheroids? Simply stating the value of the parameters s0 with no explanation is not sufficient.

      We thank the reviewer for pointing out that the physical distinction between solid-like and fluid-like spheroids was not sufficiently explained. We agree that simply stating the values of s_0 is not adequate.

      In this 3D vertex model, the target shape index s_0 controls the mechanical cost of cell rearrangements. Below the rigidity transition (s_0 < s_0^), neighbor exchanges are associated with finite energy barriers, leading to slow structural relaxation and solid-like behavior. Above the transition (s_0 > s_0^), these barriers become very small or vanish, allowing cells to readily move past one another and continuously reorganize their local neighborhood structure. The resulting tissue exhibits fluid-like behavior with efficient stress relaxation through cell rearrangements.

      This distinction was characterized in detail in Ref. [9], where the bulk 3D vertex model was shown to undergo a rigidity transition at approximately (s_0^*=5.39), based on the decay of the neighbor-overlap function and cell trajectories. The solid-like value used here lies below this transition, whereas the fluid-like value lies above it. We acknowledge that the present manuscript only briefly summarized this point, mainly in Supplementary Material A. In the revised manuscript, we will add a clearer explanation in the main text of how the target shape index controls the state of the spheroid and why the selected values correspond to solid-like and fluidlike regimes.

      (3) An explanation of how histograms in Figure 2 are calculated is necessary. Are these histograms based on one simulation or several simulations?

      We thank the reviewer for pointing out that this was not sufficiently clear. The histograms in Fig. 2 are obtained by pooling cell-level quantities from multiple independent simulations, not from a single realization. As listed in Table I, we use 30 independent realizations. We plan to state this explicitly in the revised figure caption and main text.

      (4) The experimental results are briefly mentioned, but significantly more connection between these results and the numerical results of the cell breakout model is needed.

      We agree. In the current manuscript, the experimental data are used mainly to motivate the single-cell and streaming-like breakout modes shown in Fig. 5. We plan to revise Sec. II.D and the Fig. 5 caption to make the connection more explicit: the MEF spheroid experiments show the invasion modes that motivate the model, while the extended vertex model tests minimal mechanical ingredients capable of producing analogous single-cell and follower-cell breakout.

      (5) The description of the model that incorporates variable cell-cell attachments and cell breakout is very terse and needs more detail. Moreover, while the description of the results of this model is strong, the figure that illustrates cell breakout (Figure 5) is difficult to interpret. Addressing these and other issues will make the current manuscript, which presents an interesting model and result, much stronger and easier to read.

      We thank the reviewer for this constructive assessment. We agree that the extended model with variable cell–cell attachments was described too tersely and that Fig. 5b,c was difficult to interpret in its current qualitative form.

      To make Fig. 5 more quantitative, we plan to add measurements comparing the single-cell and two-cell breakout cases. Specifically, we plan to track the displacement of the pulled boundary cell, the separation between this leader cell and its neighboring/follower cell, and the distance between the follower cell and the remaining spheroid as the fiber target length is decreased.

    1. Author response:

      The following is the authors’ response to the current reviews.

      Public Review:

      Reviewer #1 (Public review):

      Suggestions to clarify the study:

      In the revised version, the authors carefully consider these suggestions and provide further details, clarifications and even some new results. Regarding the question of how infection of a cell with one virus could lead to lower probability for a secondary infection, I think that it is possible that infected cells activate antiviral programs that lead, for example, to lower expression of surface receptors. This has been considered at least in hepatitis C virus infection. However, this is a minor point.

      Yes, the possibility that infection of a cell by a virion would reduce chance of infection by another virion was allowed in our model. However, such as a process will not result in apparent cooperativity (n>1) in our model, and thus, is irrelevant to the issue of apparent cooperativity we identified.

      Reviewer #2 (Public review):

      In their article, Peterson et al. wanted to show to what extent the classical "single hit" model of virion infection, where always the same quantity of virion is required to infect a cell, does not match with empirical observations based on human cytomegalovirus in vitro infection model, and how this would have practical impacts in experimental protocols.

      Strengths:

      The use of a very simple and robust experimental assay, where they infected cells with serially diluted virions and measured the proportion of infected cells with flow cytometry. This convincingly showed how the proportion of infected cells differed from a "single hit" model which they simulated using a simple mathematical model ("power-law model"), and better fitted a model where virions need to cooperate to infect cells.

      The use of different cell types and virus strains, which allows to draw some generalizations.

      The exploration of the mechanisms that could explain this apparent cooperation, using biologically plausible simulations.

      The practical consequences that this phenomenon has for lab virologists as well as modelers.

      Thank you.

      Weaknesses:

      The impossibility to discriminate between biological mechanisms is an important limitation of this study and calls for developing experimental designs able to further understand this question.

      The outcome of the virion clumping remains highly sensitive to the choice of the clumps size distribution, which is itself very complicated to estimate, especially at high dilution.

      The impossibility to directly fit the mathematical models to the data limit them to a qualitative discussion.

      Overall, this work is very valuable as it raises the general question of how the estimate of infectivity can be biased if extrapolated from a single virus titer assay. The observation that HCMV virions often cooperate and that this cooperation varies between context seems robust. The putative biological explanations would require further exploration.

      This topic is very well known in the case of segmented viruses and the semi-infectious particles, leading to the idea of studying "sociovirology", but to my knowledge this is the first time that it was explored for a non-segmented virus, and in the context of MOI estimation.

      Thank you. We would note, however, that inability to discriminate between alternative models is not a weakness per se. It shows that our work goes beyond a somewhat typical approach in mathematical modeling to offer a single explanation for a phenomenon in question (rather than focusing on discriminating between alternatives that is often hard to do).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) I now understand better the graphical abstract. I think my eye was too much attracted by the increase in specific infectivity that you see for more than 1 genome/cell, which is not the point of your paper. I am wondering if you should not guide even more the reader, by pointing out that the fact that the initial decline in specific infectivity represents apparent cooperativity.

      Let’s hope that the readers are smart enough to understand what to focus their eyes on. At the end, this is a graphical abstract that is not supposed to have too much text explaining where to look.

      (2) For your one-inflated geometric distribution, I agree that the estimations would remain very hypothetical because you would have to make many assumptions, however I think a hurdle model where you would fit the P(clump size = 1)=f1 and P(clump size = (i) following a one-truncated geometric distribution would be more appropriate because it would lead to a distribution closer to your PDF from figure S11C.

      The issue is that our data are not in clump sizes but in diameter of the clump D. This is why we opted for using a mixture of continuous distributions, not a mixture of discrete distributions. We are sharing the DLS data, so others are welcome to do another try of fitting other types of distribution to the data.

      (3) For the DLS data, I understand your choice to include all the datapoints, however I find the interpretation confusing: if I understand correctly, you consider that f1, the fraction of the smaller distribution, represents clumps of one virion. However, its median size is 10 times smaller than a virion. So, the number of clumps with one virion would be overestimated. I think it would be helpful for the reader to clarify this aspect, either in the results around lines 503-512, or in the discussion. Could it be that at higher dilution, what is represented by this smaller distribution would almost only be debris because the virions are so rare?

      When fitting a mixture of two log-normal distributions f<sub>1</sub> represents the proportion of clumps of larger size (as was described in the materials and methods). The actual estimated value of f<sub>1</sub> is not highly relevant in calculating change in PDF of the distribution only for D>=d (230nm) as shown in Suppl Fig S11C. But we now realize that this variable f<sub>1</sub> may be confused with a variable f<sub>1</sub> used to denote the fraction of clumps with virion size=1 (in Fig 5C). We now mention that in the caption of Supp Fig S10.

      (4) For the dashed diagonal lines of fig 2, what I don't understand is the choice of the intercept that seems a bit random. I was wondering if it would not be more helpful to make it so that the dashed line intersects the observation for 1 genome/cell, which could then be interpreted as a deviation from the "single hit" model extrapolated outside of 1 genome/cell?

      The diagonal lines in Fig 2 are exactly the same in ALL panels, as are the x/y axes ranges; the slope of the line (equals to 1) allows visually to see when the regression (shown by think black lines) deviates from slope=1, i.e., indicates apparent cooperativity. We will keep the lines are they are. Thank you for the suggestion, though.


      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      In this paper, the authors conduct both experiments and modeling of human cytomegalovirus (HCMV) infection in vitro to study how the infectivity of the virus (measured by cell infection) scales with the viral concentration in the inoculum. A naïve thought would be that this is linear in the sense that doubling the virus concentration (and thus the total virus) in the inoculum would lead to doubling the fraction of infected cells. However, the authors show convincingly that this is not the case for HCMV, using multiple strains, two different target cells, and repeated experiments. In fact, they find that for some regimens (inoculum concentration), infected cells increase faster than the concentration of the inoculum, which they term "apparent cooperativity". The authors then provided possible explanations for this phenomenon and constructed mathematical models and simulations to implement these explanations. They show that these ideas do help explain the cooperativity, but they can't be conclusive as to what the correct explanation is. In any case, this advances our knowledge of the system, and it is very important when quantitative experiments involving MOI are performed.

      Strengths:

      Careful experiments using state-of-the-art methodologies and advancing multiple competing models to explain the data.

      Weaknesses:

      There are minor weaknesses in explaining the implementation of the model. However, some specific assumptions, which to this reviewer were unclear, could have a substantial impact on the results. For example, whether cell infection is independent or not. This is expanded below.

      Suggestions to clarify the study:

      (1) Mathematically, it is clear what "increase linearly" or "increase faster than linearly" (e.g., line 94) means. However, it may be confusing for some readers to then look at plots such as in Figure 2, which appear linear (but on the log-log scale) and about which the authors also say (line 326) "data best matching the linear relationship on a log-log scale".

      This is a good point. We included a clarification to indicate that linear on the log-log scale relationship does not imply linear relationship on the linear-linear scale. We wrote:

      “Because most data did not exhibit a linear relationship between virion concentration and infection probability we fitted the models to subsets of data best matching a linear relationship on a log-log scale. Note that linear relationship on log-log scale may still be nonlinear (on linear-linear scale) when n!=1.”

      (2) One of the main issues that is unclear to me is whether the authors assume that cell infection is independent of other cells. This could be a very important issue affecting their results, both when analyzing the experimental data and running the simulations. One possible outcome of infection could be the generation of innate mediators that could protect (alter the resistance) of nearby cells. I can imagine two opposite results of this: i) one possibility is that resistance would lead to lower infection frequencies and this would result in apparent sub-linear infection (contrary to the observations); or ii) inoculums with more virus lead to faster infection, which doesn't allow enough time for the "resistance" (innate effect) to spread (potentially leading to results similar to the observations, supra-linear infection).

      In our models we assumed cells to be independent of each other (see also responses to other similar points). Because we measure infection in individual cells, assuming cells are independent is a reasonable first approximation. However, the reviewer makes an excellent point that there may be some between-cell signaling happening in the culture that “alerts” or “conditions” cells to change their “resistance”. It is also possible that at higher genome/cell numbers, exposure of cells to virions or virion debris may change the state of cells in the culture, and more cells become “susceptible” to infection. This is a good point that we now list in Limitations subsection of Discussion; it is a good hypothesis to test in our future experiments. We write:

      “Accrued damage model is also consistent with the idea that at higher genome/cell values, the inoculum itself (including cell and/or virion debris) may impact overall susceptibility of all cells in the well, for example, making them more susceptible to infection. It may be expected, though, that exposing cells to debris would increase cell resistance to infection; this would result in n < 1 that we did not observe at small genomes/cell values.”

      (3) Another unclear aspect of cell infection is whether each cell only has one chance to be infected or multiple chances, i.e., do the authors run the simulation once over all the cells or more times?

      Each cell has only one chance to be infected. Algorithm 1 clearly states that; we will add an extra sentence in “Agent-based simulations” to indicate this point.

      (4) On the other hand, the authors address the complementary issue of the virus acting independently or not, with their clumping model (which includes nice experimental measurements). However, it was unclear to me what the assumption of the simulation is in this case. In the case of infection by a clump of virus or "viral compensation", when infection is successful (the cell becomes infected), how many viruses "disappear" and what happens to the rest? For example, one of the viruses of the clump is removed by infection, but the others are free to participate in another clump, or they also disappear. The only thing I found about this is the caption of Figure S10, and it seems to indicate that only the infected virus is removed. However, a typical assumption, I think, is that viruses aggregate to improve infection, but then the whole aggregate participates in infection of a single cell, and those viruses in the clump can't participate in other infections. Viral cooperativity with higher inocula in this case would be, perhaps, the result of larger numbers of clumps for higher inocula. This seems in agreement with Figure S8, but was a little unclear in the interpretation provided.

      This is a good point. We did not remove the clump if one of the virions in the clump manages to infect a cell, and indeed, this could be the reason why in some simulations we observe apparent cooperativity when modeling viral clumping. We have explored this in the revision and found that it does not really impact how infection rate scales with the genomes/cell (e.g., see Suppl Fig S8).

      (5) In algorithm 1, how does P_i, as defined, relate to equation 1?

      These are unrelated because eqn.(1) is a phenomenological model that links infection per cell to genomes per cell. P_i in algorithm 1 is “physics-inspired” potential barrier.

      (6) In line 228, and several other places (e.g., caption of Table S2), the authors refer to the probability of a single genome infecting a cell p(1)=exp(-lambda), but shouldn't it be p(1)=1-exp(-lambda) according to equation 1?

      Indeed, it was a typo, p(1)=1-exp(-lambda) per eqn 1. Thank you, it has been corrected in the revised paper.

      (7) In line 304, the accrued damage hypothesis is defined, but it is stated as a triggering of an antiviral response; one would assume that exposure to a virion should increase the resistance to infection. Otherwise, the authors are saying that evolution has come up with intracellular viral resistance mechanisms that are detrimental to the cell. As I mentioned above, this could also be a mechanism for non-independent cell infection. For example, infected cells signal to neighboring cells to "become resistance" to infection. This would also provide a mechanism for saturation at high levels.

      We do not know how exposure of a cell to one virion would change its “antiviral state”, i.e., to become more or less resistant to the next infection. If a cell becomes more resistant, there is no possibility to observe apparent cooperativity in infection of cells, so this hypothesis cannot explain our observations with n>1. Whether this mechanism plays a role in saturation of cell infection rate at lower than 1 value when genome/cell is large is unclear but is a possibility. We added this point to Discussion in revision (see our text above that includes this point).

      (8) In Figure 3, and likely other places, t-tests are used for comparisons, but with only an n=5 (experiments). Many would prefer a non-parametric test.

      We repeated the analyses in Fig 3 with Mann-Whitney test, results were the same, so we would like to keep results from the t-test in the paper.

      Reviewer #1 (Recommendations for the authors):

      (1) The strains of HCMV used have a fluorescent reporter "in place of the US11 gene". Can you provide a brief comment on whether and how this gene deletion affects HCMV replication?

      US11 is a resident ER protein that is considered an "immune evasion factor". It promotes ERAD of MHC I and has no observable effect on replication of HCMV in cultured cells (Berger 2000 JVI, Wiertz 1996 Cell). We now add this information in Materials and methods section of the paper. We write:

      “All BAC clones were modified to express green fluorescent protein (GFP) or the monomeric red fluorescent protein mCherry (mCherry) with En passant recombineering by replacing US11 with the eGFP or mCherry gene, respectively. US11 is a resident ER protein that is considered an “immune evasion factor”. It promotes ERAD of MHC I and has no observable effect on replication of HCMV in cultured cells [27, 28]. Infectious HCMV was recovered by electroporation of BAC-DNA into MRC5 cells which were then co-cultured with either HFFCs (TB and TR) or HFF-tet cells (ME).”

      (2) I didn't understand what the section "Virus titer assays" refers to. When was this used? How or why is this different from the "Virus stock dilution and dose-response assay"? Also in this section, you refer to NHDF cells - can you provide more information about these? And how does a different type of cell affect the titer assay (here measured as infected cells), since this is one of the main points of your paper?

      Apologies for the confusion. In Ryckman lab we routinely generate viral stock and titrate it using a specific cell type, Normal (or neonatal) Human Dermal Fibroblasts (NHDF). This way, the titer of the stock is consistent between experiments by different researchers in the lab. We then use standard 10-fold dilutions to define the number of infectious units per mL of the stock. We now name this subsection as “Quantification of viral stock infectivity using standard 10-fold dilutions”. After the stock was quantified, we then used that stock in our actual experiments with very small dilution factor df that allowed us to detect deviations of the rate of infection from single hit model.

      (3) In many places, "powerlaw" is written. This is usually written as two words, "power law".

      Because powerlaw comes together with “model”, we decided to use “power-law model”.

      (4) Line 75: "have" instead of "has"?

      (5) Line 84: "with" repeated.

      Corrected, thank you.

      (6) Line 116: This section "Cell lines" seems to describe three cell lines, "HFF cells and MRC5 cells" and then "EC" cells.

      HFF cells are fibroblasts used in our main experiments and MRC5 cells are another type of fibroblasts. We used MRC5 cells in the first step of recovering infection HCMV from BAC DNA (electroporation). We clarified this in Materials and methods. We write:

      “Cell lines. Human foreskin fibroblast cells (HFFCs or fibroblasts) and MRC5 cells (also fibroblasts) were cultured in Dulbecco’s modified Eagle’s medium (DMEM, Sigma) supplemented with 5% heat-inactivated fetal bovine serum (FBS, Rocky Mountain Biologicals, Missoula, MT, USA) and 5%Fetalgro® (Rocky Mountain Biologicals, Missoula, MT, USA). We used MRC5 cells in the first step of recovering infection HCMV from BAC DNA (electroporation). For main experiments we used HFFCs as fibroblasts. Human retinal pigment epithelial cells (ECs or ARPE-19, American Type Culture Collection, Manassas, VA, USA) were cultured in a 1:1 mixture of DMEM and Ham’s F-12 medium (DMEM:F-12, Gibco) and supplemented with 10% FBS.”

      (7) Line 188: Because the virus is double-stranded, do you have to divide the qPCR result by 2 to get genomes?

      This is typically accounted for in our calculations of genome/cell.

      (8) Line 200: Typically, one would write "500g" and not "500xg".

      Corrected.

      (9) Line 248: It would be clearer to write "cell type C different from cell type C2".

      Here C and C_2 refer to actual numbers of cell in the titration/growth experiments, so it is comparing numbers, not cell types. We kept the relationship as it is.

      (10) Definition of cell class: what is n in p_n, the total number of cells, or are these divided into n classes of resistance?

      This part was incorrectly copied from an earlier version, both cell resistance and virion infectivity was sampled from normal distributions with different mean and variances (see Table 1). We corrected the text to reflect this.

      (11) Line 272 to 273: Something seems to be missing, as the change of line doesn't make sense.

      Thank you. Edited to improve readability. Now it reads

      “Clumping hypothesis. In the basic model the number of virions a given cell is exposed to follows a Poisson distribution. However, it is well recognized that as virions are produced by infected cells, they may form clumps/aggregates; the number of virions per clump/aggregate may deviate from, for example, the Poisson distribution [33].”

      (12) Line 283: How lambda is chosen is not indicated here, only later (line 424), but at this point, one can confuse it with lambda in equation 1. Is it the same? It also doesn't seem to be indicated in your Table 1.

      The mean of the Poisson distribution in clump simulations lambda is not the same as lambda in eqn 1; we re-named the mean of Poisson distribution as lambda_c which is estimated by fitting a Poisson distribution to clump size distribution estimated from DLS experiments. Because it was dependent on the virus stock dilution, it is not listed in Table 1. However, we did perform additional simulations assuming lambda_c=2 (Suppl Fig S10).

      (13) Equation 6: I understand that you mostly used kappa=0, but in equation 6, would it be positive or negative (if not zero)?

      We probably expect kappa to be negative but we did not fully explore this extension of the model.

      (14) Line 350: Instead of "infection rates" would "infection frequencies" be better?

      We agree. Changed (also changed in the sentence above that line).

      (15) Line 366: I found this sentence a bit awkward.

      We edited it to the best of our ability to improve it.

      “Importantly, for most HCMV strain-target cell combinations we estimated n>1 (Figure 2 and Supplemental Table S2). With n>1 increase in virion concentration (i.e., higher genomes/cell values) results in a higher than linear increase in the probability of a cell to be infected (eqn. (1)) indicating cooperation between virions at infecting cells. We call this phenomenon “apparent cooperativity”.

      (16) Figure 2, panel L: I wonder if it would be better to include the panel with the name of the experiment, but no data. Currently, it takes a while to find what you are talking about in panel L (or at the very least, indicate the panel in the caption).

      Changed

      (17) Figure 2: When you say that experiments were done at least twice, are you referring to the GFP and mCherry versions of the experiment, or replicates within each of those fluorescent labels?

      Replicates with each of those labels.

      (18) Figure 3: What is the number on top of the black bars? I think it is the average of the paired fold change. Is this right? Why, in panel E, is it 1.32 when only one goes up?

      Yes, fold change. Indeed, 1.32 was a typo, it is 0.70, thank you for noting.

      (19) Line 408: delete the word "there".

      Done. Thank you.

      (20) Line 412: Instead of "The", it should be "Then".

      Done. Thank you.

      Reviewer #2 (Public review):

      In their article, Peterson et al. wanted to show to what extent the classical "single hit" model of virion infection, where one virion is required to infect a cell, does not match empirical observations based on human cytomegalovirus in vitro infection model, and how this would have practical impacts in experimental protocols.

      They first used a very simple experimental assay, where they infected cells with serially diluted virions and measured the proportion of infected cells with flow cytometry. From this, they could elegantly show how the proportion of infected cells differed from a "single hit" model, which they simulated using a simple mathematical model ("powerlaw model"), and better fit a model where virions need to cooperate to infect cells. They then explore which mechanism could explain this apparent cooperation:

      (1) Stochasticity alone cannot explain the results, although I am unsure how generalizable the results are, because the mathematical model chosen cannot, by design, explain such observations only by stochasticity.

      Our null model simulations are not just about stochasticity; they also include variability in virion infectivity and cell resistance to infection. We agree that simulations cannot truly prove that such variability cannot result in apparent cooperativity; however, we also provide a mathematical proof that increase in frequency of infected cells should be linear with virion concentration at small genome/cell numbers.

      (2) Virion clumping seemed not to be enough either to generally explain such a pattern. For that, they first use a mathematical model showing that the apparent cooperation would be small. However, I am unsure how extreme the scenario of simulated virion clumping is. They then used dynamic light scattering to measure the distribution of the sizes of clumps. From these estimates, they show that virion clumps cannot reproduce the observed virion cooperation in serial dilution assays. However, the authors remain unprecise on how the uncertainty of these clumps' size distribution would impact the results, as most clumps have a size smaller than a single virion, leaving therefore a limited number of clumps truly containing virions.

      As we stated in the paper, clumping may explain apparent cooperativity in simulations depending on how stock dilution impacts distribution of virions/clump. This could be explored further, however, better experimental measurements of virions/clump would be highly informative (but we do not have resources to do these experiments at present). Our point is that the degree of apparent cooperativity is dependent on the target cell used (n is smaller on epithelial cells than on fibroblasts) that is difficult to explain by clumping which is a virion property. Per comment by reviewer 1, we have done more analyses of the clumping model to investigate importance of clump removal per successful infection on the detected degree of apparent cooperativity. We found that it was not critical to our conclusions (Suppl Fig S8).

      The two models remain unidentifiable from each other but could explain the apparent virion cooperativity: either due to an increase in susceptibility of the cell each time a virion tries to infect it, or due to viral compensation, where lesser fit viruses are able to infect cells in co-infection with a better fit virion. Unfortunately, the authors here do not attempt to fit their mathematical model to the experimental data but only show that theoretical models and experimental data generate similar patterns regarding virion apparent cooperation.

      In the revision we now provide examples of our earlier simulations that “match” experimental data with a relatively high degree of apparent cooperativity (Supp Fig S9).

      Finally, the authors show that this virions cooperation could make the relationship between the estimated multiplicity of infection and viruses/cell deviate from the 1:1 relationship. Consequently, the dilution of a virion stock would lead to an even stronger decrease in infectivity, as more diluted virions can cooperate less for infection.

      Overall, this work is very valuable as it raises the general question of how the estimate of infectivity can be biased if extrapolated from a single virus titer assay. The observation that HCMV virions often cooperate and that this cooperation varies between contexts seems robust. The putative biological explanations would require further exploration.

      This topic is very well known in the case of segmented viruses and the semi-infectious particles, leading to the idea of studying "sociovirology", but to my knowledge, this is the first time that it was explored for a nonsegmented virus, and in the context of MOI estimation.

      Thank you.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      Two aspects of the work would benefit from further thought:

      (1) The simulation of virion clumps: in both cases (Poisson distribution or one-inflated geometric distribution), the proportion of clumps containing more than one virion will be small. For the Poisson distribution, as you fit the powerlaw model on the range of genomes/cell < ~ 3 genomes/cell (Figure 4B). I wonder to what extent this explains the sudden rise in infections/cells you observe above that limit. It would be interesting to plot the (cumulative) distribution of the clump sizes at different dilution levels to have a better idea.

      The reviewer has a good eye, indeed, the relationship between infection frequency and genomes/cell is linear up to a point, and we believe the inflection point reflects the genomes/cell values when clumps contain more than 1 virion. Here is the results of simulations with distribution of virions/clump plotted:

      Similarly, for the one-inflated geometric distribution, the proportion of clumps of size 1 is the sum of two events: f1, plus 1-f1 times the probability that the geometric distribution is zero, if I follow the methods on lines 287-294. I wonder if this is appropriate regarding the estimates made with the DLC. In particular, Figure 5C shows that the proportion of clumps of size 1 is more than ~ half of all the clumps, and does not seem to be the same distribution as the estimates made on Figure S9C. Maybe a hurdle model would be more appropriate?

      This is a fair point. In our analyses we found that modeling clump size distribution is tricky and required various assumptions. The issue with the DLS data is that we do not really know the distribution of intact virions per clump so how to relate the size of the clump to the number of virions in a clump is wide-open; we explored several possibilities and found that the answer (whether clumping results in apparent cooperativity) depends on assumptions of how clumps are modelled (e.g., compare Fig 4B and Suppl. Fig S11). Hurdle model is not appropriate for clumps because by our definition of a clump, it must have at least 1 virion. Our key observation, however, is that the degree of apparent cooperativity depends on the target cell type – and thus should be independent of virion clumping (unless there is viral cooperativity in the clumps). Overall, we decided that exploring more clumping models would take extra effort, but it is unclear if it brings any benefits to our conclusions.

      The analysis of the clump size distribution using dynamic light scattering, in Figure S8. If I interpret correctly, events with size < 230 nm should be excluded as they do not represent clumps of virions but rather media impurities or cell debris. Therefore, I don't understand the choice of fitting the whole set with a combination of two normal distributions, as even the larger normal distribution covers clumps < 230 nm. If the f1 indicated here is the one used in the methods line 287-294, this is then wrong because it does not represent the fraction of clumps of size 1, but rather debris.

      We used two normal (on log-scale) distributions when quantifying clump distribution data (Supp Fig S10) to avoid sub-selection of the data; in this way, two distribution fit the whole dataset with excellent quality. An alternative approach would be to sub-select data with size >=230nm and fit a normal (or similar) distribution of the clumps; such an approach may generate biases and/or unreliable estimates at high dilutions due to small number of clumps with large size (e.g., see Supp Fig S10S-X). In our simulations to model clump distribution and infection (Fig 5) we attempted to simulate the estimated clump size distribution (Suppl Fig S11C) only approximately. Again, because in our measurements we don’t really know the number of virions per clump, efforts to model exactly clump size distribution, we believe, are not going to give full answers.

      (2) Figure 4 and results lines 419-465: Why didn't you try to fit the different models to the data, instead of qualitatively comparing the estimate of n in the simulations with arbitrary parameters to the one for empirical data? Your models match the expectation of virion cooperation by design, so they are not more convincing for a virologist than logical non-quantitative reasoning. They would be of stronger evidence in my opinion if you could show how well they fit the data. You could then directly compare the different models' fits using goodness-of-fit metrics and decide whether one is better than another or if they all explain equally well the observations.

      Well, we have 11 different relationships between infection rate and genome/cell, finding parameter combinations that would match all the data with at least 2 alternative models seems excessive at present but it is a good direction as we get extra funding to continue this work. It is also difficult to extensively search for the parameter values that would result in a perfect fit of the stochastic simulations to data since the methods of fitting agent-based models to data are not fully developed. However, following this suggestion we now show results of simulations for the two alternative models (accrued damage and viral compensation) that we believe do match experimental data somewhat (see new Suppl Fig S9).

      Minor comments:

      (1) Graphical abstract: This requires more context as it is too rough here to help me understand the general idea of the paper. Plus, why does specific infectivity first decrease with genome/cell?

      We added few elements to the graphical abstract including the strain and target cell used. The decrease in specific infectivity at lower genome/cell is due to apparent cooperativity.

      (2) Equation (7): It would be beneficial for the reader if the reasoning behind the likelihood computation were further described.

      This is a relatively standard approach to model/estimate parameters of a binary outcome, e.g., see Wikipedia: https://en.wikipedia.org/wiki/Logistic_regression

      (3) Line 352-357: could the drop in infectivity also be enhanced/explained by increased cell mortality? Did you gate on cell viability during FCM?

      The infection rate was measured in live cells only, so increased cell mortality may be an explanation.

      (4) Figure 2: I don't understand the dashed diagonal lines: what do they represent exactly? Especially, wouldn't the single-hit model depend on p(1), in which case it should vary by cell x virus?

      As the caption to Figure 2 clearly states, diagonal dashed lines show the slope =1 (i.e, single hit model), so one would be able compare how far the data and/or model fit line deviate from 1. The note for p(1) in panel A is to illustrate how p(1) is calculated; obviously it varies by the strain-cell combination as is indicated in Suppl. Tab S2).

      (5) Fig3G: Is it not surprising to find a positive relationship between p(1) and n? I would have intuitively expected that the stricter the environment is, the more cooperation you observe. But maybe these viruses did not evolve in this context, and therefore, this relationship is different from what you expect from an evolutionary optimum.

      Well, we simply don’t know. The relationship simply suggests that there is connection between infectivity of a single virion and the degree of apparent cooperativity. We are not certain what is the context in which these viruses have evolved.

      (6) Flow cytometry assay: could it be possible that cells infected by more virions generate more fluorescent proteins and are therefore less likely to be false negatives? Maybe you could compare the fluorescence intensity distribution among infected cells in the context of low MOI vs high MOI?

      This is an interesting point. From presented flow cytometry plots (e.g., Suppl Fig S3), the MFI for infected cells does not seem to depend on the dilution (or genome/cell).

      (7) Figure S9B: I did not understand this figure. Are the axes labels correct? How is it possible to have less than 1 virion/well?

      The y axis shows a scaled number calculated from integrating estimated clump size distribution, we assume 1 “scaled” virion/well at highest virion/cell values. With scaling, yes, it is possible to have less than 1 virion/well.

      Reviewer #3 (Public review):

      Summary:

      The authors dilute fluorescent HCMV stocks in small steps (df ≈ 1.3-1.5) across 23 points, quantify infections by flow cytometry at 3 dpi, and fit a power-law model to estimate a cooperativity parameter n (n > 1 indicates apparent cooperativity). They compare fibroblasts vs epithelial cells and multiple strains/reporters, and explore alternative mechanisms (clumping, accrued damage, viral compensation) via analytical modeling and stochastic simulations. They discuss implications for titer/MOI estimation and suggest a method for detecting "apparent cooperativity," noting that for viruses showing this behavior, MOI estimation may be biased.

      Strengths:

      (1) High-resolution titration & rigor: The small-step dilution design (23 serial dilutions; tailored df) improves dose-response resolution beyond conventional 10× series.

      (2) Clear quantitative signal: Multiple strain-cell pairs show n > 1, with appropriate model fitting and visualization of the linear regime on log-log axes.

      (3) Mechanistic exploration: Side-by-side modeling of clumping vs accrued damage vs compensation frames testable hypotheses for cooperativity.

      Thank you.

      Weaknesses:

      (1) Secondary infection control: The authors argue that 3 dpi largely avoids progeny-mediated secondary infection; this claim should be strengthened (e.g., entry inhibitors/control infections) or add sensitivity checks showing results are robust to a small secondary-infection contribution.

      This is an important point. We do believe that the current knowledge about HCMV virion production time – it takes 3-4 days to make virions per multiple papers (see Fig 7 in Vonka and Benyesh-Melnick JB 1966; Fig 3B in Stanton et al JCI 2010; and Fig 1A in Li et al. PNAS 2015) – is sufficient to justify our experimental design but we do agree that an additional control to block novel infections with would be useful. We had previously performed experiments with a HCMV TB-gL-KO that cannot make infectious virions (but the stock virions can be made from complemented target cells). We will investigate if our titration experiments with this virus strain have sufficient resolution to detect apparent cooperativity. However, at present we do not have the resources to perform novel experiments.

      (2) Discriminating mechanisms: At present, simulations cannot distinguish between accrued damage and viral compensation. The authors should propose or add a decisive experiment (e.g., dual-color coinfection to quantify true coinfection rates versus "priming" without coinfection; timed sequential inocula) and outline expected signatures for each mechanism.

      Excellent suggestion. Because infection of a cell is a result of the joint viral infectivity and cell resistance, it may be hard to discriminate between these alternatives unless we specify them as particular molecular mechanisms. But we tried our and listed potential future experiments in the revised version of the paper. Specifically, we write:

      “Second, while we have proposed alternative mechanisms that may result in apparent cooperativity, at present we could not discriminate between these alternatives, in part, because the models lacked specifics – e.g., if virions interacting with a cell reduce its resistance to infection, what does it mean exactly [12]? If virions in a collection augment their infectivity (which may be expected for segmented viruses), how does that viral compensation actually work? Designing experiments that would discriminate between these alternatives would require focusing on a specific mechanism. For example, it may be that that the initiation of gene expression is difficult but is more efficient when there are more virions bringing in more tegument transactivators like pp72/ppUL35 [59]. Alternatively, it may be that there is a bona fide resistance mechanism at play here (e.g. “interferon”) that is antagonized by a viral tegument protein (like TRS1/IRS1 that acts against PKR and 2’5’OAS) [60]. Accrued damage model is also consistent with the idea that at higher genome/cell values, the inoculum itself (including cell and/or virion debris) may impact overall susceptibility of all cells in the well, for example, making them more susceptible to infection. It may be expected, though, that exposing cells to debris would increase cell resistance to infection; this would result in n < 1 that we did not observe at small genomes/cell values. Addressing these hypotheses is an area of future research that will require funding.”

      (3) Decline at high genomes/cell: Several datasets show a downturn at high input. Hypotheses should be provided (cytotoxicity, receptor depletion, and measurement ceiling) and any supportive controls.

      Another good point. We do not have a good explanation, but we do not believe this is because of saturation of available target cells. It seemed to only happen (or was most pronounced) with the ME stocks, which are typically lower in titer and so the higher MOI were nearly undiluted stock. It may be the effect of the conditioned medium. Or perhaps there are non-infectious particles like dense bodies (enveloped particles that lack a capsid and genome) and non-infectious, enveloped particles (NIEPs) that compete for receptors or otherwise damage cells and these don’t get diluted out at the higher doses. We included the point about cell death in Discussion of the revised version of the paper. Specifically, we write:

      “We also do not have a clear explanation of why infection frequency declines at high genomes/cell values for some strain-cell combinations (e.g., Figure 2A, C, D, I, J). Because we measured cell infection in live cells, increase in cell death at higher genomes/cell values may result in the decrease in the number of viable cells.”

      (4) Include experimental data: In Figure 6, please include the experimentally measured titers (IU/mL), if available.

      This is a model-simulated scenario, and as such, there is no measured titers.

      (5) MOI guidance: The practical guidance is important; please add a short "best-practice box" (how to determine titer at multiple genomes/cell and cell densities; when single-hit assumptions fail) for end-users.

      Good suggestion. We now include best-practice box using guidelines developed in Ryckman lab over the years in the revised version of the paper. This is how it reads:

      “Match viral titration methods to the experiment as far as possible. This includes using the same dilution of the viral stock, the cell type, duration of inoculation, and readout of infection.

      When possible, determine the degree of apparent cooperativity (“n”-value, eqn. (1)) for each virus strain/cell type pair being studied.

      If n= 1 (no cooperativity), it is reasonable to calculate experimental MOI based on stock infectivity value determined from a convenient stock dilution.

      If n > 1 or unknown, then stock infectivity should be determined at a dilution resulting in an MOI as close as possible to the desired experimental MOI. Alternatively, the inoculum size can be empirically determined to yield the desired number of infected cells. In these ways different virus/cell type pairs can be compared more fairly.

      Box 1: Recommendations on titrating viral stocks and on performing experiments when comparing different viral strains.”

      Reviewer #3 (Recommendations for the authors):

      FROM PUBLIC REVIEWS (2) Discriminating mechanisms: At present, simulations cannot distinguish between accrued damage and viral compensation. The authors should propose or add a decisive experiment (e.g., dual-color coinfection to quantify true coinfection rates versus "priming" without coinfection; timed sequential inocula) and outline expected signatures for each mechanism.

      This is a good point but to propose a good experiment we need to narrow down the “generic” mechanism to specific processes/genes. We put forward some ideas but clearly more work is needed here:

      “Second, while we have proposed alternative mechanisms that may result in apparent cooperativity, at present we could not discriminate between these alternatives, in part, because the models lacked specifics – e.g., if virions interacting with a cell reduce its resistance to infection, what does it mean exactly [12]? If virions in a collection augment their infectivity (which may be expected for segmented viruses), how does that viral compensation actually work? Designing experiments that would discriminate between these alternatives would require focusing on a specific mechanism. For example, it may be that that the initiation of gene expression is just difficult but is more efficient when there are more virions bringing in more tegument transactivators like pp72/ppUL35 [59]. Alternatively, it may be that there is a bona fide resistance mechanism at play here (e.g. “interferon”) that is antagonized by a viral tegument protein (like TRS1/IRS1 that acts against PKR and 2’5’OAS) [60]. Accrued damage model is also consistent with the idea that at higher genome/cell, the inoculum itself (including cell and/or virion debris) may impact overall susceptibility of all cells in culture, for example, making them more susceptible to infection. It may be expected, though, that exposing cells to debris would increase cell resistance to infection; this would result in n < 1 that we did not observe at small genomes/cell values. Addressing these hypotheses is an area of future research that will require funding.”

      (1) Methods transparency: Include raw spreadsheets or tables of dilution factors and per-well genome estimates used for Figure 1A; this will help reproducibility of the df = 1.3-1.5 pipeline.

      Provided as supplemental xlsx file.

      (2) Epithelial vs fibroblast contrast: Since n is lower on epithelial cells, expand on cell-intrinsic barriers that could dampen apparent cooperativity, and if this argues against simple clumping.

      Indeed, this is our point that we raised in Discussion. Since ECs show lower n than fibroblasts, this observation argues against clumps. Going forward the contrast between cell types will be an approach to understand mechanism. One difference is entry pathways, the ECs involve endocytosis and endosome acidification whereas the fibroblasts do not. There are clearly different receptors involved also, although they are not clearly characterized. One recent report that might be relevant is Ohman 2024 PNAS that shows the gH/gL/UL128-131 complex (aka, "pentamer") is not just dispensable for entry into fibroblasts, but inhibitory. They suggest that the pentamer might bind to a receptor on fibroblasts that activates a pathways that acts against viral IE expression, It could be that in this situation, more virions are really helpful to overcome that block, whatever it is. We now update this point in Discussion.

      (3) Visualization: In Figure 2, consider showing confidence bands for the fitted slope (n) within the colored fit window and reporting n {plus minus} SE in the panels.

      Because we used custom scripts to fit models to data, showing bands of model predictions was a bit complex and would interfere with data points. But we now show 95% Cis for the estimated value n (that are listed in Suppl. Tab S2).

      (4) Symbols: Define all symbols (e.g., V₀, n) on first use in the main text, not only in Methods.

      Done.

      (5) Plot axes check: Explain non-uniform axis labeling ("genomes/cell," "infections/cell").

      This comment was unclear – which labels were not “uniform”? Genomes/cell indicate the expected number of genomes (or virions) that a cell is on average exposed to, infections/cell indicates the probability that a cell actually gets infected.

      (6) Confidence interval for estimated parameters: Figure 3 A-C, please report estimated parameter intervals.

      These are listed in Suppl. Tab S2. Putting Cis for all estimates would clutter the figure making it hard to tell which CIs are for which estimate. But we put the Cis for estimated parameter n in Figure 2.

    1. Author response:

      Reviewer #1 (Public Review):

      This study by Charendoff et al provides interesting observations related to global histone hypermethylation in host cells, during Chlamydia trachomatis infections. The core observation they report is that the host histones are highly hypermethylated during infection, and this appears to be an amplifying effect due to continuous inhibition of demethylases, in part due to a metabolic shift in the host where succinate amounts (which inhibit demethylases) increases. The authors claim specifically due to the bacteria, since antibiotic treatment prevents histone hypermethylation (but leaves you wondering about cause/consequence correlations).

      The core observation of hyper methylation is very interesting, and well documented. There are a number of points to consider though in order to fully substantiate the findings, and close out loose ends. My comments are broad - and built around the interpretations (vs the data presented).

      (1) Related to observations coming Fig 1C etc, and connecting to Fig 3 - the hyper methylation appears to be across different protein arg/lys residues - and is not histone specific. So, is it just a consequence of high SAM pools and flux in infected cells? i.e. the bacterial infection increases SAM pools in cells, and provides an increase in substrate pools for the methyltransferases, leading to protein hyper methylation. The approach used here only measures steady-state SAM amounts (and not SAM flux or utilisation).

      For example, reduced SAM amounts in nuclei could be due to increased utilisation of SAM. The experiments done with the demethylase does not actually answer this question - if you decrease demethylase activity, you will get an increase in net methylation. The authors see an increase in net methylation in the infected cells - this would suggest that in addition (or perhaps primarily) to reduced demethylase activity, there could be much higher SAM utilisation/flux. Again, the over expression of JMJ proteins does not resolve this problem.

      This is an important point. Indeed, one limitation of the initial version of the paper was that we had measured SAM concentration only at one time point (40 hpi) and on the whole population. During revision we used a ratiometric sensor to measure SAM concentration in cells (PMID 34937909). We observed cell-to-cell heterogeneity in SAM levels in HeLa cells, as previously reported in other cell lines. Chlamydia inclusions develop asynchronously, which allows to observe, 40 hpi, a continuum of early (low bacterial load) to late (high bacterial load) stages of infection. We observed no correlation between bacterial load and SAM level, and SAM levels were globally similar when comparing infected and non-infected cells. This experiment strongly supports the hypothesis that protein hypermethylation is not due to an increase in SAM during infection. The data were added in the New Fig. 3. Note that the former Fig. 3 is now split into New Fig. 3 and New Fig. 4.

      (2) Adding to this - what happens to SAM pools in the cells treated with the inhibitors? This actually may not look like the slightly reduced SAM pool observed in infected cell nuclei. Also, what is the SAM/SAH ratio (a very useful indicator of methylation activity).

      Based on the high cell-to-cell heterogeneity of SAM levels observed with the ratiometric probe, we reasoned that measuring SAM/SAH ratio without single cell resolution would not bring crucial information. Also, the discrepancy between data displayed in new Fig. 3A (nuclear extracts) and 3C (live cell imaging) indicate that SAM might be less stable in cellular extracts from infected cells compared to non-infected ones, which would complicate the interpretation of the data. Therefore, we did not implement LC-MS/MS on nuclear extracts to measure SAM/SAH ratio.  

      (3) There is a correlation/implication issue here in Fig 2 - cells with C. trachoma's infection show hyper methylation. But these are the only cells with high C. trachomatis. So it is a bit ingenious to say that histone hyper methylation correlates with bacterial proliferation. The cells without bacteria don't have hyper methylation - and that does not have anything to do with the bacterial proliferation.

      In Fig. 2B, we compared the methylation signal within the population of infected cells only (excluding the uninfected cells). We edited the text to clarify this point. “We observed that, within the population of infected cells, the sum intensity of the mCherry signal was higher in cells that displayed hypermethylation of H3K9me3 than in cells with low level of H3K9me3, indicating that histone hypermethylation correlated with bacterial load (Fig. 2B).”

      (4) The claim that demethylase activity is down in infected cells again comes primarily from the increased succinate (2-fold) amounts in infected nuclei - and then correlated with experiments where succinate, (permeable) a-KG are supplemented in excess. While I personally like the hypothesis that the hypermethylation might be a result of an imbalance in cofactors (succinate vs a-KG) in infected cells, the data presented is very premature to make that conclusion. Again, steady state measurements of only succinate cannot provide a clear answer to that question. For example, is there a clear allocation/flux difference (between a-KG, and leading out to glutamate/glutamine, vs flux through the TCA and increased succinate accumulation? Is there a bottleneck/build-up of succinate in cells that might lead to the increase in nuclei? This also opens another direction of possible regulation - increased histone succinylation. When you see a large increase in succinate in the nucleus, before looking at demethylase activity - it becomes obvious if succinate itself increases histone succinylation (through HATs).

      Our work confirms the accumulation of succinate in cells infected by C. trachomatis, previously reported in Rother et al 2018. The reason for this accumulation remains to be investigated in detail. We have previously shown that OxPhos is relatively stable in infected cells (PMID 35931114), indicating that the flux through the TCA of the eukaryotic host proceeds normally. As mentioned in our discussion, the TCA of the bacteria is disrupted with several enzymes missing, although not in the step immediately downstream of succinate/fumarate production. Still, synthesis of succinate and fumarate (fumarate accumulation was observed in the Rother 2018 study) by bacterial enzymes might contribute to their accumulation in infected cells. The approach we chose to measure methylation at the proteome level is not suitable to look for histone succinylation, because of the diversity of post translational modifications on histones, which occur in combinations. However, following on this reviewer’s comment, we reanalysed the proteomic data to compare protein succinylation levels in infected and non-infected samples. We detected 41 succinylated peptides in the infected samples, against 23 in the uninfected samples. For many of these, we did not have quantitative data in all condition and only one protein, transportin 1 (TNPO1), reached statistical significance, with a 4-fold increase in succinylation in infected samples. Thus, while essentially qualitative, this analysis fully supports the hypothesis that succinate accumulates in infected cells. These data were added to Table S1 and to the result section.

      (5) What might the authors hypothesise about why this hyper methylation happens? It appears in some ways that hyper methylation happens - potentially due to a metabolic bottleneck that the bacteria triggers (and there is a build-up of SAM and/or succinate, and altered flux out of a-kg). The methylation is just a visible outcome - but may not be central to pathogenesis or viability.

      We discussed this question in the penultimate paragraph of the discussion by giving some elements of answer to the question: “Does it benefit the host or the bacteria? ». In our study, we showed that protein hypermethylation affected the transcriptional response of the host. We did not investigate whether the activity of some of the host proteins engaged in the response to infection were affected. It might be the case, considering that methylation is a common PTM regulating protein’s activity. Still, we agree with this reviewer that hypermethylation might not be central to pathogenesis or viability. Addressing this question would require a complex model in which protein methylation levels could be controlled experimentally.  

      Reviewer #2 (Public Review):

      Strengths:

      (1) Because the study compares genuinely infected cells with uninfected cells within the same infected cell population, it enables a clearer and more rigorous comparison.

      (2) By using multiple Chlamydia species and cells from multiple host species (human and mouse), and obtaining consistent findings across these systems, the study demonstrates the generality of bacterium-induced epigenomic alterations.

      (3) The study shows that the epigenomic changes are caused by reduced activity of JMJC domain-containing lysine demethylases, demonstrating through multiple complementary approaches-including the use of a demethylase inhibitor, overexpression of target-specific demethylases, and analysis from the perspective of cofactors required for JMJC domain-containing demethylases-that decreased lysine demethylase activity constitutes the molecular mechanism underlying the increased H3 methylation levels induced by Chlamydia infection.

      (4) By performing ChIP-seq analyses of H3K4me3 and H3K9me3, the study clearly delineates, on a genome-wide scale, how infection leads to increased levels of these epigenomic marks.

      Weakness:

      (1) Reduction of cofactors such as Fe2+ or a-KG decreases the activity of JMJC-domaincontaining lysine demethylases (thereby directly affecting histone H3 lysine methylation). However, these cofactors are also involved in the activities of other epigenetic regulators, such as TET enzymes that contribute to DNA demethylation and SIRT family proteins that mediate histone deacetylation. Therefore, it cannot be excluded that modulation of these factors indirectly leads to the changes in H3 lysine methylation dynamics targeted in this study.

      Indeed, reduction of the concentration of Fe2+ and aKG is expected to have other consequences in addition to the inhibition of JMJC-domain containing lysine demethylases on which we focus in this study. As a matter of fact, we reported a decrease in the methylation level of host DNA in infected cells, and we brought some elements that might explain the discrepancy between DNA and histone methylation status in the discussion (e.g., infected cells display enhanced expression of GADD45, which recruit TET enzymes and thus facilitate DNA demethylation). This example illustrates the complexity of host/pathogen interplay, which affect many parameters simultaneously. Indeed, we cannot rule out that modulation of enzymatic activities other than JMJC-domain containing lysine demethylase contribute significantly to the hypermethylation phenotype.

      (2) Related to point 1, although overexpression of JMJC-type demethylases has been shown to reduce the Chlamydia infection-induced increase in H3 lysine methylation, it is well known that over production of these enzymes, while target-specific, also leads to a genome-wide reduction of lysine methylation. Thus, a decrease in lysine methylation upon expression of these demethylases does not necessarily demonstrate that the infection-induced increase in H3 lysine methylation is caused by impaired JMJC-type demethylase activity.

      We fully agree. We included this experiment to show that increasing the expression of one demethylase only restored demethylation of its cognate target. This support the hypothesis that if the hypermethylation is due to poor demethylase activity, it is likely that several demethylases show impaired activity (as opposed to a scenario in which failure of activity of a single demethylase would indirectly affect all other methylation marks).  

      Reviewer #3 (Public Review):

      In this manuscript, the authors explore a molecular basis for hypermethylation of histones in epithelial cells infected with the obligate intracellular bacterial pathogen Chlamydia trachomatis. This is of particular interest given that Chlamydia is known to drastically alter host cell gene transcription, and histone hypermethylation would suggest a new way by which Chlamydia interferes with gene expression of its host. Histone methylation was previously implicated in the introduction of dsDNA breaks in infected cells, and the chlamydial effector NUE was reported to methylate histones, but the role of this modification in dictating host cell gene transcription has been unexplored. The authors use a suite of tools to approach this question, including various -omics techniques, genetic approaches, and biochemical assays. Overall, the manuscript provides many interesting pieces of data, though some of them are difficult to reconcile, which may reflect methodological hurdles that are not fully addressed in the current version of the manuscript. My major concerns regard the rationale/interpretation for various mechanistic experiments and that the heterogeneity of the histone hypermethylation phenotype is not addressed which I believe may explain some apparent inconsistencies in the results.

      We thank this reviewer for insightful comments. We address these two major concerns during revision and bring some elements in our responses below.

      Using an immunofluorescent approach, the authors show that a subpopulation of the nuclei in Chlamydia-infected cells (~10-20%) exhibit high amounts of methylated histone species. This occurs during the late stages of infection, near the time when Chlamydia would lyse the host cell and positively correlates with bacterial burden.

      Accordingly, halting chlamydial growth blocks the onset of histone hypermethylation. Exogenously supplying cofactors for histone demethylases, the low activity of which is implicated in the histone hypermethylation phenotype, reduces histone hypermethylation. In general, these data are compelling and raise interesting questions about the role of histone methylation in governing chlamydial egress from infected cells. Interestingly, these behaviors seem to arise independently of NUE, the secreted chlamydial histone methyltransferase, supporting the notion that a metabolic reprogramming may underlie the hypermethylation phenomenon.

      As noted above, the authors propose that hypermethylation arises due to decreased demethylase activity in infected cells. However, the data do not conclusively support this interpretation. For example, the approaches used to probe demethylase activity rely on (i) a direct biochemical measure of demethylase activity, (ii), pharmacological inhibition of demethylase, and (iii) heterologous expression of a specific demethylase. With the exception of (i), these approaches would be expected to alter histone methylation regardless of the source. That is, inhibition of demethylases should increase histone methylation regardless of whether the source of methylation is increased methylase or decreased demethylase activity. Similarly, overexpression of a demethylase would be expected to reduce cognate histone methylation arising either from increased methylase or decreased demethylase activity.

      We agree with the reviewer’s comments. The experiment using pharmacological inhibitors (ii) show that infected cells are sensitized to these inhibitors but doesn’t provide direct mechanistic insight. The experiment using heterologous expression of demethylases (iii) was included to show that increasing the expression of one demethylase only restored demethylation of its cognate target. This supports the hypothesis that several demethylases show impaired activity (as opposed to a scenario in which failure of activity of a single demethylase would indirectly affect all other methylation marks).  

      The most direct evidence for impaired demethylase activity come from the direct measure of demethylation of H3K4me3 in nuclear extract (i). It is strengthened by indirect evidence that metabolite concentrations hinder demethylase activities late in infection: 1/ iron and DMKG supply diminish hypermethylation of histone lysine residues 2/ succinate levels (a competitor of aKG) are two-fold higher in nuclei isolated from infected cells. This latter finding was confirmed during revision as we identified more succinylated proteins in infected samples compared to non-infected ones.

      We also considered the possibility that infected cells displayed increased histone methyl transferase (HMT) activity. This would be compatible with decrease KDM activity and could contribute to the histone hypermethylation. Unfortunately, this hypothesis cannot be tested directly (as we did for the measure of H3K4me3 demethylation activity). Indeed, SAM is notoriously labile and in vitro assays to measure HMT require to add exogenous SAM to cell extracts to detect any HMT activity, which would not allow us to test activity based on endogenous SAM levels.

      Instead, we used a ratiometric sensor to measure SAM concentration in cells (PMID 34937909). Chlamydia inclusions develop asynchronously, which allows to observe, 40 hpi, a continuum of early (low bacterial load) to late (high bacterial load) stages of infection. There was no correlation between bacterial load and SAM level, and this level was globally similar when comparing infected and non-infected cells. This experiment supports our hypothesis that protein hypermethylation is not due to an increase in SAM during infection.

      This experiment was also very interesting because it revealed a high cell-to-cell heterogeneity in SAM levels in HeLa cells. Thus, in some cells, SAM might be limiting, which could explain why only a fraction of cells display histone hypermethylation.

      Still, we cannot fully rule out the possibility that increase in SAM availability late in the infectious cycle in some cells, and is immediately consumed through protein methylation, resulting in no net [SAM] increase. The discussion was expanded to take these comments into consideration.

      Altogether, we think that the evidence of decrease KDM activities in infected cells late in infection are strong. Our data do not rule out the possibility that additional mechanisms may contribute.

      Moreover, the authors report that the effect of the demethylase inhibitor on histone hypermethylation is significantly potentiated by infection, suggesting that infected cells have greater methylase activity than uninfected cells, because the latter barely respond to the presence of demethylase inhibitor. In other words, a dramatic increase in histone methylation in the presence of demethylase inhibitor is most parsimoniously explained by increased methylation (no longer being removed by demethylase), not decreased demethylation (which would be analogous to treatment with demethylase inhibitor). The authors do not directly assay methylase activity. These concerns extend to the rationale used to justify experiments with infected mice, which the authors treat with the demethylase inhibitor.

      The observation that the same concentration of JIB-04 leads to an increase of histone methylation in infected cells and not in non-infected cells, is coherent with the data showing that aKG or iron supply diminish histone hypermethylation in infected cells. Indeed, the inhibitor is taken up similarly by infected and uninfected cells but the potency of the inhibitor will depend partly on levels of iron, aKG and succinate found in the cellular milieu so same concentration of inhibitor may inhibit demethylase activity in cells with higher succinate and/or low aKG and low iron but fail to inhibit demethylase activity in cells with higher iron or aKG or lower succinate. In other words, high iron, high aKG or low succinate will “buffer” JIB-04 and make it less potent since JIB-04 partly acts by competing with the iron (competitively) and the aKG (mixed competitive inhibition) PMID 23792809. The same phenomenon is expected for SD70 and TACH101 that share aspects of the mode of action of JIB-04 regarding partly competing for aKG and/or iron in the catalytic site.

      The authors perform experiments to characterize the consequence of hypermethylation genome-wide. Because the authors do not enrich for those cells which exhibit histone hypermethylation, the results reflect the mixed population, and therefore presumably dilute out important signal related to the phenomena under investigation. For example, the proteomic analysis of post-translational modifications identifies only one methylated histone species, whereas the immunofluorescent approach shows consistent effects across five different methylated histone species. Moreover, the chromatin immunoprecipitation analysis indicates that there is unexpectedly a lower density of methylated histones at regions which are also enriched in uninfected cells. The authors argue that this suggests increased methylation is happening "outside" of these histone-dense regions, but direct evidence in support of this claim is lacking.

      The caveat of bulk analyses as opposed to single cell resolution is indeed important to consider when analysing the chIP-seq data and we emphasized this point in the revised manuscript. We could have sorted the cells with high bacterial burden; this would probably have given stronger differences between the two samples. Still, the change in distribution of H3K4me3 in infected samples was very clear and statistically significant. A change in H3K9me3 distribution would be more difficult to catch, as the mark is more widespread.

      In sum, this paper provides compelling evidence in support of the notion that histones are hypermethylated at various residues late in chlamydial infection, that this process is modulated by known cofactors of demethylases, and is the result of high levels of bacterial replication in the cell. That histone hypermethylation governs host gene transcription during chlamydial infection suggests a relatively novel mechanism by which Chlamydia subverts the host cell to establish a replicative niche or egress to infect a new cell. The information obtained regarding the methylation status of host proteins and host gene transcription controlled by a metabolic cofactor during infection will be a useful resource for other researchers. However, in the current version of the manuscript, the mechanistic basis for these behaviors is relatively unclear.

      We thank this reviewer for constructive feedback. We believe that the mechanistic conclusions of our report have been strengthened during revision with additional experiments and text clarification.

    1. Author response:

      We thank the editor and reviewers for the positive comments and critical feedback on our manuscript. We are currently preparing revisions to address the critiques provided by reviewer 2, which focused primarily on growth experiments performed with ∆ACP and ∆FabD P. falciparum parasites in minimal lipid conditions. We note that the major conclusions of our manuscript regarding an essential, FASII-independent function for ACP in apicoplast biogenesis do not require or rely on these experiments in minimal lipid conditions.

      Nevertheless, we believe that these observations have value and agree that they contrast with similar experiments reported in the Amiar et al. 2020 study referenced by the reviewer. We note that this prior study (and others cited by the reviewer) primarily focused on the related apicomplexan parasite, Toxoplasma gondii. We fully agree that available evidence in these and other papers supports a key, fitness-conferring role for FASII activity in growth of T. gondii parasites, including possible expanded functions for ACP that may differ from P. falciparum. We will revise our manuscript to clarify that our results only apply to P. falciparum. We note that our minimal lipid growth experiments with P. falciparum utilized culture conditions and concentrations that appear identical to those reported in the Amiar et al. 2020 study. Nevertheless, we agree with the reviewer that additional experiments will be required to fully test and understand FASII functions in asexual blood-stage malaria parasites, including possible functions in low-lipid conditions. We plan to revise our manuscript to clarify this and other points, and we will include expanded responses to the reviewer critiques.

    1. Author response:

      We thank the reviewers for their positive and constructive feedback and for the careful reading of our manuscript.

      We plan to address the reviewers’ comments and, specifically, to more thoroughly compare movement-associated activity with optogenetic stimulation of the locus coeruleus (LC), with new experiments, clarifications, and additional analyses.

      (1) We plan to perform new experiments using two-photon imaging of noradrenaline (NA) sensors in head-fixed mice during both optogenetic LC stimulation and spontaneous movement. This will, if successful, allow us to directly compare the spatial and temporal structure of NA release across conditions, and to quantify NA amplitude during locomotion versus LC stimulation.

      (2) We will analyze existing NA fiber photometry data for movement-related NA release and compare it to release evoked by LC stimulation.

      (3) In general, we plan to more prominently highlight the limitations of our study that were brought up by the reviewers. In particular, we will expand our discussion of other neuromodulatory systems and their interactions with the LC-NA system, and will tone down conclusions of our study if they cannot be supported by the additional planned experiments and analyses.

      Finally, a reviewer suggested the additional experiment to inhibit LC while performing two-photon imaging in head-fixed animals. These experiments have, due to their technical complexity, a low likelihood of success. In addition, recent work from the lab of Emily Macé already performs LC inhibition during functional recordings (doi: 10.64898/2026.03.06.710089). This work supports our interpretation that the contribution of LC-evoked NA release does not dominate movement-related signals. We will discuss these recent findings in the revised version of our manuscript.

      Together, we believe that these planned experiments, analyses, and revisions will address all main concerns raised by the reviewers.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      (1) Presentation of Figures in the Response Letter

      I would like to note that the figures included in the response letter would benefit from improved organization. For example, Author response image 1 lacks clarity for experimental conditions. From the response letter, my understanding is that a "Labeling rate index", Rg−Rn, was calculated to represent the difference in the rate of increase in labeling between neurons and glial across two time intervals based on experiments shown in Figure 2-figure supplement 1C and G. It seems that a mean convergence index was calculated for each experimental condition at each time point for glial and neurons, and then the differences in mean convergence index increase between time intervals were calculated for glial and neurons. The legend needs more detail to enhance clarity.

      Yes, the “labeling rate index” (Rg−Rn) corresponds exactly to the reviewer’s understanding. Specifically, it quantifies the difference between neurons and glia in the increase of the mean convergence index across two defined time intervals, calculated separately for each experimental condition based on the experiments shown in Figure 2–figure supplement 1C and G.

      To improve clarity, we have substantially revised the figure legend to explicitly describe (i) the definition of labeling rate, (ii) how the mean convergence index was computed for neurons and glia at each time point, (iii) how changes across time intervals were derived, and (iv) how to calculate the labeling rate index. In addition, we have moved this analysis to Figure 2-figure supplement 2 and cited it in Line 191.

      Furthermore, the manuscript should clearly distinguish between figures generated from re-analysis of existing data and those based on newly conducted experiments. This distinction should be explicitly stated in the figure legends and/or main text.

      I recommend that all response figures containing data integral to the authors' rebuttal be properly integrated into the manuscript's existing supplementary figure set, rather than remaining isolated in the response document. This would enhance clarity and ensure that key supporting data are fully accessible to readers. For instance, Author response image 1 can be integrated with Figure 2-figure supplement.

      We appreciate the reviewers’ valuable suggestions. We have revised the figure legends and/or corresponding main text to clearly distinguish figures derived from re-analysis of existing data from those based on newly conducted experiments. In addition, all response figures containing data integral to our rebuttal have now been integrated into the current manuscript’s supplementary figure set.

      Specifically, Author response images 1 and 3 have been incorporated into Figure 2–figure supplement 2 and Figure 2–figure supplement 3, respectively; Author response image 2 has been incorporated into Figure 1–figure supplement 2. Author response image 4 has been incorporated into Figure 1,2–figure supplement 1. These changes improve clarity and ensure that all supporting data are readily accessible to readers.

      (2) Glial Cell Labeling and Specificity of Trans-Synaptic Spread

      The authors provided a comprehensive and well-reasoned response to the concern regarding the labeling of radial glial cells. The inclusion of a dedicated section in the revised Discussion and response figures (possibly to be integrated with supplementary figures), strengthens the manuscript.

      The authors have made an interesting observation in Author response image 2 that glial labeling was frequently observed near the soma and dendrites of starter cells, suggesting that transneuronal labeled glial cells may be synaptically associated with the starter neurons. Also astroglia starter cells lead to infection of nearby TVA-negative astroglia, suggesting astroglia-to- astroglia transmission.

      I find the response scientifically satisfactory and appreciate the authors' transparency in addressing the limitations of their approach.

      We thank the reviewer for the positive and thoughtful evaluation. As suggested, we have integrated the revised Discussion and the corresponding response figures into the main text and the supplementary figure set, ensuring that these observations and their interpretation are clearly presented and readily accessible to readers.

      (3) Temperature Effects and Larval Viability

      The authors' justification for raising larvae at 36C to improve labeling efficiency is reasonable. The supporting data indicating minimal impact on larval viability within the experimental timeframe are convincing. Referencing prior behavioral studies and including survival data under controlled conditions adds credibility to their claims. I find this issue satisfactorily addressed.

      We thank the reviewer for this positive and constructive evaluation.

      (4) Viral Toxicity and Dosage Considerations, Secondary Starter Cells

      The authors present a well-reasoned explanation that viral cytotoxicity is primarily driven by replication and not by viral titer or injection volume. However, the inclusion of experimental data directly testing the effects of higher titer or volume on starter cell viability would have strengthened this point, particularly since such tests are relatively straightforward to perform.

      We agree with the reviewer that directly testing the effects of viral titer and injection volume on starter cell viability would further strengthen this point. In practice, we have already used the highest CVS virus titer that could be reliably generated in our system. Therefore, we tested injection volumes of up to 20 nl and observed no detectable effect on starter cell survival, whereas higher injection volumes resulted in deformation of the larval brain, precluding their use.

      Although not shown as a separate figure, these data informed our interpretation of viral toxicity, which is now described more clearly in the revised Discussion. We hope that this explanation and the clarified discussion adequately address the reviewer’s concern.

      Regarding the potential contribution of secondary starter cells, the authors provide a convincing rationale for why such effects are unlikely under their sparse labeling conditions. However, in cases where TVA and G are broadly expressed-such as under the vglut2a promoter, as shown in Author response image 2 it would be valuable to directly evaluate this possibility experimentally. While the authors' interpretation is reasonable, empirical validation would further strengthen their conclusions.

      We appreciate the reviewer’s interest in experimentally evaluating the potential contribution of secondary starter cells under conditions of broad TVA and G expression. In response, we performed additional viral tracing experiments in which TVA and G were driven by the excitatory neuronal marker vglut2a to achieve broad helper expression.

      As shown in a representative case (Author response image 1), newly appearing tdTomato<sup>+</sup> neurons were observed at the later time (6 vs. 3 dpi, circles), many of which were spatially separated from EGFP<sup>+</sup>/tdTomato<sup>+</sup> starter neurons identified at the early time point (3 dpi, dashed circles). Notably, a subset of these newly labeled tdTomato<sup>+</sup> neurons colocalized with EGFP (6 vs. 3 dpi, dashed cyan circles). These new EGFP<sup>+</sup>/tdTomato<sup>+</sup> neurons may represent secondary starter cells or delayed infection of initially targeted starters. Interpretation of tdTomato<sup>+</sup>-only neurons (6 dpi, gray circles) is further complicated by variability in projection distance and synaptic strength, as short-range secondary-order (or multi-level) inputs and long-range first-order inputs may be labeled within similar time windows. In addition, in the presence of multiple primary or secondary starter neurons, unambiguous assignment of labeled inputs to specific starters remains challenging, even with high-temporal-resolution imaging.

      Owing to these constraints, empirical identification of secondary (or multi-level) connections is not readily achievable with the current tracing strategy. A potential solution would be to combine pan-neuronal helper expression with spatiotemporally controlled activation, for example, through a transgenic line enabling light-inducible helper expression (e.g., G protein). Such an approach would enable delayed and cell-specific initiation of secondary (or multi-level) starters, thereby temporally separating long-range first-order inputs from multi-step circuit propagation and permitting input tracing of targeted cells, ultimately improving the spatiotemporal resolution of circuit mapping.

      We have incorporated a dedicated section in the revised Discussion to clarify the applicable scenarios, limitations, and future directions of this viral tracing strategy in zebrafish.

      Author response image 1.

      Recombinant RV-based viral tracing under broad helper expression conditions.

      Time-lapse (3 and 6 dpi) confocal images of the larval hindbrain showing recombinant RV-based viral tracing under broad helper expression (TVA and G, green) via vglut2a promoter-driven UGNT, following posterior hindbrain infection with CVSdG-tdTomato[EnvA] (magenta). Dashed circles, areas enriched with EGFP<sup>+</sup>/tdTomato<sup>+</sup> neurons; gray circles, areas enriched with tdTomato<sup>+</sup>-only neurons; dashed white lines, hindbrain boundaries. C, caudal; R, rostral. Scale bars, 20 μm.

      Reviewer #2 (Public review):

      The study by Chen, Deng et al. aims to develop an efficient viral transneuronal tracing method that allows efficient retrograde tracing in the larval zebrafish. The authors utilize pseudotyped-rabies virus that can be targeted to specific cell types using the EnvA-TvA systems. Pseudotyped rabies virus has been used extensively in rodent models and, in recent years, has begun to be developed for use in adult zebrafish. However, compared to rodents, the efficiency of spread in adult zebrafish is very low (~one upstream neuron labeled per starter cell). Additionally, there is limited evidence of retrograde tracing with pseudotyped rabies in the larval stage, which is the stage when most functional neural imaging studies are done in the field. In this study, the authors systematically optimized several parameters of rabies tracing, including different rabies virus strains, glycoprotein types, temperatures, expression construct designs, and elimination of glial labeling. The optimal configurations developed by the authors are up to 5-10 fold higher than more typically used configurations.

      The results are convincing and support the conclusions. There are some additional changes that are recommended:

      (1) The new data included in the response to reviewer's letter are important to support the main conclusions and should be included in the manuscript.

      We agree with the reviewer that the new data provided in the response are important for supporting the main conclusions. Accordingly, we have now incorporated all four figures from the response into the supplementary figure set of the revised manuscript and added the corresponding descriptions and discussion to the main text where appropriate.

      (2) Line 357-362: This section should include all of the response letter figures and associated details. Additionally, the Author response image 3 is at odds with Fig 2-supplement 1G. In Author response image 3, ~75% of glial cells labeled at 4 dpi loses their fluorescence by 10 dpi. However, Figure 2-supplement 1G shows that glial overall labeling increases ~2 fold from 4 dpi to 10 dpi. This would suggest that the de novo labeling rate for glia is much higher than the net labeling rate calculated from the convergence index. The authors should clarify these findings.

      We agree with the reviewer that the original section at Lines 357-362 should cite the relevant figures and include the associated details. We have now relocated this content to the Results section and incorporated the corresponding figures and descriptions.

      In addition, we fully agree with the reviewer’s interpretation regarding the apparent discrepancy between the high loss rate of early-labeled glial cells (previously Author response image 3, now Figure 2—figure supplement 3) and the net increase in total glial labeling (Figure 2—figure supplement 1G). This pattern indicates that the net convergence index underestimates the true rate of de novo glial infection, as early labeled glial cells progressively lose detectable fluorescence while overall glial labeling continues to increase, implying ongoing de novo infection events outpace this loss. We have clarified this point in the Results section.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The new data included in the response to reviewer letter are important to support the main conclusions and should be included in the manuscript.

      This recommendation echoes the point raised in Reviewer #2’s Public Comment #1. As detailed in our response there, all new data originally included in the response letter have now been fully integrated into the manuscript’s supplementary figure set, with corresponding descriptions added to the main text.

      Line 357-362: This section should include all of the Author response images and associated details. Additionally, Author response image 3 is at odds with Fig 2-supplement 1G. In Author response image 3, ~75% of glial cells labeled at 4 dpi loses their fluorescence by 10 dpi. However, Figure 2-supplement 1G shows that glial overall labeling increases ~2 fold from 4 dpi to 10 dpi. This would suggest that the de novo labeling rate for glia is much higher than the net labeling rate calculated from the convergence index. The authors should clarify these findings.

      This recommendation echoes the concern raised in Reviewer #2’s Public Comment #2 regarding the apparent discrepancy between glial cell loss and the net increase in glial labeling. Please refer to our response to that comment for a detailed explanation. Briefly, we clarify that the continued increase in overall glial labeling despite substantial loss of early-labeled glia indicates a high rate of ongoing de novo infection that is not captured by net convergence index measurements alone. The relevant figure and associated details, including this clarification, have now been incorporated into the revised main text.

      Data and description for response letter Figure 4 should be quantified and added to the manuscript.

      Across nine infected larvae examined, initial infection was consistently restricted to TVA-positive astroglia, typically involving a single starter glial cell per larva. No viral spread was observed in three larvae injected with SADdG-mCherry[EnvA], whereas astroglia-to-astroglia transmission was detected in three of six larvae injected with CVSdG-tdTomato[EnvA]. Importantly, no neuronal labeling was observed in any of the experiments. These quantitative data and descriptions, originally presented as Author response image 4, have now been incorporated into the main text as Figure 1,2–figure supplement 1).

    1. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment:

      This study reports the important finding that the dynamin inhibitor Dyngo-4a broadly affects lipid packing and plasma membrane dynamics, independently of its action on dynamin. While solid computational, biophysical, and cell-based evidence supports this conclusion, there is incomplete support for the authors' main claim on the role of lipid packing in caveolae internalization, as the causal relationship remains unclear and direct analyses are lacking. With stronger evidence, this work would be of significant interest to cell biologists, biophysicists, and chemists interested in membrane remodeling and drug-membrane interactions.

      We are thankful for the very positive feedback and enthusiasm for our work and sincerely thank all the reviewers for their time, their constructive criticism and valuable comments. Based on this, we have revised our manuscript as detailed below in the point-by-point response where the responses to reviewers’ comments are indicated in blue font. Text edits in the revised manuscript are indicated in red font.

      We agree that providing sufficient evidence for inhibition of caveolae endocytosis by Dyngo-4a is critical and have therefore worked hard on identifying suitable assays that enable conclusive experiments as described below. We have now added a new figure with data that we think firmly supports our statement that caveolae internalization is restricted by Dyngo-4a. Additionally, EM images and quantifications of caveola morphology with or without treatment has been added within the same figure. Taken together, we believe that we have provided strong data to support this main claim and challenged this hypothesis as far as current methodology allows. Therefore, we hope that the revised manuscript warrants a new eLife assessment and we would like this to be the version of accord for the publication in eLife.

      Point-by-point response to reviewers comments

      Reviewer #1 (Public review):

      The authors use Dyngo-4a, a known Dynamin inhibitor to test its influence on caveolar assembly and surface mobility. They investigate whether it incorporates into membranes with Quartz-Crystal Microbalance, they investigate how it is organized in membranes using simulations. Finally, they use lipid-packing sensitive dyes to investigate lipid packing in the presence of Dyngo-4a, membrane stiffness using AFM and membrane undulation using fluorescence microscopy. They also use a measure they call "caveola duration time" to claim that something happens to caveolae after Dyngo-4a addition and using this parameter, they do indeed see an increase in it in response to Dyngo-4a, which is reduced back to the baseline after addition of cholesterol. 

      Overall, the authors claim: 1) Dyngo-4a inserts into the membrane and this 2) results in "a dramatic dynamin-independent inhibition of caveola scission". 3) Dyngo-4a was inserted and positioned at the level of cholesterol in the bilayer and 4) Dyngo-4a-treatment resulted in decreased lipid packing in the outer leaflet of the plasma membrane 5) but Dyngo-4a did not affect caveola morphology, caveolae-associated proteins, or the overall membrane stiffness 6) acute addition of cholesterol counteracts the block in caveola scission caused by Dyngo-4a. 

      Overall, in this reviewers opinion, claims 1, 3, 4, 5 are well-supported by the presented data from electron and live cell microscopy, QCM-D and AFM.

      We thank the reviewer for these positive and encouraging words and believe that the new experiments added to the manuscript has provided strong evidence that caveola internalization is greatly inhibited by Dyngo-4a (see below).

      However, there is no convincing assay for caveolar endocytosis presented besides the "caveola duration" which although unclearly described seems to be the time it takes in imaging until a caveolae is not picked up by the tracking software anymore in TIRF microscopy. Since the main claim of the paper is a mechanism of caveolar endocytosis being blocked by Dyngo-4a, a true caveolar internalization assay is required to make this claim. This means either the intracellular detection of not surface connected caveolar cargo or the quantification of caveolar movement from TIRF into epifluorescence detection in the fluorescence microscope. Otherwise, the authors could remove the claim and just claim that caveolar mobility is influenced.

      We thank the reviewer and agree that this is a very important point to verify. Therefore, we have worked hard to quantify the endocytosis of caveolae in thin sections of MEF cells using transmission electron microscopy. By incubating cells with externally added HRP for two-minutes followed by washing, vesicles internalized during this period can be contrasted and distinguished from surface associated vesicles. Sections were quantified by counting both surface-associated and internalized caveolae and CCVs (see figure below). Surface associated caveolae and CCVs can be distinguished based on size and shape for CCV the presence of a coat, but the number of vesicles per image is very low because a cross section has to go right through the vesicle. Furthermore, although internalized caveolae and CCVs can be differentiated by size, it is much harder to separate these from other vesicles, tubules and tubular endosomes positive for HRP.  We detect an approximate 50% reduction in internalized caveolae and CCVs (ie. containing the internalized marker) in Dyngo-4a cells, which confirms that internalization is impaired following Dyngo-4a treatment. Yet, CCV endocytosis was simultaneously confirmed by Tfn uptake assay to be reduced by a greater extent, approximately 95%. We believe that this discrepancy in numbers is due to the low frequency of counted vesicles per section and the difficulties in distinguishing different internalized vesicles and endosomal tubules making a robust quantification of endocytic events difficult. It is also important to note that the EM assay relies on structural criteria to identify only the budded CCVs and caveolae containing the internalized marker, in transit to the early endosome. Other labeled structures are excluded. In contrast, uptake of Tfn into endosomes would also be measured by the light microscopy assay. Therefore, we have chosen not to include these data in the revised manuscript.

      Author response image 1

      Instead, we have developed a new assay in which we can quantify internalization in whole cells and clearly separate internalized caveolae from those that are surface associated or have fused with endosomal structures. For this we use the HeLa FlpIn Cav1-GFP cells which are induced to express Cav1-GFP at endogenous levels to label caveolae. The cells are incubated for five minutes with fluorescent CTxB known to be internalized by caveolae (but also via other mechanisms). To be able to separate internalized caveolae from early endosomes, cells were fixed and labelled with antibodies against the marker EEA1.  Cells were analyzed by fluorescence microscopy and confocal z-stacks of entire cells were recorded. The data was analyzed by software to identify only the caveolae that were positive for CTxB but negative for EEA1. The results from quantification showed a very clear inhibition in the number of internalized caveolae in Dyngo-4a treated cells in comparison to control cells. These data have been included in the manuscript as an important new figure 2 together with TEM data where we quantify the morphology of surface associated caveolae with or without Dyngo-4a treatment. We have also extensively edited the text in the results section to describe these new data and to convey that Dyngo-4a indeed affects internalization. We are very happy to have established means to address this important point by extending the current methodology and tools. Together with the TIRF data and FRAP data we believe that we have provided strong data for this claim and challenged our hypothesis as far as current methodology allows.

      Significance: 

      A number of small molecule inhibitors for the GTPase dynamics exist, that are commonly used tools in the investigation of endocytosis. This goes as far that the use of some of these inhibitors alone is considered in some publications as sufficient to declare a process to be dynamin-dependent. However, this is not correct, as there are considerable off-target effects, including the inhibition of caveolar internalization by a dynamin-independent mechanism. This is important, as for example the influence of dynamin small molecule inhibitors on chemotherapy resistance is currently investigated (see for example Tremblay et al., Nature Communications, 2020). The investigation of the true effect of small molecules discovered as and used as specific inhibitors and their offside effects is extremely important and this reviewer applauds the effort. It is important that inhibitors are not used alone, but other means of targeting a mechanism are exploited as well in functional studies. The audience here thus is besides membrane biophysicists interested in the immediate effect of the small molecule Dyngo-4a also cell biologists and everyone using dynamic inhibitors to investigate cellular function. 

      Thank you for the comments. We very much appreciate the interest and enthusiasm of the reviewer for our work. This has inspired and supported us to perform additional work for the revision of our manuscript.

      Reviewer #2 (Public review): 

      In this manuscript, the authors probe the mechanisms by which Dyngo-4a, a dynamin inhibitor used to block endocytosis, disrupts caveolae dynamics. They provide compelling evidence that Dyngo-4a inhibits caveolae dynamics and endocytosis (as well as several other aspects of plasma membrane dynamics) by a dynamin-independent mechanism. They also provide strong computational and experimental data showing that Dyngo-4a inserts into membranes and decreases lipid packing in the outer leaflet of the plasma membrane. Finally, they demonstrate that the addition of excess cholesterol to cells reverses the effects of Dyngo-4a on caveolae dynamics, presumably by reversing lipid packing defects. Based on these findings they conclude that lipid packing regulates caveolae dynamics and endocytosis in a cholesterol-dependent manner. 

      This work should be of value to cell biologists interested in plasma membrane remodeling and membrane trafficking, biophysicists that study small molecule/membrane interactions and membrane remodeling processes, and chemists interested in designing drugs to target membrane trafficking machinery and pathways. 

      This work addresses the important topic of how a widely used endocytic inhibitor actually works. In the process of addressing this question, the authors uncover unexpected connections between how lipids are packed in cell membranes and membrane dynamics. The methods are appropriate and many of the claims made in this work are well supported by data.

      We very much appreciate the thorough review and very positive feedback constructive critique and thank the reviewer for the time spent on our manuscript.

      Weaknesses: 

      I appreciate that the manuscript has already gone through one round of revisions and that many of the concerns from the previous reviewers appear to have been addressed. However, as an interested reader, I would like to offer several additional comments for the authors to consider. 

      (1) It is not clear based on the data presented whether the effects of Dyngo-4a on lipid packing give rise to defects in caveolae dynamics or if these effects are merely correlated. To show this more definitively, one might expect additional experimental approaches to be used to perturb lipid packing. I appreciate this is probably beyond the scope of the current study. However, it seems important for the manuscript to be clear about how far this interpretation can be pushed in the absence of additional independent lines of evidence.

      We are very proud of the direct experimental support of the effect on lipid packing that we have performed using incorporation of extra cholesterol to the membrane which supports these effects are not merely correlated. Unfortunately, specifically perturbing lipid packing in other ways and conclusively interpreting such data is not uncomplicated. We agree that data and conclusions should be further challenged but we believe that this goes beyond the scope of this manuscript.

      (2) On a related note, it is not obvious how changes in lipid packing in the outer leaflet could impact caveolae dynamics. It would be helpful to include a cartoon illustrating how this might work.

      Thank you for pointing out this important aspect. We have elaborated on this within the discussion and referred to our recently published perspective article in Nature Cell Biology ('A lipid-centric view of endocytosis by caveolae' Parton, Kozlov and Lundmark DOI: 10.1038/s41556-026-01945-5) where this topic is extensively discussed. In short, insertion of the 8S disc in the inner leaflet of the PM replaces approximately 250 lipids and spans the entire thickness of the leaflet. The insertion of the flat, hydrophobic phase of the 8S disc, that faces the outer leaflet, results in a differential contact energy favoring the uneven packing of lipids and preferred accumulation of cholesterol in the PM of mammalian cells. Increased cholesterol content in the PM leads to more tilt and splay and hence curvature generation and, if not constrained by EHD2, scission. Thus, the distinct lipid packing of cholesterol and sphingomyelin opposite the Cav1 complex is key to drive curvature generation and internalization of caveolae.

      We agree that a schematic figure could be nice to illustrate how packing affects caveolae internalization. However, we realized that providing a comprehensible concept this would require an extensive figure with vast discussions in the text. Therefore, we have chosen not to include this here, but refer to the figures in Parton et al. Nature Cell Biology DOI: 10.1038/s41556-026-01945-5

      (3) The authors note that Dyngo-4a inhibits several dynamic processes including generalized plasma membrane mobility (Fig 4A&B), transferrin uptake (Fig S4C), and fusion of fusogenic liposomes (Fig S4G). This clearly indicates there is a major disruption of the plasma membrane going on here that is not limited to caveolae. They go on to show that the addition of cholesterol reverses the effects of Dyngo-4a on caveolae dynamics. However, they do not discuss whether adding back cholesterol has similar effects on plasma membrane mobility and transferrin uptake. This information could help to further pinpoint whether the mechanisms of action are shared, and if the role of cholesterol is more general in controlling these events or is instead specific to caveolae. 

      Yes, this is correct, and we agree that this important finding leads to many follow up questions on the mechanism of action of Dyngo-4a on cellular processes. Yet, to dissect the mechanism for all these processes goes way beyond the scope and our resources for this manuscript.

      (4) In Fig 4C, the morphology of the neck region of the Dyngo-4a treated caveolae structure appears to be "pinched" compared to the control. I appreciate that more EM studies are underway. It would be useful to specifically compare the morphology of the caveolae as part of those studies.

      Thanks, this is a relevant and interesting question. In the revised manuscript, we have therefore performed and included extra quantitative EM data addressing the morphology of caveolae. Based on this we conclude that there is no statistically significant difference in the height, width or neck diameter of caveolae treated with Dyngo-4a in comparison to control cells. When analyzing the ratio of height, width and neck diameter of each caveolae, there is a trend in that neck diameter is increased in Dyngo-4a-treated cells. These data have been included in the new figure 2 A-B and discussed in the text.

      (5) In Line 91, a statement is made that 8S complex formation requires cholesterol. This is debatable, as they appear to form in E. coli in the absence of cholesterol (reference 14).

      Thank you, we have clarified that this statement is referring to mammalian cells.

      Some minor spelling errors include: 

      Line 66 generrating

      Line 182 signigicantly 

      Line 197 treatmend 

      Line 347 succefully 

      These errors have been corrected

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      One criticism that I would still make in the revised version of the paper concerns the description of the behavior of the two monkeys which is still minimal, while acknowledging differences in their choice and RT performance that reflect "individual differences in sensitivity to motion stimulus and a common heuristic-based satisficing strategy". This sentence is not clear to me. Moreover, the potential consequences of these differences on neuronal activity are only considered in the cluster analysis done for each of the two animals separately and for which it turns out there is no notable difference.

      We have revised the text to emphasize the key, common feature of their behavior and refer readers interested in variability across sessions and individuals to our previous study: “Both monkeys showed consistent biases toward the large-reward choice (Figure 1B, C). Details of their performance, including variations across sessions and individuals, have been reported in a previous study (Fan et al., 2018).”

      Given that both monkeys’ choices and RT showed clear and consistent coherence and reward dependencies, and that the clustering analysis were consistent across the two monkeys, we believe that our analyses presented here are appropriate. Future work is needed to examine if and how STN contributes to more nuanced aspects of behavioral variability.

      Compared to the first version of the paper, the cluster analysis in this revised version yields three distinct populations instead of the previous four. While the authors suggest that these subpopulations play important roles in encoding different aspects of decision-making, the identification of three rather than four subpopulations seems to me an important update that warrants discussion.

      The clustering results are slightly different because, following suggestions from the first round of reviews, we now use more principled approaches for selecting neurons and computing the clusters. The primary difference is that Clusters 1 and 3 in the original manuscript have mostly been merged into one cluster (new Cluster 3). We updated the text to note that our use of three clusters depends on our choice of clustering cutoff and continue to emphasize that the clusters are consistent across monkeys and clustering techniques: In Results: “Inspection of the dendrogram (hierarchical cluster tree) suggested that our STN samples can be reasonably grouped into three clusters, although other groupings are possible using different clustering cutoffs (Figure 5-S1).” In Discussion: “Furthermore, our clustering analysis aimed to identify common activity profiles in the STN population, while leaving behind many neurons that either did not show consistent task-related modulation or had less common activity profiles (e.g., those that were far from others in the vector space and those with too infrequent occurrence to form detectable clusters). More work is needed to continue to refine our understanding of the specific computational contributions of the STN to decision formation.”

      Finally, I think it would have been interesting to identify the level of collinearity in the model proposed by the authors (equation 7). Indeed, one can expect significant collinearity between some of the proposed explanatory factors of neuronal activity, such as choice and coherence level, for example.

      The reviewer is correct that choice and coherence are correlated with the formulation of Eq. 7. However, such collinearity does not seem to bias the regression results (Author response image 1). We have performed simulations with different modulation strengths and noise levels (A and C) and observed generally good recoverability of the ground-truth regression coefficients (red: unity-slope lines), despite the strong correlation between choice and coherence for one choice (B).

      Author response image 1.

      Similarly, for the analysis relating neuron activity to decision evaluation signals (p 16), firing rates calculated using sliding averages with 1-ms steps are compared, but the method does not specify controls for multiple comparisons or for non-independent data.

      We have made multiple comparison corrections using the Benjamini and Hochberg procedure and updated the relevant text in Methods, Results, and Abstract accordingly.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the paper, the authors propose a new RNA velocity method, TSvelo, which predicts the transcription rate linearly based on the expression of RNA levels of transcription factors. This framework is an extension of its recent work TFvelo by including unspliced reads and designing a coherent neuralODE framework. Improved performance was demonstrated in six diverse datasets.

      Strengths:

      Overall, this method introduces innovative solutions to link cell differentiation and gene regulation, with a balance between model complexity (neuralODE) and interpretability (raw gene space).

      We thank the reviewer for the positive evaluation of our work and for recognizing the novelty of the proposed framework. We appreciate the reviewer’s summary highlighting that TSvelo extends our previous method TFvelo by incorporating unspliced reads and introducing a coherent neuralODE framework to model transcription dynamics.

      We are encouraged that the reviewer recognizes the potential of our approach to link cell differentiation with gene regulatory mechanisms, while maintaining a balance between model expressiveness and interpretability in the gene expression space. In the revised manuscript, we have further clarified several methodological details and strengthened the presentation to better highlight these aspects.

      Weaknesses:

      While it seems to provide convincing results, there are multiple technical concerns for the authors to clarify and double-check.

      (1) The authors should clarify and discuss the TF-target map: here, the TF-target genes map is predefined by the TF binding's ChIP-seq data. This annotation is largely incomplete and mostly compiled from a set of bulk tissues. Therefore, for a certain population, the TF-target relation may change. This requires clarification and discussion, possibly exploring how to address this in the model. In addition, a regulon database could be added, e.g., DoRothEA?

      We thank the reviewer for this important comment. The TF–target maps used in TSvelo (e.g., derived from ChIP-seq-based resources such as ENCODE) reflect aggregated TF binding evidence collected across diverse bulk cell types and experimental conditions. As such, they are inherently incomplete and do not capture fully context-specific regulatory activity in a given primary tissue. In TSvelo, we therefore do not treat these annotations as fixed or cell-type-specific ground truth regulatory relationships. Instead, they are used as a permissive prior that encodes a broad set of potential regulatory interactions.

      Within the TSvelo framework, the contribution of each TF–target interaction is learned from data through weight estimation, allowing the model to down-weight or effectively ignore prior edges that are inconsistent with the observed single-cell expression dynamics. This design enables TSvelo to remain robust even when the prior TF–target map is noisy, incomplete, or derived from heterogeneous bulk contexts.

      Following the reviewer’s suggestion, we additionally incorporated the DoRothEA regulon database as an alternative prior with confidence-level filtering. We further performed ablation studies on the pancreas dataset and the gastrulation erythroid dataset using different TF–target resources, including ChEA, ENCODE, and their combinations with DoRothEA.

      The results on the pancreas dataset and the gastrulation erythroid dataset are shown in Figure S13 and Figure S14 respectively, which come up with the same conclusion. We observed highly consistent results across most TF–target prior combinations, including ChEA, ENCODE, ChEA+ENCODE, ChEA+DoRothEA, ENCODE+DoRothEA, and ChEA+ENCODE+DoRothEA. Using the pancreas dataset as example, the mean velocity consistency ranged from 0.985 to 0.995, the mean in-cluster coherence ranged from 0.983 to 0.992, and the mean cross-boundary direction correctness ranged from 0.719 to 0.740 across all settings. These consistently high and tightly bounded metrics indicate that TSvelo is largely insensitive to the specific choice of TF–target prior.

      The only configuration showing reduced stability was the use of DoRothEA alone, particularly in terms of cross-boundary direction correctness. This is likely due to its comparatively limited coverage of TF–target interactions. For instance, in the pancreas dataset, only 81 out of 2000 highly variable genes (HVGs) could be associated with TFs based on DoRothEA, corresponding to 102 TF–target links in total, which may restrict downstream regulatory modeling. In contrast, ChEA covered 1793 genes with 13,976 TF–target links, and ENCODE covered 1854 genes with 33,076 links. These results further suggest that integrating multiple TF–target resources could improve performance, likely due to increased coverage and complementary regulatory information.

      We further acknowledge that regulatory interactions are inherently context-dependent, and that no static TF–target resource can fully capture tissue-specific regulatory programs. In the revised Discussion, we explicitly clarify this limitation and highlight that incorporating context-specific regulatory data (e.g., single-cell chromatin accessibility or perturbation-based regulatory maps) represents an important direction for future improvement.

      (2) The authors should clarify how example genes are selected. This is particularly unclear in Figure 2d.

      We thank the reviewer for raising this point. The example genes shown in Fig. 2d were selected to illustrate representative scenarios where our method provides advantages, particularly cases in which the unspliced–spliced 2D phase portrait exhibits mixed or overlapping patterns that are difficult to model using conventional RNA velocity approaches. These examples are therefore intended to demonstrate the types of transcriptional dynamics that TSvelo is designed to better capture.

      To avoid the impression of selective presentation, we note that our conclusions are based on systematic evaluation across all genes and datasets. Additional visualizations for a broader set of genes on this dataset are provided in Fig. S3. We have clarified the example gene selection criteria in the revised manuscript.

      (3) The authors should clarify confidence in the statement in lines 179-180, that ANXA4 should initially decrease. This is particularly concerning, as TSvelo didn't capture the cell cycle transitions well during the initial part.

      We thank the reviewer for raising this point. The statement that ANXA4 initially decreases is based on the observed expression pattern in the dataset rather than on cell-cycle–related dynamics inferred by the model. Specifically, ANXA4 shows higher expression in Ductal cells compared to Ngn3 EP cells, and Ductal represents an earlier stage in the developmental trajectory. Therefore, along the Ductal to Ngn3 EP transition, ANXA4 naturally exhibits an initial decrease in expression. We have clarified this point in the revised manuscript.

      (4) A support reference should be added for the statement in line 260 that "neuron migrations are inside-out manner". There is no reference supporting this, and this statement is critical for the model assessment.

      We thank the reviewer for this suggestion. This pattern has been reported in previous studies [1,2], which have been added into the revised manuscript.

      To Improve clarity, we have also revised the statement in the manuscript as follows:

      “During cortical development, neurons follow an inside-out layering pattern in which earlier-born neurons populate the deep cortical layers, whereas later-born neurons migrate past them to occupy more superficial layers.”

      (1) Nadarajah, B., Parnavelas, J. Modes of neuronal migration in the developing cerebral cortex. Nat Rev Neurosci 3, 423–432 (2002).

      (2) Li, C., Virgilio, M.C., Collins, K.L. et al. Multi-omic single-cell velocity models epigenome–transcriptome interactions and improves cell fate prediction. Nat Biotechnol 41, 387–398 (2023).

      (5) The comparison to scMultiomics data is particularly interesting, as MultiVelo uses ATAC data to predict the transcription rate. It would be very insightful to add a direct comparison of the estimated transcription rate between using ATAC and directly using TFs' RNA expressions.

      We thank the reviewer for suggesting this highly interesting comparison between ATAC-derived regulatory activity and TF RNA-based proxies for transcription rate estimation.

      We have conducted the requested analysis by computing gene-wise chrome accessibility rate used in MultiVelo and the learned transcription rate from TSvelo, and evaluated their correlation across genes. As shown in Figure S15, the two estimates exhibit almost no global correlation across genes, indicating that they capture substantially different aspects of regulatory information.

      This discrepancy is not unexpected and reflects the fundamental differences between these modalities. scATAC-seq measures chromatin accessibility, which provides a proxy for cis-regulatory potential of genomic regions. However, ATAC signals are inherently sparse and often exhibit a near-binary structure, limiting their ability to directly capture fine-grained temporal regulatory dynamics. In contrast, TF RNA expression reflects downstream transcriptional output, which is shaped by multiple regulatory layers, including post-transcriptional regulation, protein activity, temporal delays, and indirect regulation through intermediate transcriptional or signaling pathways. As a result, these two modalities are expected to capture complementary but not directly comparable aspects of gene regulation.

      Overall, this result suggests that ATAC-based and TF RNA-based signals capture distinct aspects of gene regulation. This further implies that integrating both modalities may be beneficial for future models that aim to more comprehensively characterize transcriptional regulation. We have added this discussion to the supplementary information.

      (6) In Figure 6g, it should be clarified how the lineage was determined. Did the authors use the LARRY barcodes, predicted cell fate, or any other methods? Here, the best way is probably using the LARRY barcodes for individual clones.

      We thank the reviewer for this suggestion. The lineage assignment used in Fig. 6g is described in the Methods section (“Lineage segmentation and pseudotime initialization”). Briefly, lineages are inferred from the transcriptomic structure of the data by performing Leiden clustering followed by PAGA-based connectivity analysis. Starting from an initial Leiden cluster, the filtered PAGA graph defines the shortest paths to other clusters, which are considered as the detected lineages, and diffusion pseudotime (DPT) is then used to initialize pseudotime along each lineage. Thus, in this analysis lineages are determined from the expression-derived trajectory structure. We have clarified this point in the revised manuscript and refer readers to the Methods section.

      Reviewer #2 (Public review):

      Summary:

      Li et al. propose TSvelo, a computational framework for RNA velocity inference that models transcriptional regulation and gene-specific splicing using a neural ODE approach. The method is intended to improve trajectory reconstruction and capture dynamic gene expression changes in scRNA-seq data. However, the manuscript in its current form falls short in several critical areas, including rigorous validation, quantitative benchmarking, clarity of definitions, proper use of prior knowledge, and interpretive caution. Many of the authors' claims are not fully supported by the evidence.

      We thank the reviewer for the careful evaluation of our manuscript and for the constructive comments. We appreciate the concerns regarding validation, benchmarking, methodological clarity, and interpretation. In the revised manuscript, we have carefully addressed these points by adding additional analyses, clarifying methodological details, and moderating several claims to ensure they are fully supported by the data. Detailed responses to each comment are provided below.

      Major comments:

      (1) Modeling comments

      (a) Lines 512-513: How does the U-to-S delay validate the accuracy of pseudotime? Using only a single gene as an example is not sufficient for "validation."

      We thank the reviewer for this important clarification. In the revised manuscript, we have rephrased this part to clarify that Fig. 1a serves only as an illustrative example showing the U-to-S delay for a single gene. Accordingly, we have corrected our statement to indicate that the U-to-S delay is used to infer trajectory orientation, rather than to validate the accuracy of pseudotime.

      In addition, we have expanded the description to explain that U-to-S delay signals are aggregated across all genes to provide a more robust and comprehensive assessment for this purpose. Additional analysis is provided in our response to the next comment.

      (b) Lines 512-518: The authors propose a strategy for selecting the initial state, but do not benchmark how accurate this selection procedure is, nor do they provide sufficient rationale. While some genes may indeed exhibit U-to-S delay during lineage differentiation, why does the highest U-to-S delay score indicate the correct initiation states? Please provide mathematical justification and demonstrate accuracy beyond using a single gene example. Maybe a simulation with ground truth could help here, too.

      We thank the reviewer for this insightful comment. In the revised manuscript, we have clarified both the intuition and justification of this approach. Briefly, along a correctly oriented trajectory, unspliced (U) expression is expected to precede spliced (S) expression due to transcriptional dynamics. Ideally, this U-to-S delay would be observable at the level of individual genes. However, due to the high noise inherent in scRNA-seq data, such delays are often not consistently detectable on a per-gene basis. To address this, we aggregate U-to-S delay signals across all genes and determine the lineage orientation by maximizing a global delay score. Under this criterion, the cluster from which all outgoing lineages exhibit the highest aggregated U-to-S delay is inferred to correspond to the initial state.

      We emphasize that this approach relies on genome-wide aggregation rather than any single gene. Moreover, the same strategy is applied uniformly across all six datasets using identical parameter settings, demonstrating its robustness and stability. To further address the reviewer’s concern, we additionally present the U-to-S delay scores for each Leiden cluster when treated as the initial state across all datasets (Author response image 1). The results on all datasets suggest that the highest U-to-S delay scores can be used to detect the initial cluster.

      Author response image 1.

      The U-to-S delay scores for each Leiden cluster when treated as the initial state across all datasets.

      Following your suggestions, we also add a simulation study. We generated synthetic single-cell RNA velocity datasets using a mechanistic transcriptional dynamics model with one or multiple developmental branches. The system included 200 genes, among which 30 were designated as transcription factors (TFs).

      For each branch, we independently sampled a TF–target regulatory matrix W ϵ R<sup>30×200</sup> from a standard normal distribution to simulate distinct GRN structures. Gene expression dynamics were modeled using a coupled ordinary differential equation (ODE) system describing unspliced and spliced RNA abundances:

      where u and s denote unspliced and spliced RNA levels, respectively. The transcription rate α was computed as a nonlinear function of TF expression, defined as a weighted sum of spliced TF abundance, followed by clipping to ensure bounded activation.

      Each branch is initialized from the same randomly sampled initial condition drawn from a gamma distribution, allowing controlled divergence of trajectories driven solely by branch-specific regulatory programs.

      To simulate observed sequencing counts, we introduced technical noise by scaling latent expression levels with cell-specific library sizes drawn from a log-normal distribution. The resulting expression counts were generated using a negative binomial sampling model:

      where θ controls over dispersion, with smaller values corresponding to higher noise levels. The final datasets consist of paired unspliced (U) and spliced (S) count matrices with realistic transcriptional stochasticity and branching gene regulatory dynamics. For each branch, cells were further divided into three developmental stages for downstream analysis.

      We evaluated TSvelo on multiple simulated datasets with varying numbers of branches and noise levels. There are two or three branches start from the same root cell groups in these datasets (Branch 1: stage 0 - stage 1 - stage 2. Branch 2: stage 0 - stage 3 - stage 4. Branch 3: stage 0 - stage 5 - stage 6). The results of initial state identification based on the unspliced-to-spliced (U-to-S) delay, along with the corresponding 2D velocity stream visualizations, are presented in Supplementary Figure S1. These results demonstrate that the U-to-S delay–based initialization is robust and consistently identifies cells corresponding to the earliest developmental stage (“stage 0”) across different simulation settings. All additional results have been included in the Supplementary Information.

      (c) Equation (8): The formulation looks to be incorrect. If $$W \in \mathbb{R}^{G\times G}$$ and $$W' - \Gamma' \in \mathbb{R}^{K\times K}$$, how can they be aligned within the same row? Please clarify.

      We thank the reviewer for pointing this out. This was a typographical error in the manuscript. In the third line of Equation (8), the term should be W’ instead of W. We have corrected this in the revised manuscript to ensure dimensional consistency.

      (d) The use of prior knowledge graphs from ENCODE or ChEA to constrain regulation raises concerns. Much of the regulatory information in these databases comes from cell lines. How can such cell-line-based regulation be reliably applied to primary tissues, as is done throughout the manuscript? Additional experiments are needed to test the robustness of TSvelo with respect to prior knowledge.

      We thank the reviewer for this important comment. In TSvelo, TF–target networks from resources such as ENCODE and ChEA are incorporated as priors that guide the model toward biologically plausible regulatory structures. Importantly, the contribution of each TF–target interaction is learned from the data, allowing the model to down-weight or override potentially inaccurate or context-mismatched regulatory links. By aggregating signals across a large number of genes, the model further reduces sensitivity to noise and incompleteness in any single prior network.

      To evaluate robustness with respect to prior knowledge, we incorporated the DoRothEA regulon resource as an alternative TF–target prior with confidence-level filtering. We further performed ablation studies on the pancreas dataset and the gastrulation erythroid dataset using different TF–target resources, including ChEA, ENCODE, and their combinations with DoRothEA.

      The results on the pancreas dataset and the gastrulation erythroid dataset are shown in Figure S13 and Figure S14 respectively, which come up with the same conclusion. We observed highly consistent results across most TF–target prior combinations, including ChEA, ENCODE, ChEA+ENCODE, ChEA+DoRothEA, ENCODE+DoRothEA, and ChEA+ENCODE+DoRothEA. Using the pancreas dataset as example, the mean velocity consistency ranged from 0.985 to 0.995, the mean in-cluster coherence ranged from 0.983 to 0.992, and the mean cross-boundary direction correctness ranged from 0.719 to 0.740 across all settings. These consistently high and tightly bounded metrics indicate that TSvelo is largely insensitive to the specific choice of TF–target prior. Notably, these results further suggest that even when the underlying regulatory resources differ in origin (e.g., cell-line-derived vs. curated or aggregated datasets), the inferred dynamics remain stable.

      The only configuration showing reduced stability was the use of DoRothEA alone, particularly for cross-boundary direction correctness. This is likely due to its comparatively limited coverage of TF–target interactions. For instance, in the pancreas dataset, only 81 out of 2000 highly variable genes (HVGs) could be associated with TFs based on DoRothEA, corresponding to 102 TF–target links in total, which may limit downstream regulatory modeling. In contrast, ChEA covered 1793 genes with 13,976 TF–target links, and ENCODE covered 1854 genes with 33,076 links. These results further suggest that integrating multiple TF–target resources can improve performance, likely due to increased coverage and complementary regulatory information.

      We agree that regulatory interactions derived from resources such as ENCODE and ChEA may not fully generalize to primary tissues due to their context-dependent nature. In the revised Discussion, we explicitly clarify this limitation, particularly their inability to capture tissue-specific regulatory programs. We further highlight that incorporating context-specific regulatory data, such as single-cell chromatin accessibility or perturbation-based regulatory maps, represents an important direction for future improvement.

      (e) Lines 579-580: How is the grid search performed? More methodological details are required. If an existing method was used, please provide a citation.

      The grid search for the time step means that the model evaluates the loss in equation (10) across all candidate values of t<sub>step</sub> in the set {0,1,2,...,999}. This strategy was originally adopted in scVelo for optimizing the time step parameter. We have now added the corresponding citation to scVelo in the revised manuscript.

      (2) Application on pancreatic endocrine datasets

      (a) Lines 140-141: What is the definition of the final pseudotime-fitted time t or velocity pseudotime?

      There is no distinction between “final pseudotime”, “fitted time t” and “velocity pseudotime”. All of them refer to the same quantity in our framework. To eliminate any potential ambiguity, we have standardized the terminology by replacing “final pseudotime” with “pseudotime”.

      (b) Lines 143-144: The use of the velocity consistency metric to benchmark methods in multi-lineage datasets is incorrect. In multi-lineage differentiation systems, cells (e.g., those in fate priming stages) may inherently show inconsistency in their velocity. Thus, it is difficult to distinguish inconsistency caused by estimation error from that arising from biological signals. Velocity consistency metrics are only appropriate in systems with unidirectional trajectories (e.g., cell cycling). The abnormally high consistency values here raise concerns about whether the estimated velocities meaningfully capture lineage differences.

      We thank the reviewer for raising this important point regarding the use of the velocity consistency metric in multi-lineage systems. Velocity consistency was initially introduced by scVelo [1] and implemented as scvelo.velocity_confidence() in its package. Velocity consistency provides one of the few widely adopted quantitative criteria for benchmarking RNA velocities [2]. We agree that it is especially suitable for single-lineage processes. For datasets with clear multi-lineage differentiation (Fig. 5 and Fig. 6), we do not use this metric, precisely to avoid the issue highlighted by the reviewer.

      However, the pancreatic endocrine dataset (Fig. 2) exhibits minimal branching, making velocity consistency be more appropriate. As introduced by veloVI study, RNA velocities are supposed to change smoothly over the phenotypic manifold [3]. Higher consistency indicates that neighboring cells show compatible velocity directions, reflecting stable and coherence of the inferred velocity field. Additionally, multiple previous studies used velocity consistency to evaluate model performance on this pancreas dataset [2,3,4], providing a standard point of comparison.

      To better address your concerns, we have replaced the corresponding panel in Fig. 2 of the main text with an evaluation of cell-type separability in both the traditional 2D (unspliced–spliced) phase portrait and the learned 3D (α–unspliced–spliced) phase portrait by TSvelo (Author response image 4 in our response to your subsequent question). We appreciate your suggestions, as the comparison more clearly highlights the novelty and contribution of TSvelo and helps explain its improved performance. Now, the velocity consistency panel has been moved to the Supplementary Information. In addition, we have added a clearer explanation of the cross-boundary correctness metric in the revised manuscript.

      (1) Bergen, V., Lange, M., Peidli, S., Wolf, F. A., & Theis, F. J. (2020). Generalizing RNA velocity to transient cell states through dynamical modeling. Nature Biotechnology, 38(12), 1408-1414.

      (2) Luo, Y., Ren, J., Yang, Q. ... & Li, Q. (2026). Benchmarking RNA velocity methods across 17 independent studies, Cell Reports Methods, 101367.

      (3) Gayoso, A., Weiler, P., Lotfollahi, M., Klein, D., Hong, J., Streets, A., ... & Yosef, N. (2024). Deep generative modeling of transcriptional dynamics for RNA velocity analysis in single cells. Nature Methods, 21(1), 50-59.

      (4) Li, J., Pan, X., Yuan, Y., & Shen, H. B. (2024). TFvelo: gene regulation inspired RNA velocity estimation. Nature Communications, 15(1), 1387.

      (c) The improvement of TSvelo over other methods in terms of cross-boundary direction correctness looks marginal; a statistical test would help to assess its significance.

      We thank the reviewer for this insightful comment. In the revised manuscript, we have added statistical tests for evaluated metrics, including velocity consistency, cross-boundary direction correctness, and in-cluster coherence.

      As shown in Author response image 2, TSvelo significantly outperforms all baseline methods in terms of velocity consistency across both datasets. For in-cluster coherence, TSvelo achieves significantly better performance on the gastrulation (erythroid) dataset, while on the pancreas dataset it performs comparably to the best-performing baselines (UniTVelo and TFvelo) and significantly outperforms several competing methods, including CellDancer, Dynamo, and scVelo.

      For cross-boundary direction correctness, TSvelo shows consistent improvements in mean performance on the pancreas dataset (Author response image 3), and significantly outperforms Dynamo and scVelo on the gastrulation dataset. Although not all pairwise comparisons on cross-boundary direction correctness reach statistical significance, this is likely influenced by the limited number of independent samples (n = 7 and n = 4 for the two datasets, respectively), which reduces statistical power for detecting differences. Importantly, TSvelo still achieves the best average performance among all methods, indicating a consistent overall trend in favor of TSvelo.

      We have added these results into the revised manuscript.

      Author response image 2.

      The quantitative comparison between TSvelo and baseline approaches on the pancreas dataset (panel a) and the gastrulation erythroid dataset (panel b). In each plot, methods are ranked in descending order of their mean values. Numbers at the bottom indicate the sample size for each metric. Significance is determined using a one-sided Mann–Whitney U test. *****, ***, ** and * represent p < 0.00001, 0.0001 ≤ p < 0.001, 0.001 ≤ p < 0.01, and 0.01 ≤ p < 0.05, respectively.

      Author response image 3.

      The comparison of mean cross-boundary direction correctness on the pancreas dataset.

      (d) Lines 177-178: Based on the figure, TSvelo does not appear to clearly distinguish cell types. A quantitative metric, such as Adjusted Rand Index (ARI), should be provided.

      We thank the reviewer for this helpful suggestion. To quantitatively assess whether TSvelo can distinguish cell types, we evaluated the separability of cell-type labels in both the 2D (unspliced–spliced) phase portrait adopted by previous RNA velocity approaches, and the 3D (α–unspliced–spliced, α denotes the transcriptional rate) phase portrait introduced by TSvelo.

      Specifically, we evaluated how well the embedding preserves cell-type information using a k-nearest neighbors (kNN) classification accuracy with 5-fold cross-validation. Given an embedding matrix in 2D or 3D space (X 𝛜 ℝ<sup>n*d</sup>, where n is the number of cells and d is 2 or 3) and corresponding cell-type labels (y 𝛜 {1, … ,C}, we partition the data into five folds. For each fold (k), a kNN classifier with K = 5, denoted asf<sup>(k)</sup>, is trained on the training subset and evaluated on the held-out test subset. The classification accuracy for the k-th fold is defined as ℝ

      where n<sub>k</sub> is the number of samples in the test set and 1(.)is the indicator function. The final score is obtained by averaging across all folds:

      This metric directly assesses whether cells of the same type are positioned close to each other in the embedding space, and is widely used to quantify representation quality.

      Using this evaluation, we observed that the 3D phase portrait consistently achieves significantly higher accuracy than the 2D phase portrait (Author response image 4). The improvement is highly statistically significant (one-sided Mann–Whitney U test, p-value = 4.37 × 10<sup>-10</sup>), demonstrating that the 3D representation provides substantially better separation of cell types.

      We have added these quantitative results to the revised manuscript to complement the visual evidence and to clarify that TSvelo effectively distinguishes cell types in the learned representation.

      Author response image 4.

      The evaluation of the separability of cell-type labels in both the 2D (unspliced–spliced) phase portrait and the 3D (α–unspliced–spliced) phase portrait for the pancreas dataset.

      (e) Lines 179-183: The claim that traditional methods cannot capture dynamics in the unspliced-spliced phase portrait is vague. What specific aspect is not captured-the fitted values or something else? Evidence is lacking. Please provide a detailed explanation and quantitative metrics to support this claim.

      We thank the reviewer for this important comment. We have revised the text to more clearly illustrate this point using representative example genes as follows: “For instance, ANXA4 shows higher expression in Ductal cells compared to Ngn3 low EP cells, which mean its expression pattern exhibits an initial decrease followed by an increase. Such dynamics are not easily captured in the conventional unspliced–spliced phase portrait used by previous approaches, as many baseline methods implicitly assume a decreasing–then–increasing expression pattern. By comparison, TSvelo can still fit such expression pattern by using additional information from the 3D phase portrait.”

      In addition, we also clarify that the 2D u–s representation has limited capacity to separate heterogeneous dynamic cell states, which can affect downstream velocity field estimation. In the conventional 2D u–s phase portrait, cells from different dynamic regimes may overlap in the same region of the embedding space. This overlap reduces the identifiability of underlying transcriptional states and makes the inferred local dynamics more ambiguous. In contrast, TSvelo introduces an additional latent variable α, forming a 3D (α, u, s) phase portrait, which helps disentangle these mixed trajectories and yields a more structured and separable representation of cell dynamics. We have provided quantitative evidence in the previous response (Author response image 4). Briefly, the proposed 3D representation achieves consistently higher kNN classification accuracy (5-fold cross-validation, k=5) for cell state identification compared to the 2D u–s embedding.

      (3) Application to gastrulation erythroid datasets

      (a) Lines 191-194: The observation that velocity genes are enriched for erythropoiesis-related pathways is trivial, since the analysis is restricted to highly variable genes (HVGs) from an erythropoiesis dataset. This enrichment is expected and therefore not informative.

      We thank the reviewer for this comment and agree that such enrichment is expected given the use of HVGs from an erythropoiesis dataset. This analysis was included only as a preliminary sanity check to support the plausibility of the inferred velocity genes, rather than as a main result. We have accordingly simplified the description and clarified that this analysis serves only as a preliminary check in the revised manuscript.

      (b) Lines 227-228: It remains unclear how TSvelo "accurately captures the dynamics." What is the definition of dynamics in this context? Figure 3g shows unspliced/spliced vs. fitted time plots and phase portraits, but without a quantitative definition or measure, the claim of superiority cannot be supported. Visualization of a single gene is insufficient; a systematic and quantitative analysis is needed.

      We thank the reviewer for this important comment. We have revised the text to more clearly illustrate this point using representative example genes as follows: “For HSP90AB1, which exhibits a counter-clockwise pattern in the unspliced–spliced phase portrait, in contrast to the clockwise dynamics typically assumed by most baseline approaches, it is difficult for previous methods to capture this behavior, whereas TSvelo can still faithfully model such patterns. For genes such as RPS26, which have critical roles in the development in blood progenitors to erythroid40, the unspliced-spliced data is so noisy that cells of different types overlap in phase portrait. TSvelo can still captures the gene dynamics and reveals differences in transcription rates across cell types.”

      In addition, we explicitly emphasize the role of the 3D (α, u, s) phase portrait, which provides a more structured and separable representation of transcriptional states compared to the conventional 2D u–s space. This improved representation is the key factor underlying the advantages of TSvelo in modeling transcriptional processes. In the conventional 2D u–s phase portrait, cells from different transcriptional states may overlap, leading to reduced separability. In contrast, introducing the latent variable α expands the representation to a 3D space, which helps disentangle these mixed states and yields a clearer phase structure. Similar to our previous response in Author response image 4, we provide quantitative evidence on this gastrulation erythroid dataset in Figure S7, showing that the 3D representation achieves consistently higher kNN classification accuracy for cell state separation compared to the 2D u–s embedding (one-sided Mann–Whitney U test, p-value = 0.002).

      (4) Application to the mouse brain and other datasets

      (a) Lines 280-281: The authors cannot claim that velocity streams are smoother in TSvelo than in Multivelo based solely on 2D visualization. Similarly, claiming that one model predicts the correct differentiation trajectory from a 2D projection is over-interpretation, as has been discussed in prior literature see PMID: 37885016.

      We thank the reviewer for this important comment. Consistent with other RNA velocity studies, TSvelo employs the 2D UMAP stream plot for visualizing the results. We agree that conclusions based solely on 2D visualizations may lead to over-interpretation. Our intention was to provide an intuitive visualization rather than a rigorous quantitative comparison. Accordingly, we have revised the text to avoid making definitive claims about smoothness or correctness of differentiation trajectories based solely on 2D projections.

      (b) Lines 304-306: Beyond transcriptional signal estimation, how is regulation inferred solely from scRNA-seq data validated, especially compared with scATAC-seq data? Are there cases where transcriptome-based regulatory inference is supported by epigenomic evidence, thereby demonstrating TSvelo's GRN inference accuracy?

      We thank the reviewer for this important question regarding the validation of regulatory inference derived from scRNA-seq data and its comparison to scATAC-seq-based evidence.

      We would like to first clarify the scope of TSvelo. Similar to existing RNA velocity methods, the primary goal of TSvelo is to model transcriptional dynamics and accurately infer cell state transitions and cell fate trajectories. In this context, gene regulatory information is not inferred de novo from data, but incorporated as prior knowledge from curated TF–target databases to guide and constrain the dynamics modeling process, as described in our Introduction.

      We have conducted the requested analysis by computing gene-wise chrome accessibility rate used in MultiVelo and the learned transcription rate from TSvelo, and evaluated their correlation across genes. As shown in Figure S15, the two estimates exhibit almost no global correlation across genes, indicating that they capture substantially different aspects of regulatory information.

      This discrepancy is not unexpected and reflects the fundamental differences between these modalities. scATAC-seq measures chromatin accessibility, which provides a proxy for cis-regulatory potential of genomic regions. In contrast, TF RNA expression reflects downstream transcriptional output, which is shaped by multiple regulatory layers, including post-transcriptional regulation, protein activity, temporal delays, and indirect regulation through intermediate transcriptional or signaling pathways. As a result, these two modalities are expected to capture complementary but not directly comparable aspects of gene regulation.

      We acknowledge that scATAC-seq provides valuable complementary information on chromatin accessibility and regulatory potential, and will consider incorporating matched multi-omics data in future work. In the revised manuscript, we further clarify that TSvelo is an RNA velocity method that incorporates prior knowledge from curated TF–target databases, and we have added a discussion on the potential use of scATAC-seq data for future extension of our framework.

      (c) The claim that TSvelo can model multi-lineage datasets hinges on its use of PAGA for lineage segmentation, followed by independent modeling of dynamics within each subset. However, the procedure for merging results across subsets remains unclear.

      We thank the reviewer for pointing out that the merging step was not sufficiently described. After modeling dynamics independently within each lineage-specific subset, TSvelo integrates the results via a weighted aggregation procedure at the cell level.

      For each cell and each inferred quantity (e.g., velocity or other dynamic variables), we collect the estimates obtained from different lineage-specific models and combine them using a weighted average. The weights are defined by the size of each lineage, reflecting its statistical support. We have clarified details about this merging procedure in the Methods section.

      This aggregation reconciles multiple lineage-specific estimates for the same cell into a single value and mitigates discontinuities that could arise from directly combining independent lineage analyses. The resulting values define a unified set of dynamics for each cell across lineages.

      Reviewer #3 (Public review):

      Despite the abundance of RNA velocity tools, there are still major limitations, and there is strong skepticism about the results these methods lead to. In this paper, the authors try to address some limitations of current RNA velocity approaches by proposing a unified framework to jointly infer transcriptional and splicing dynamics. The method is then benchmarked on 6 real datasets against the most popular RNA velocity tools.

      While the approach has the potential to be of interest for the field, and may present improvements compared to existing approaches, there are some major limitations that should be addressed, particularly concerning the benchmark (see major comment 1).

      Major comments:

      (1) My main criticism concerns the benchmarking: real data lack a ground truth, and are absolutely not ideal for comparing methods, because one can only speculate what results appear to be more plausible.

      A solid and extensive simulation study, which covers various scenarios and possibly distinct data-generating models, is needed for comparing approaches. The authors should check, for example, the simulation studies in the BayVel approach (Section 4, BayVel: A Bayesian Framework for RNA Velocity Estimation in Single-Cell Transcriptomics). Clearly, all methods should be included in the simulation.

      Following your recommendation, we have added the simulation analysis to compare TSvelo with existing RNA velocity approaches. We generated synthetic single-cell RNA velocity datasets using a mechanistic transcriptional dynamics model with one or multiple developmental branches. The system included 200 genes, among which 30 were designated as transcription factors (TFs).

      For each branch, we independently sampled a TF–target regulatory matrix W ϵ ℝ<sup>30×200</sup> from a standard normal distribution to simulate distinct GRN structures. Gene expression dynamics were modeled using a coupled ordinary differential equation (ODE) system describing unspliced and spliced RNA abundances:

      where u and s denote unspliced and spliced RNA levels, respectively. The transcription rate α was computed as a nonlinear function of TF expression, defined as a weighted sum of spliced TF abundance, followed by clipping to ensure bounded activation.

      Each branch is initialized from the same randomly sampled initial condition drawn from a gamma distribution, allowing controlled divergence of trajectories driven solely by branch-specific regulatory programs.

      To simulate observed sequencing counts, we introduced technical noise by scaling latent expression levels with cell-specific library sizes drawn from a log-normal distribution. The resulting expression counts were generated using a negative binomial sampling model:

      where θ controls over dispersion, with smaller values corresponding to higher noise levels. The final datasets consist of paired unspliced (U) and spliced (S) count matrices with realistic transcriptional stochasticity and branching gene regulatory dynamics. For each branch, cells were further divided into three developmental stages for downstream analysis.

      We evaluated TSvelo and those splicing-based RNA velocity approaches on multiple simulated datasets with varying numbers of branches and noise levels. There are one, two or three branches start from the same cell group in these datasets (Branch 1: stage 0 - stage 1 - stage 2. Branch 2: stage 0 - stage 3 - stage 4. Branch 3: stage 0 - stage 5 - stage 6). We primarily assessed performance using the cross-boundary direction correctness (CBDir) metric, as it directly evaluates inferred trajectories against ground-truth cell stage annotations, which have been widely adopted in RNA velocity studies such as VeloAE and UniTvelo. In detail, Cross-boundary direction correctness assesses the accuracy of transitions from a source cluster to a target cluster by examining the boundary cells, and requires ground truth annotations. We directly run the function unitvelo.evaluate() provided in UniTVelo to obtain the Cross-boundary direction correctness. In detail, the CBDir is calculated as follows:

      where θ controls over dispersion, with smaller values corresponding to higher noise levels. The final datasets consist of paired unspliced (U) and spliced (S) count matrices with realistic transcriptional stochasticity and branching gene regulatory dynamics. For each branch, cells were further divided into three developmental stages for downstream analysis.

      where C<sub>A</sub> denotes the set of cells in the target cluster A, and N(c) represents the neighboring cells of a given cell c v<sub>c</sub> and x<sub>c</sub> denote the low-dimensional velocity and state vectors of cell c, respectively, and x<sub>c’</sub> denotes the state vector of its neighboring cell.

      As shown in Figure S2, TSvelo consistently achieves the highest accuracy across all simulation settings, particularly in scenarios with complex branching structures, which pose significant challenges for baseline methods.

      (2) Related to the above: since a ground truth is missing, the real data analyses need to be interpreted with caution. I recommend avoiding strong statements, such as "successfully captures the correct gene dynamics", or "accurately infer", in favour of milder statements supported by the data, such as "... aligns with the biological processes described" (as in page 12), or "results are compatible with current biological knowledge", etc...

      We thank the reviewer for this helpful comment. We agree that analyses on real datasets should be interpreted with appropriate caution because definitive ground truth is typically unavailable. Following the reviewer’s suggestion, we have revised the wording throughout the manuscript to avoid overly strong claims. For example, statements such as “successfully captures the correct gene dynamics” and “accurately infer” have been replaced with more cautious descriptions such as “consistent with known biological processes”.

      (3) Many methods perform RNA velocity analyses. While there is a brief description, I think it'd be useful to have a schematic summary (e.g., via a Table) of the main conceptual, mathematical, and computational characteristics of each approach.

      We thank the reviewer for this insightful suggestion. We agree that a structured summary of existing RNA velocity methods would improve clarity and accessibility. We have added a new summary table (Table S1) that systematically compares representative RNA velocity approaches in the supplementary information.

      (4) Related to the above: I struggled to identify the main conceptual novelty of TSvelo, compared to existing approaches. I recommend explaining this aspect more extensively.

      We thank the reviewer for this insightful comment. We agree that the conceptual novelty of TSvelo can be more clearly articulated.

      In the revised manuscript, we have expanded the discussion at the beginning of the Results section to explicitly highlight the key distinctions between TSvelo and existing approaches. Specifically, we now clarify that most existing RNA velocity methods predominantly focus on splicing dynamics and typically operate in a gene-wise manner, without capturing coordinated dynamics across genes. In contrast, TSvelo models the full cascade of transcriptional regulation, transcription, and splicing within a unified framework, and estimates RNA velocity jointly across all genes, thereby capturing their coordinated dynamics at the system level.

      (5) A computational benchmark is missing; I'd appreciate seeing the runtime and memory cost of all methods in a couple of datasets.

      We thank the reviewer for this helpful suggestion regarding computational benchmarking. In the revised manuscript, we have added a systematic comparison of runtime and GPU memory usage across TSvelo and ba methods using simulated datasets of increasing scale (600, 1200, and 1800 cells) on our NVIDIA GeForce RTX 3090 device with 24 GB memory.

      Table S2 shows differences in computational efficiency and resource requirements among methods. Specifically, classical methods such as scVelo and Dynamo exhibit very fast runtimes (10–24 seconds) and do not rely on GPU acceleration, reflecting their relatively lightweight modeling strategies. In contrast, deep learning–based approaches, including UniTVelo, cellDancer, and TSvelo, have higher computational costs due to their increased model complexity.

      TSvelo exhibits a stable GPU memory footprint (~1.26 GB) across different dataset sizes, indicating that its memory usage is primarily determined by model architecture rather than the number of cells. This level of memory consumption is well within the capacity of modern GPUs and does not pose practical limitations. In terms of runtime, TSvelo scales approximately linearly with dataset size. The higher computational cost of TSvelo is mainly due to its EM-style optimization procedure, where each M-step also involves multiple optimization updates to infer gene regulatory effects in a global model. This design enables TSvelo to explicitly incorporate regulatory priors and jointly model gene interactions, which is not supported by these baseline methods.

      To further improve runtime efficiency, TSvelo allows flexible control of the number of EM iterations. As shown in Figure S16 and Table S3, we evaluated performance under different iteration settings on the simulation dataset. The early stopping strategy employed in the EM framework of TSvelo, which will stop modeling if the loss is not further reduced in the last 3 iterations. Results show that convergence is typically achieved within 3 iterations for this dataset, and increasing the maximum number of iterations beyond this does not further change the results. Notably, even a single iteration already yields competitive performance, likely benefiting from the strong initialization based on unspliced-to-spliced temporal delay.

      Overall, these results highlight a trade-off between computational efficiency and modeling expressiveness. While TSvelo is more computationally demanding than classical approaches, it provides a more flexible framework for incorporating regulatory information and capturing complex gene interactions, which we believe justifies the additional computational cost in scenarios requiring accurate dynamical inference.

      (6) I think BayVel (mentioned above) should be added to the list of competing methods (both in the text and in the benchmarks). The package can be found here: https://github.com/elenasabbioni/BayVel_pkgJulia.

      We thank the reviewer for suggesting BayVel and for providing the repository link. We carefully review the available resources, including both the BayVel_pkgJulia and the BayVel_notebooks, and we appreciate the authors’ efforts in making their code and data publicly available.

      We note that BayVel repositories primarily provide scripts and data for reproducing the figures and results reported in their manuscript. However, at present, the available resources do not yet provide a complete guideline or standardized pipeline for applying BayVel to new datasets. To ensure a fair and reproducible comparison, we therefore tend to use BayVel results officially provided by the authors. We are grateful that the BayVel results on the pancreas dataset is released at BayVel_notebooks page: https://github.com/elenasabbioni/BayVel_notebooks/tree/main/real%20data/Pancreas/moments/output.

      Based on these results, we conducted comparisons across all methods on the pancreas dataset, with quantitative evaluations shown in Author response image 55. In each plot, methods are ranked in descending order of their mean values. Numbers at the bottom indicate the sample size for each metric. Statistical significance is assessed using a one-sided Mann–Whitney U test, where *****, ***, **, and * denote p < 0.00001, 0.0001 ≤ p < 0.001, 0.001 ≤ p < 0.01, and 0.01 ≤ p < 0.05, respectively.

      BayVel has now been included in the Introduction, and corresponding comparisons have been added in the revised manuscript.

      Author response image 5.

      The quantitative comparison between TSvelo and baseline approaches on the pancreas dataset. In each plot, methods are ranked in descending order of their mean values. Numbers at the bottom indicate the sample size for each metric. Significance is determined using a one-sided Mann–Whitney U test. *****, ****,***, ** and * represent p < 0.00001, 0.00001 ≤ p < 0.0001, 0.0001 ≤ p < 0.001, 0.001 ≤ p < 0.01, and 0.01 ≤ p < 0.05, respectively.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please carefully proofread the text. Some typos:

      (1) Line 110: differentia -> differential.

      (2) Line 280: ".," to be corrected.

      (3) Line 566: optimize -> optimizes.

      We thank the reviewer for carefully proofreading the manuscript and for pointing out these typographical errors. We have corrected the identified typos in the revised manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Regarding Major Comment 1 in the Public Review, I contacted BayVel authors, who told me that they'll upload all their scripts here within a few days: https://github.com/elenasabbioni/BayVel_notebooks

      Thank you very much for reaching out to the BayVel authors. We sincerely appreciate the BayVel authors’ efforts to make their scripts and results publicly available through BayVel_notebooks. We believe this is a valuable contribution that will greatly benefit the community.

      We have followed the repository and have now included BayVel in the revised manuscript, with corresponding comparisons added to both the main text and the benchmarking results.

      (2) Page 9 mentions "consistency", "coherence", and "correctness". Instead of these qualitative (and potentially subjective) evaluations, I'd appreciate using quantitative metrics or visual descriptions when differences are visually clear.

      We thank the reviewer for this insightful comment. The terms “velocity consistency,” “in-cluster coherence,” and “cross-boundary correctness” used in our manuscript are not intended as subjective descriptions. They correspond to commonly used evaluation criteria in this field and have been adopted as quantitative metrics in previous studies, such as VeloAE[1] and UniTVelo[2]. We have incorporated the following updated definition into the Methods section.

      (1) Velocity consistency (VCon). We used the scvelo.velocity_confidence() function from scVelo to evaluate velocity consistency, interpreting the results as a measure of how consistent velocities are within neighboring cells. Velocity consistency is especially suitable for evaluating the RNA velocity modeling on single lineage. For each cell , the velocity consistency is calculated as follows:

      Where N (c) represents the neighboring cells of a given cell c v<sub>c</sub> v<sub>c’</sub> denote the low-dimensional velocity vectors of cell cand its neighboring cell c’.

      (2) Cross-boundary direction correctness (CBDir). Cross-boundary direction correctness assesses the accuracy of transitions from a source cluster to a target cluster by examining the boundary cells, and requires ground truth annotations. We directly run the function unitvelo.evaluate() provided in UniTVelo to obtain the Cross-boundary direction correctness. In detail, the CBDir is calculated as follows:

      Where C<sub>A</sub> denotes the set of cells in the target cluster A, and represents the neighboring cells of a given cell c v<sub>c</sub> v<sub>c’</sub> denote the low-dimensional velocity and state vectors of cell cand its neighboring cell c’.

      (3) Within-cluster velocity coherence (ICCoh). Within-cluster velocity coherence measures the coherence of velocities within a single cluster using a cosine similarity score between cell velocities. We applied the function unitvelo.evaluate() provided by UniTVelo to directly compute the within-cluster velocity coherence. Using the same notation as defined above, the CBDir is calculated as follows:

      (1) Qiao, C. & Huang, Y. Representation learning of RNA velocity reveals robust cell transitions. Proceedings of the National Academy of Sciences 118, e2105859118 (2021).

      (2) Gao, M., Qiao, C. & Huang, Y. UniTVelo: temporally unified RNA velocity reinforces single-cell trajectory inference. Nature Communications 13, 6586 (2022).

      (3) At page 3, some objects are not defined after formula (3):

      ReLU finction, and w_gi

      Additionally, parenthesis of ReLU function should be bigger.

      We thank the reviewer for pointing this out. In the revised manuscript, we have explicitly defined the ReLU activation function and clarified that w<sub>gi</sub> represents the regulatory weight of TF i on the target gene g. In addition, we have adjusted the formatting of Eq. (3) by enlarging the parentheses in the ReLU function to improve readability.

    1. Author response:

      The following is the authors’ response to the original reviews.

      General recommendations (from the Reviewing Editor):

      The reviewers discussed the revision at length, and all were appreciative of the revisions to the paper. Nonetheless, they agreed that the evidence against alternative hypotheses was not yet decisive, and it may not be possible to provide the evidence needed given the difficulty of acquiring this data. Thus they feel that a more nuanced interpretation of the data and tempering of the conclusions is necessary. These points are described in more detail in the reviewer-specific comments in the Public reviews.

      We thank the editor and the reviewers for their constructive discussion. In this revision, we have adopted these recommendations: we have tempered our conclusions and removed binary framing, taking into consideration that other alternative explanations might exist. We have also expanded the Discussion to consider additional potential mechanisms and added corresponding limitations. We also changed the paper title to avoid strong inference; the new title is “Evidence that humans underestimate body mass in microgravity: kinematic signatures in reaching movements during spaceflight”.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors have conducted substantial additional analyses to address the reviewers' comments. However, several key points still require attention. I was unable to see the correspondence between the model predictions and the data in the added quantitative analysis. In the rebuttal letter, the delta peak speed time displays values in the range of [20, 30] ms, whereas the data were negative for the 45{degree sign} direction. Should the reader directly compare panel B of Figure 6 with Figure 1E? The correspondence between the model and the data should be made more apparent in Figure 6. Furthermore, the rebuttal states that a quantitative prediction was not expected, yet it subsequently argues that there was a quantitative match. Overall, this response remains unclear.

      We thank the reviewer raising the question about Figure 6B. We would like to clarify that the phrase "quantitative match" in the summary of our previous rebuttal letter was a wording error; in fact, the subsequent detailed responses consistently and correctly described the comparison as qualitative. We apologize for the confusion this may have caused, and address this point below.

      First, we have revised the manuscript to clarify this point. We have added the following statement: "We note that these correlations evaluate the directional trend rather than the absolute magnitude of the effects; a precise quantitative match is not expected given the simplifications of the two-joint arm model." in the main text.

      Second, we have replaced Figure 6 with a revised version that presents model-predicted Δ values and experimentally observed Δ values side by side, allowing for a more intuitive visual comparison. As shown in the updated figure, the directional trends are broadly consistent amplitude changes and timing shifts are rank-ordered by movement direction in both model and data while the absolute magnitudes do not precisely match. We believe this layout makes the intended comparison more transparent.

      As discussed in our previous response, as noted above, a precise quantitative match is not expected given our model's simplifications, and this level of qualitative comparison is consistent with established practice in similar modeling studies (e.g., Gaveau et al., 2016).

      Regarding the negative Δ peak speed time at 45°: as shown in our statistical analyses (Figure 4A, Figure 5F), there was no significant timing change at 45°. The negative value reflects a small, non-significant mean difference. The key pattern that timing advance increases for directions associated with higher effective inertia holds for the 90° and 135° directions, which is the directional trend our analysis was designed to capture.

      A follow-up question concerns the argument about strategic slowing. The authors argue that this explanation can be rejected because the timing of peak speed should be delayed, contrary to the data. However, there appears to be a sign difference between the model and the data for the 45{degree sign} direction, which means that it was delayed in this case. Did I understand correctly? In that regard, I believe that the hypothesis of strategic slowing cannot yet be firmly rejected and the discussion should more clearly indicate that this argument is based on some, but not all, directions.

      I agree with the authors on the importance of the mass underestimation hypothesis, and I am not particularly committed to the strategic slowing explanation, but I do not see a strong argument against it. If the conclusion relies on the sign of the delta peak speed, then the authors' claims are not valid across all directions, and greater caution in the interpretation and discussion is warranted. Regarding the peak acceleration time, I would be hesitant to draw firm conclusions based on differences smaller than 10 ms (Figures R3 and 6D).

      The authors state in the rebuttal that the two hypotheses are competing. This is not accurate, as they are not mutually exclusive and could even vary as a function of movement direction. The abstract also claims that the data "refutes" strategic slowing, which I believe is too strong. The main issue is that, based on the authors' revised manuscript, the lack of quantitative agreement between the model and the data for the mass underestimation hypothesis is considered acceptable because a precise quantitative match is not expected, and the predictions overall agree for some (though not all) directions and phases (excluding post-in). That is reasonable, but by the same logic, the small differences between the model prediction and the strategic slowing hypothesis should not be taken as firm evidence against it, as the authors seem to suggest. In practice, I recommend a more transparent and cautious interpretation to avoid giving readers the false impression that the evidence is decisive. The mass underestimation hypothesis is clearly supported, but the remaining aspects are less clear, and several features of the data remain unexplained.

      We thank the reviewer for this critical assessment. We acknowledge that our previous framing was too binary, and we agree that strategic slowing and mass underestimation are not mutually exclusive. We would like to clarify our view: we did not find evidence supporting strategic slowing (e.g., slower reaction times, symmetric velocity/acceleration peaks), whereas we did find evidence supporting mass underestimation (asymmetric peaks, unchanged reaction times, more sub movements). This is not a case of rejecting one hypothesis to affirm the other; our data simply do not support one while providing positive evidence for the other. We do not rule out the possibility that both mechanisms could operate together, though we note that our data did not reveal evidence supporting strategic slowing in the current reaching task.

      We also agree that the lack of significant timing changes at 45° limits the scope of our argument against strategic slowing in that direction. However, the null result at 45° likewise cannot serve as positive evidence for strategic slowing either. As discussed in our previous revision and in Discussion, this null effect may arise because 45° reaches are predominantly single-joint (evidenced by curvature patterns characteristic), making them less suitable for modeling with a simplified two-link arm model than the 90° and 135° directions.

      In line with these considerations, we have made the following revisions to the manuscript:

      (1) We have removed binary framing throughout, replacing claims of mutual exclusivity or outright rejection of strategic slowing with more measured language. For example, "refutes" in the abstract has been changed to "These findings provide support for the body mass underestimation hypothesis while being inconsistent with the strategic slowing hypothesis." The two hypotheses are no longer presented as mutually exclusive, and strategic slowing is now characterized as insufficient to fully explain the direction-dependent pattern, rather than ruled out entirely.

      (2) We have revised the conclusion. The concluding paragraph no longer presents an either-or outcome. We describe the direction-dependent under-actuation pattern, note that it strongly supports mass underestimation while not being readily explained by a uniform strategic adjustment, and acknowledge that other factors may also contribute. A new limitation paragraph discusses the simplified nature of our model and acknowledges that other neurophysiological and biomechanical factors cannot be excluded.

      Reviewer #2 (Public review):

      This study explores the underlying causes of the generalized movement slowness observed in astronauts in weightlessness compared to their performance on Earth. The authors argue that this movement slowness stems from an underestimation of mass rather than a deliberate reduction in speed for enhanced stability and safety.

      Overall, this is a fascinating and well-written work. The kinematic analysis is thorough and comprehensive. The design of the study is solid, the collected dataset is rare, and the model adds confidence to the proposed conclusions.

      Compared to the previous version, the authors have thoroughly addressed my concerns. The model is now clear and well-articulated, and alternative hypotheses have been ruled out convincingly. The paper is improved and suitable for publication in my opinion, making a significant contribution to the field.

      Strengths:

      Comprehensive analysis of a unique data set of reaching movement in microgravity

      Use of a sensible and well-thought experimental approach

      State-of-the-art analyses of main kinematic parameter

      Computational model simulations of arm reaching to test alternative hypotheses and support the mass underestimation one

      This work has no major weakness as it stands, and the discussion provides a fair evaluation of the findings and conclusions.

      We thank the reviewer for the supportive feedback, and we are grateful for the earlier comments that helped us improve the manuscript.

      Reviewer #3 (Public review):

      Summary:

      The authors describe an interesting study of arm movements carried out in weightlessness after a prolonged exposure to the so-called microgravity conditions of orbital spaceflight. Subjects performed radial point-to-point motions of the fingertip on a touch pad. The authors note a reduction in movement speed in weightlessness, which they hypothesize could be due to either an overall strategy of lowering movement speed to better accommodate the instability of the body in weightlessness or an underestimation of body mass. They conclude for the latter, mainly based on two effects. One, slowing in weightlessness is greater for movement directions with higher effective mass at the end effector of the arm. Two, they present evidence for increased number of corrective submovements in weightlessness. They contend that this provides conclusive evidence to accept the hypothesis of an underestimation of body mass.

      Strengths:

      In my opinion, the study provides a valuable contribution, the theoretical aspects are well presented through simulations, the statistical analyses are meticulous, the applicable literature is comprehensively considered and cited and the manuscript is well written.

      Weaknesses:

      I nevertheless am of the opinion that the interpretation of the observations leaves room for other possible explanations of the observed phenomenon, thus weakening the strength of the arguments.

      To strengthen the conclusions, I feel that the following points would need to be addressed:

      We thank the reviewer for the insightful critique and constructive suggestions. Following the reviewer's advice, we have re-framed our Introduction and Discussion to present mass underestimation as a plausible mechanism identified by our simplified model, while explicitly acknowledging other potential factors. Below we address each point in detail.

      (1) The authors model the movement control through equations that derive the input control variable in terms of the force acting on the hand and treating the arm as a second-order low pass filter (Eq. 13). Underestimation of the mass in the computation of a feedforward command would lead to a lower-than-expected displacement to that command. But it is not clear if and how the authors account for a potential modification of the time constants of the 2nd order system. The CNS does not effectuate movements with pure torque generators. Muscles have elastic properties that depend on their tonic excitation level, reflex feedback and other parameters. Indeed, Fisk et al.* showed variations of movement characteristics consistent with lower muscle tone, lower bandwidth and lower damping ratio in 0g compared to 1g. Could the variations in the response to the initial feedforward command be explained by a misrepresentation of the limbs damping and natural frequency, leading to greater uncertainty to the consequences of the initial command. This would still be an argument for un-adapted feedforward control of the movement, leading to the need for more corrective movements. But it would not necessarily reflect an underestimation of body mass.

      *Fisk, J. O. H. N., Lackner, J. R., & DiZio, P. A. U. L. (1993). Gravitoinertial force level influences arm movement control. Journal of neurophysiology, 69(2), 504-511.

      While the authors attempt to differentiate their study from previous studies where limb neuromechanical impedance was shown to be modified in weightlessness by emphasizing that in the current study the movements were rapid and the initial movement is "feedforward". But this incorrectly implies that the limb's mechanical response to the motor command is determined only by active feedback mechanisms. In fact:

      (a) All commands to the muscle pass through the motor neurons. These neurons receive descending activations related not only to the volitional movement, but also to the dynamic state of the body and the influence of other sensory inputs, including the vestibular system. A decrease in descending influences from the vestibular organs will lower the background sensitivity to all other neural influences on the motor neuron. Thus, the motor neuron may be less sensitive to the other volitional and reflexive synaptic inputs that it may receive.

      (b) Muscle tone plays a significant role in determining the force and the time course of the muscle contraction. In a weightless environment, where tonic muscle activity is likely to be reduced, there is the distinct possibility that muscles will react more slowly and with lower amplitude to an otherwise equivalent descending motor command, particularly in the initial moments before spinal reflexes come into play. These, and other neuronal mechanisms could lead to the "under-actuation" effect observed in the current study, without necessarily being reflective of an underestimation of mass per se.

      The reviewer raises an important point that the observed underactuation may not necessarily reflect mass underestimation per se. It could also arise from changes in the time constants of the control system, tonic muscle activation levels, vestibular descending inputs, or altered spinal reflex gains. We agree that our simplified model does not capture these neuromuscular factors, and we have made several revisions to address this concern.

      In the Discussion (paragraph 4), we have added a new substantive section discussing how reduced tonic muscle activity, diminished vestibular inputs to motor neurons, and altered muscle activation dynamics (Fisk et al., 1993) may contribute to the observed under-actuation independently of mass misestimation. We argue that while these factors likely affect motor output, they would be expected to produce a relatively uniform effect across movement directions, as tonic muscle activation and vestibular descending inputs are not specific to a particular reaching direction. In contrast, the direction-dependent pattern of our results with greater effects for directions involving higher effective mass is more naturally explained by a misrepresentation of inertial properties than by a uniform change in neuromuscular excitability. Nevertheless, we explicitly acknowledge that these mechanisms may act in concert with mass underestimation, and that our current data cannot fully disentangle them.

      Additionally, the paragraph discussing proprioceptive mechanisms (paragraph 6 of Discussion) now opens with the conditional framing "If mass underestimation contributes to the observed underactuation," and closes by noting that the same proprioceptive degradation could affect motor output through other pathways such as reducing tonic muscle activation or altering spinal reflex gains independent of any explicit misrepresentation of body mass.

      We have also added a new limitation (the fourth in the Limitations section) explicitly acknowledging that our model treats muscles as ideal torque generators and does not capture potential changes in muscle activation dynamics, damping, or reflex gains that may occur in microgravity. Future studies combining detailed musculoskeletal modeling with direct measurements of muscle activation, joint impedance, and trunk kinematics would be needed to distinguish between mass underestimation and other sources of underactuation.

      That said, the assumption of relatively preserved muscle properties is partly supported by the available evidence. A systematic review of simulated microgravity studies found that upper limb maximal voluntary contraction remained mostly unchanged for up to 45 days of unloading, and that upper limb muscles declined substantially more slowly than lower limb and trunk muscles (Winnard et al., 2019). A more recent review similarly reported that upper limb muscle outcomes are less affected by microgravity exposure (Bosutti et al., 2025). This is also consistent with our own unpublished observations in Chinese astronauts, which did not indicate an obvious decline in upper limb force output. While these findings do not rule out subtler changes in muscle tone or activation dynamics, they suggest that gross alterations in upper limb neuromuscular capacity are unlikely to be the primary driver of the underactuation we observed.

      Refs.

      Winnard, A., Scott, J., Waters, N., Vance, M., & Caplan, N. (2019). Effect of time on human muscle outcomes during simulated microgravity exposure without countermeasures—systematic review. Frontiers in physiology, 10, 1046.

      Bosutti, A., Ganse, B., Maffiuletti, N. A., Wüst, R. C., Strijkers, G. J., Sanderson, A., & Degens, H. (2025). Microgravity‐induced changes in skeletal muscle and possible countermeasures: What we can learn from bed rest and human space studies. Experimental Physiology.

      (2) The subject's body in weightless is much more sensitive to reaction forces in interactions with the environment in the absence of the anchoring effect of gravity pushing the body into the floor and in the absence of anticipatory postural adjustments that typically accompany upper-limb motions in Earth gravity in order to maintain an upright posture. The authors dismiss this possibility because the taikonauts were asked to stabilize their bodies with the contralateral hand. But the authors present no evidence that this was sufficient to maintain the shoulder and trunk at a strictly constant position, as is supposed by the simplified biomechanical model used in their optimal control framework. Indeed, a small backward motion of the shoulder would result in a smaller acceleration of the fingertip and a smaller extent of the initial ballistic motion of the hand with respect to the measurement device (the tablet), consistent with the observations reported in the study. Note that stability of the base might explain why 45º movements were apparently less affected in weightlessness, according to many of the reported analyses, including those related to corrective movements (Fig. 5 B, C, F; Fig. 6D), than the other two directions. If the trunk is being stabilized by the left arm, the same reaction forces on the trunk due to the acceleration of the hand will result in less effective torque on the trunk, given that the reaction forces act with a much smaller moment arm with respect to the left shoulder (the hand movement axis passes approximately through the left shoulder for the 45º target) compared to either the forward or rightward motions of the hand.

      The reviewer raises an important point about the potential influence of reaction forces on trunk and shoulder stability in microgravity. We have revised the relevant Discussion paragraph to address this concern more thoroughly.

      We would like to clarify that, in addition to stabilizing the body with the left hand grasping a fixed bar, the taikonauts’ feet were also constrained with foot straps, providing multi-point stabilization. Furthermore, the reviewer's trunk displacement hypothesis predicts that the 45° direction should be systematically less affected across all kinematic measures. However, while 45° did not show significant changes in the timing of kinematics peaks, it did show significant changes in movement duration, peak acceleration, and peak speed comparable to the other directions. This dissociation is difficult to reconcile with a uniform trunk displacement artifact, but is consistent with a direction-dependent inertial effect.

      We acknowledge that we did not directly measure trunk or shoulder kinematics, highlight that we did our best to provide multi-point stabilization in our setup, and we have added this as a limitation in the revised Discussion.

      (3) The above is exacerbated by potential changes in the frictional forces between the fingertip and the tablet. The movements were measured by having the subjects slide their finger on the surface of a touch screen. In weightlessness, the implications of this contact can be expected to be quite different than on the ground. While these forces may be low on Earth, the fact is that we do not know what forces the taikonauts used on orbit. In weightlessness, the taikonauts would need to actively press downward to maintain contact with the screen, while on Earth gravity will do the work. The tangential forces that resist movement due to friction might therefore be different in 0g. . Indeed, given the increased instability of the body and the increased uncertainty of movement direction of the hand, taikonauts may have been induced to apply greater forces against the tablet in order to maintain contact in weightlessness, which would in turn slow the motion of the finger on the table and increase the reaction forces acting on the trunk. This could be particularly relevant given that the effect of friction would interact with the limb in a direction-dependent fashion, given the anisotropy of the equivalent mass at the fingertip evoked by the authors

      We agree that in microgravity, taikonauts must actively press on the screen to maintain contact, potentially altering normal forces and thus friction compared to ground conditions. We have acknowledged this point in the revised Discussion. However, we note several reasons why friction is unlikely to be the dominant factor. First, the tablet uses a capacitive touchscreen, which registers touch through changes in electrical capacitance and does not require substantial normal force to maintain contact. Second, typical tangential friction forces during touchscreen interaction range from 0.1 to 0.5 N (Ayyildiz et al., 2018), which are small compared to the 10–15 N required to accelerate the arm during reaching. Third, touchscreen performance has been shown to be largely unaffected during long-duration spaceflight (Holden et al., 2022). Lastly but importantly, the friction hypothesis does not readily account for the direction-specific pattern of effects we observed. While we cannot exclude a contribution of altered friction, particularly in interaction with the direction-dependent effective mass, its magnitude makes it unlikely to account for the observed kinematic changes.

      Ref:

      Ayyildiz, M., Scaraggi, M., Sirin, O., Basdogan, C., & Persson, B. N. J. (2018). Contact mechanics between the human finger and a touchscreen under electroadhesion. Proceedings of the National Academy of Sciences of the United States of America, 115(50), 12668–12673.

      Holden, K., Greene, M., Vincent Cross, E., Sandor, A., Thompson, S., Feiveson, A., & Munson, B. (2023). Effects of long-duration microgravity and gravitational transitions on fine motor skills. Human Factors, 65(6), 1046-1058.

      I feel that the authors have done an admirable job of exploring the how to explain the modifications to movement kinematics that they observed on orbit within the constraints of the optimal control theory applied to a simplified model of the human motor system. While I fully appreciate the value of such models to provide insights into question of human sensorimotor behaviour, to draw firm conclusions on what humans are actually experiencing based only on manipulations of the computational model, without testing the model's implicit assumptions and without considering the actual neurophysiological and biomechanical mechanisms, can be misleading. One way to do this could be to examine these questions through extensions to the model used in the simulations (changing activation dynamics of the torque generators, allowing for potential motion backward motion of the shoulder and trunk, etc.). A better solution would be to emulate the physiological and biomechanical conditions on Earth (supporting the arm against gravity to reduce muscle tone, placing the subject on a moveable base that requires that the body be stabilized with the other hand) in order to distinguish the hypothesis of an underestimation of mass vs. other potential sources of under-actuation and other potential effects of weightlessness on the body.

      In sum, my opinion is that the authors are relying too much on a theoretical model as a ground truth and thus overstate their conclusions. But to provide a convincing argument that humans truly underestimate mass in weightlessness, they should consider more judiciously the neurophysiology and biomechanics that fall outside the purview of the simplified model that they have chosen. If a more thorough assessment of this nature is not possible, then I would argue that a more measured conclusion of the paper should be 1) that the authors observed modifications to movement kinematics in weightlessness consistent with an under-actuation for the intended motion, 2) that a simplified model of human physiology and biomechanics that incorporates principles of optimal control suggest that the source of this under-actuation might be an underestimation of mass in the computation of an appropriate feedforward motor command, and 3) that other potential neurophysiological or biomechanical effects cannot be excluded due to limitations of the computational model.

      We appreciate the reviewer's thoughtful assessment. We fully agree that a simplified computational model should not be treated as ground truth, and that the neurophysiology and biomechanics beyond the computational model must be carefully considered.

      As detailed in our responses above, we have substantially revised the Discussion to address each of these concerns—including new discussions of neuromuscular factors, more balanced treatment of trunk stability and friction, conditional framing of the mass underestimation interpretation, and a new limitation on model simplifications. The conclusion has been restructured following the reviewer's recommended framework.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      If possible and allowed, the authors are strongly encouraged to consider sharing this unique dataset. Making the data publicly available alongside the paper could foster future studies and further accelerate research in this area.

      We sincerely thank the reviewer for this suggestion. The ground control data and all analysis code will be made publicly available alongside the Version of Record.

      However, unfortunately, the raw in-flight data from the taikonaut cohort cannot be made publicly available due to confidentiality regulations of China's manned space program; access for scientific research requires approval from the China Astronaut Research and Training Center and can be requested through the corresponding author.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The objective of this study was to infer the population dynamics (rates of differentiation, division and loss) and lineage relationships of NK cell subsets during an acute immune response and under homeostatic conditions.

      Strengths:

      A rich dataset and a detailed analysis of a particular class of stochastic models.

      Weaknesses: (relating to initial submission)

      The stochastic models used are quite simple; each population is considered homogeneous with first-order rates of division, death, and differentiation. In Markov process models such as these there is no dependence of cellular behavior on its history of divisions. In recent years models of clonal expansion and diversification, in the settings of T and B cells, have progressed beyond this picture. So I was a little surprised that there was no mention of the literature exploring the role of replicative history in differentiation (e.g. Bresser Nat Imm 2022), nor of the notion of family 'division destinies' (either in division number, or the time spent proliferating, as described by the Cyton and Cyton2 models developed by Hodgkin and collaborators; e.g. Heinzel Nat Imm 2017). The emerging view is that variability in clone (family) size arises may arise predominantly from the signals delivered at activation, which dictate each precursor's subsequent degree of expansion, rather than from the fluctuations deriving from division and death modeled as Poisson processes.

      As you pointed out, the Gerlach and Buchholz Science papers showed evidence for highly skewed distributions of family sizes, and correlations between family size and phenotypic composition. Is it possible that your observed correlations could arise if the propensity for immature CD27+ cells to differentiate into mature CD27- cells increases with division number? The relative frequency of the two populations would then also be impacted by differences in the division rates of each subset - one would need to explore this. But depending on the dependence of the differentiation rate on division number, there may be parameter regimes (and timepoints) at which the more differentiated cells can predominate within large clones even if they divide more slowly than their immature precursors. One might not then be able to rule out the two-state model. I would like to see a discussion or rebuttal of these issues.

      Comments on revisions:

      (1) The authors have put in a lot of effort to address the reviews and have explored alternative models carefully.

      We appreciate the reviewers’ comments.

      (2) In the sections relating to homeostasis and the endogenous response, as far as I can tell you are estimating net growth rates (the k parameters) throughout - this is to be expected if you're working with just cell numbers and no information relating to proliferation. In these sections there are many places where you refer to proliferation rates and death rates when I think you just mean net positive or net negative growth rates. It's important to be precise about this even if the language can get a bit repetitive. (These net rates of growth or loss relate to clonal rather than cellular dynamics, which may be worth explaining). Later, you do use data relating to dead cells, which in principle can be used to get independent measures of death rates, but these data were not used in the fitting.

      We have modified the main text to address the comment.

      (3) There is so much evidence that T and B cell differentiation are often contingent on division that it would be very reasonable to consider it as a possibility for NK cells too. (Differentiation could be asymmetric, as you explored, or simply symmetric with some probability per division). These processes can be cast into simple ODE models but no longer allow you to aggregate division and death rates - so for parameter estimation you need to add measures of proliferation (Ki67 or similar) or death. This may be worth some discussion?

      We have modified the main text (lines 242-245) to address the comment.

      Reviewer #2 (Public review):

      Summary:

      Wethington et al. investigated the mechanistic principles underlying antigen-specific proliferation and memory formation in mouse natural killer (NK) cells following exposure to mouse cytomegalovirus (MCMV), a phenomenon predominantly associated with CD8+ T cells. Using a stochastic modeling approach, the authors aimed to develop a quantitative model of NK cell clonal dynamics during MCMV infection. Starting from a single immature Ly49+CD27+ NK cell, a two-state linear model (with a death variant) explained the negative correlation between clone size at 8 dpi and the CD27+ fraction, but failed to reproduce the first and second moments of CD27+ and CD27− NK cell populations at 8 dpi. To address this limitation, the authors added an intermediate maturation state, yielding a three-stage model (CD27+Ly6C− → CD27−Ly6C− → CD27−Ly6C+) that fits the first and second moments under two constraints: CD27+ NK cells proliferate faster than CD27− NK cells, and clone size is negatively correlated with the CD27+ fraction (upper bound of −0.2). The model predicts high proliferation in the intermediate state and high death in mature CD27−Ly6C+ cells, and it was validated using Adams et al. (2021) NK reporter mice tracking CD27+/− populations after tamoxifen, allowing discrimination between bone marrow-derived and pre-existing peripheral NK cells. To test the prediction that mature CD27− NK cells have a higher death rate, the authors measured Ly49H+ NK cell viability in the mouse spleen at different time points post-MCMV infection. Data confirmed lower viability of mature (CD27−) than immature (CD27+) cells during days 4-8 post-infection, and a model variant supported that higher CD27− death increases their proportion in the dead cell compartment. Altogether, the authors propose a three-stage quantitative model of antigen-specific expansion and maturation of naïve Ly49H+ NK cells with the trajectory CD27+Ly6C− (immature) → CD27−Ly6C− (mature I) → CD27−Ly6C+ (mature II), highlighting high proliferation in the mature I state and increased death in the mature II state.

      Strengths:

      Models explaining correlations and first and second moments, supported by analytical investigations, stochastic simulations, and model selection, identify key processes in antigen-specific NK expansion and maturation. The work distinguishes expansion, contraction, and memory in NK cells from CD8+ T cells and informs NK therapy development.

      Weaknesses (relating to initial submission):

      The conclusions of this paper are largely supported by the available data. However, a comparative analysis with more recent works in the field would be desirable. Clarifications:

      (1) Initial Conditions and Grassmann Data: The Grassmann data is used solely as a constraint, while the simulated values of CD27+/CD27− cells could have been directly fitted to the Grassmann data, which assumes a 1:1 ratio of CD27+/CD27− at t = 0. This would allow an alternative initial condition rather than starting from a single CD27+ cell.

      (2) Correlation Coefficients in the Three-State Model: Although the parameter scan of the three-stage model (Figure 2) demonstrates the potential for negative correlations between colony size and the fraction of CD27+ cells, the calculated correlation coefficients using the fitted parameter values are not shown. Including these would validate that the fitted parameters lie in the negative-correlation regime.

      (3) Viability Dynamics and Adaptive Response: The authors measured the time evolution of CD27+/− dynamics and viability over 30 days post-infection (Figure 4). It would be valuable to test whether the three-state model can reproduce the adaptive response of CD27− cells to MCMV infection, particularly the observed drop in CD27− viability at 5 dpi and its rebound at 8 dpi. Demonstrating this would test whether the model can simultaneously explain viability dynamics and moment dynamics, and would enable sensitivity analysis of CD27− viability with respect to model parameters.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor points:

      (1) line 175 - Here I think you have only ruled out the two state model with no death, and not the two state model in general?

      Edited the sentence to address the comment.

      (2) Figures 2 and 5 - the phenotypes (CD27+ Ly6C-, etc.) should be clearly labeled above each cell type. Fig 1 could be improved in the same way.

      Done.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Kashiwagi et al. undertook a population analysis of dendritic spine nanostructure applied to the objective grouping of 8 mouse models of neuropsychiatric disorders. They report that spine morphology in cultured hippocampal neurons shows a higher similarity among schizophrenia mouse models (compared with autism spectrum disorder (ASD) mouse models), and identify an effect of Ecrg4 (encoding small secretory peptides) on spine dynamics and shape in these models.

      Strengths:

      The study developed a method for objectively comparing spine properties in primary hippocampal neuron cultures from 8 mouse models of psychiatric disorders at the population level using high-resolution structured illumination microscopy (SIM) imaging. This novel technique identified two distinct groups of mouse models according to the population-level spine properties: those with ASD-related gene mutations and those with schizophreniarelated gene mutations. Functional studies, including gene knockdown and overexpression experiments, identified an effect of Ecrg4 on the spine phenotype of the schizophrenia model mice.

      We thank the reviewer for finding our strategy novel and useful for identifying molecules associated with the spine phenotype in schizophrenia-related mouse models.

      Weaknesses:

      The main weakness is that the study is wholly in vitro, using cultured hippocampal neurons. The authors present this as an advantage, however, arguing that spine morphology as measured in a reduced culture system can demonstrate direct effects of gene mutations on neuronal phenotypes in the absence of indirect influences from non-neuronal cells or specific environments.

      We appreciate this reviewer's concern about the limitation of cultured hippocampal neurons in extracting disease-related spine phenotypes. While we fully recognize this limitation, we consider that this in vitro system has several advantages that contribute to translational research on mental disorders.

      First, our culture system has been shown to support the development of spine morphology similar to that of the hippocampal CA1 excitatory synapse in vivo. High-resolution imaging techniques confirmed that the in vitro spine structure was highly preserved compared with in vivo preparations (Kashiwagi et al., Nature Communications, 2019). The present study used the same culture system and SIM imaging. Therefore, the difference we detected in samples derived from disease models is likely to reflect impairment of molecular mechanisms underlying native structural development in vivo.

      Second, super-resolution imaging of thousands of spines in tissue preparations under precisely controlled conditions cannot be practically applied using currently available techniques. The advantage of our imaging and analytical pipeline is its reproducibility, which enabled us to compare the spine population data from eight different mouse models without normalization.

      Third, a reduced culture system can demonstrate the direct effects of gene mutations on synapse phenotypes, independent of environmental influences. This property is highly advantageous for screening chemical compounds that rescue spine phenotypes. Neuronal firing patterns and receptor functions can also be easily controlled in a culture system. The difference in spine structure between ASD- and schizophrenia-related mouse models is valuable information to establish a drug screening system.

      Fourth, establishing an in vitro system for evaluating synapse phenotypes could reduce the need for animal experiments. Researchers should be aware of the 3Rs principles. In the future, combined with differentiation techniques for human iPS cells, our in vitro approach will enable the evaluation of disease-related spine phenotypes without the need for animal experiments. The effort to establish a reliable culture system should not be eliminated.

      We modified our text to have a balanced discussion on both advantages and disadvantages of the in vitro culture system in the study of mental disorder mouse models, as follows:

      "Finally, while the spine phenotype identified in the human postmortem brain undoubtedly resulted from complex interactions among genetic background, environmental influences, and regulation by non-neuronal cells, data from pure neuronal cultures are more likely to reflect the direct effects of schizophrenia-related gene mutations on synaptic functions. This property may be advantageous for identifying synaptic molecules that regulate synapse phenotypes in schizophrenia-related mouse models. However, the phenotype observed in the culture system requires confirmation using in vivo experiments of mouse models or human tissue samples. Efficient in vitro screening combined with reliable in vivo evaluation of synapses will facilitate translational research on mental disorders."

      Another weakness is that CaMKIIαK42R/K42R mutant mice are presented as a schizophrenia model, the authors justifying this by saying that "CaMKII-related signaling pathway disruption has been implicated in the working memory deficits found in schizophrenia patients". Since mutations in CAMK2A cause autosomal dominant intellectual developmental disorder-53 (OMIM 617798) and autosomal recessive intellectual developmental disorder-63 (OMIM 618095), and mice carrying the CAMK2A E183V mutation exhibit ASD-related synaptic and behavioral phenotypes (PMID: 28130356), I think it's stretching credibility to refer to the CaMKIIαK42R/K42R mice as a schizophrenia model.

      We agree with this reviewer that CAMK2A mutations in humans are linked to multiple mental disorders, including developmental disorders, ASD, and schizophrenia. Association of gene mutations with the categories of mental disorders is not straightforward, as the symptoms of these disorders also overlap with each other. For the CaMKIIα K42R/K42R mutant, we considered the following points in its characterization as a model of mental disorder. Analysis of CaMKIIα +/- mice in Dr. Tsuyoshi Miyakawa's lab has provided evidence for the reduced CaMKIIα in schizophrenia-related phenotypes (Yamasaki et al., Mol Brain 2008; Frankland et al., Mol Brain Editorial 2008). It is also known that the CaMKIIα R8H mutation in the kinase domain is linked to schizophrenia (Brown et al., 2021). Both CaMKIIα R8H and CaMKIIα K42R mutations are located in the N-terminal domain and eliminate kinase activity. On the other hand, the representative CaMKIIα E183V mutation identified in ASD patients exhibits unique characteristics, including reduced kinase activity, decreased protein stability and expression levels, and disrupted interactions with ASD-associated proteins such as Shank3 (Stephenson et al., 2017). Importantly, reduced dendritic spines in neurons expressing CaMKIIα E183V is a property opposite to that of the CaMKIIα K42R/K42R mutant, which showed increased spine density (Koeberle et al. 2017).

      References related to this discussion.

      (1) Yamasaki et al., Mol Brain. 2008 DOI: 10.1186/1756-6606-1-6

      (2) Frankland et al. Mol Brain. 2008 DOI: 10.1186/1756-6606-1-5

      (3) Stephenson et al., J Neurosci. 2017 DOI: 10.1523/JNEUROSCI.2068-16.2017

      (4) Koeberle et al. Sci Rep. 2017 DOI: 10.1038/s41598-017-13728-y

      (5) Brown et al., iScience. 2021 DOI: 10.1016/j.isci.2021.103184

      We fully agree with the reviewer that different CAMK2A mutations likely cause distinct phenotypes observed in the broad spectrum of mental disorders. In the revised manuscript, we include a discussion of the relevant literature to categorize this mouse model appropriately.

      "CaMKII-related signaling pathway disruption has been implicated in the working memory deficits found in schizophrenia patients [45,46]. CAMK2A mutations in humans are linked to multiple mental disorders, including developmental disorders, ASD, and schizophrenia [47]. The K42R mutation of CAMK2A does not correspond to any known human genetic variant, but the CAMK2A R8H mutation is linked to schizophrenia [48]. Both R8H and K42R mutations in the N-terminal domain of CaMKIIα eliminate kinase activity; these mutations may have a similar impact on human mental disorders."

      Although the manuscript is largely well written, there are some instances of ambiguous/unspecific language. This extends to the title (Decoding Spine Nanostructure in Mental Disorders Reveals a Schizophrenia-1 Linked Role for Ecrg4), which gives no indication that the work was in vitro on cultured neurons derived from mouse models.

      We appreciate the reviewer for pointing out the lack of information about the experimental system in the title of this manuscript. According to the suggestion of the reviewer, we modified the title as "Decoding spine nanostructure in cultured neurons derived from mouse models of mental disorder reveals a schizophrenia-linked role for Ecrg4".

      Reviewer #2 (Public review):

      Okabe and colleagues build on a super-resolution-based technique that they have previously developed in cultured hippocampal neurons, improving the pipeline and using it to analyze spine nanostructure differences across 8 different mouse lines with mutations in autism or schizophrenia (Sz) risk genes/pathways. It is a worthy goal to try to use multiple models to examine potential convergent (or not) phenotypes, and the authors have made a good selection of models. They identify some key differences between the autism versus the Sz risk gene models, primarily that dendritic spines are smaller in Sz models and (mostly) larger in autism risk gene models. They then focus on three models (2 Sz - 22q11.2 deletion, Setd1a; 1 ASD - Nlgn3) for time-lapse imaging of spine dynamics, and together with computational modelling provide a mechanistic rationale for the smaller spines in Sz risk models. Bulk RNA sequencing of all 8 model cultures identifies several differentially expressed genes, which they go on to test in cultures, finding that ecgr4 is upregulated in several Sz models and its misexpression recapitulates spine dynamics changes seen in the Sz mutants, while knockdown rescues spine dynamics changes in the Sz mutants. Overall, these have the potential to be very interesting findings and useful for the field. However, I do have a number of major concerns.

      We thank the reviewer for evaluating our findings as potentially very interesting and useful.

      (1) The main finding of spine nanostructure changes is done by carrying out a PCA on various structural parameters, creating spine density plots across PC1 and PC2, and then subtracting the WT density plot from the mutant. Then, spines in the areas with obvious differences only are analyzed, from which they derive the finding that, for example, spine sizes are smaller. However, this seems a circular approach. It is like first identifying where there might be a difference in the data, then only analyzing that part of the data. I welcome input from a statistician, but to me, this is at best unconventional and potentially misleading. I assume the overall means are not different (although this should be included), but could they look at the distribution of sizes and see if these are shifted?

      We appreciate the reviewer's concern regarding our analysis of spine population data. The intention of pre-selecting the areas showing differences between wild-type and mutant was to make a direct comparison between two subareas (one is enriched with wild-type spines and the other is enriched with mutant spines) and clarify that the spines of schizophreniarelated mouse models were smaller than wild-type spines. Conventional methods of comparing the total spine population using simple size parameters are not useful for this purpose, as shown in Supplementary Figure 2.

      To clarify the reviewer's concern, we revised the analysis of the spine population data for both Figure 3 and Figure 8.

      Figure 3: We first divided the feature space projected onto PC1 and PC2 into four areas with distinct structural properties: (1) small and short, (2) small and long, (3) large and short, and (4) large and long. Next, we calculated the normalized spine counts in the four areas for both wild-type and mutant spines and obtained the relative ratio (mutant/wild-type) for each area. As we performed three independent SIM imaging experiments (in one, we imaged both wild type and mutant culture dishes prepared from the same pregnant mouse), there are three independent datasets from 8 mouse models.

      We found that the spine ratio (mutant/wild-type) only in area 2 (small and long spines) differed significantly between genotypes. This result is shown in Fig. 3 and explained in the text. The spine ratios in areas 1 and 3 did not show a clear relationship to the genotypes, while the ratio in area 4 showed the opposite trend to that in area 2. The opposite trend between areas 2 and 4 indicates enrichment of both small and long spines in schizophrenia-related mouse models, consistent with our previous analysis.

      Figure 8: In this analysis, we aimed to evaluate the rescue effect of Ecrg4 shRNA relative to that of control shRNA. If Ecrg4 shRNA is effective, the spine population enriched in the control shRNA condition should be reduced in the Ecrg4 shRNA condition. To confirm this point in the revised manuscript, we first defined areas in the projected PC1-PC2 plane showing either enrichment or depletion of spines in the control shRNA condition (spine numbers increasing or decreasing by more than 3 × SD). We next measured the difference in spine numbers between the control and Ecrg4 shRNA conditions in either enriched or depleted areas. The expectation is that Ecrg4 shRNA treatment reduces the extent of both enrichment and depletion. The effect was significant in both the 22qdel and Setd1a mouse models, as indicated by permutation tests. This analysis was explained in the revised manuscript.

      (2) Despite extracting 64 parameters describing spine structure, only 5 of these seemed to be used for the PCA. It should be possible to use all parameters and show the same results. More information on PC1 and PC2 would be helpful, given that the rest of the paper is based on these - what features are they related to?

      We thank the reviewer for the advice on providing the rationale for parameter selection in PCA. We divided spines into 160-nm segments along their long axis, and the spine segments were used to calculate the 64 parameters, which include volume of each spine segment (20 segments), convex hull volume of each spine segment (20 segments), and convex hull ratio of each spine segment (20 segments). As most spines are shorter than 0.16 × 20 =3.2 μm, these segment-related parameters contain a large fraction of zero values, which affect the proper calculation of principal components. Therefore, we selected two parameters that reflect the principal structural features (length and volume), together with three other parameters that were mutually independent and also independent from the first two parameters (pairwise correlation coefficients < 0.3). These selection criteria were described in the original manuscript. We also confirmed that PCA using all 64 parameters yields a cross correlation map similar to that shown in Fig. 2B.

      Author response image 1.

      We provided additional information in the Materials and Methods section of the revised manuscript.

      As described previously, the pattern of four areas with distinct spine structures (1. small and short, 2. small and long, 3. large and short, 4. large and long) supports the idea that the PC1PC2 plane reflects the relationship between spine volume and length (Fig. 3A and B).

      These specific features could then be analyzed in the full dataset, without doing the cherry picking above.

      We provided the dataset for the relative enrichment of spine counts across four areas of the PC1-PC2 plane in Fig. 3A and B. This analysis provides a comprehensive view of spine population properties related to spine volume and length, without relying on a pre-set region of interest.

      It would also be helpful to demonstrate whether PC1 and 2 differ across groups - for example, the authors could break their WT data into 2 subsets and repeat the analysis.

      We noticed differences in the pattern of spine distribution across the PC1-PC2 planes in each experiment. The subtraction of the distributional data between wild-type and mutant samples effectively cancels out such differences. In general, the difference between two wild-type samples is smaller than that between wild-type and mutant samples, as shown in Author response image 2.

      Author response image 2.

      We added a description of variation across groups to the revised manuscript.

      (3) Throughout the paper, the 'n' used for statistical analysis is often spine, which is not appropriate. At a minimum, cell should be used, but ideally a nested mixed model, which would take into account factors like cell, culture, and animal, would be preferable. Also, all of these factors should be listed, with sufficient independent cultures.

      We agree that nested mixed models are more appropriate for evaluating genotype effects in most of our datasets. We confirm that the results of statistical analysis using nested mixed models were consistent with our previous conclusions in most cases.

      Figure 3: We performed three independent primary cultures of embryonic hippocampal tissue with genotypes of both wild-type and mutant from the same pregnant mice for each mouse model. In our new Figure 3, each data point represents an independent culture experiment, and group comparisons were performed using one-way ANOVA followed by Tukey's post hoc test. In this analysis, statistical analysis using neurons as units of 'n' is not possible, as the number of spines measured from a single neuron is insufficient to generate the density map shown in Figure 3. The statistical analysis was described in the revised text. The details of experimental conditions related to Figure 3 are provided in Supplementary Table 1.

      Figure 5A-C: We analyzed spine turnover rate using a linear mixed-effects model with genotype as a fixed effect and plate, cell, and dendrite as nested random effects. In both 22q deletion model and Setd1a model, there were significant effects of genotype (F(1,25) = 5.79, p = 0.024 for 22q deletion model and F(1,22) = 7.33, p = 0.013 for Setd1a model). In contrast, Nlgn3 mutant neurons did not show a significant difference (F(1,14) = 1.35, p = 0.26). This analysis was described in the revised text.

      Figure 5D-F: Spine lifetime was analyzed using a linear mixed-effects model accounting for the hierarchical structure of the data (spines nested within dendrites, cells, and culture plates). The analysis revealed a significant effect of genotype in both 22q deletion mutant and Setd1a mutant (22qdel mutant; F(1,336) =5.33, p=0.022, Setd1a mutant; F(1,282)=6.38, p=0.012 ). The neurons of both mutants exhibited significantly longer spine lifetimes compared with wild-type neurons (22qdel mutant; ratio = 1.28, 95% CI 1.04–1.58, Setd1a mutant; ratio = 1.35, 95% CI 1.07–1.70). In contrast, Nlg3 mutation did not significantly alter spine lifetime (ratio = 0.86, 95% CI 0.61–1.22; F(1,220)=0.69, p=0.41). This analysis was described in the revised text.

      Figure 5G-I: Spine volume trajectories were analyzed using linear mixed-effects models incorporating nested random effects (spine/dendrite/cell/culture plate) to account for the hierarchical structure of the data. In the 22q deletion model, newly formed spines were significantly smaller than those in wild-type neurons (genotype effect: p < 0.001). The spines in Setd1a mutant neurons also displayed significantly smaller volume than those in wild-type neurons (p < 10<sup>-7</sup>). There were also differences in the temporal profiles of spine growth in these two mutants (p < 0.001). In contrast, newly formed spines in the Nlgn3 mutant neurons were significantly larger than those in wild-type neurons (p < 10<sup>-4</sup>) with preserved time-course of spine growth. This analysis was described in the revised text.

      Figure 5J-L: Similar analyses using linear mixed-effects models incorporating nested random effects (spine within dendrite within cell within culture plate) identified significantly smaller initial spine size in the 22q deletion model (p < 10<sup>⁻6</sup>), while no significant differences in the initial spine volume were found for Setd1a mutants. The temporal trajectories of spine shrinkage before their loss were also not significantly altered in both 22qdel and Setd1a mutants. The Nlg3 mutant showed a significantly different time-course of spine shrinkage (p < 0.05), while the initial spine size was not altered. This analysis was described in the revised text.

      Figure 7A overexpression dataset: We analyzed plate-averaged lifetime values using a linear mixed-effects model with treatment as a fixed effect. There exists a significant main effect of treatment (F(3,8) = 4.59, p = 0.038), with post hoc examination showing a significant increase in lifetime by Ecrg4 overexpression (β = 0.49 ± 0.16 SE, t(8) = 3.16, p = 0.013). Figure 7A shRNA dataset: We also applied a linear mixed-effects model for plate-averaged lifetime values with treatment as a fixed effect. The analysis revealed no significant effect of treatment (F(2,6) = 0.29, p = 0.76).

      The analyses of overexpression and shRNA datasets were described in the revised text.

      Figure 8: As in Figure 3, we performed three independent primary cultures of embryonic hippocampal tissue with genotypes of both wild-type and mutant from the same pregnant mice for each mouse model. The culture plates were transfected with either a control shRNA or an Ecrg4 shRNA construct. Each data point represents an independent culture experiment, and the effect of Ecrg4 shRNA relative to that of control shRNA was evaluated using a permutation test. The data analysis was described in the revised text. The details of experimental conditions related to Figure 8 are provided in Supplementary Table 1.

      (4) The authors should confirm that all mutants are also on the C57BL/6J background, and clarify whether control cultures are from littermates (this would be important). Also, are control versus mutant cultures done simultaneously? There can be significant batch effects with cultures.

      The mutant mice we used in this study are on C57BL/6J or C57BL/6N background. It is known that C57BL/6J or C57BL/6N mice exhibit distinct phenotypes across a range of physiological, biochemical, and behavioral systems. However, it is less likely that our analysis is affected by differences between C57BL/6J and C57BL/6N, as we compared wild-type and mutant littermates on the same genetic background. This experimental design can also reduce the batch effects with different culture preparations. This point was described in the revised text.

      (5) The spine analysis uses cultures from 18-22 DIV - this is quite a large range. It would be worth checking whether age is a confounder or correlated with any parameters / principal components.

      We described in the method sections that culture samples were processed for imaging at 18-22 DIV. However, all the SIM imaging experiments for eight mutant mouse models were performed on samples fixed at DIV 19. The wide range of imaging experiments (DIV 18-22) includes test samples we used to optimize imaging conditions. In the revised manuscript, we specified the timing of SIM imaging.

      (6) The computational modelling is interesting, but again, I am concerned about some circularity. Parameter optimization was used to identify the best fit model that replicated the spine turnover rates, so it is somewhat circular to say that this matched the observations when one of these is the turnover rate.

      We appreciate the reviewer's comment on some circularity of the argument. We agree that the turnover rate is already incorporated into the simulation model and is not an appropriate criterion for the evaluation. We modified the text accordingly.

      It is more convincing for spine density and size, but why not go back and test whether parameter differences are actually seen - for example, it would be possible to extract the probability of nascent spine loss, etc.

      We thank the reviewer for giving this important suggestion. The probability of nascent spine loss is an important parameter, and we initially attempted to estimate it from the original data set. However, the upper limit of our time-lapse imaging is 24 h, which is insufficient to distinguish stable and nascent spines clearly. The difficulty of extracting all the necessary parameters for spine remodeling is our motivation for starting this computational modelling.

      More compelling would be to repeat the experiments and see if the model still fits the data. In the interpretation (line 314-318) it is stated that '... reduced spine maturation rate can account for the three key properties of schizophrenia-related spines...', which is interesting if true, but it has just been stated that the probability of spine destabilization is also higher in mutants (line 303) - the authors should test whether if the latter is set to be the same as controls whether all the findings are replicated.

      As suggested by the reviewer, we set the probability of spine destabilization equal across wild-type and mutant models and repeated the simulations. The results indicate that this modification has small effects on spine density (0.61 vs 0.62), spine turnover rate (0.22 vs 0.21), fraction of small spines (0.21 vs 0.20), and mean spine size (0.37 vs 0.36). We described this point in the revised manuscript.

      (7) No validation for overexpression or knockdown is shown, although it is mentioned in the methods - please include.

      As suggested by the reviewer, we validated overexpression and knockdown. The results are summarized in Supplementary Figure 8.

      Supplementary Figure 8A-C shows the immunocytochemistry of anti-Ecrg4, anti-Cip4, and anti-NPAS4 for the confirmation of overexpression of these molecules.

      Supplementary Figure 8D-E shows the confirmation of the appropriate size of exogenously expressed Ecrg4, Cip4, and NPAS4 by immunoblotting. (previous Supplementary Figure 10F is now Supplementary Figure 8E).

      Supplementary Figure 8F-H indicates the efficient knockdown of exogenously expressed Met-GFP, ARHGAP15-GFP, and Ecrg4-HA by respective shRNA constructs in COS-7 cells. (previous Supplementary Figure 10G is now Supplementary Figure 8H)

      Also, for the knockdown, a scrambled shRNA control would be preferable.

      We used Stealth RNAi Negative Control Duplexes (Invitrogen) as the shRNA control in this study. To confirm that this RNAi sequence does not affect spine turnover, we performed timelapse imaging of neurons transfected with GFP alone or with GFP and the Stealth RNAi Negative Control. No detectable change in spine turnover was observed (Supplementary Figure 8I), indicating that this RNAi control sequence is suitable for our study.

      (8) The finding regarding ecgr4 is interesting, but showing that some ecgr4 is expressed at boutons and spines and some in DCVs is not enough evidence to suggest that actively involved in the regulation of synapse formation and maturation (line 356).

      To reveal the active roles of Ecrg4 in spine regulation, we exogenously applied a synthetic Ecrg4 peptide to wild-type neurons and monitored both spine density and turnover rate after Ecrg4 application. The Ecrg4 application increased the spine turnover rate, whereas samples treated with the scrambled peptide did not. This result supports the active role of Ecrg4 in regulating spine turnover. The data were added as Supplementary Figures 9F and G.

      (9) The same caveats that apply to the analysis also apply to the ecgr4 rescue. In addition, while for 22q the control shRNA mutant vs WT looks vaguely like Figure 2, setd1a looks completely different.

      We thank the reviewer for pointing out the apparent difference in the pattern of spine population data between Figure 2 and Figure 8. We performed SIM analysis using DiI-labeled neurons in Figure 2, whereas the data in Figure 8 are derived from GFP-expressing neurons. The images of cell-surface labeling and cytoplasmic labeling cannot be analyzed in the same way, as it is necessary to adjust parameters in SIM image processing and PCA-based dimensional reduction. Consequently, the distribution of the spine population projected onto the PC1-PC2 plane differs between DiI-labeled neurons and GFP-expressing neurons. To facilitate the comparison of PCA analysis applied to GFP-expressing neurons, we replaced the weight matrix for GFP-expressing neurons with that previously calculated for the DiIlabeled neurons. This adjustment increased the similarity of the data distributions shown in Figures 2 and 8. The explanation for the different patterns in the spine population map between Figure 2 and Figure 8 was added to the revised text. The related explanation for the data processing was described in the Materials and Methods.

      And if rescued, surely shRNA in the mutant should now resemble control in WT, so there shouldn't be big differences, but in fact, there are just as many differences as comparing mutant vs wild-type? Plus, for spine features, they only compare mutant rescue with mutant control, but this is not ideal - something more like a 2-way ANOVA is really needed. Maybe input from a statistician might be useful here?

      We appreciate the reviewer's important comment and agree that the analytical approach used in the original manuscript was not optimal. We therefore revised our analysis to examine whether the difference observed between wild-type and mutant neurons was reduced by suppression of Ecrg4 expression.

      To this end, we first identified two regions in the PC1–PC2 plane where mutant spines were either enriched or depleted relative to wild-type neurons (Areas A and B). We then counted the number of spines located in Areas A and B in control shRNA-treated mutant neurons (normalized spine counts XA and XB). Next, we quantified spine counts in the same areas using data from Ecrg4-suppressed mutant neurons (normalized spine counts YA and YB). If XA > YA and XB < YB, suppression of Ecrg4 would indicate a shift toward rescue of the phenotype observed in control shRNA-treated mutant neurons. Indeed, the datasets were consistent with this shift in relative spine counts.

      To determine whether these differences exceeded those expected from random variation in spine counts, we performed a permutation test. Specifically, spine identities were randomly shuffled between the two conditions while preserving the total number of spines in each dataset. The observed differences were then compared with the distribution obtained from the permuted datasets to assess statistical significance.

      We found that all three culture replicates showed statistical significance in both areas A and B for both the 22qdel and Setd1a mutations. This analysis is described in the Result section.

      (10) Although this is a study entirely focused on spine changes in mouse models for Sz, there is no discussion (or citation) of the various studies that have examined this in the literature. For example, for Setd1a, smaller spines or reduced spine densities have been described in various papers (Mukai et al, Neuron 2019; Chen et al, Sci Adv 2022; Nagahama et al, Cell Rep 2020).

      We appreciate the reviewer's suggestion to include a discussion of schizophrenia-related mouse models. We added more information related to the Setd1a mouse model to the Discussion section.

      "Population-level spine properties were more homogeneous in schizophrenia models (those with gene mutations implicated in schizophrenia) than in the other 4 models studied, in part due to a shared tendency for smaller spines. This observation is consistent with previous studies on Setd1a mutant mice, which showed reduced spine width, decreased mushroomtype spines, and lower spine density in the prefrontal cortex [43,56,57]. In contrast to these findings, several previous studies reported reduced numbers of small spines in the postmortem cortical tissues of schizophrenia patients [22,58]. "

      (11) There is a conceptual problem with the models if being used to differentiate autism risk from Sz risk genes. It is difficult to find good mouse models for Sz, so the choice of 22q11.2del and Setd1a haploinsufficiency is completely reasonable. However, these are both syndromic. 22qdel syndrome involves multiple issues, including hearing loss, delayed development, and learning disabilities, and is associated with autism (20% have autism, as compared to 25% with Sz). Similarly, Setd1a is also strongly associated with autism as well as Sz (and also involves global developmental delay and intellectual disability). While I think this is still the best we can do, and it is reasonable to say that these models show biased risk for these developmental disorders, it definitely can't be used as an explanation for the higher variability seen in the autism risk models.

      We appreciate the reviewer's suggestion for more careful consideration of the interpretation of phenotypes in mouse models, with regard to their relation to clinical phenotypes in human patients. According to the suggestion of the reviewer, we modified the relevant text as follows:

      "The nanoscale features of dendritic spines in ASD-associated mouse models were more variable than those in schizophrenia-associated mouse models. This difference may be related to the broader clinical spectrum of ASD, which ranges from mild impairments in social skills to severe intellectual disability. The four ASD-associated mouse models examined in this study, Nlgn3<sup>R451C/(y or R451C) , Syngap1<sup>+/-</sup>, POGZ<sup>Q1038R/+</sup>, and 15q11-13<sup>dup/+</sup>, may represent subgroups with different levels of hippocampal dysfunction. Among the four ASD-associated mouse models, 15q11-13<sup>dup/+</sup> showed population-level spine properties closer to those of the schizophrenia models. To understand this similarity, further analysis of neural circuit changes in both ASD- and schizophrenia-associated mouse models will be necessary. Analysis of the relationships between rare genetic variants and synapse phenotypes in mouse models may contribute to their eventual categorization. This information should be useful to understand the underlying mechanisms of the broader clinical spectrum of ASD."

      (12) I am not convinced that using dissociated cultures is 'more likely to reflect the direct impact of schizophrenia-related gene mutations on synaptic properties' - first, cultures do have non-neuronal cells, although here glial proliferation was arrested at 2 days, glia will be present with the protocol used (or if not, this needs demonstrating).

      In our culture system, the density of non-neuronal cells is low, and most neurons are not in direct contact with non-neuronal cells. We reported this method in Nat. Neurosci. 1999, where we utilized this culture system to visualize GFP-tagged PSD-95 in neurons using recombinant adenovirus. Because recombinant adenovirus shows higher infection efficiency in glial cells, it was essential for us to establish a culture condition that isolates neurons from glial cells.

      Second, activity levels will affect spine size, and activity patterns are very abnormal in dissociated cultures, so it is very possible that spine changes may not translate into in vivo scenarios. Overall, it is a weakness that the dissociated culture system has been used, which is not to say that it is not useful, and from a technical and practical perspective, there are good justifications.

      We appreciate the reviewer's comment on the advantages and disadvantages of using an in vitro culture system. This comment aligns with the first reviewer's. We modified our text to have a balanced discussion on the role of the in vitro culture system in the study of mental disorder mouse models as follows:

      "Finally, while the spine phenotype identified in the human postmortem brain undoubtedly resulted from complex interactions among genetic background, environmental influences, and regulation by non-neuronal cells, data from pure neuronal cultures are more likely to reflect the direct effects of schizophrenia-related gene mutations on synaptic functions. This property may be advantageous for identifying synaptic molecules that regulate synapse phenotypes in schizophrenia-related mouse models. However, the phenotype observed in the culture system requires confirmation using in vivo experiments of mouse models or human tissue samples. Efficient in vitro screening combined with reliable in vivo evaluation of synapses will facilitate translational research on mental disorders."

      (13) As a minor comment, the spine time-lapse imaging is a strength of the paper. I wonder about the interpretation of Figure 5. For example, the results in Figure 5G and J look as if they may be more that the spines grow to a smaller size and start from a smaller size, rather than necessarily the rate of growth.

      We thank the reviewer for the insightful comment. In the revised manuscript, we analyze the time-lapse data using linear mixed-effects models incorporating nested random effects (spine/dendrite/cell/culture plate). This analysis suggested the difference in the initial size of spines. This point is described in the revised manuscript as follows:

      "Schizophrenia-associated mouse models showed higher similarity in spine morphology, driven by reduced size and growth of nascent spines."

      "We further compared the initial increase in spine volume between genotypes (Figure 5G-I). Linear mixed-effects models incorporating nested random effects revealed significantly smaller initial spine volumes in both 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> models (genotype effect: p < 0.001 for 22q11.2<sup>del/+</sup> and p < 10<sup>-7</sup> for Setd1a<sup>+/-</sup>). The spines in both mutants also displayed a significant reduction in spine volume increase (p < 0.001). In contrast, newly formed spines in the Nlgn3<sup>R451C/(y or R451C)</sup> neurons were significantly larger than those in wild-type neurons (p < 10<sup>-4</sup>) with preserved time-course of spine growth.”

      We tested whether the initial size difference in spines can be incorporated into the computational simulation. However, due to the large variability in the initial spine size, it was difficult to perform parameter optimization in the model with additional factors. Therefore, we did not further pursue this possibility in this revision. This point is described in the revised text.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The manuscript would be strengthened if the following issues were adequately addressed:

      (1) It would be helpful to know more about the in/ex vivo dendritic spine phenotype of the mouse models of neuropsychiatric disorders, to allow readers to judge whether and how the in vitro spine phenotype in hippocampal neuronal cultures overlaps with/replicates the spine phenotype within the mouse brain.

      We appreciate this comment, but our currently available data is insufficient to specify the difference between in vitro and in vivo spine phenotypes. Our previous study, published in Nature. Comm. (2019), provided data showing that the overall distribution of spine size is similar between in vivo and in vitro conditions in the mouse hippocampus.

      (2) Although the manuscript is largely well written, there are instances of ambiguous language, particularly when describing the spine phenotypes. For example, we are told that "ASD mouse models showed a tendency of decreasing spine subpopulation with small volumes." This description and other examples should be expressed more clearly.

      Following the reviewer's suggestions, we revised the text to improve clarity. We modified the sentence "ASD mouse models showed a tendency of decreasing spine subpopulation with small volumes" to "ASD-related mouse models showed an opposite spine phenotype."To avoid possible confusion for readers, we have revised several sentences in the text to clarify the intended meaning.

      Also, I question whether the word "decoding", meaning to convert (a coded message) into intelligible language, is the most appropriate for the title and abstract.

      The original meaning of the word "decoding" is the conversion of a coded message into an intelligible form; however, in this study, we use the term in a broader sense, referring to the extraction of latent population-level properties of dendritic spines from multidimensional structural parameters. We believe this usage is consistent with its common use in neuroscience and systems biology, where "decoding" often refers to inferring underlying biological states or information from complex datasets.

      (3) The authors should reconsider whether CaMKIIαK42R/K42R mice should be described as a schizophrenia model, when mutations in CAMK2A are known to cause autosomal dominant intellectual developmental disorder-53 (OMIM 617798) and autosomal recessive intellectual developmental disorder-63 (OMIM 618095), and mice carrying the CAMK2A E183V mutation exhibit ASD-related synaptic and behavioral phenotypes (PMID: 28130356).

      We provided a detailed answer to this question in the previous part of the rebuttal.

      (4) The title doesn't adequately summarise the contents of the manuscript. It should mention mice/mouse models and cultured neurons.

      We also responded to this request in the previous part of the rebuttal.

      Reviewer #2 (Recommendations for the authors):

      (1) Please provide a supplementary table with all DEGs. Also, DEGs are listed if present in 'more than 2' models - does this mean they had to be in 3 or more? Please clarify.

      According to the reviewer's suggestion, we added data on DEGs shared by >2 mouse models in Supplementary Figure 7. We also added Supplementary Tables 2 and 3 for all DEGs. The phrase "in more than 2 models" means "in 3 or 4 models".

      (2) There are several references to 'schizophrenia mouse models' - it is worth rephrasing this to make clear that these are not mice with schizophrenia.

      We replaced the expression "schizophrenia (or ASD) mouse models" with "schizophrenia (or ASD)-associated mouse models" or similar appropriate wording throughout the manuscript.

      (3) Line 66: 'a recent...' - 2014 is not really recent.

      We removed the word "recent" from the sentence.

      (4) Figure S1: The legend says A-D, but they are not on the figure. Also, make clear whether this data is only WT data - it seems to be from disorder models, with 4 colors for each model - please clarify.

      We changed the sentence from "shown as A to D" to "shown as A to C". The datasets in Supplementary Figure 1 are wild-type only. Each graph uses four colors to represent wildtype data from four imaging datasets obtained from different mouse models. Graphs A to C correspond to spine length, surface area, and volume, respectively.

      (5) Methods, line 680-4: More detail here would be helpful.

      We added more explanation for the generation of subtraction maps.

      (6) Line 193: Make it clear this is hippocampal in the main text.

      We added "cultures of embryonic hippocampi" to the text.

      (7) Figure 5, D-F: Make clear that these are transient spines (as per main text)

      We added "Lifetimes of transient spines" to both the main text and figure legend.

      (8) Figure 6B: More detail is needed; no idea what this is - no axis label. D - also not clear what numbers on the y-axis mean. E - color scale??

      We added details to the figure legend, the axis labels for Figures 6B and 6D, and the color scale for Figure 6E.

      (9) Supplementary Figure 9 - not clear what matrices are actually showing, nor what the scale refers to - is this the number of shared DEGs? If so, please make it clearer.

      The matrices show the shared DEG numbers, as shown in their titles. The scale indicates DEG numbers. We added the explanation of the color code to the figure legend.

      (10) Please make clear in the main text that ecgr4 affected the turnover rate. It would be good to measure other parameters as well.

      We added the phrase "a significant increase in spine turnover rate by Ecrg4 overexpression" to the main text.

      (11) Figure 7: Suggest to label C on images as well, so obvious which is GFP/anti-HA overlay (and respective colors) and which is anti-HA staining.

      We added the labels with respective colors to Figure 7.

      (12) Ecgr4 is a precursor protein that is cleaved to produce several hormone-like peptides. Where is the HA tag - so which cleavage products will it label? Any antibodies that work in immunocytochem?

      HA tag was attached to the C-terminal domain. We predict that anti-HA binds to four cleavage products (the full-length Ecrg4, Augurin, Argilin, and Δ16). Among several commercially available antibodies, only the SIGMA product could detect cells expressing Ecrg4-HA by immunocytochemistry.

      (13) Supplementary Figure 10: Synaptosome would be a good addition.

      We isolated the fraction of synaptosomes using Syn-PER™ Synaptic Protein Extraction Reagent in Supplementary Figure 9A. We added this explanation to the Materials and Methods section.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      Strengths of this paper include the important question addressed and the elegant and innovative combination of methods, which led to clear insights into the sensory biology of self-righting, and that will be useful for others in the field. This is a substantial contribution to understanding how animals correct their body position. The manuscript is very clearly written and couched in interesting biology.

      Limitations:

      (1.1) The interpretation of functional experiments is complicated by the proposed excitatory and inhibitory roles of dorsal and ventral sensory neuron activity, respectively. So, while silencing of an excitatory (dorsal) element might slow righting, silencing of inputs that inhibit righting could speed the behavior. Silencing them together, as is done here, could nullify or mask important D-V-specific roles. Selective manipulation of cells along the D-V axis could help address this caveat.

      We highly appreciate the thoughtful comments by Rev1 pointing out the relative simplicity of our current inferences regarding the role of dorsal vs. ventral substrate contact, and agree with the suggestion that cells along the DV axis could have diverse roles in their contribution to self-righting. In this context, we wish to point out two aspects, one theoretical and one practical. Regarding theory, our view is that this may not be a simple case of “excitation vs. inhibition”, but rather one in which the coordinated and dynamic activity of distributed sensory neurons promotes differential action selection in alignment with environmental conditions – a framework that could involve many different behaviours with a still uncertain level of granularity (e.g., is self-righting different if the larva is rotated to 160º instead of exactly 180º?). Regarding the practical aspect, while this area represents a fascinating point for future investigation, it is currently limited by technological development, particularly in the context of this study where a relatively low-cost implementation has been used to probe the AP axis. Investigation of the DV axis would require further technological development, since optogenetic light would need to be precisely delivered from the side rather than from underneath, with a greater degree of resolution compared to the AP axis given the much smaller width of the larva (~120-140µm) relative to its length (~550-600µm). Therefore, whilst we appreciate these comments and suggestion, we believe this line of experiments is ideal for a follow-up investigation, rather than being implemented in the current study.

      (1.2) Prior studies from the authors implicated daIV neurons in the righting response. One of the main advances of the current manuscript is the clever demonstration of region-specific roles of sensory input. However, this is only confirmed with a general md driver, 190(2)80, and not with the subsetspecific Gal4, so it is not clear if daIV sensory neurons are also acting in a regionally-specific manner along the A-P axis.

      To address this interesting and important comment by Rev1 we have carried out a new experiment using an alternative driver to 109(2)80-Gal4 and testing the impact of these manipulations on larval behaviour. The revised version of our MS includes a new figure Supp Fig S3 which shows self-righting times when using the ppk-Gal4 driver with the opto-axial technique. As observed with the 109(2)80-Gal4 driver, self-righting was delayed in anterior but not posterior inhibition conditions, suggesting the daIV neurons act in a region-specific manner to trigger postural control behaviour.

      We have also conducted a head casting analysis in the ppk domain; in another new figure, Supp Fig S7, we also show that head casting behaviour is also increased in the same manner as with the 109(2)80-Gal4 driver.

      These new panels and figures are cited within the sub sections entitled “Optogenetic inhibition of anterior but not posterior multidendritic neurons delays self-righting” and “Inhibition of anterior multidendritic neurons is associated with increased head casting during self-righting”, on pages 25 and 28, respectively. We are grateful to Rev1 for this suggestion, which we consider qualitatively improves our paper.

      (1.3) The manuscript is narrowly focused on sensory neurons that initiate righting, which limits the advance given the known roles for daIV neurons in righting. With the suite of innovative new tools, there is a missed opportunity to gain a more general understanding of how sensory neurons contribute to the righting response, including promoting and inhibiting righting in different regions of the larva, as well as aspects of proprioceptive sensing that could be necessary for righting and account for some of the observed effects of 109(2)80.

      Once again, we appreciate this interesting comment by Rev1. We feel our study provides novelty in understanding how sensory neurons in different body regions contribute to the induction of the behaviour. We developed new technology to show that the activity of anterior sensory neurons is essential for normal righting and inhibiting this activity leads to a switch to a different behavioural regime. We feel this represents a substantial advancement in our understanding of how this behaviour is initiated that has not been previously described. Whilst we also appreciate there is likely to be a substantial role of proprioception in self-righting behaviour, our work here focuses on the external stimuli that elicit self-righting, as a detailed understanding of proprioception would be out of scope and require the development of further techniques to manipulate and measure larval posture. As detailed in the above comment, we feel that the more targeted investigation of daIV neurons can also shed some light on the cell-type specificity and inputs to the self-righting induction process.

      (1.4) Although the authors observe an influence of Hox genes in righting, the possible mechanisms are not pursued, resulting in an unsatisfying conclusion that these genes are somehow involved in a certain region-specific behavior by their region-specific expression. Are the cells properly maintained upon knockdown? Are axon or dendrite morphologies of the cells disrupted upon knockdown?

      We agree with this comment in that further investigating the effects of Hox expression on localised aspects of the sensory system poses an interesting line of investigation. Indeed, we are currently conducting a full scale analysis of Hox gene effects across the sensory field. As things stands, it is not clear how Hox gene expression could affect local sensory processes, a mechanism which could involve morphological changes, changes in neuronal excitability (e.g. due to changes in channel expression), synapse formation and/or efficiency, cell development and identity, and/or combinations of these effects, amongst other possibilities. It is clear that a complete and satisfying investigation of this mechanism for each of the Hox genes would pose a substantial amount of work so, while we acknowledge the merit of Rev1’s comment, we consider that adding a cellular-mechanistic analysis of Hox effects is out of scope for the present study and shall constitute a central matter for a followup study emerging from current projects. We think that our data on Hox expression/function as reported here should serve to open up the analysis of genetic regulation of local sensory function, an area in which we are currently working very actively.

      (1.5) There could be many reasons for delays in righting behavior in the various manipulations, including ineffective sensory 'triggering', incoherent muscle contraction patterns, initiation of inappropriate behaviors that interfere with righting sequencing, and deficits in sensing body position. The authors show that delays in righting upon silencing of 109(2)80 are caused by a switch to head casting behavior. Is this also the case for silencing of daIV neurons, Hox RNAi experiments, and silencing of CO neurons? Does daIII silencing reduce head casting to lead to faster righting responses?

      This is an insightful comment. In the revised version of the manuscript, we do indeed show that anterior inhibition of daIV neurons leads to the same head casting behaviour as with the 109(2)80 domain, which we interpret as an inability of the larvae to sense the underlying substrate (see page 28). We hope the new data addresses this comment, at least to an extent. While we acknowledge it would also be insightful to run this behavioural analysis for other experimental conditions, such as the daIII inhibition and Hox RNAi lines, these experiments pose a specific technical difficulty: the behavioural analysis relies on a deep neural network (DNN) which was trained solely on recordings of the opto-axial technique, meaning it does not translate well to other experimental situations. This problem is further compounded by the use of L1 larvae, which means recording resolution is insufficient to accurately define the body landmarks used in the posture tracking at a smaller scale. Therefore, the recourse for identifying behavioural changes is manual observation, which we feel is too inconsistent to address a quantitative question like this.

      (1.6) 109(2)80 is expressed in a number of central neurons, so at least some of the righting phenotype with this line could be due to silenced neurons in the CNS. This should at least be acknowledged in the manuscript and controlled for, if possible, with other Gal4 lines.

      We thank the reviewer for making this interesting comment. We have added a phrase to the section “Conditional inhibition of multidendritic neurons delays self-righting” (p21) which acknowledges the presence of 109(2)80 expression in the CNS (as reported by Hughes and Thomas). We agree that ideally, a variety of sensory Gal4 lines would be used to check for consistency of the effects. However, it is also important to note that 109(2)80 is one of the only available Gal4 lines with near sole md neuron expression, as other Gal4s also drive expression strongly in external sensory cells for example. Thus, re-running experiments with these other lines – which would involve a substantial investment of time and resources – would not be an ideal strategy. We feel that the new observation of (very) similar axial results using the ppk-Gal4, which does express solely in the daIV neurons, better helps to confirm the specificity of the findings to multidendritic neurons.

      Other points:

      (1.7) Interpretation of roles of Hox gene expression and function in righting response should consider previous data on Hox expression and function in multidendritic neurons reported by Parrish et al. Genes and Development, 2007.

      We thank Rev1 for pointing out this study, which is definitively important to discuss given our results on Hox genes. To address this gap, we have added an additional paragraph in the Discussion (p37) to discuss the documented effects of Hox genes on da neuron dendritic morphology and how our results can be interpreted in light of this.

      (1.8) The daIII silencing phenotype could conceivably be explained if these neurons act as the ventral inhibitors. Do the authors have evidence for or against such roles?

      This is another interesting suggestion. If the daIII neurons were to fulfil this role, then in theory, their inhibition would result in self-righting behaviour under conditions of combined dorsal and ventral substrate contact. This is not an experiment we performed, so we are currently unable to confirm or rule out this possibility. However, we note from casual observation that daIII inhibition does not cause larvae to spontaneously self-right. As mentioned above, our view is not one in which the system has “dorsal/ventral stimulators/inhibitors” for a given behaviour, but that action selection proceeds according to a coordination of many (dynamic) contextual clues. Given the new results with the axial inhibition of daIV neurons (see above) it might be more parsimonious to suggest that these “tiling” neurons are primarily responsible for detecting substrate contact around the full circumference of the animal, rather than this involving different cell types according to the different sides of the body.

      Reviewer #2 (Public review):

      Strengths:

      The work of Roseby et al. does what it says on the tin. The experimental design is elegant, introducing innovative methods that will likely benefit the fly behavior community, and the results are robustly supported, without overstatement.

      Weaknesses:

      The manuscript is clearly written, flows smoothly, and features well-designed experiments. Nevertheless, there are areas that could be improved. Below is a list of suggestions and questions that, if addressed, would strengthen this work:

      (2.1) Figure 1A illustrates the sequence of self-righting behavior in a first instar larva, while the experiments in the same figure are performed on third instar larvae. It would be helpful to clarify whether the sequence of self-righting movements differs between larval stages. Later on in the manuscript, experiments are conducted on first instar larvae without explanation for the choice of stage. Providing the rationale for using different larval stages would improve clarity.

      This is a very interesting point raised by Rev2. Most of our previous work on self-righting (e.g. PicaoOsorio et al. 2015 Science; Picao-Osorio, Baldaia et al. 2017 Genetics; Klann et al. 2021 Journal of Neuroscience) was focused on the first instar larva (L1) because this early stage: (i) represents the simplest form of all larval stages, (ii) allows meaningful comparisons with late embryonic processes guiding the development and physiology of the nervous system, (iii) captures the system in a relatively naïve state, that had limited if any exposure to external stimuli. Although these attributes remain valid for the investigation of the sensory stimuli that trigger self-righting, the implementation of the necessary regional physical measurements and manipulations used in this study (surface contact, opto-axial technique, deep neural network analysis) would be impossible to implement in the early forms of the larva simply due to its reduced size. Due to this, we employed L3s, which due to their larger dimensions enabled the development and use of the sophisticated regional stimulation techniques reported here. Yet, as Rev2 rightly points out, we return to the late embryo and early L1 at the point of conducting gene expression analyses as these are optimised for those early stages. The selection of larval stage according to experiment relies on the fact that all forms of the larva display self-righting (Issa, Picao-Osorio, et al. 2019 Current Biology), that SR does not differ according to larval stage and that the characterisation of the structure of the nervous system across larval stages has shown a large level of similarity and consistent topographically arranged connectivity between identified neurons (Gerhard et al. 2017 eLife).

      (2.2) What was the genotype of the larvae used for the initial behavioral characterization (Figure 1)? It is assumed they were wild type or w1118, but this should be stated explicitly. This also raises the question of whether different wild-type strains exhibit this behavior consistently or if there is variability among them. Has this been tested?

      Thank you to the reviewer for pointing this out. The genotype for Figure 1 was w<sup>1118</sup>; this has now been added to the figure legend and the results section – thank you to Rev2 for pointing this out. Although in this study we did not explicitly compare self-righting (SR) performance in wild type/control genotypes (as we are internally consistent in using w<sup>1118</sup>) based on previous data collected in our lab we know that self-righting times are similar and very consistent amongst inbred control lines such as w<sup>1118</sup>, yw, and Oregon Red. Furthermore, we can also add that when comparing SR times between these inbred populations with a highly polymorphic outbred Drosophila population (Martins et al. 2013 PLoS Pathogens) we observed that their SR time (i.e. 6.14s ± 1.06) was not significantly different from the inbred lines (p<0.05, U test) (Picao-Osorio, J. 2014 Doctoral Thesis, Chapter 4, p112).

      (2.3) Could the observed slight leftward bias in movement angles of the tail (Figure 1I and S1) be related to the experimental setup, for example, the way water is added during the unlocking procedure? It would be helpful to include some speculation on whether the authors believe this preference to be endogenous or potentially a technical artifact.

      This is an interesting comment, and we recognise that lateral manipulation biases in self-righting could indeed reflect experimental limitations or biological tendencies. At this point we cannot interpret these results as formal evidence of chirality, given that they may reflect subtle aspects of the micromanipulation of specimens. We are currently developing a motorised platform to conduct self-righting tests, which when fully developed, should help addressing the chirality question.

      (2.4) The genotype of the larvae used for Figure 2 experiments is missing.

      Thank you for pointing this out. These were again w<sup>1118</sup> larvae; this detail has now been added to the figure legend and the main text.

      (2.5) The experiment shown in Figure 2E-G reports the proportion of larvae exhibiting self-righting behavior. Is the self-righting speed comparable to that measured using the setup in Figure 1?

      Thank you for pointing this out. We have now added average self-righting times to the figure legends of figures 1 and 2. The self-righting times across for the dorsal + ventral contact conditions was notably longer than dorsal-only cases, which were also slightly longer than the “standard” case. This is perhaps to be expected, as the larvae are encountering unusual and ambiguous situations. We suggest the extra time could reflect an additional decision-making step or action flip-flopping process, or simply physical constraints on the movement (for example, not being able to use some parts of the body).

      (2.6) Line 496 states: "However, the effect size was smaller than that for the entire multidendritic population, suggesting neurons other than the daIVs are important for self-righting". Although I agree that this is the more parsimonious hypothesis, an alternative interpretation of the observed phenomenon could be that the effect is not due to the involvement of other neuronal populations, but rather to stronger Gal4 expression in daIVs with the general driver compared to the specific one. Have the authors (or someone else) measured or compared the relative strengths of these two drivers?

      We agree with this suggestion and to address this concern, we have added as part of our new figure Supp. Fig. S3, a dedicated panel S3C showing fluorescence measurements from ddaC using the 109(2)80-Gal4 and ppk-Gal4 lines. We found no difference in tdTomato fluorescence intensity, suggesting equal expression strength across the two Gal4 drivers. Our new results for axial daIV inhibition are also consistent with this effect size difference, further suggesting that inhibition of all md neurons poses stronger challenges for self-righting compared to the daIV neurons alone.

      (2.7) Is there a way to quantify or semi-quantify the expression of the Hox genes shown in Figure 6A? Also, was this experiment performed more than once (are there any technical replicates?), or was the amount of RNA material insufficient to allow replication?

      Unfortunately, we only had limited amounts of mRNA extracted from FACS-sorted 109(2)80>GFP cells to feed our reverse transcriptase reactions and used much of these samples for the experiment reported. After Rev2 suggestion we went back to our freezers, recovered traces of the samples used in the original experiment, and attempted a new amplification; despite this effort, this new experiment was unsuccessful. We feel that the main point deduced from the original experiment is valid in that we obtained amplicons of the expected size for all the Hox transcripts analysed and that for those cases in which we observed biological effects – i.e. Antp and Abd-B – we corroborated protein expression in the 109(2)80 domain using immunohistochemistry. We are currently expanding this project examining the roles of all Hox genes across the entire sensory system and shall report the expression patterns of all Hox genes in each of the subcomponents of the sensory system the future.

      (2.8) Since RNAi constructs can sometimes produce off-target effects, it is generally advisable to use more than one RNAi line per gene, targeting different regions. Given that Hox genes have been extensively studied, the RNAis used in Figure 6B are likely already characterized. If this were the case, it would strengthen the data to mention it explicitly and provide references documenting the specificity and knockdown efficiency of the Hox gene RNAis employed. For example, does Antp RNAi expression in the 109(2)80 domain decrease Antp protein levels in multidendritic anterior neurons in immunofluorescence assays?

      We used the TRiP RNAi lines, specifically the Valium10 selection available from the Bloomington Stock Centre. Unfortunately, there is not much information on how specific the Hox RNAi lines areor whether their might have off-target effects.

      (2.9) In addition to increasing self-righting time, does Antp downregulation also affect head casting behavior or head movement speed? A more detailed behavioral characterization of this genetic manipulation could help clarify how closely it relates to the behavioral phenotypes described in the previous experiments.

      This would be interesting line of investigation. As described in a previous comment, this is currently unfeasible for us given some important differences between experiments including larval stage and recording conditions. We have added some speculative comments to the manuscript describing the larval behaviour under Hox RNAi.

      (2.10) Does down-regulation of Antp in the daIV domain also increase self-righting time?

      Given the new results with axial effects of daIV neurons, we also sought to address this point with a new series of experiments expressing Hox RNAi constructs in the ppk-Gal4 domain. The new data is shown in a new figure (Figure S8) displaying self-righting times for ppk-Gal4-Hox-RNAi. Interestingly, we found no effect of any RNAi expression on self-righting times, suggesting that md types other than daIVs are under Hox regulation that is important for self-righting.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The reviewers were enthusiastic about the value and quality of this study by Roseby and colleagues. There were two main issues that emerged from the reviews that we're highlighting for the authors to address, should they choose to:

      (1) A little more cell-type resolution of the anterior region

      The anterior region includes a lot of sensory neurons that may be contributing to the effect. Some sensory neurons (e.g., daIV) have been implicated in righting - are these the ones carrying the anterior signal? Are dorsal sensory neurons promoting righting and ventral ones stalling it?

      We are not suggesting a complete sensory-neuron mapping in the anterior region. Instead, we propose the authors conduct a focused check: repeat the axial inhibition with a daIV-specific driver (same photomask assay) to show the A-P effect within the implicated class, and, if possible, replicate one key result with an alternative broad md driver to address Gal4 strength/off-target expression.

      As mentioned above (see Rev1 comment) we have indeed carried out a new experiment using an alternative driver to 109(2)80-Gal4 and testing the impact of these manipulations on larval behaviour. The revised version of our MS includes a new figure Supp Fig S3 which shows self-righting times when using the ppk-Gal4 driver with the opto-axial technique. As with the 109(2)80-Gal4 driver, self-righting was delayed in anterior but not posterior inhibition conditions, suggesting the daIV neurons specifically act in a region-specific manner to trigger postural control behaviour.

      Furthermore, in another new figure, Supp Fig S7, we show that head casting behaviour is also increased in the same manner as with the 109(2)80-Gal4 driver. These new panels and figures are cited within the sub-sections entitled “Optogenetic inhibition of anterior but not posterior multidendritic neurons delays self-righting” and “Inhibition of anterior multidendritic neurons is associated with increased head casting during self-righting”, on pages 25 and 28, respectively. We are grateful to R1 for this suggestion, which we consider qualitatively improves the quality of our paper.

      (2) The Hox section to strengthen this section, we recommend:

      (a) Confirm specificity/efficacy of knockdown (e.g., Antp protein reduction in targeted md neurons and a second RNAi line if available).

      This is a reasonable comment. For our experiments, we selected a UAS-Antp<sup>RNAi</sup> line (Bloomington #27675) given that this construct has been: (i) utilised in several previous studies as the main and single line to interfere with Anpt expression (e.g. Baek et al. 2013 Development; Paul et al. 2021 Nature Comms) and (ii) shown to display a consistent reduction in Antp protein levels of approximately 50% (see Poliacikova et al. 2024 Science Adv.). Furthermore, previous work comparing #27675 with other UAS-Antp<sup>RNAi</sup> lines has demonstrated that all available lines lead to a similar level of reduction in protein expression, although the #27675 line exhibits the most consistent effects (lower variability) (Poliacikova et al. 2024 Science Adv.). Unfortunately, at this point in time, we do not have the capacity to conduct new experiments with other RNAi lines, but consider that the information and arguments mentioned above should be reassuring about our choice of a reasonable and previously validated method to interfere with Antp expression.

      (b) Perform one temporal control (GAL80^ts) or a simple rescue, to separate developmental vs acute roles.

      This is a good and interesting suggestion, but we consider that the discrimination between developmental and physiological effects falls outside the scope of this study. Indeed, experiments of this kind are currently being conducted in our lab as part of a wider examination of Hox gene roles in the sensory system.

      (c) Place the results clearly in the context of prior work (e.g., Parrish 2007), so the mechanism isn't left hanging.

      This is an important point, and we have now done this. Many thanks for pointing this out.

      Reviewer #1 (Recommendations for the authors):

      (1.1) A Gal4 line for the pannier dorsal specification gene shows expression in dorsal sensory neurons, as described in Galindo et al., Development, 2023, and could help tease apart dorsal v. ventral contributions.

      This is an interesting suggestion. However, we understand that the pannier (pnr) Gal4 line mentioned in Galindo et al. 2023 is an enhancer trap inserted in the pnr locus which drives expression in neural as well as non-neural tissues such as the embryonic dorsal ectoderm (see: Calleja et al. 1996 Development; Stronach et al. 2014 Genetics). Although, as Rev1 rightly indicates, this line also labels dorsal cluster sensory neurons, including ddaC (cIV) and ddaF (cIII) neurons the fact that the line displays expression in non-neural tissues makes its use in behavioural experiments difficult as non-neural effects might affect the behavioural patterns studied. A possible way to instrument the pnrGal4 tool into behavioural analyses might involve the creation of the necessary variants to implement a split-Gal4 approach, but this, we believe, unfortunately falls out of the scope of this study.

      (1.2) Potential roles for daII neurons and daI neurons are not examined. Drivers have been described for daII neurons, and there are drivers that will target a majority of proprioceptive md neurons, so these could be examined to complete the analysis started here.

      This is another interesting suggestion by Rev1, but we consider that the fine-grain mapping of effects mediated by sensory neuron sub-clases falls outside the scope of this study aimed at mapping sensory regional effects on self-righting. This does not take the merit of the suggestion away, and indeed, experiments of this kind are currently being conducted in our lab as part of a comprehensive examination of Hox gene roles in the sensory system.

      (1.3) To account for 109(2)80 off targets, the authors could consider other lines that silence most or all md neurons (clh201-Gal4; 5-40-Gal4; 21-7-Gal4) that could at least have different central offtargets. Some other lines are broad somatosensory system drivers but sensory-specific (pebbledGal4).

      This is an interesting comment, and so are the suggestions made. Although to include this kind of verification would be interesting, when carrying out our experiments, we did not observe any central expression at all. Also, to repeat all our experiments in which we use the established and validated 109(2) 80 line using instead these four Gal4 lines, is unfortunately out of scope for us at this point in time. We will nonetheless consider these comments by Rev1 in future extensions of our work.

      (1.4) There is a typo on line 481; it should be "other".

      We are grateful to R1 for pointing this out. This has now been amended

      Reviewer #2 (Recommendations for the authors):

      (2.1) Lines 91-92 cite references describing self-righting behavior across different animal groups, which is illustrated in Figure 1B. It would be helpful to indicate these references directly in the figure. For example, instead of using dots to denote their presence (which are, in a way, redundant since the behavior is reported in all groups), numbers or letters could be used to refer to the specific papers describing them.

      Thank you for this suggestion. We have now replaced the original dots by an abridged citation of a key paper providing evidence in that specific animal group, e.g. Smith, et al. 1997; Rogers et al. 2015

      (2.2) In Figure 1A, the diagrams illustrate the two large dorsal tracheae, which nicely indicate the larva's orientation. However, since they are drawn in a very light gray, they can be difficult to distinguish without zooming in. It might improve clarity if the tracheae were made slightly more prominent.

      Thank you for this suggestion. We have now implemented this change.

      (2.3) In Figure 1E, the dotted line and green bar mark the segment of the recording corresponding to self-righting, which is then quantified in Figure 1G. Was the same procedure applied when analyzing tail speed, or was it limited to head speed? Figure 1F does not show a dotted line or green bar, which is confusing; it would be helpful to clarify the reason for this discrepancy. Also, in Figure 1G, there is an inset showing photos of the movement sequence with the green bar and the caption 'Trimmed to SR sequence,' which implies to me that for tail speed, the 0.75-1 segment of the recording was also used for quantification. I suggest adding the dotted line and green bar to Figure 1F and removing this inset from Figure 1G, as it appears quite small and disrupts the layout of the figure. If it is retained, the figure legend should explicitly refer to the inset.

      Thank you for pointing this out. We have amended these figures as suggested.

      (2.4) In Figures 1 and 2, the box plots include the individual data points, whereas Figures 3 and S2 do not. For data transparency, it would be important to show the individual measurements here as well. I strongly recommend adding them to the figure, or alternatively providing a clear rationale in the text for not doing so.

      Thank you for mentioning this. The reason data points are not shown in Fig 3 or S2 is because the variance extends the scale and compresses the box making it illegible. To make this clear we now explain this in the figure legends.

      (2.5) In Figures 4 and 5, the distribution of self-righting times from the optogenetic inhibition experiments is shown using bar graphs rather than box plots, as in the previous figures. This choice obscures the data distribution, since all bars reach down to zero. Replacing the bar graphs in Figures 4 and 5 with box plots would more clearly convey the experimental results.

      We thak Rev2 for this comment, which gives us an opportunity to clarify the matter. Distributions of SR times are drawn with bars because we compare means +/- variance in the analysis, and not medians +/- IQR as is done in the other experiments. The choice of visualisation reflects the analysis, which is what is recommended by statisticians. Plus, we also show the individual observations, meaning the distribution can be observed. We hope that it is now clear that we are not obscuring any distributions.

      (2.6) Figure 6 would benefit from some reorganization. Panel A is very small and dense with information, making it difficult to interpret without significant zooming. In particular, the FACS graph is nearly impossible to read, as the axes remain unclear even when enlarged. It might be best to either remove this graph and replace it with a cartoon version of FACS-sorted populations, and reorganize the figure to ensure legibility. Additionally, the current layout progresses from the bottom up, which takes time to follow. Comprehension could be improved if the sequence began with the larva dissection placed in the top left area of the figure, where readers typically look first (I appreciate that this is mentioned in the figure legend; however, a different layout might present the information more effectively).

      We appreciate the constructive spirit of this comment and have indeed considered Rev2 suggestions including drafting new layouts of this figure. After all this experimentation, we remain of the view that the original presentation is probably the best trade-off between size and clarity, offering more space for the appreciation of confocal imaging and its interpretation.

      Minor corrections:

      (1) Throughout the text, the word Drosophila appears sometimes in italics and sometimes in regular font; please standardize its formatting for consistency.

      Amended

      (2) Line 179: the use of three hyphens in the sentence "minimum --- in all cases < 30 s --- to avoid larval desiccation" is unusual; exchanging them for commas or brackets is advised.

      Amended

      (3) Line 183: in w1118, the numbers are usually in superscript (not subscript), and the w should be italicized.

      Amended

      (4) In line 783, there is an incorrect space between "is" and the comma in "...repertoire, which is , in...".

      Amended

      (5) In Figure 2G, the left panel appears partially cut off, which makes the text at the edges difficult to read. It might help to adjust the panel so that all labels are fully visible.

      Done

      (6) In the current version of the manuscript, Figure 5 is presented before Figure 4, which is confusing.

      This has been amended.

      (7) Two videos are included in the supplementary material, but I could not find any reference to them in the main text of the manuscript.

      This has been amended.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      (1) Research scope

      The results primarily focus on mutations in ZNF217, ZNF703, and ZNF750, with limited correlation analyses between mutations and gene expression. The rationale for focusing only on these genes is unclear. Given the availability of large breast cancer cohorts such as TCGA and METABRIC, the authors should compare their mutation profiles with these datasets. Beyond European and U.S. cohorts, sequencing data from multiple countries, including a recent Nigerian breast cancer study (doi: 10.1038/s41467-021-27079-w), should also be considered. Since whole-exome sequencing was performed, it is unclear why only four genes were highlighted, and why comparisons to previous literature were not included.

      We have significantly strengthened the biological and clinical rationale for focusing on these three genes in the Introduction. Specifically, we now clearly justify their selection based on distinct functional roles: ZNF217 (oncogene, 20q13 amplification); ZNF703 (luminal subtype oncogenic driver); ZNF750 (tumor suppressor involved in differentiation). We have also explicitly define the knowledge gap: lack of mutation and expression data for these genes in African populations, particularly Kenyan cohorts.

      Importantly, we have now incorporated comparative analysis with TCGA data in the Results. This include; A new section on “Recurrent mutations and comparison with TCGA”; a new table, “Table 6” and a curated dataset, “Supplementary Table S4”

      (2) Language and Style Issues

      There are many typos and clear errors in the main text (e.g. (ref)).

      Additionally, several statements read unnaturally. For example:

      "Investigators uncovered 170 mutations ..." should instead be phrased as "We identified 170 mutations ...."

      "The research team ..." should be rephrased as "Our team ...."

      The manuscript has undergone comprehensive language editing throughout the revised draft.

      (3) Methods and Data Analysis Details

      The methods section is vague, with general descriptions rather than specific details of data processing and analysis. The authors should provide:

      (a) Parameters used for trimming, mapping, and variant calling (rather than referencing another paper such as Tang et al. 2023).

      (b) Statistical methods for somatic mutation/SNP detection.

      (c) Details of RNA purification and RNA-seq library preparation.

      Without these details, the reproducibility of the study is limited.

      We have fully revised and substantially expanded the Methods section to improve clarity, transparency, and reproducibility. In the revised manuscript, we now provide explicit details of all key analytical steps. These include quality control procedures using FastQC and MultiQC, as well as read trimming parameters implemented in Trimmomatic (leading and trailing quality <3, sliding window 4:15, and minimum read length of 36 bp). We also clearly describe alignment of reads to the hg38 reference genome using BWA-MEM, followed by somatic variant calling using MuTect2 in paired tumor–normal mode with incorporation of a Panel of Normals (PON). Variant filtering criteria are now explicitly stated, including minimum read depth (≥10), base quality (≥20), and variant allele fraction (≥0.05), and functional annotation was performed using VEP (v108).

      In addition, we have included details on variant validation through visualization in the Integrative Genomics Viewer (IGV), as well as RNA-seq processing steps using STAR for alignment, featureCounts for quantification, and DESeq2 for normalization and differential expression analysis. Statistical analyses are now clearly described, including the use of paired tests and Benjamini–Hochberg correction for multiple testing. Collectively, these additions directly address the reviewer’s concerns by ensuring that all analytical procedures are transparently reported and fully reproducible.

      (4) Data Reporting

      This study has the potential to provide a valuable resource for the field. However, data-sharing plans are unclear. The authors should:

      (a) Deposit sequencing data in a public repository.

      (b) Provide supplementary tables listing all detected mutations and all differentially expressed genes (DEGs).

      (c) Clarify whether raw or adjusted p-values were used for DEG analysis.

      (d) Perform DEG analyses stratified by breast cancer subtypes, since differential expression was observed by HER2 status, and some zinc finger proteins are known to be enriched in luminal subtypes.

      We have improved data transparency and reporting in the revised manuscript. All sequencing data are now publicly available, with whole-exome sequencing (WES) data deposited in the Sequence Read Archive (SRA; PRJNA913947) and RNA-seq data available in the Gene Expression Omnibus (GEO; GSE225846). In addition, we have provided comprehensive Supplementary Materials to support reproducibility and facilitate further analysis, including detailed mutation summaries (Table S1), mutation positions (Table S2), amino acid changes (Table S3), the curated TCGA comparison dataset (Table S4), protein domain annotations (Table S5), and the combined gene expression and clinical dataset (Table S6).

      We have also clarified key aspects of the statistical analysis, including the use of Benjamini–Hochberg adjusted p-values and the thresholds applied for significance. Furthermore, in response to reviewer comments regarding subtype-specific analyses, we have explicitly addressed in the Discussion why subtype-stratified differential expression analysis was not performed, noting that the limited sample size would reduce statistical power and increase the risk of overinterpretation. Together, these revisions enhance the transparency, accessibility, and interpretability of the study.

      (5) Mutation Analysis

      Visualizations of mutation distribution across protein domains would greatly strengthen interpretation. Comparing mutation distribution and frequency with published datasets would also contextualize the findings.

      We have substantially enhanced the mutation analysis by incorporating several new figures and complementary analyses that provide deeper biological interpretation. Specifically, we added Figure 1 to summarize mutation burden, coding consequences, and prevalence; Figure 2 to illustrate the nucleotide substitution spectrum; Figure 3 to map mutations across protein domains; Figure 4 to assess functional enrichment and mutation composition; and Figure 5 to highlight recurrent mutations.

      Reviewer #2 (Public review):

      Weaknesses:

      The current cohort size is relatively small to reach significant findings, and targeted exploration on ZNF family without emphasizing the reason or clinical significance hinders the overall significance of the entire work.

      We acknowledge the limitation posed by the relatively small cohort size and have addressed this concern in several ways in the revised manuscript. First, we have explicitly stated this limitation in the Discussion section. We have also reframed the study as a pilot and population-specific exploratory analysis to better reflect its scope. To strengthen the overall significance, we integrated both mutation and gene expression data, incorporated comparisons with TCGA datasets, and emphasized the importance of African-specific genomic insights. Importantly, we highlight that this study provides novel data from an underrepresented population, which represents a key contribution to the field.

      Reviewer #3 (Public review):

      Weaknesses:

      The author has enhanced the descriptive depth of the study by adding details on mutations, expression subgroup analyses, and functional annotations but has not addressed the core weaknesses of small cohort size and lack of functional validation. While the revised version is more comprehensive in cataloging molecular alterations, it remains confined to descriptive analysis, with no substantial improvement in the reliability or generalizability of its conclusions.

      We have addressed this concern by clearly acknowledging the key limitations of the study, including the absence of functional validation, the relatively small sample size, and the limited generalizability of the findings. In response, we have refined our interpretation to avoid causal claims and instead present the results as hypothesis-generating. We have also expanded the Discussion to include future research directions, recommending functional validation studies, multi-omics approaches, and validation in larger, more diverse cohorts.

      In addition, we have strengthened the robustness of the study by incorporating comparisons with TCGA data, providing more detailed mutation classification, and integrating genomic and transcriptomic analyses. Beyond addressing reviewer comments, we have further improved the manuscript by reorganizing the Results section to follow a clear and logical flow—from mutation burden and spectrum to protein-level distribution, functional enrichment, recurrent mutations, and TCGA comparison. We have also improved figure quality and labeling to meet journal standards, added clear and consistent figure captions, and ensured alignment between the text, figures, and tables throughout the manuscript.

      We sincerely thank the reviewers for their valuable feedback, which has significantly improved the quality and rigor of this work.

    1. Author response:

      We thank the Editor and the Reviewers for their detailed and constructive feedback. We look forward to submitting a revised version of the manuscript that addresses their comments and suggestions, with a special focus on clarifying the assumptions and implications of our analysis. In particular, we will aim to demonstrate that (i) many of our qualitative findings -- and even some quantitative results -- extend beyond the simplest two-resource case considered in the main text, and (ii) that they can also be generalized to account for simple forms of cross-feeding. We hope that these changes will help to illustrate the broader applicability of our underlying mathematical framework.

    1. Author response:

      We thank the editors and reviewers for their thoughtful and constructive evaluation of our manuscript. We are pleased that the reviewers found the study valuable and the evidence supporting a role for Yme1 in MDC formation solid. As described below, we plan to modify the manuscript to clarify the lipid model, better explain the relationship between Ups-family proteins and MICOS, distinguish MDC formation from Atg32-dependent mitophagy, clarify metabolic conditions, add statistical analyses where missing, and strengthen Yme1 validation with immunoblotting.

      eLife Assessment

      This valuable study demonstrates that the inner membrane protease YME1 contributes to the formation of mitochondrial-derived compartments in yeast through the modulation of both the lipid transporter UPS2 and the MICOS complex. The evidence supporting this model is solid, although this manuscript could be improved by providing additional evidence supporting the independent roles for UPS2 and MICOS regulation in this process. This work will be of interest to cell biologists, biochemists, and geneticists interested in understanding the molecular basis of mitochondrial regulation and function.

      We appreciate this positive assessment and agree that the roles of Ups-family lipid transport and MICOS in MDC regulation could be expanded further. This will be an important topic for future studies, especially with regard to how MICOS contributes to MDC formation. In the current revision, we will add new genetic data focused on PA-linked lipid metabolism through the yeast Pah1/Lipin pathway, which we think will help strengthen and clarify the lipid arm of the model. Our current interpretation is that Yme1-regulated Ups-family lipid transport and MICOS may both influence a shared mitochondrial membrane state that permits MDC formation. This interpretation is consistent with our genetic data and with known connections between Ups proteins, MICOS, and mitochondrial membrane organization.

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Balasubramaniam and colleagues continue this group's efforts to understand mitochondrial-derived compartments (MDCs) that bud off from yeast mitochondria in response to metabolic stress. In a previous genetic screen, they identified Ups lipid transfer proteins and the AAA-protease Yme1 as components that modulate MDC formation. In this study, the authors link these observations by showing that Yme1 modulates levels of Ups1, Ups2, as well as MICOS complex members in the mitochondrial proteome. Using genetic approaches, they then show that Yme1's role on MDCs is dependent on its catalytic activity (via an inactive mutant) and that YME1 shows genetic interactions with UPS1/2 and MIC10/MIC60. The overall model is that Yme1 activity responds to metabolic cues and acts via proteolysis of these two distinct mitochondrial machineries to regulate MDC biogenesis.

      Strengths:

      The strengths of the study are its integration of mitochondrial proteomics with strong genetic approaches, as well as synergy with the authors' previous studies on the role of lipids in MD genesis. The work is overall well carried-out and experiments are thoughtfully discussed.

      Weaknesses:

      The major weaknesses are a lack of mechanistic resolution surrounding the model, e.g., proposed or tested mechanisms by which Yme1 activity is regulated by metabolic cues, or how Ups1/2 activity and the MICOS contribute to MDC generation. The authors acknowledge these as open questions, but addressing them would still enhance the significance of the study.

      We thank the reviewer for the positive assessment, and we agree that the upstream regulation of this response remains an important open question. Yme1-dependent MDC regulation could involve changes in Yme1 activity, substrate accessibility, or broader changes in mitochondrial lipid and protein organization. Fully resolving how metabolic state gates this response will require future work, likely outside the scope of the current study.

      We also agree that the manuscript would benefit from a more developed discussion of how lipid changes could contribute to MDC formation. Our prior work showed that reduced mitochondrial PE promotes MDC formation, whereas cardiolipin is required for MDC biogenesis (Xiao et al., 2024). We proposed that reduced PE changes the membrane environment of mitochondrial outer membrane proteins, potentially affecting their stability, abundance, insertion, or lateral organization within the membrane. Such changes could increase the pool of proteins available for sorting into MDCs or make the outer membrane more permissive for domain formation. In the revision, we will connect this model more directly to Yme1-dependent regulation of Ups-family lipid transport.

      We will also expand the model to incorporate PA-linked metabolism. We did not initially focus heavily on Ups1 because complete loss of UPS1, or loss of downstream cardiolipin synthesis through CRD1, blocks MDC formation because cardiolipin is required. Thus, complete disruption of Ups1-dependent lipid transport may obscure the effects of more moderate changes in PA flux. To address this, we will include additional lipid measurements and new genetic data targeting PA metabolism through the yeast Pah1/Lipin pathway. Because Pah1 converts PA to DAG, this provides a way to alter PA-linked metabolism without simply eliminating cardiolipin synthesis. Our new data suggest that PA accumulation or altered PA-linked lipid flux may also promote MDC formation. Together, these findings support a broader model in which reduced PE and increased PA alter both the organization of OMM proteins and the physical properties of the membrane, including curvature and domain formation, thereby creating a membrane state that is more permissive for MDC biogenesis.

      Reviewer #2 (Public review):

      In this manuscript, the authors report a novel regulation of the outer mitochondrial membrane remodeling domains called mitochondria-derived compartments, MDCs. The team has previously established the main principles behind this recently identified quality control pathway, but the mechanisms that control MDCs formation remain incompletely understood. Using the baker's yeast model, the authors identify the conserved mitochondrial protease Yme1 as a crucial factor that regulates MDC formation. Mechanistically, Yme1's proteolytic function controls the levels of Ups1 and Ups2 lipid transfer proteins and the components of the membrane organizing complex called MICOS, thus providing a plausible model as to how Yme1-dependent proteolysis permits MDC formation through the removal of lipid and MICOS-dependent constraints. Finally, the authors show that this Yme1-mediated activity is also defined by metabolic conditions. In principle, this study is interesting and novel, and holds potential to provide new insights into the regulation of the MDC pathway that emerged as a new fundamental mitochondrial quality control mechanism. However, the following points should be carefully addressed.

      Major points:

      (1) Yme1 has been previously shown to regulate mitochondria-specific autophagy through Atg32 processing. Given the high similarity of the MDC pathway to piecemeal autophagy and the fact that both pathways share some of the core components, the authors should address the involvement of Atg32 in their model. It would also be important to include a brief discussion addressing the differences between piecemeal autophagy and the MDC pathway.

      We agree that this is an important point. The reason we did not focus on Atg32 in the current manuscript is that we previously investigated the relationship between MDC formation and Atg32-dependent mitophagy and found that Atg32 is dispensable for MDC formation (Hughes et al., 2016). Based on that result, we do not anticipate that Atg32 is required for the Yme1-dependent MDC phenotypes described here. This is also consistent with the different growth conditions associated with these pathways: Atg32-dependent mitophagy is stimulated under respiratory or post-diauxic conditions, whereas MDCs do not form under the respiratory conditions that stimulate Atg32-dependent mitophagy (Hughes et al., 2016; Raghuram and Hughes, 2024).

      We will clarify this distinction in the revised manuscript. In addition, to be thorough, we plan to generate and test the Atg32-GFP variant previously shown to block Yme1-dependent Atg32 processing and mitophagy (Wang et al., 2013). This will allow us to test directly whether preventing Yme1-dependent Atg32 cleavage affects MDC formation. If successful and interpretable, we will include these data in the revised manuscript.

      (2) The Rpt3 (P215L) expression experiment is interesting, but appears to be somewhat superficial due to the unclear mechanism by which the mitochondrial network morphology is restored in these cells. Could this result be replicated in the dnm1∆ mgm1∆ double deletion mutant, which is a well-established model for mitochondrial network restoration?

      We agree that the Rpt3(P215L) experiment is best viewed as a morphology control. The purpose was to test whether abnormal mitochondrial morphology alone explains the MDC defect in yme1Δ cells. Because Rpt3(P215L) improved mitochondrial morphology but did not restore MDC formation, we interpret this as evidence that morphology alone is not sufficient.

      We attempted to generate the requested dnm1Δ mgm1Δ yme1Δ triple-mutant combination, but that strain combination has not been viable in our hands. However, we do have dnm1Δ data showing that altering mitochondrial structure can rescue some morphological features but does not restore MDC formation in yme1Δ cells. We will include these data where appropriate and clarify that this experiment is intended as a morphology control.

      (3) Figure 3E. The changes in PE levels appear to be minor. While statistically significant, the observed differences may not be physiologically relevant. More in-depth lipidomic analysis data should be presented to substantiate the authors' argument and better address the questions at hand. Related to that, could PE or PA supplementation stimulate MDC formation?

      We agree that additional lipid data would strengthen this part of the manuscript. We initially streamlined the lipid section because we had previously examined the lipid requirements for MDC formation in detail, showing that reduced mitochondrial PE can promote MDC formation, whereas cardiolipin is required (Xiao et al., 2024). However, the current study would benefit from a broader analysis of the lipid changes associated with Yme1-dependent regulation.

      In the revision, we will expand the lipid data to include additional lipid species and incorporate these results into the model. We will also add new genetic data targeting PA metabolism through the yeast Pah1/Lipin pathway. Together, these data suggest that PA accumulation or altered PA-linked lipid flux may also contribute to MDC formation. This supports a broader lipid-balance or lipid-shunting model in which reduced PE, increased PA, or altered lipid distribution between mitochondrial membranes could influence OMM remodeling through effects on membrane curvature, OMM protein organization, or mitochondrial membrane contacts.

      We agree that direct PE or PA supplementation would be a valuable experiment. We have attempted lipid supplementation but have not been able to deliver these lipids effectively to yeast cells in a way that produces interpretable results. We are therefore focusing on lipid profiling and genetic approaches that alter lipid metabolism inside the cell.

      (4) The connection between rapamycin treatment and Yme1-regulated MDC formation is unclear and puzzling and needs to be explained better.

      We agree that this connection is not fully clear. In this manuscript, rapamycin is used primarily as a robust MDC-inducing condition. Our data do not define the full pathway connecting TORC1 inhibition to Yme1-dependent mitochondrial remodeling.

      In the revision, we will either clarify this point or reduce the emphasis on rapamycin as a mechanistic entry point. Our current interpretation is that rapamycin creates a metabolic/mitochondrial state in which Yme1-dependent remodeling of lipid and membrane-organization pathways becomes important for MDC formation. Whether this involves direct regulation of Yme1, altered substrate availability, altered membrane composition, or a combination of these remains open.

      (5) The MICOS complex is clearly involved in the regulation of MDC, but the manuscript misses the mark on providing compelling evidence and a clear explanation as to how MICOS contributes to said regulation.

      We agree that the mechanism by which MICOS regulates MDC formation remains an important open question and will be a major focus of future work. Our current data show that MICOS perturbation can partially restore MDC formation in yme1Δ cells, supporting a role for MICOS in this pathway. This analysis was motivated in part by the incomplete genetic suppression achieved through the lipid pathway alone, which suggested that additional Yme1-regulated factors contribute to MDC formation.

      MICOS therefore represents a strong candidate for this additional regulatory input. However, defining whether MICOS acts through lipid distribution, OMM-IMM organization, membrane architecture, or another mechanism will require a deeper investigation than is possible within the scope of the current study. We will clarify this point in the revised manuscript and present the current findings as the beginning of a broader investigation into how MICOS contributes to MDC biogenesis.

      Minor points:

      (1) The authors should discuss potential reasons for the dramatically different rates of MDC formation in the S288C and W303 background cells. Does this have anything to do with generally more robust mitochondrial functions in the latter cells?

      We agree this is worth discussing. One likely explanation is that the difference reflects broader differences in mitochondrial activity and metabolic state between these strain backgrounds. We and others have shown that W303 cells have more robust respiratory mitochondrial function than BY/S288C-derived cells, and in our hands W303 also shows lower MDC formation. This fits our broader model that MDCs are favored in glucose-grown or metabolically perturbed cells and do not form under respiratory conditions (Raghuram and Hughes, 2024). We do not yet know the genetic basis for this difference, so we will present this as an interesting future direction.

      (2) Proper statistical analyses should be provided for all the graphs presented.

      We will add statistical analyses where missing.

      (3) The authors should include Yme1 immunoblots to confirm the identity of strains being studied and validate the presence or overexpression of Yme1 and its catalytic mutant in their experiments.

      We agree that direct validation of Yme1 protein levels will strengthen the manuscript. Our quantitative mitochondrial proteomics already confirms strong depletion of Yme1 in yme1Δ cells, and we will also include quantitative proteomics showing increased Yme1 abundance in the overexpression strain. In addition, we have now obtained a Yme1 antibody from a colleague and will include immunoblots validating Yme1 loss, re-expression, catalytic mutant expression, and overexpression where appropriate.

      Reviewer #3 (Public review):

      Summary:

      Since describing MDCs over a decade ago, the lab of the corresponding author, Hughes, has been at the forefront of further characterizing these structures. Here, they follow up on recent work (PMID: 38497895), where a screen identified Yme1 as a potential regulator of MDCs. After confirming that Yme1-ko prevents MDCs that are usually induced via various established treatments (Rapamycin, cycloheximide, Concanavalin A), the authors confirmed that the proteolytic activity of Yme1 is required. Next, using proteomics, they identified how loss of Yme1 impacts the mitochondrial proteome with and without Rapamycin treatment to induce MDCs. From this result and based on insight from other published data implicating lipids, the focused initially on the lipid transfer protein Usp2, a known target of Yme1. Here, they showed that loss of Usp2 could partially rescue MDC formation in Yme1-ko cells. To look for other Yme1 targets that might also be involved in MDC formation, next, they investigated the MICOS complex, which was also notable in their proteomics data. They then showed that inhibiting MICOS also partially restored MDC formation in Yme1-ko cells. They then tested the combined effects of Usp2 and MDC inhibition on MDCs, which was limited by the fact that the combination of full MICOS disruption, Usp2-KO, and Yme1-KO was not viable. To circumvent this limitation, they investigated the knockout of individual MICOS subunits in combination with Usp2 and/or Yme1. Finally, they showed that growth conditions also mediate MDC formation in the context of Yme1 overexpression. In rich media, Yme1 overexpression induces MDCs on its own. However, this induction is lost upon amino acid starvation, suggesting that there are still other as-yet-unidentified factors regulating the formation of MDCs.

      Strengths:

      The authors use unbiased approaches and genetic models to begin unraveling a novel regulatory role of Yme1 in the formation of MDCs.

      Weaknesses:

      (1) The authors find both Ups1 and Ups2 in their screens, but only focus on Ups2 in this paper. It would be good to know why they did not also investigate Ups1, and its other protease Atp23, which could potentially act similarly to Yme1, or even rescue the loss of Yme1.

      We agree that Ups1 and Atp23 are important to consider. We initially focused on Ups2 because its deletion partially restores MDC formation in yme1Δ cells and because of its connection to mitochondrial PE synthesis, which we had previously shown to regulate MDC formation (Xiao et al., 2024). Ups1 is more difficult to assess genetically because complete loss of UPS1, or of downstream cardiolipin synthesis through CRD1, blocks MDC formation due to the requirement for cardiolipin. Thus, an ups1Δ phenotype cannot readily reveal whether a more moderate reduction in Ups1 activity, and the resulting accumulation or redistribution of PA, might promote MDC formation.

      In the revision, we will explain this rationale and include new genetic data targeting PA metabolism through the yeast Pah1/Lipin pathway. This provides a way to test the contribution of PA accumulation without simultaneously eliminating cardiolipin synthesis, and our initial results support a role for PA-linked lipid remodeling in partially bypassing the requirement for Yme1. We will also discuss Atp23 as a potentially important regulator of Ups1 and PA metabolism. A full investigation of Atp23 will be an important direction for future work.

      (2) I'm not convinced that the data support the notion that Usp2 and MICOS have distinct effects on MDCs. In Figure S3C-D, there is no statistical analysis to indicate whether the small differences between the MICOS-ko and the double knockout are significant. If MICOS-ko and Ups2-ko were acting through different mechanisms, one would expect their combination to be additive; this does not appear to be the case, as both single deletions and the double deletion all cause similar levels of MDCs (~30-40%). Rather, this result is what you would expect if they were working through the same mechanism. There also does not appear to be an additive effect in Figure 4F-G, when using the mic60-ko rather than the complete MICOS-ko. In this regard, the authors note in their discussion that 'loss of MICOS may disrupt membrane associations or alter lipid distribution between mitochondrial subcompartments' (lines 390-392). The latter situation seems like it would be the same mechanism as Usp2 and would more accurately explain their findings.

      This is a very good point, and we agree with the reviewer’s interpretation. The lack of strong additivity is consistent with Ups2 and MICOS acting within the same pathway or converging on a shared mechanism, rather than representing two separate mechanisms of MDC regulation. We did not intend to imply that these must be independent pathways. In the revised manuscript, we will ensure that the text reflects this interpretation and will add statistical analyses to the relevant comparisons.

      (3) The manuscript is missing key data confirming the re-expression or overexpression of Yme1 protein (Figure 1 E/G and Figure 5A). It is important to know the relative levels of expression of the re-expressed proteins to each other and to endogenous Yme1.

      We agree that direct validation of Yme1 protein levels is important. Our quantitative mitochondrial proteomics already confirms strong depletion of Yme1 in yme1Δ cells, and we will also include quantitative proteomics showing increased Yme1 abundance in the overexpression strain. In addition, we have now obtained a Yme1 antibody from a colleague and will add immunoblots validating Yme1 loss, re-expression, catalytic mutant expression, and overexpression.

      (4) Some clarification of the details for metabolically restrictive conditions would be helpful.

      Thanks for this suggestion. We will clarify these conditions throughout the manuscript and figure legends and will define exactly what we mean by low-amino-acid, amino-acid-free, synthetic, and rich media conditions. More broadly, MDC formation is strongly influenced by media composition and mitochondrial metabolic state. MDCs form less efficiently in synthetic media and do not form under conditions that promote respiratory mitochondrial function (Raghuram and Hughes, 2024).

      (5) Beyond just the presence/absence of MDCs, does more detailed quantification of their size/shape reveal any subtle differences between conditions?

      This is an interesting question. In our hands, MDC size and shape are variable and appear strongly influenced by mitochondrial fission/fusion state. Conditions that favor more fused mitochondrial networks can produce larger MDC-like structures, whereas fragmented networks can produce smaller structures. So far, we have not found a simple size or shape metric that explains the Yme1/Ups2/MICOS phenotypes better than MDC frequency.

      We will clarify this point in the revised manuscript and avoid implying that MDC frequency captures every possible morphological difference. More detailed morphometric analysis of MDC size, topology, and maturation state will be an important future direction, especially as we connect lipid remodeling to membrane curvature and MDC biogenesis.

      References

      Hughes, A.L., Hughes, C.E., Henderson, K.A., Yazvenko, N., and Gottschling, D.E. 2016. Selective sorting and destruction of mitochondrial membrane proteins in aged yeast. eLife. 5. doi: 10.7554/eLife.13943.

      Raghuram, N., and Hughes, A.L. 2024. Amino acids trigger MDC-dependent mitochondrial remodeling by altering mitochondrial function. bioRxiv. 2024.07.09.602707. doi: 10.1101/2024.07.09.602707.

      Wang, K., Jin, M., Liu, X., and Klionsky, D.J. 2013. Proteolytic processing of Atg32 by the mitochondrial i-AAA protease Yme1 regulates mitophagy. Autophagy. 9(11):1828–1836. doi: 10.4161/auto.26281.

      Xiao, T., English, A.M., Wilson, Z.N., Maschek, J.A., Cox, J.E., and Hughes, A.L. 2024. The phospholipids cardiolipin and phosphatidylethanolamine differentially regulate MDC biogenesis. Journal of Cell Biology. 223(5). doi: 10.1083/jcb.202302069.

    1. Author response:

      eLife Assessment

      This important study investigates the peptide-binding principles of promiscuous chicken MHC molecules. The data from crystallography, mass spectrometry, and modeling are convincing. However, the presentation would benefit from streamlining and clear links between data and conclusions. This paper will be of broad interest to immunologists and those interested in vaccine development.

      Overall, we are delighted and grateful to the eLife editors and the two reviewers for the careful and thoughtful assessments and reviews of our paper. We are glad that the strengths of the paper were apparent and appreciated. And of course, every paper has weaknesses, especially for a story as complex as this one.

      We are making only minor changes in our revision, so we would be happy if the editors decide to evaluate the revised manuscript without involving the reviewers further.

      Before answering the comments and questions directly, perhaps a few points would help clarify why the paper is as it is.

      First, the experiments cover over three decades of work, with the first gas phase sequencing results done in 1992. Unlike some of the chicken class I alleles which immediately gave completely clear stringent motifs (B4, B12 and B15 in Wallny et al 2006 PNAS, B19 in Han et al 2023 J Immunol), we harvested nothing but confusion from the B21 class I results (Fig. 1). Initially, we thought that the lack of a clear motif for B21 was due to multiple well-expressed class I molecules but only one dominantly-expressed class I molecule was found (Wallny et al 2006 PNAS, Shaw et al 2007 J Immunol) and, to our surprise, bacterially-expressed BF2*21:01 heavy chain and b2-microglobulin refolded with two synthetic peptides without sequence in common, and the crystal structures showed that this molecule remodeled the binding site to accommodate two such disparate peptides (Koch et al 2008 Immunity). This was the beginning of our understanding of the spectrum of class I alleles from promiscuous generalists to fastidious specialists, which we have explored in a series of further papers (in particular, Chappell et al 2015 eLife, Tresgaskes et al 2016 PNAS, Kaufman 2018 Trends Immunol, Tregaskes and Kaufman 2022 Mol Immunol).

      Second, over these many years, we continued to explore the binding properties of BF2*21:01 in ever more detail, resulting in the current manuscript. We learned only slowly how to probe this unexpected promiscuity, unprecedented in the MHC literature, so that the experiments proceeded with our best understanding at the time, including taking advantage of new approaches as they become available. Each experiment built on the previous set of experiments and each brought us closer to an understanding.

      Third, having amassed a collection of data, we chose eLIFE exactly because it allows us to present the entire story from beginning to end without compromise, not just the highlights with the major points illustrated by a few main figures and with the supporting data in many supplementary figures. We include all the data, because it is all part of the story, and so interested researchers to look at the data from their own perspective. Although mostly we provide bar graphs, we include the raw data (or close to it) for the final experiments (illustrated by Figs. 10 and 18) in the single supplementary data spreadsheet, so these can be assessed easily by others in the field, perhaps using approaches that we may not feel competent to perform.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Combining in vitro refolding, SEC-based assembly assays, peptide-library screening, MALDI-TOF, LC-MS/MS, structural analysis and immunopeptidomics, this manuscript investigates the peptide-binding principles of the promiscuous chicken MHC-I molecule BF2*21:01.

      Strengths:

      Although the peptide motif of BF2*21:01 is highly complex, this manuscript identified several principles, including a preference for 10-mer peptides, co-variation between P2 and Pc-2, effects of P3 and Pc-3, and a strong cellular preference for Leu at Pc. The results are important for avian MHC biology and poultry vaccine epitope prediction.

      Weaknesses:

      The manuscript is sometimes difficult to follow because the authors present a large amount of peptide-library, structural and immunopeptidomics data. without always clearly explaining how these datasets support the proposed simplifying principles.

      We are delighted and grateful to the reviewer 1 for the careful and thoughtful comments and questions concerning our manuscript. We are glad that the strengths of the paper were apparent and appreciated, and acknowledge the weaknesses that come with such a complex story with experiments performed over decades.

      Major Issues - Points Requiring Clarification or Additional Support:

      (1) (Line 282-301, 537-545)

      The immunopeptidomics conclusions are mainly based on one B21 cell line with one biological replicate and at least two technical replicates. Given the complexity of the BF2*21:01 peptide repertoire, this is a major limitation. The authors should either provide additional biological replicates or clearly state this limitation in the Abstract, Results and Discussion.

      This limitation is clearly stated in lines 537-545, as part of a paragraph covering the various ways in which the data presented in this manuscript could be improved. In fact, we have performed immunopeptidomics of several different B21 cell types, with many replicates and found similar data as presented, giving us confidence in our interpretations. However, these other experiments belong in different stories, so it is not appropriate that the data be reported in this manuscript.

      (2) (Lines 290-313)

      The B21 cell preparations contain both BF2 and the lowly expressed BF1 molecule. Some peptides, especially 8-mers or peptides with atypical motifs, may derive from BF1*21:01. The authors should clarify how BF2*21:01-bound peptides were distinguished from possible BF1-derived peptides, or interpret the immunopeptidomics motif more cautiously. The authors should also provide or cite evidence confirming the B21 haplotype identity of the cell line and chicken materials used for immunopeptidomics.

      The concern about the contribution of BF1*21:01 to the immunopeptidomics is clearly stated in the manuscript, both lines 290-313 and as part of the paragraph describing the limitations of the experiments (lines 542-543). In fact, the expression of BF1 molecules has long been known to be less than 10% of BF2 molecules at the RNA level, and much less at the protein level (Wallny et al 2006 PNAS, Shaw et al 2007 J Immunol). The proportion of 8mers identified by immunopeptidomics is also low (Fig. 14), and it is not impossible that most 8mers are due to BF1*21:01. We have used assembly assays with peptide libraries, immunopeptidomics and a crystal structure to determine the peptide motif for typical BF1 molecules, of which BF1*21:01 is one and found it may contribute to 8mer peptides but very seldom to longer peptides. This work is unpublished but gives us confidence that the characteristics of BF2*21:01 are not misrepresented by the data in this manuscript.

      The sources of the chicken samples and the cell lines are described in detail under Materials and Methods (lines 577-590), citing relevant publications. 

      (3) (Lines 217-221, 243-253)

      The authors acknowledge that MALDI-TOF cannot reliably distinguish peptide combinations with identical or similar masses, nor determine residue positions in some cases. Therefore, MALDI-TOF results should not be over-interpreted as precise evidence for residue preference. The authors should clearly indicate which conclusions are supported by LC-MS/MS.

      As described, the experiments follow each other in temporal sequence, so that we started with single peptides, then peptide libraries that varied in one position, then peptide libraries that varied in two positions first analysed by MALDI-TOF and later by LC-MS/MS. The final experiment (Fig. 10, with the original data in the supplementary spreadsheet) directly compares MALDI-TOF and LC-MS/MS results for six peptide libraries, so that the strength of the evidence for residue preference is clear. Throughout the manuscript, we do our best to not to overstate conclusions based on the data of any particular experiment.

      (4) (Lines 297-301, 316-330)

      The authors suggest that longer peptides may bulge in the middle or extend out of the groove at the C-terminal end. The rationale for the C-terminal extension is not clearly explained. Why is the C-terminal extension considered rather than the N-terminal extension? If the binding register is uncertain, long peptides should be analyzed separately from canonical-length peptides.

      When the first sequence of a chicken class I cDNA was determined, an immediate mystery was why one of the so-called invariant residues that coordinate the N- and C-termini of the bound peptide is not conserved (Kaufman et al 1992 J Immunol). In fact, this residue Tyr at position 86 in HLA-A2 and the equivalent position in all mammalian classical class I molecules is an Arg in the classical class I molecules of all non-mammalian vertebrates and is common with class II molecules (Kaufman et al 1995 Semin Immunol). Similar to class II molecules, this Arg in chicken class I molecules allows the peptide to extend out of the C-terminus, as shown by a crystal structure (Xiao et al 2018 J Immunol). The concern that we might be misidentifying the C-terminal amino acid was the basis for the analysis in Figs. 23 and 24, but in the absence of crystal structures, we are not able to provide a final answer this question. Perhaps relevant is the fact that a chicken class II molecule can bind exactly the same peptide in two conformations, one with a canonical 9mer core and the other with an unexpected 10mer core (Goryanin et al 2026 J Virol).

      By contrast, N-terminal extensions are only found for some class I alleles and thus far depend on the substitution of small amino acid sidechains for W166 (Li et al 2011 J Virol for bovine, Ma et al 2020 J Immunol for Xenopus, Wei et al 2022 J Immunol for ovine). Thus far, no chicken BF2 sequences have this substitution, consonant with the many crystal structures, including those for BF2*21:01 (Koch et al 2008 Immunity, Chappell et al 2015 eLlife, this manuscript). However, in unpublished data, we find that most BF1 sequences have sequence differences that could allow N-terminal extensions, although we have no crystal structures to support this possibility.

      (5) (Lines 406-439)

      In vitro assembly assays show that several hydrophobic residues can be tolerated at Pc, whereas immunopeptidomics shows a strong Leu preference at this position. The authors should clarify whether this Leu preference reflects intrinsic BF2*21:01 binding specificity, TAP-mediated peptide transport, antigen processing, peptide loading, or a cell-line-specific effect. Additional experimental support, such as TAP transport analysis, would strengthen this conclusion.

      The preference for Leu at the final position of the peptide by immunopeptidomics of the B21 cell line is strong but not absolute and is certainly affected at the least by the length of the peptide (Figs. 23 and 24). Unpublished immunopeptidomics results (mentioned above) show that this is not a cell line-specific result. The evidence from assembly assays of various peptides is that several hydrophobic amino acids are tolerated with sufficient stability of BF2*21:01 that they are detected in the assay (Figs. 3, 5, 9 and 10). Thermostability assays (Fig. 6) show that peptides with these same hydrophobic amino acids are stable to at least body temperature of chickens. These experiments show that such stability is peptide-dependent (that is, whether a particular amino acid is tolerated depends on the stability conferred by the rest of the peptide). Finally, peptide translocation assays using B21 cells have been done (Tregaskes et al 2016 PNAS) and show that peptides with several hydrophobic amino acids can be pumped into the lumen of the endoplasmic reticulum. However, the assays are with single synthetic peptides, so the data are not extensive enough to separate the effects of the final amino acid from the rest of the peptide. Certainly, peptides with amino acids other than Leu at the C-terminus can be translocated. So, it is not yet clear at which point the preference for Leu at the C-terminus of the peptide arises.

      (6) (Lines 172-178, 243-279, 442-457)

      The structural analysis explains some residue combinations, such as Arg at P2 with Glu at Pc-2 or Trp at Pc. However, the structural interpretation is not fully integrated with the large-scale peptide library and immunopeptidomics results. Representative high- and low-frequency combinations should be discussed structurally.

      Six crystal structures show that BF2*21:02 remodels the binding to accommodate a variety of anchor residues (Koch et al 2008 Immunity, Chappel et al 2015 eLife). These crystal structures are representative of sequences found by the immunopeptidomics from very frequent (H-E at roughly 15% 8-12mers) to moderately frequent (E-L at roughly 6% 8-12mers) to infrequent (N-F, A-D and E-D at roughly 1.5%, 1.6% and 0.7% 8-12mers) based on Fig. 18. All but one of the structures has Leu at the C-terminus, with the last one having Val which is found but not frequently by immunopeptidomics.

      Similar numbers are found by LC-MS/MS of double-substitution libraries of the two original peptide sequences in Fig. 10 with H-E found frequently (8.1% in P390, 3.8% in P498) and the others infrequently (0.1, 0.9, 1.0, 0.3% in P390, 0, 1.4, 1.0, 0.3% in P498), as calculated from the numbers in the Supplementary data spreadsheet. As discussed in the manuscript, for single-substitution peptide libraries of the two original peptides, Ile/Leu at the C-terminus was very frequent but at the same or slightly less level as Phe, with Met less frequent and Val even less so (Fig. 7).

      In addition, there are two more structures along with models explicitly testing some substitutions (Fig. 5). Attempting more current modelling approaches, we found AlphaFold 3 was unable to correctly predict most of the conformations that are found in the crystal structures of BF2*21:01, so we don’t feel confident in using them to predict unknown structures of this kind.

      (7) The inference of co-variation between P2 and Pc-2, as well as the modulatory effects of P3 and Pc-3, should be better explained. At present, some conclusions appear to be based mainly on residue-frequency patterns, and the logical connection between these observations and the proposed binding principles is not always clear. Statistical analyses, such as mutual information, chi-square tests or permutation tests, and representative structural explanations would strengthen this conclusion.

      We endeavored to do our best to explain the data, our interpretations and our reasoning, so we apologise if we have not managed to be as clear as might be desired. We have included as close to raw data as possible for the LC-MS/MS and MALDI-TOF (Fig. 10) and for the immunopeptidomics (Fig. 14 and 18) in the Supplementary Data spreadsheet, exactly so that competent practitioners can carry out further analyses (including the sophisticated statistical tests mentioned).

      Reviewer #2 (Public review):

      Summary:

      The study presents an in-depth analysis of the peptide repertoire bound by a promiscuous chicken MHC molecule using mass spectrometry, x-ray crystallography and modelling. While the MHC can bind a very diverse set of peptides, the authors have found some new rules that govern peptide binding to this MHC that could help to build a predictive model to study the repertoire of pathogen-derived peptides.

      Strengths:

      The study uses a range of well performed experiment across multiple techniques and provides an in-depth analysis of the peptide repertoire, including peptide sequences, length, preferred residues, stability and MHC presentation.

      Weaknesses:

      The data overall support the analysis and conclusion well. The only caveat is linked to Figure 4, which does not describe the stability of the peptide-MHC complex, but instead shows refold yield, and the two are not always linked.

      We are grateful for the clear understanding of the strengths of the work. With regards to Fig. 4, we agree with the reviewer that there are differences in refold yield but that measure may not be correlated with stability of the peptide-MHC complex. However, we were basing our interpretation of stability on the position and quality of the monomer peak, as illustrated by the trace in Fig. 2, in which a sharp peak at the monomer position represents a stable complex (as seen for the 10 and 11mer peptides) and later peaks represent unstable complexes falling apart during the chromatography (as seen for the 7, 8 and 9mer peptides).

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript entitled "Essential function reflected in the phylodynamics of a multigene family - the pir genes of malaria parasites" by Jackson and colleagues investigates the global phylogeny of pir genes across 14 Plasmodium species and one Hepatocystis species. The authors also focus on the functional characterization of the conserved ortholog pirC1 and claim that pirC1 is not the founder of the family and that it plays an essential role in blood-stage growth.

      Strengths:

      Overall, the manuscript is well written and interesting, as it combines comparative genomics and evolutionary analysis with functional experiments. The phylogenetic analysis is rigorous and represents a major strength of the manuscript.

      Weaknesses:

      The general conclusions regarding the potential function of this gene family are not fully supported by the data presented. The manuscript moves too quickly from growth phenotype and localization studies to a specific mechanistic model. The discussion argues that PIRC1 may be involved in nutrient acquisition, host sensing, or metabolic support, but the data provided do not directly support these functions, and the manuscript in its present form remains speculative. Although the manuscript includes some experimental results, it lacks direct mechanistic validation of the specific functions of the pir genes, including pirC1. In its current form, the study does not yet establish a definitive role for pirC1 in metabolic processes.

      The reviewer is correct that there is no definitive proof for the function of the PIRC1 protein. We speculate that this protein is involved in a metabolic process based on mutant phenotype – small, poorly developed parasites that do not produce the same amount of DNA as wildtype parasites (and hence likely fewer merozoites). That this occurs in an in vitro culture of Plasmodium knowlesi rules out a role in the interaction with the host organism, such as sequestration or facilitating passage through the spleen. The localization of the protein outside of the parasite is consistent with a role in nutrient uptake, but we agree that additional experiments are required to determine the role of the protein definitively. We aim to look at the differences in the transcriptome and the metabolome to gain more insight into the pirC1 phenotype; this should reveal metabolic deficiencies in the mutant parasite.

      Reviewer #2 (Public review):

      Summary:

      This is an extensive study using phylogenetic comparison across multiple plasmodium species to gain new insights in relation to their evolutionary pathways and the potential function of pir. In addition to establishing a framework to identify related orthologues across species as well as expanding paralogues families within a species, the work also focuses on understanding loss and gain of different PIRs and how this indicates a relative lack of functional constraints and essentiality for most members of the gene family.

      The authors provide evidence that at least pirC has a conserved function and plays an important role in parasite growth in multiple species.

      While this study represents a significant effort and does provide interesting new insights that would help our understanding of this complex gene family in the future, it has a number of limitations.

      Strengths:

      Extensive and thorough phylogenetic analysis that is supported by some biological validation. Provides an indication that the PIR gene family has limited biological constraints and evolved independently across different species, leading to rapid expansion and deletion of orthologous groups. Identified pirC as a functional and important member of the family that is conserved across the species.

      Weaknesses:

      The phylogenetic tree is based on a truncated sequence that focuses on the more conserved parts of the pir sequence. This could potentially lead to missing the key functional drivers of evolution. The biological validation of the role of pirC has some inconsistencies that need to be addressed.

      The reviewer is correct. We do not use the repetitive parts of the pir gene sequences for the phylogeny. We define these as the ‘distal variable’ and ‘proximal’ domains of the protein in Fig. S1, results text and supplementary results. We remove these parts from the alignment because they are only nominally homologous (they cannot be aligned) and so break the basic assumption of phylogenetic analysis. Amino acid repeats evolve quickly and are homoplasic (their similarities do not reflect ancestry) so omitting them is correct and makes the phylogeny more reliable. While these features do not contribute to the phylogenetic estimate, we propose in the results text and Fig. S3, in agreement with the reviewer, that they are an important demonstration of how pirs have differentiated and what is different between the subfamilies. The reviewer is also correct that we have considered the whole gene sequence when comparing Alphafold predictions and in selection analyses of closely related sequences (in these cases, the repeat sequences can be aligned).

      A structural prediction for the sequence used in the alignment would mostly reflect the distal conserved domain but would be misleading because the alignment combines conserved regions that are not physically attached in reality. We will clarify these points.

      Reviewer #3 (Public review):

      This paper aims to classify, from an evolutionary perspective, the multigene family PIR found in malaria parasites infecting rodents and Old World monkeys, and to link this classification to functional diversification. The authors also hypothesize that PIR members conserved across species play important roles in parasite survival, and seek to clarify their functions.

      To achieve these aims, the authors comprehensively analyze the evolution of PIR genes using genomic and transcriptomic information from many malaria parasite species. They focus on PIRC1, a member conserved across species, and attempt to clarify its function in rodent and simian malaria parasites by examining the phenotypes of parasites in which the corresponding genetic locus has been disrupted. They also attempt to determine its localization using PIRC1 tagged with an epitope sequence. However, although the locus-disrupted parasites appear to show an approximately 50% reduction in growth rate, this effect seems to be overestimated. Another weakness is that the cause of the reduced growth rate has not been clarified. The localization analysis also remains insufficiently conclusive.

      Therefore, I consider that the first half of the paper, consisting of the bioinformatics analyses, achieves the objective of comprehensively summarizing PIR and may become a reference paper for discussing the evolution and function of the PIR gene family. On the other hand, regarding the function of PIRC1, no clear conclusion can be drawn from the results presented, and several additional experiments are necessary.

      My major comments are as follows.

      (1) The claim that the failure of eight disruption attempts indicates that pirC1 is essential is too strong.

      Lines 319-321: The authors argue that a total of eight failed attempts to disrupt the pirC1 locus using two different construct designs suggest that pirC1 is essential in P. berghei. However, the failure of these attempts could also reflect technical issues with the construct design itself, such as the length of the homologous regions used for recombination, which are approximately 650 bp. Therefore, it is an overstatement to conclude that "pirC1 is essential for P. berghei blood-stage growth." Given that parasites with disruption of the corresponding locus could be obtained in both P. chabaudi and P. knowlesi, a more appropriate statement would be that "pirC1 is important for P. berghei blood-stage growth."

      It is correct that we cannot rule out that the inability to delete the pirC1 gene is Plasmodium berghei is unrelated to an essential function. We are happy to change the text to the suggested description.

      (2) The data on the mCherry-expressing P. berghei line shown in Supplementary Figure 11 are insufficient.

      (a) Panel C: Southern blot analysis

      To conclusively identify the lower band in panel C as chromosome 1, additional probes specific to genes located on chromosomes 1 and 2 would be required. In addition, a parental parasite control should also be included. The Southern blot image of the parental parasite should show only a single band at the higher position, with no band at the lower position. Probes specific to chromosomes 1 and 2 would help demonstrate that the lower band corresponds to chromosome 1, rather than chromosome 2.

      To this end, the authors could describe the result as follows:

      "In the parental parasite, only a single band corresponding to chromosome 7 was detected, indicating that the smaller chromosome was genetically modified. The size of the lower band detected with the dhfr probe was identical to that of the band detected with the control chromosome 1 probe, but distinct from that detected with the chromosome 2 probe, indicating that chromosome 1 was modified."

      That said, this chromosome-level Southern blot analysis is not sufficient to demonstrate that the target PBANKA_0100500 locus was specifically modified. The authors should provide more direct evidence showing that the PBANKA_0100500 locus, rather than another genomic locus, was modified. For example, Southern blot analysis after restriction enzyme digestion would provide more definitive evidence. Diagnostic PCR may also provide more specific evidence.

      Although we are confident that the parasites has been modified in the expected way, we are planning to generate PCR data confirming that the mCherry tag is correctly integrated into PBANKA_010050.

      (b) Panel D: Flow cytometry analysis

      To allow a more accurate interpretation of the percentage of mCherry-positive cells, flow cytometry data for the parental parasite line should also be presented.

      We will repeat the flow cytometry experiments and include a wildtype strain in the analysis.

      (3) There are unclear points in the PCR results shown in Supplementary Figure 12.

      Supplementary Figure 12: In panel B, a PCR product should also be amplified from dPCHAS_0101200 using the P1-P3 primer pair. Why is this band absent? The authors should provide the uncropped electrophoresis image so that the larger band can be seen. In addition, if labels 1 and 2 indicate independent clones, this should be stated in the figure legend.

      We will gladly supply the full, uncropped electrophoresis image and we will clarify what the numbers indicate in the legend.

      (4) The growth rates of P. chabaudi and P. knowlesi parasites with disruption of the PIRC1 gene locus should be quantitatively analyzed.

      The growth rates of P. chabaudi and P. knowlesi are described only qualitatively, but they should be evaluated quantitatively. In Figure 4A, the parasitemia of wild-type P. chabaudi increases from approximately 6.1% on day 6 to approximately 15.6% on day 8, corresponding to a 3.8-fold increase. However, because parasite growth may already be affected by immune-mediated suppression at this stage, this value should be regarded as a minimum estimate. In contrast, the mutant increases from approximately 3.2% on day 8 to approximately 6.8% on day 10, corresponding to a 2.1-fold increase. Based on these values, the daily growth rate of the mutant appears to be reduced to at least approximately 56% of that of the wild type. Similarly, from the growth curve of P. knowlesi in Fig. 5A, the DMSO-treated group appears to increase approximately two-fold per day, whereas the rapamycin-treated group increases only approximately one-fold per day. Thus, P. knowlesi also appears to show an approximately 50% reduction in growth rate. Taken together, both P. chabaudi and P. knowlesi appear to reproducibly show an approximately 50% reduction in growth capacity. A reduction of this magnitude is difficult to describe as a "severe growth defect"; a more appropriate wording would be simply that the parasites "showed a growth defect." In addition, the terms "a severe growth defect" and "essential" appear to be overstated throughout the manuscript, and the wording should be toned down. Finally, I recommend presenting Figure 4A and Figure 5A on a logarithmic scale so that the trend in growth rates can be more intuitively appreciated from the graphs.

      It should be possible to determine the growth rate of the wildtype and mutant P. knowlesi parasites. In addition, we can change the text to reflect that although there is a growth phenotype in the two species in which we obtained mutants, the parasites do have the capacity to replicate. Note that in the case of P. knowlesi, the parasites numbers in vitro do not increase, hence any additional factors that decrease the growth rate, such as immune system and spleen, will lower the reproductive rate further and render the mutant parasite unable to proliferate.

      (5) The evidence that disruption of the PIRC1 gene locus in P. knowlesi does not affect erythrocyte invasion is weak.

      The authors describe that "the developmental cycle of the parasites lacking PIRCl is slightly longer than that of parasites that produce PIRCl (line 383-384)," and appear to support this interpretation with data showing that "mutant parasites are significantly smaller than wild-type parasites (line 414)" and that "the DNA content in ML10-arrested parasites lacking PIRCl is lower than that of DMSO-treated parasites (line 417-418)" at 24 hours after invasion. However, a slightly longer developmental cycle alone does not seem sufficient to explain a 50% growth reduction.

      I think the erythrocyte invasion capacity has not been quantitatively evaluated, and therefore, the evidence supporting the conclusion that the phenotype of P. knowlesi parasites with disruption of the PIRC1 gene locus is unrelated to erythrocyte invasion is weak. The authors should assess invasion efficiency using purified merozoites. For P. chabaudi, it should also be possible to apply an in vitro or in vivo erythrocyte invasion assay similar to that used for other rodent malaria parasites, and this should be evaluated as well.

      We can further investigate the invasion phenotype of the mutant P. knowlesi parasites. The presence of a clear phenotype during the intraerythrocytic stage indicates that the protein also has a role after invasion, but we agree that determining the effect on invasion directly will be useful.

      Alternatively, the reduced DNA content in ML10-arrested parasites lacking PIRC1 (lines 416-417) could suggest that the number of merozoites formed per schizont may be reduced. To clarify this point, the authors should assess whether the number of merozoites per schizont is altered in P. knowlesi (and P. chabaudi parasites lacking PIRC1).

      We aim to count merozoites and the level of invasion, which will allow us to determine the reproductive rate of the mutant parasites.

      (7) The authors propose the possibility that PIRC1 expressed in merozoites is released after invasion; however, the evidence that PIRC1 localizes to intracellular organelles is weak.

      Line 333: "a peripheral pattern around the parasite" is indicative of parasite plasma membrane, PV, or PVM. ", indicative of a parasitophorous vacuole (PV) or parasitophorous vacuole membrane (PVM) location" should be amended to ", indicative of parasite plasma membrane, a parasitophorous vacuole (PV) or parasitophorous vacuole membrane (PVM) location". In the Figure S14 image, red signals are uniformly detected from the merozoites formed in the schizont stage parasite (not really microorganelle patterns), but not from the PVM surrounding the schizont, suggesting parasite plasma membrane localization, not PVM. I agree that the signal is detected from the compartments extending into the iRBC cytosol, which may be difficult to explain if it is located on the parasite plasma membrane, but how frequently were such images seen?

      To determine the localization of the protein in the merozoite, we will image P. knowlesi merozoites.

      Figure 4D. In the images of liver-stage schizonts, AMA1 does not appear to localize to the micronemes in mature merozoites, suggesting this image is an immature schizont. Although PIRC1 appears to be expressed in liver-stage schizonts, it is difficult to clearly determine whether it localizes to intracellular organelles or to the parasite plasma membrane.

      This is a valuable comment. It is difficult to impossible to determine the exact localization of the protein at this stage, irrespective of the exact stage of the parasite. It is clear from the images is that the protein is not secreted at this stage. The main aim of the experiment was to determine whether the protein is produced by the parasite during the liver stage, which the results confirm.

      To clarify the above points, the authors should examine whether PIRC1 is detected in intracellular organelles or around the merozoites by analyzing its localization in purified merozoites.

      This we aim to do.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable manuscript presents an open-source and low-cost acoustic system for quantifying biting and chewing in mice. The approach is carefully validated against human observers, demonstrating strong methodological reliability and enabling high-resolution analysis of feeding microstructure. The tool has broad relevance for studies of appetite circuits and pharmacological interventions. An important contribution is the identification of previously unrecognized "meal-related" neurons in the lateral hypothalamus, providing novel biological insight into solid food consumption. While the support for the methodological advances is compelling and robust, some circuit-level conclusions are preliminary or incomplete, relying on small pilot samples and manual classification, and should be interpreted with caution. This paper will be of interest to those interested in ingestive behavior and/or the hypothalamus.

      We thank the reviewers for their careful reading and constructive comments, which have substantially strengthened the manuscript. In the revised version, we have addressed every suggestion and introduced the following major additions: New experiments. We added one additional Vglut2 mouse to the calcium imaging cohort, achieving 386 neurons (Figure 8), and three naive Vgat mice with unilateral DREADD injections (Supplementary Fig. 5-1). New analyses. We performed ROC analyses on all feeding- and licking-related responses of n = 79 LH GABAergic and n = 386 LH glutamatergic neurons (Figures 7D-F and 8D-F). We also characterized the robustness of the Crunchometer to additive white-noise injection (Supplementary Fig. 1-2). New supplementary material. Three new supplementary figures have been added in total (Supplementary Figs. 1-2, 5-1, and 6-1). Supplementary Fig. 6-1 provides instructions for building a 1-Hz pulse generator that blinks an LED in synchrony with the video. Software improvements. We upgraded the original MATLAB scripts to an App GUI version, migrated the full codebase from MATLAB to Python, and packaged it as fully standalone executables for macOS (Apple Silicon) and Windows both of which run without a MATLAB license.

      Our point-by-point responses to the reviewers' comments are in red below. Deletions are omitted for brevity. We hope that the revisions fully address the points raised and render the manuscript suitable for publication.

      Public Reviews:

      Reviewer #1 (Public review):

      This is an interesting and valuable paper by Gil-Lievana, Arroyo et al. that presents an open-source method (the "Crunchometer") for quantifying biting and chewing behavior in mice using audio detection. The work addresses an important and unmet need in the field: quantitative measures of feeding behavior with solid foods, since most prior approaches have been limited to liquids. The authors make a clear and compelling case for why this problem is important, and I fully agree with their motivation.

      The system is carefully validated against human-scored video data and is shown to be at least as accurate, and in some cases more accurate, than human observers. This is a major strength of the study. I also particularly appreciate the demonstration of the technology in the context of LHA circuitry, which nicely illustrates its utility and importance for mechanistic studies of feeding. I also appreciate the ability to readily time-lock neural data to individual crunches. Overall, the manuscript is well-executed and represents a useful contribution to the field.

      We thank you for your appreciation of the Crunchometer and its alignment with ephys:

      To further facilitate alignment with neuronal activity, we have now also included a schematic diagram of the pulse generator used to blink an LED in synchronization with the video (see the new Supplementary Fig. 6-1).

      The comments I have are largely minor and should be straightforward to address:

      (1) The authors should report sample sizes for all mouse cohorts, either alongside the statistics or in the figure legends for mean data.

      We apologize for this oversight. We have now included all sample sizes in the figure captions.

      (2) Clarification is needed as to whether crunch detection fidelity is influenced by the hardness or softness of the food. The focus here is on standard pellets, with some additional high-fat pellet data, but it would be useful to know how generalizable the method is across different textures.

      We thank the reviewer for this important observation. Because the Crunchometer depends on bites generating an audible acoustic signal, food hardness directly impacts detection fidelity. Hard, brittle foods are readily detected, whereas soft foods such as jelly, pudding, or peanut butter are unlikely to produce a reliably detectable signal. This is a genuine scope limitation of the method, and we now make it explicit in the manuscript (see below).

      Regarding the two diets used in our study, Chow and HFD pellets differ only slightly in consistency, with HFD being marginally softer. These differences proved too subtle to separate acoustically: the intensity (dB) and spectral content of bites on the two diets were closely overlapping. Accordingly, when we trained an SVM on audio features alone, it could not reliably discriminate Chow from HFD bites.

      Importantly, the Crunchometer does not need to resolve food identity from sound, because audio and video play complementary roles in the system: the acoustic channel confirms that a bite occurred, while the mouse's position within the food-specific ROI determines which food was consumed. This division of labor is what allows per-diet attribution despite acoustically similar pellets.

      We have added to the Result section:

      “The Crunchometer, therefore, does not need to infer food identity acoustically: audio confirms that a bite occurred, and the mouse's position within a food-specific ROI identifies which food was consumed. This design enables per-diet attribution even for pellets with indistinguishable crunch signatures.”

      We fully agree with the reviewer that the study of solid-food consumption should not be restricted to standard murine diets. Foods with naturalistic textures, for example, the Granny Smith apple, chocolate, and salted peanuts used by O'Connell et al. (2025), span a much wider range of hardness and elasticity than Chow vs. HFD, and would likely generate more clearly differentiated acoustic signatures. We hypothesize that the Crunchometer could generalize to such foods to the extent that each food produces a clear and distinct acoustic pattern, and even where acoustic signatures overlap, ROI-based spatial attribution would continue to resolve food identity as long as each food is presented at a separate, trackable location.

      To make this scope explicit for readers, we have added the following clarification to the Behavioral Protocol section:

      "Our study is limited to the acoustic detection of standard Chow and HFD pellets, both of which exhibit a firm, brittle consistency. Future work should evaluate the fidelity of the Crunchometer across a broader range of food textures, encompassing varying degrees of hardness and elasticity, as explored by O'Connell et al. (2025)."

      (3) The authors should comment on how susceptible the Crunchometer is to background noise. For example, how well does it perform in the presence of white noise, experimenter movement, or other task-related sounds?

      We thank the reviewer for this valuable comment. The Crunchometer performs reliably in controlled, low-noise environments, but like any acoustic detection system, it is vulnerable to interference from sounds whose spectral content overlaps with the bite-related frequency band (500–950 Hz). To quantify this vulnerability, we stress-tested both the threshold-based and SVM-based detection methods by adding white noise to the original audio recordings at progressively decreasing amplitudes and measuring how detection performance degraded as the signal-to-noise ratio decreased. We found that the threshold-based method was more robust to white-noise contamination than the SVM-based method, maintaining acceptable detection performance at lower SNR values before degrading [see the new Supplementary Fig. 1-2].

      First, the white noise amplitude is generated as follows:

      Where L<sub>𝑛𝑜𝑖𝑠𝑒</sub> is the desired amplitude of the White Noise in dB. Then, the audio signal was range-normalized to its absolute maximum value, and the white noise was added with its desired amplitude, as shown by the following formula:

      (4) Chemogenetic activation of LHA GABAergic neurons is used. DREADD-based activation may strongly drive these neurons in a way that is not directly comparable to optogenetic or more physiological manipulations. While I do not think additional experiments are required, it would strengthen the discussion to briefly acknowledge this limitation.

      We thank the Reviewer for this thoughtful observation, which we agree with. Chemogenetic activation of LHA GABAergic neurons via DREADDs does not reproduce the physiological firing dynamics of these neurons along several dimensions: it imposes a sustained, tonic drive lasting hours after CNO administration; it likely produces firing rates above the endogenous range; and it lacks the fine temporal structure, phasic bursts, behaviorally- phased locked activity that these neurons exhibit during natural feeding episodes.

      We recognize, however, that this limitation is not unique to chemogenetics. Optogenetic approaches likewise fail to reproduce endogenous activity, as they impose synchronous, high-frequency activation patterns on a single cell type that are unlikely to occur under physiological conditions. Moreover, as we previously described in a phenomenon our laboratory termed optoception (Luis-Islas et al., 2022), optogenetic stimulation can itself generate signals perceptible to the animal, adding a further interpretive caveat. Thus, both techniques depart from physiological activity.

      For these reasons, we interpret our findings as evidence that activation of LHA GABAergic neurons is sufficient to drive the observed behavioral effects, without claiming that the endogenous firing pattern encodes these behaviors in the same manner or with the same dynamics imposed by our manipulation. We have now added a brief statement to the Discussion acknowledging this limitation explicitly:

      “A methodological consideration is that chemogenetic activation via DREADDs imposes a sustained, supra-physiological drive that does not reproduce the temporal structure of endogenous LHA GABAergic activity during feeding; optogenetic manipulations share analogous limitations (see optoception; Luis-Islas et al., 2022). Our findings, therefore, establish that activation of this neuronal population is sufficient to produce uncontrolled feeding and gnawing, without implying that its endogenous firing encodes them in the same manner.”

      Reviewer #2 (Public review):

      Summary:

      This manuscript introduces the Crunchometer, a low-cost, open-source acoustic platform for monitoring the microstructure of solid food intake in mice. The Crunchometer is designed to overcome the limitations of existing methods for studying feeding behavior in rodents. The goal was to provide a tool that could precisely capture the microstructure of solid food intake, something often overlooked in favor of liquid-based assays, while being affordable, scalable, and compatible with neural recording techniques. By doing so, the authors aimed to enable detailed analysis of how physiological states, drugs, and specific neural circuits shape naturalistic feeding behaviors.

      Strengths:

      The study's strengths lie in its clear innovation, methodological rigor in validation against human annotation, and demonstration of broad utility across behavioral and neuroscience paradigms. The approach addresses a significant methodological gap in the field by moving beyond liquid-based feeding assays and provides an accessible tool for precisely dissecting ingestive behavior. The system is validated across multiple contexts, including physiological state (fed vs. fasted), pharmacological manipulation (semaglutide), and circuit-level interventions (chemogenetic activation of LH neurons), and is further shown to integrate seamlessly with both electrophysiology and calcium imaging.

      (1) Introduces a low-cost, open-source acoustic tool for measuring solid food intake, filling a critical gap left by expensive and proprietary systems.

      (2) Makes the method easily adoptable across labs with detailed setup instructions and shared benchmark datasets.

      (3) Provides high temporal precision for detecting bite events compared to human observers.

      (4) Successfully distinguishes feeding microstructure (bites, bouts, IBIs, gnawing vs.

      consumption) with greater objectivity than manual annotation.

      (5) Demonstrates compatibility with electrophysiology and calcium imaging, enabling fine-scale alignment of neural activity with feeding behavior.

      (6) Effectively discriminates between fed vs. fasted states, validating physiological sensitivity.

      (7) Captures the pharmacological effects of semaglutide, although this is really just reduced feeding and associated readouts (bouts, latency, etc).

      (8) Has potential to distinguish consummatory vs. non-consummatory behaviors (e.g., food spillage, gnawing); however, the current SVM model struggles to separate biting from gnawing due to similar acoustic profiles, and manual validation is still required.

      (9) Provides potential for closed-loop experiments.

      Weaknesses:

      Several limitations temper the strength of the conclusions: the supervised classifier still requires manual correction for gnawing, generalizability across different setups is limited, and the neuroscience findings, particularly calcium imaging of GABAergic and glutamatergic neurons, are based on small pilot samples. These issues do not undermine the value of the tool, but mean that the neural circuit findings should be interpreted as preliminary.

      We sincerely thank the Reviewer for the careful and generous reading of our manuscript, and particularly for recognizing the methodological gap that the Crunchometer seeks to fill. We appreciate the acknowledgment that the tool's validation spans physiological, pharmacological, and circuit-level contexts, and that its integration with electrophysiology and calcium imaging was considered seamless. The Reviewer has also accurately identified the three main limitations of the current version of the platform, which we address in turn below:

      (1) The supervised SVM classifier still requires manual correction for gnawing.

      We agree with the Reviewer. The acoustic signatures of biting (consummatory) and gnawing (non-consummatory manipulation of the pellet) share overlapping linear spectrotemporal features that our SVM exploits for discrimination. This overlap reflects a genuine biomechanical similarity (both involve incisor contact with the pellet surface) rather than a shortcoming of the classifier per se. In ongoing work toward Crunchometer 2.0, we are addressing these limitations. The Crunchometer 2.0 will incorporate more sophisticated deep learning algorithms, such as ResNet, to better exploit non-linear features. Also, we are currently collecting a larger database of bite, gnawing, and environmental noise sounds across different setups, microphones, and conditions to build a more robust dataset for training new AI algorithms that can discriminate between gnawing and biting and generalize more robustly across microphones and behavioral setups. This effort will also be important for developing a closed-loop version of the Crunchometer to detect bites in real time and trigger an actuator (e.g., a laser). But we agree that, for the present manuscript, gnawing classification remains the weakest link in the pipeline.

      Nevertheless, we think that having a human in the loop is an advantage (not a disadvantage) of the equipment, as it improves the quality of database curation. No matter how sophisticated future algorithms become, human intervention will remain essential. To this end, we have now developed a human-validation GUI that further facilitates human revision of snippets through an intuitive, easy workflow, reducing human effort (Author response image 1).

      Author response image 1.

      The visual validator GUI allows a human to verify and reclassify snippets into the correct category in a friendly interface.

      (1) Generalizability across different setups is limited.

      This is a fair concern and one we have taken seriously, as noted above, and one we have already recognized. The acoustic signal captured by the Crunchometer is inherently sensitive to the geometry and material of the box, microphone placement, the ambient noise floor of the vivarium or experimental room, and the hardness of the specific pellet batch. To mitigate this, we have 1) released the full hardware specifications and bill of materials so that other laboratories can reproduce the acquisition geometry, and 2) provided the benchmark dataset and trained classifier weights so that groups using comparable setups can deploy the tool directly. We have already acknowledged that the SVM does not always generalize across setups. In this regard, we have now shown that the threshold method is more resistant to white-noise contamination (see new Supplementary Fig. 1–2) and, in our experience in the lab, it performs robustly across multiple setups and conditions we have tested. More importantly, improved algorithms are currently under development in our laboratory.

      (1) Some neuroscience findings (calcium imaging of GABAergic vs. glutamatergic neurons) are based on small pilot samples (n=2 mice per condition), limiting generalizability.

      (3) The neuroscience findings (calcium imaging of GABAergic and glutamatergic LH neurons) are based on small pilot samples.

      The Reviewer is correct, and we appreciate the comment. As noted in the manuscript, we explicitly state in the Results and Discussion that these findings are presented as preliminary. As the Reviewer noted, these findings do not undermine the value of the Crunchometer; we fully agree. The calcium imaging experiments were designed as a proof-of-concept to demonstrate that the temporal precision of the Crunchometer is sufficient to align neural activity with individual bite events, rather than as a definitive circuit-level characterization of LH GABAergic and glutamatergic populations during feeding. Nevertheless, we have now increased the number of Vglut2 mice by 1, bringing the total number of glutamatergic neurons to 386. We have now also performed a formal quantification of all the experiments recorded in Vgat (n=2, three sessions, 79 neurons) and Vglut2 (n=3, 6 sessions, 386 neurons). This new formal analysis uncovers neurons selectively tuned to liquid, solid, and both food types. A fully powered characterization of these two populations is underway in our laboratory, once funding arrives in the lab, and will be reported in a dedicated follow-up study.

      (2) Chemogenetic and pharmacological experiments used small cohorts, raising statistical power concerns.

      The chemogenetic experiments were conducted with a modest sample size (n = 4 bilaterally infected mice). Nevertheless, the data revealed a robust, reproducible behavioral effect consistent across all four subjects. The primary aim of this study was to illustrate the potential utility of the Crunchometer using complementary experimental approaches, including chemogenetic activation of GABAergic neurons in the lateral hypothalamic area (LHA). To further address this concern, we have now included three additional transgenic mice with unilateral infections and obtained results comparable to those of the bilateral condition. These new data are presented in a new supplementary figure comparing unilateral and bilateral infections (Supplementary Fig. 5-1). Notably, chemogenetic activation of LHA GABAergic neurons promoted eating-related consummatory behaviors to a similar extent under both unilateral and bilateral DREADD activation. Accordingly, we have now added the following text to the Results section:

      “Notably, unilateral DREADD infections in other naïve n=3 Vgat-cre mice yielded results comparable to bilateral infections. While the effect size was slightly reduced with unilateral administration, the difference between the two delivery methods was not statistically significant (Supplementary Fig. 5-1)”

      (3) Correlation with actual food intake is modest and sometimes less accurate than human observers.

      We agree that this result highlights the complexity of feeding behavior, influenced by factors such as hoarding and spillage. The threshold method detects feeding behavior solely based on the magnitude of bite-related sounds (e.g., when the mouse bites the pellet close to the microphone), whereas human observers incorporate additional visual information to infer feeding behavior even in the absence of detectable chewing sounds, introducing variability in detection criteria. Although the number of bouts identified by the Threshold method was comparable to those annotated by human observers, the estimated duration (Bout Size) of those detections differed. This discrepancy likely reflects some inconsistency in the detection criteria among human observers and delays in identifying the onset. Moreover, instances of mice chewing pellets without consuming them (i.e., spillage) were observed. These events were often misclassified as feeding bouts, resulting in false positives for both the threshold method and human observers.

      (4) Sensitive to hoarding behavior, which can reduce detection accuracy and requires manual correction for misclassifications (e.g., tail movements, non-food noises). However, these limitations are discussed and not ignored.

      We thank the reviewer for this constructive comment and for acknowledging that we explicitly discuss these limitations rather than overlook them. Indeed, gnawing and hoarding behaviors (together with tail movements and non-food noises) are factors that can reduce the accuracy of feeding detection. Even using the Crunchometer, an accurate measurement of solid-food consumption therefore remains challenging, which further supports the inclusion of a human-in-the-loop step to ensure a high-quality, well-curated database. Accordingly, we have added the following sentence to the Result section:

      "This human validation was essential for ensuring the high fidelity of our behavioral database and mitigating the inherent limitations of automated classification."

      Conclusion:

      Overall, this is an exciting and impactful methodological advance that will likely be widely adopted in the field. I recommend minor revisions to clarify the limits of classifier generalizability, better contextualize the small-sample neuroscience findings as pilot data, and discuss future directions (e.g., real-time closed-loop applications).

      We thank you for your constructive comments.

      Reviewer #3 (Public review):

      Summary:

      The manuscript provides detailed information on the construction of open-source systems to monitor ingestive behavior with low-cost equipment. Overall, this is a welcome addition to the arsenal of equipment that could be used to make measurements. The authors show interesting applications with data that reveal important neurophysiological properties of neurons in the lateral hypothalamus. The identification of previously unknown "meal-related" neurons in the LH highlights the utility of the device and is a novel insight that should spark further investigation on the LH. This manuscript and videos provide a wealth of useful information that should be a must-read for anyone in the ingestive behavior or hypothalamus fields.

      A scholarly introduction to the history and utility of various ways feeding is measured in rodents is provided. One point - the microstructure of eating solid food - has been studied extensively (for one of many studies, see https://doi.org/10.1371/journal.pone.0246569 ). However, I agree that the crunchometer will allow for more people to access recordings during food intake and temporally lock consummatory behavior to neural activity.

      Apologize for this oversight. This is indeed an important reference for the microstructure of eating solid food in a social context. We have now included it in the Introduction of this reference “Food intake in social contexts is a more ethologically valid model, in which radio-frequency identification (RFID) transponders enable the simultaneous assessment of feeding behavior across multiple mice in a single box (Rathod and Fulvio, 2021)”

      Questions on results:

      (1) It is unclear why 10% sucrose solution was used as a liquid instead of water, given that the study is focusing on the solid food source.

      One motivation for using sucrose rather than water alone was to create a highly palatable environment and to test whether mice would prefer palatable liquid sucrose over HFD. However, the choice of liquid stimulus will ultimately depend on the end user and the specific experimental conditions of each lab implementing the Crunchometer. Future versions of the apparatus could also incorporate multiple sippers to deliver several tastants alongside solid food.

      (2) It is unclear how essential the human verification is in the pipeline - results for Figure 1 keep referring to the verification as essential. Is that dispensable once the ML algorithms have been trained?

      Human validation, also referred to as a human-in-the-loop approach, is a deliberate design feature of the Crunchometer rather than a limitation (also see answer to Reviewer 2). The outputs of machine-learning algorithms, no matter how accurate, require expert corroboration to confirm or reject the specific behaviors under study, particularly when the behavioral repertoire is as heterogeneous as feeding (which encompasses sniffing, gnawing, biting, hoarding, and manipulating the food item). For this reason, we view human oversight as a safeguard for scientific rigor that remains valuable even as more advanced algorithms (e.g., deep learning and convolutional neural networks) are incorporated into future versions of the pipeline. As noted above, we have implemented a graphical user interface (GUI) that enables batch sorting and rapid inspection of multiple snippets (using a photographic montage view strategy), substantially reducing manual curation time.

      (3) The ability to extrapolate food quantity consumed is limited, with high variability. This limitation does not undercut the utility of the crunchometer, but should be highlighted as one of the parameters that are not suitable for this system. This limitation should be added to the limitations section.

      We thank the reviewer for this constructive observation. We fully agree that, although the Crunchometer reliably detects feeding events and their temporal microstructure (bouts, meals, and latencies), extrapolating absolute food quantity consumed from acoustic signals is indirect and carries substantial variability and should not be the primary readout for studies that require precise gravimetric measurements. As recommended, we have now explicitly listed this limitation in the Limitations section of the Discussion:

      "While the Crunchometer provides accurate temporal detection of bites and feeding microstructure, the estimation of absolute food mass consumed from bite-related acoustic signals shows considerable variability across trials and subjects. This limitation arises from individual differences in gnawing patterns, food fragmentation, and hoarding behavior. Accordingly, the Crunchometer is best suited for analyses of feeding dynamics and behavioral microstructure, whereas studies requiring precise quantification of ingested mass should complement the system with direct gravimetric measurements for example, real-time weighing of feeders."

      (4) The ability to discriminate between gnawing and consummatory behavior is a strength (Figure 5), and these findings are important. However, it is unclear what can be made of mice that have 'gnawing' behavior in the fasted state (like in Figure 3). It seems they would need to be eliminated from the analysis with this tool?

      We apologize for this misunderstanding. We have now more clearly indicated in Figure 3A that the cumulative feeding time reflects only Chow and HFD feeding bouts, excluding gnawing.

      We now state: “The lower panel shows the cumulative feeding time (only for Chow and HFD pellets, gnawing is excluded) over a two-hour session for the fed (green) and fasted (purple) groups (n = 6 mice).”

      Under normal physiological conditions, gnawing is an infrequent behavior in rodents. In our study, however, its frequency increased in the fasted state a change possibly attributable to heightened stress. This behavior was further exacerbated by chemogenetic manipulation, driving it to non-physiological levels.

      (5) Why is there a post-semaglutide fed group and not a fasted group in Figure 4? It seems both would have been interesting, as one could expect an effect on feeding even 24h after semaglutide treatment. This would help parse the preference better because the animals eat such a small amount of semaglutide, that it is hard to compare to the fasted condition with saline treatment.

      We thank the reviewer for this insightful suggestion. It would have been interesting to include a fasted post-semaglutide group, as it could provide relevant information about the lasting effect of an acute administration of semaglutide. However, we decided not to include this additional experimental condition because the semaglutide fasted mice displayed a markedly reduced food intake during the experimental session. An additional post-semaglutide fasted session would have required a prolonged food restriction (at least 24 hours), which we consider an unnecessarily stressful condition for the mice. Therefore, we decided to feed the mice once the experiment was completed. Nevertheless, we believe that comparing the food intake (grams) between the fed group shown in Figure 3C and the post-semaglutide fed group reported in Figure 4D provides insight into the lasting effect of semaglutide. The comparison reveals a remarkable reduction of food intake in the post-sem fed mice relative to the fed group, suggesting that the acute administration of semaglutide suppresses the feeding behavior for up to 24 hours.

      (6) The identification of 'meal-related' neurons in the LH is another strength of the manuscript. Although there is currently insufficient data, could similar recordings be used to give a neurophysiological definition of a 'meal' duration/size? Typically, these were somewhat arbitrarily defined behaviorally. Having a neural correlate to a 'meal' would be a powerful tool for understanding how meals are involved in overall caloric intake.

      We thank the reviewer for this insightful suggestion. We agree that the traditional behavioral criteria for defining meals, typically derived from log-survivor analyses of inter-pellet or inter-lick intervals, are operationally useful but ultimately arbitrary, and that a neurophysiologically grounded definition would be a valuable complement for the field.

      Our current dataset was not designed to formally establish such a definition, and we want to be cautious about the logic of the problem: validating a neural criterion solely against the behavioral one it would replace is circular. A genuinely neural definition of a meal would need to be anchored to independent criteria, for example, its ability to predict the latency and size of the subsequent meal, its correspondence with post-prandial satiety markers, or its response to anorectic agents such as GLP-1 receptor agonists. This is a methodologically nontrivial undertaking that we believe deserves a dedicated follow-up study.

      As preliminary evidence that such a problem is tractable, we note that the meal-related LH neurons identified here display sustained activity with onset and offset dynamics that broadly parallel the behaviorally defined meal boundaries (Figure 6), suggesting that meal structure is reliably encoded at the population level. A related approach, using neural activity to segment ingestive behavior at finer temporal scales, has been successful in our previous work on licking microstructure in the nucleus accumbens (Tellez, et al. 2012), and we consider the present findings a natural extension of that line of research to the larger meal timescale.

      (7) The conclusion in the title of Figure 8 is premature, given the pilot nature and small number of neurons and mice sampled.

      We appreciate this comment and agree with the reviewer. Accordingly, we have performed additional experiments on the Vglut2 glutamatergic population, in some cases using three-plane recordings, which substantially increased the yield to 386 glutamatergic neurons. As the reviewer anticipated, we observed a broad diversity of response profiles in this population, including neurons selective for liquid licking, for solid food intake, and for both food types. We also formally quantified these responses using ROC analysis, applying the same procedure to the Vgat GABAergic neurons (n = 79). These new findings have been incorporated into the revised manuscript (Results and Discussion). We thank the reviewer for prompting this extension of the analysis (see Manuscript).

      Conclusion:

      Overall, this report on the Crunchometer is well done and provides a valuable tool for all who study food intake and the behaviors around food intake. Clarification or answers to the points above will only further the utility and understanding of the tool for the research community. I am excited to see the future utility of this tool in emerging research.

      We sincerely thank the Reviewer for these kind and encouraging words, and for the constructive feedback provided throughout the review. The clarifications and additional analyses prompted by these comments have substantially improved the manuscript, and we share the Reviewer's enthusiasm about the potential of the Crunchometer to contribute to future research on feeding behavior.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) The authors have done a phenomenal job with the Introduction, highlighting the need for this tool, citing the history of feeding measurement systems and their relative strengths and weaknesses.

      Thank you for your comment; we greatly appreciate your positive feedback.

      (2) A limitation of Automated Pellet Dispensers is the possibility that the animals fail to consume the pellet after it has been retrieved from and registered by the device, potentially constraining accuracy.

      We address this issue in the Introduction, specifically, we wrote:

      “Current methods to monitor feeding behavior could be classified into four different classes…3) Automated Pellet Dispensers: Often integrated into operant conditioning chambers, these devices provide a controlled way of delivering food pellets. While devices like the open-source Feeding Experimentation Device (FED3) (Ali and Kravitz, 2018; Matikainen-Ankney et al., 2021), a pellet dispenser, are useful for measuring reinforcement, they alter the natural feeding patterns of mice, for example, requiring a simple action, such as a nose-poke can reduce overeating and weight gain in mice (Barrett et al., 2025). A further limitation is that FED3 may overestimate consumption if an animal retrieves and registers a pellet without actually consuming it. A significant strength of this method is its ability to enable closed-loop optogenetic stimulation concurrent with neuronal recordings.”

      (3) I really appreciate the data in Figure 2G, where they displayed the results of an "outlier" animal, as behavior is extremely variable, and it's useful to see how this system deals with the variability of the subjects. This is again highlighted by mouse number 5 in Figure 3A, which exhibited profound gnawing behavior.

      We thank the reviewer for this positive comment. Our decision to include the outlier animal in Fig. 2G and to report the atypical gnawing behavior of mouse 5 in Fig. 3A reflects a deliberate commitment to documenting inter-individual variability, which we consider a core strength rather than a limitation of behavioral work. We believe that such cases are particularly informative for evaluating the robustness of automated monitoring systems under behavioral-lab conditions.

      (4) It would be useful to know if the mice had prior exposure to HFD, as I found it surprising that many animals consumed the chow at all, sometimes completely ignoring the HFD (fasted mouse 3). I only ask because in our experience, mice with constant exposure to both HFD and chow predominantly, if not always, consume the HFD over chow. This could have something to do with the way the food substrates are presented in this chamber.

      We thank the reviewer for this point. Mice in this experiment did receive prior exposure to both Chow and HFD during the habituation phase, with at least two 30-min sessions in the experimental chamber with both diets available (no video was collected at this stage). The Chow and HFD feeders were identical in geometry, position, and accessibility, so we do not consider either environmental novelty or spatial bias to be the main driver of the pattern. Rather, we interpret the strong chow preference of fasted mouse 3 as a case of residual neophobia toward the HFD pellet. Since performing these experiments, we have refined our habituation protocol: pre-exposing animals to a single HFD pellet in their home cage, a familiar and safe environment, prior to any chamber session, greatly mitigates HFD neophobia in our hands. Familiarity with the novel food in a safe context thus appears to be the critical factor, rather than the duration of exposure in the experimental chamber. We have added this refinement to the Methods as a recommendation for future users of the Crunchometer.

      “Behavioral protocol. All mice were habituated to the Crunchometer for 2 days before the recording session. Each habituation session lasted 30 minutes, during which two food pellets were placed in the chamber: one standard Chow pellet (LabDiet 5008) and one highly palatable high-fat diet (HFD) pellet (Research Diet, D12451). As a practical note, we recommend allowing the HFD to equilibrate to room temperature before the experiment and pre-exposing mice to a single HFD pellet in their home cage to attenuate neophobia prior to testing.”

      (5) The authors claim saline or semaglutide was administered immediately before the start of the behavioral experiment, but given the time it takes for this drug to blunt appetite, I was somewhat surprised it led to such a rapid decrease in both chow and HFD intake. Could the authors comment on this? How quickly do these animals experience the malaise associated with these drugs? Also, this dose seems to be on the very high side, so I imagine it's making the animals feel quite sick and is probably a big reason why the effects last so long into the post-sem measurements. Was bodyweight tracked across this treatment? I'm not so convinced that sema treatment led to a loss of strong HFD preference, as the chow intake was already very low to begin with, and as mentioned above, it looks like the drug just led to a cessation of all intake. I'd just tamp down this claim of preference switch. It clearly reduced intake of both substrates, it's just harder to detect for the chow because it was already so low to begin with.

      Thank you for these comments. We agree with the Reviewer and have toned down the claim regarding a switch in HFD/chow preference. In the revised Results section, we now explicitly acknowledge that further characterization is needed using chronic semaglutide treatment. Specifically, we added the following sentence:

      "Future studies should use the Crunchometer to characterize changes in HFD/chow preference during 24-h monitoring under chronic semaglutide treatment."

      In addition, we administered a single subcutaneous dose of semaglutide at 30 nmol/kg (0.123 mg/kg), following the protocol described by Zhang et al. (2023). In their study, pharmacokinetic analyses showed that plasma concentrations, measured by an ELISA assay that immunoreacts with both growth differentiation factor 15 (GDF15) and the intact N-terminal region of glucagon-like peptide-1 (GLP-1), increased shortly after administration of the 30 nmol/kg dose in C57BL/6 mice. Peak plasma concentration (Cmax = 43.1 nmol/L) was reached at 6.7 hours (Tmax), and levels returned to baseline by 24 hours post-administration, indicating complete drug clearance. Although this dose is relatively high, it was intentionally selected to produce a robust acute response from a single administration, as our objective was to assess the drug’s effects within a short, 2-hour observational window. Under these conditions, we observed a rapid reduction in food intake immediately following the onset of Crunchometer recording. While we do not exclude the possibility that these effects could be more pronounced over longer observation periods or with chronic dosing regimens, our study was strictly limited to a single acute exposure.

      Although semaglutide is known to suppress food intake through multiple mechanisms, including stress and malaise measured by Conditioned Taste Aversion and release of stress hormones (Teixidor-Deulofeu et al., 2025), we do not believe that discomfort or malaise played a significant role in our study. While the mice did reduce their food intake during semaglutide administration, this reduction persisted for at least 24 hours after the final dose—at which point the drug was no longer present—suggesting a satiety-driven effect rather than one mediated by aversion. In this sense, previous studies have demonstrated that semaglutide continues to suppress food intake even when the aversive pathway mediated by Area Postrema GLP1R neurons is inhibited. Although blocking this pathway reduces flavour aversion, the anorexic effect remains, indicating that suppression of intake can be driven by satiety independently of nausea or malaise (Huang et al., 2024). In summary, although we selected a relatively high dose to ensure a detectable acute effect within our experimental window, this choice was grounded in previously published data, and our findings are consistent with established mechanisms of action for semaglutide.

      Additionally, body weight data have now been included in Figure 4D. We observed a similar body weight loss of approximately 5% on the first day of drug administration, consistent with the findings reported by Zhang et al. (2023).

      (6) The authors demonstrate that CNO administration prompted significant increase in liquid sugar intake in the last panel of Figure 5F as a confirmation that LH GABAergic neurons are implicated in processing reward, however given the above results it seems likely that these mice will drink anything including water (when not thirsty, thus in a non-rewarding scenario) or possibly aversive agents like quinine.

      This is an interesting question, and we agree with the Reviewer. The original discovery by Jennings and Stuber showed that optogenetic activation of these GABAergic neurons induces voracious feeding and that Vgat mice kept licking for liquid rewards in an appetitive task (Jennings et al., 2015). We also acknowledge that prior work has shown LH GABAergic neuron activation can drive consumption of non-caloric and biologically irrelevant stimuli, including wood gnawing, water, or saccharin (Navarro et al., 2016). However, several lines of evidence support a role in reward/palatability processing rather than purely indiscriminate consumption. Our own lab (Garcia et al., 2021) showed that activation of LH Vgat+ neurons increased quinine intake only during water deprivation; in sated animals, activation failed to promote quinine intake. Instead, these neurons promoted overconsumption of sucrose when available, leading us to conclude that LH Vgat+ neurons increase the drive to consume the nearest food, but this drive is potentiated by the palatability of the tastant. In non-human primates, LH GABA activation drives goal-directed eating predominantly for palatable food (Ha et al., 2024), supporting a reward-related function across species. Together, these findings indicate that while LH GABAergic activation does broadly promote consumption, the selectivity toward palatable stimuli observed in Figure 5F is consistent with a reward-related function.

    1. Author response:

      The following is the authors’ response to the previous reviews

      We thank the reviewers for their careful evaluation and constructive comments throughout the two rounds of revision. We hope that the revisions have satisfactorily addressed all concerns and that the manuscript is now suitable for publication.

      This novel contribution highlights the role of this pro-inflammatory factor in the pathogenesis of and resistance to Plasmodium chabaudi infection in mice. While aspects of this response have been previously described, this study is the first to link the TNF–iNOS–HIF-1α axis to the in vivo mediation of malaria disease through its involvement in glucose metabolism. Despite well-documented metabolic alterations during malaria, including hypoglycemia and hyperlactatemia, the mechanisms underlying these changes and their relationship to host immune responses remain poorly understood. Addressing this gap is essential for elucidating how metabolic adaptation shapes disease outcomes during Plasmodium infection.

      In response to the reviewer’s comments, we have revised the Abstract, Introduction, and Discussion to clearly distinguish between:

      Previously established mechanisms (TNF–iNOS–HIF-1α–glycolysis axis), and

      The novel contribution of our study (its in vivo integration during Plasmodium infection and association with host resistance).

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      The premise of the manuscript by Matteucci et al. is interesting and elaborates a mechanism via which TNFa regulates monocyte activation and metabolism to promote murine survival during Plasmodium infection. The authors show that TNF signaling (via an unknown mechanism) induces nitrite synthesis, which (via yet an unknown mechanism), and stabilizes the transcription factor HIF1a. Furthermore, that HIF1a (via an unknown mechanism) increases GLUT1 expression and increases glycolysis in monocytes. The authors demonstrate that this metabolic rewiring towards increased glycolysis in a subset of monocytes is necessary for monocyte activation including cytokine secretion, and parasite control.

      Strengths:

      The authors provide elegant in vivo experiments to characterize metabolic consequences of Plasmodium infection, and isolate cell populations whose metabolic state is regulated downstream of TNFa. Furthermore, the authors tie together several interesting observations to propose an interesting model regarding

      Weaknesses:

      The main conclusion of this work - that "Reprogramming of host energy metabolism mediated by the TNF-iNOS-HIF1a axis plays a key role in host resistance to Plasmodium infection" is unsubstantiated. The authors show that TNFa induces GLUT1 in monocytes, but never show a direct role for GLUT1 or glucose uptake in monocytes in host resistance to infection (nor the hypoglycemia phenotype they describe).

      We thank the reviewer for this important comment and for highlighting the need to clarify the mechanistic link between TNF-driven metabolic rewiring and host resistance to Plasmodium infection. As noted in our first revision, our primary objective was to investigate how TNF integrates systemic and cellular metabolic responses during infection in vivo. We demonstrate that glucose uptake is significantly increased in spleen and liver during infection in a partially TNF-dependent manner, and that TNF promotes GLUT1 expression (main glucose transporter in immune cells) and glycolysis specifically in monocytic cells. Importantly, to directly address the role of TNF signaling in myeloid cells, we also observed the same phenotype (higher parasitemia, but absence of hypothermia and hypoglycemia) in mice with conditional deletion of TNF receptor 1 in lysozyme M–expressing cells (TNFR1^ΔLyz2) (Figure 4P–R), thereby validating in a cell-specific context the findings previously observed in mice with global TNFR1 deficiency. Together, these findings support a functional link between TNF signaling in monocytes, induction of GLUT1-dependent glucose metabolism, and the regulation of both systemic metabolic responses and host resistance during experimental malaria.

      While we agree that we do not demonstrate a cell-intrinsic role for GLUT1 in monocytes, multiple lines of evidence in our study support the functional relevance of glycolytic metabolism downstream of the TNF–iNOS–HIF-1α axis.

      (1) First, we show that Pc infection results in a marked increase in glucose uptake in the spleen and liver, but not in skeletal muscle or adipose tissues (Figure 2K), and that this effect is absent in TNFR-/- mice (Figure 2L), indicating a TNF-dependent and tissue-specific metabolic reprogramming. We have also clarified in the Discussion that this process appears to be insulin-independent and likely driven by pro-inflammatory signals.

      (2) Second, we show that the TNF–iNOS–HIF-1α axis. induces GLUT1 expression in monocytic cells (Figures 4M, 5D, 6L). This supports a model in which these cells contribute to observed systemic metabolic changes.

      (3) Third, we also observed a similar phenotype—characterized by higher parasitemia but absence of hypothermia and hypoglycaemia-in mice with conditional deletion of TNF receptor 1 in lysozyme M–expressing cells (TNFR1^ΔLyz2) (Figure 4P–R), thereby validating in a cell-specific context the findings previously observed in mice with global TNFR1 deficiency. These findings indicate that disruption of glycolysis phenocopies key aspects of the TNF-driven metabolic and immunological response to infection. 

      (4) Finally, we demonstrate that glycolytic metabolism is functionally relevant for host resistance. Pharmacological inhibition of glycolysis in vivo using 2-DG led to increased parasitemia (Figure 6O), resembling the impaired parasite control observed in HIF-1α^ΔLyz2, TNFR-/-, and iNOS-/- mice. These findings indicate that disruption of glycolysis phenocopies key aspects of the TNF–iNOS–HIF-1α axis deficiency, supporting the conclusion that this pathway is required to sustain glycolytic metabolism and effective parasite control during infection.

      About the hypoglycemia phenotype and resistance, our previous study (PMID: 29805094) demonstrates that TNF-driven inflammation regulates systemic glucose metabolism during Plasmodium chabaudi infection. We showed that infection-induced hypoglycemia correlates with TNF levels and is associated with changes in parasite development. Specifically, leukocytes primed with IFNγ display increased expression of glucose metabolism and inflammatory genes, and TNFα-induced hypoglycemia is linked to the accumulation of non-proliferative trophozoite forms, whereas parasite replication (schizogony) occurs during host feeding. These findings indicate that blood glucose availability, regulated by TNF, directly influences parasite growth dynamics and infection outcome. Although the cellular mechanisms were not addressed in that study, our current work builds on these findings by identifying the TNF-iNOS–HIF-1α axis as a driver of GLUT1-dependent glycolysis in monocytes, linking systemic metabolic changes to a cell-intrinsic mechanism that contributes to host resistance. 

      We agree that directly establishing the cell-intrinsic contribution of GLUT1 would require dedicated genetic approaches (e.g., conditional deletion in monocytes), which are beyond the scope of the present study. 

      Comments on revisions:

      The demonstration that the established TNF-iNOS-HIF-1α-glycolysis axis operates in vivo during P. chabaudi infection is valuable and relevant. However, it constitutes contextual validation and must be carefully described as such. This distinction, i.e., "what has already been shown vs. what is new" is not consistently reflected in the framing of the manuscript raising overstatement concerns. This is particularly evident in the abstract and other conclusive statements, where mechanistic novelty is implied, even when the underlying pathways/mechanisms are already known. To improve the manuscript, all sentences that refer to already established findings should be accurately described as such.

      For example, the abstract states: "Here, we show that TNF signaling hampers physical activity, food intake, and energy expenditure while enhancing glucose uptake by the liver and spleen as well as controlling parasitemia in P. chabaudi-infected mice." In this sentence, the effects of TNF signaling on physical activity, food intake, energy expenditure, glucose metabolism and control of parasitemia are unequivocally established and therefore do not, in themselves, constitute new findings. Feeding behavior, not cell-intrinsic metabolism, may drive glycemic differences.

      We thank the reviewer for this comment and for highlighting the importance of distinguishing systemic metabolic effects from cell-intrinsic mechanisms. We have now revised the manuscript to more consistently distinguish between previously established mechanisms and our novel findings, particularly in the Abstract and other summary statements, to avoid any potential overstatement.

      We also would like to emphasize that, in both the Introduction and Discussion, we explicitly acknowledge that key components of the TNF–iNOS–HIF-1α–glycolysis axis have been previously described. In the Introduction, we cite studies demonstrating that TNF can induce glucose uptake and metabolic reprogramming in immune cells (refs. 14–17), as well as the role of HIF-1α as a central regulator of glycolysis and inflammation in myeloid cells (refs. 21–28). Similarly, in the Discussion, we detail prior evidence that TNF induces iNOS-derived RNI (refs. 51–54), that RNI stabilizes HIF-1α (ref. 52), and that HIF-1α drives the expression of glycolytic genes including GLUT1 (refs. 55–57). We also cite studies showing that TNF contributes to parasite control and glucose metabolism in malaria (refs. 58–61).

      Importantly, while these pathways have been described in other contexts, their integration and functional relevance in vivo during Plasmodium infection, particularly in the context of host systemic metabolism and monocytic cell function, have not been previously demonstrated. Our study addresses this gap by showing that this axis operates during P. chabaudi infection and links inflammatory signaling to both cellular metabolic reprogramming and organismal metabolic changes.

      Specifically, we demonstrate that TNF signaling drives increased glucose uptake in spleen and liver in a tissue-specific manner, promotes GLUT1 expression and glycolysis in monocytic cells, and that disruption of this axis (genetically or pharmacologically via glycolysis inhibition) impairs parasite control. In addition, we provide evidence connecting these cellular processes to systemic metabolic alterations, including hypoglycemia.

      The authors propose that TNF signaling leads to GLUT1 upregulation (in inflammatory monocytes, MO-DCs, and within the liver and spleen) during Plasmodium infection, and that this results in increased glucose uptake contributing to systemic hypoglycemia. While this is an intriguing hypothesis, we urge the authors to consider an alternative explanation that, at present, is not adequately ruled out. Given that glycemia serves as a central functional readout in the manuscript, this distinction is essential to clarify.

      The observed regulation of glycemia is likely not a direct consequence of increased glucose uptake by immune cells or by tissues but may instead reflect broader differences in disease severity across genotypes. The iNOS KO, TNFR KO, and HIF-1ΔLyz2 mice likely experience a dampened inflammatory response, which would blunt infection-induced anorexia and help preserve overall metabolic homeostasis. This alternate interpretation is supported by the authors' metabolic cage data showing increased physical activity in TNFR KO mice and the elevated food intake shown in Figure 2B.

      We thank the reviewer for this important point regarding the potential contribution of feeding behavior and systemic energy balance to the observed metabolic phenotypes. In fact, this possibility has been explicitly already incorporated into the revised manuscript. Also, we have revised the Discussion to explicitly state that the hypoglycemia observed during infection likely reflects both systemic changes in energy balance and TNF-driven metabolic reprogramming in immune cells, rather than a single isolated mechanism. Specifically, we have had already added the following statement to the Discussion:

      “Although restored physical activity, food consumption and energy expenditure in knockout mice may contribute to the observed systemic metabolic parameters by altering energy balance, these effects are not mutually exclusive with the TNF-driven, cell-intrinsic metabolic mechanisms described here”.

      In addition, we note that under naive conditions, we did not observe differences between genotypes in physical activity, food intake, energy expenditure, respiratory exchange ratio, or glycemia. These findings support that baseline metabolic parameters are comparable and that the differences observed during infection arise in the context of TNF-dependent inflammatory responses. During infection, although TNFR-deficient mice display increased food intake and activity, these differences arise in the context of altered inflammatory signaling. Therefore, rather than being mutually exclusive, behavioral and metabolic changes are likely coordinated downstream of TNF signaling.

      Furthermore, our data using pharmacological inhibition of glycolysis (2-deoxy-D-glucose) demonstrate that disruption of glycolytic metabolism results in increased parasitemia and reduced lactate levels, recapitulating key aspects of the phenotype observed in TNFR-/-, iNOS-/-, and HIF-1αΔLyz2 mice. This supports a functional role for glycolytic metabolism in host response, beyond differences in feeding behavior.

      Since anorexia and energy expenditure are tightly coupled to the inflammatory milieu, it is plausible that these behavioral and systemic differences-not monocyte nor tissue GLUT1 expression per se-are the primary contributors to the observed glycemic patterns. To support their current interpretation, the authors should perform a pair-feeding experiment in which (at least) TNFR KO mice are restricted to the same food intake as infected WT controls. This would help disentangle whether differences in glycemia truly reflect immune-driven metabolic rewiring or are secondary to differences in caloric intake.

      We thank the reviewer for this suggestion. We agree that pair-feeding experiments would provide an additional layer of control to isolate the contribution of caloric intake. However, we note that:

      (1) Baseline metabolic equivalence in naive animals argues against intrinsic differences in energy balance.

      (2) The observed phenotypes occur in the context of infection-driven inflammation, where anorexia is itself a TNF-dependent host response.

      (3) Our data support a model in which behavioral changes and metabolic rewiring are integrated components of the host response rather than independent variables.

      Importantly, our data already support a role for TNF-driven metabolic rewiring beyond feeding behavior, as inhibition of glycolysis with 2-deoxy-D-glucose recapitulates the impaired parasite control observed in genetic models. In addition, as discussed in the manuscript, systemic factors such as food intake are not mutually exclusive with cell-intrinsic metabolic mechanisms.

      We therefore consider that pair-feeding experiments are beyond the scope of the present study.

      The contribution of monocyte-specific glucose metabolism to host resistance remains unresolved.

      We appreciate the authors' effort to address the mechanistic role of glycolysis in host resistance using in vivo 2-deoxyglucose (2DG) treatment. However, I would like to point out that while this experiment is informative, it does not fully resolve the specific concern raised regarding the cell-intrinsic role of TNF-induced glycolysis in monocytes. 2DG acts systemically, inhibiting glycolysis across a wide range of cell types-including hepatocytes, endothelial cells, lymphocytes, and myeloid populations. Therefore, the observed increase in parasitemia following 2DG treatment may reflect the broad importance of glycolysis for host defense, or alternatively, may result from elevated circulating glucose levels induced by 2DG (PMID: 35841892), which could enhance parasite growth by increasing nutrient availability. Therefore, this experiment does not allow for a specific conclusion about the requirement for TNF-driven metabolic reprogramming in monocytes.

      We thank the reviewer for this comment regarding the interpretation of the 2-deoxyglucose (2DG) experiments. We agree that systemic 2DG treatment does not allow cell-specific conclusions, as it broadly inhibits glycolysis across multiple cell types. Accordingly, these data are interpreted as supporting a role for glycolysis in host defense at the organismal level, rather than as direct evidence for a monocyte-intrinsic requirement of TNF-driven metabolic reprogramming.

      At the same time, our study includes cell-specific analyses that support the engagement of this pathway in myeloid populations. In particular, we observe increased GLUT1 expression in CD11b<sup>+</sup> cells within both the liver and spleen during infection, with marked upregulation in monocyte-derived dendritic cells (MODCs). Importantly, this induction is not observed in the corresponding knockout models, supporting the idea that TNF signaling is required for this metabolic adaptation in these cells in vivo. Consistent with this, we validated that both parasitemia and systemic glucose levels in TNFR1^ΔLyz2 mice phenocopy those observed in TNFR-deficient animals, reinforcing the contribution of myeloid TNF signaling to the metabolic and disease outcomes.

      In addition, our in vitro data demonstrate increased GLUT1 expression in WT monocytes but not in cells lacking components of the TNF–iNOS–HIF-1α axis, further supporting a pathway-specific effect. Given that GLUT1 is the primary glucose transporter in immune cells, these combined in vivo and in vitro findings, together with the 2DG experiments, provide strong evidence supporting our proposed model. 

      We agree that directly establishing a monocyte-intrinsic role would require targeted genetic approaches, which are beyond the scope of the present study.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Statement

      This valuable study characterizes the emergence of the membrane-associated periodic cytoskeleton (MPS) in the axons of human motor neurons derived from induced pluripotent stem cells. Super-resolution imaging of beta-II spectrin provides convincing evidence for the patterned assembly of spectrin-poor gaps and spectrin-rich MPS in the medial region of the axons and its enhancement by the kinase inhibitor staurosporine. The data advocates against gap formation by cytoskeleton disassembly in a continuous MPS. Instead, a continuous MPS may result from nascent MPS patches and their maturation, a model that would benefit from live imaging for validation.

      (R1) We thank the reviewers and editor for their constructive and thoughtful feedback. We are pleased the reviewers found our evidence to be convincing and that our study provides a valuable framework for understanding the complex dynamics of MPS assembly.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Ever since the surprising discovery of the membrane-associated Periodic Skeleton (MPS) in axons, a significant body of published work has been aimed at trying to understand its assembly mechanism and function. Despite this, we still lack a mechanistic understanding of how this amazing structure is assembled in neuronal cells. In this article, the authors report a "gap-and-patch" pattern of labelled spectrin in iPSC-derived human motor neurons grown in culture. The mid-sections of these axons exhibit patches with reasonably well-organized MPS that are separated by gaps lacking any detectable MPS and having low spectrin content. Further, they report that the intensity modulation of spectrin is correlated with intensity modulations of tubulin as well. However, neurofilament fluorescence does not show any correlation. Using DIC imaging, the authors show that often the axonal diameter remains uniform across segments, showing a patch-gap pattern. Gaps are seen more abundantly in the midsection of the axon, with the proximal section showing continuous MPS and the distal segment showing continuous spectrin fluorescence but no organized MPS. The authors show that spectrin degradation by caspase/calpain is not responsible for gap formation, and the patches are nascent MPS domains. The gap and patch pattern increases with days in culture and can be enhanced by treating the cells using the general kinase inhibitor staurosporine. Treatment with the actin depolymerizing agent Latrunculin A reduces gap formation. The reasons for the last two observations are not well understood/explained.

      (R2) We thank the reviewer for the detailed and accurate description of the data shown and its relevance to further our understanding of MPS assembly mechanism and function.

      Strengths:

      The claims made in the paper are supported by extensive imaging work and quantification of MPS. Overall, the paper is well written and the findings are interesting. Although much of the reported data are from axons treated with staurosporine, this may be a convenient system to investigate the dynamics of MPS assembly, which is still an open question.

      (R3) We thank the reviewer for the positive comments on the manuscript and the convenience of the experimental system developed to further study the dynamics of MPS assembly. We hope others turn into motor neurons to explore cortical cytoskeleton biology and hopefully shed light into their susceptibility in various degenerative diseases.

      Weaknesses:

      Much of the analysis is on staurosporine-treated cells, and the effects of this treatment can be broad. The increase in patch-gap pattern with days in culture is intriguing, and the reason for this needs to be checked carefully. It would have been nice to have live cell data on the evolution of the patch and gap pattern using a GFP tag on spectrin. The evolution of individual patches and possible coalescence of patches can be observed even with confocal microscopy if live cell super-resolution observation is difficult.

      (R4) Because staurosporine may hit various kinases relevant to the phenomenon under study we did not elaborate too deeply on the likely targets in the discussion. We have, however, included the possibility that the relevant kinase in this matter could be PKC, in light of the new study published while our manuscript was under revision (Heller et al., 2025) (see second last paragraph in the Discussion section). Staurosporine represented a convenient initial approach that allowed us to find the phenomenon, and we are now conducting new studies dissecting the molecular pathways involved. However, the extent of such studies lies beyond the scope of the present report.

      See R16 regarding possible live-imaging experiments using tagged βII-spectrin constructs.

      Some more comments:

      (1) Axons can undergo transient beading or regularly spaced varicosity formation during media change if changes in osmolarity or chemical composition occur. Such shape modulations can induce cytoskeletal modulations as well (the authors report modulations in microtubule fluorescence). The authors mention axonal enlargements in some instances. Although they present DIC images to argue that the axons showing gaps are often tubular, possible beading artefacts need to be checked. Beading can be transient and can be checked by doing media changes while observing the axons on a microscope.

      (R5) As we acknowledge this possibility, we believe that, even if they occurred, they could not contribute to our observations of gaps-and-patches phenomenon since this latter subsisted long (hours and days) after any gross manipulation of media. Moreover fixed samples, when observed under DIC, confocal or STED did not evidence such beadings. We do refer to a characteristic local enlargement that was very localized and very low in numbers (see Fig.1C and E, and Suppl. Fig1C and E), so we don't believe these are transient, and do not resemble the structure referred to as beading. Structurally, beading is essentially different since it appears in rows of consecutive “beads” in long stretches, where round, small enlargements of axonal caliber are arranged in a consecutive manner, resembling pearls on a string. As mentioned by the reviewer, the beading phenomena can occur transiently when drastically changing media osmolarity (rarely done in cell culture manipulations) or non-tranciently when axons are undergoing degeneration. Indeed, to prevent gross changes in osmolarity, our routine fixation is a 4% PFA and 4% sucrose in PBS. In any case, we did not observe signs of beading in the cultures used for this study.

      (2) Why do microtubules appear patchy? One would imagine the microtubule lengths to be greater than the patch size and hence to be more uniform.

      (R6) Our stainings are for tubulin protein isoforms beta-III and alpha-II. That is, they would label microtubules, but free tubulin as well. Hence we don't think this is evidence for “patchy microtubules”. The slight decrease in intensity for tubulin within gaps is indeed something to investigate, and can indicate that tubulin prefers to accumulate within patches.

      (3) Why do axons with gaps increase with days in culture? If patches are nascent MPS that progressively grow, one would have expected fewer gaps with increasing days in culture. Is this indicative of some sort of degeneration of axons?

      (R7) We agree with the apparent discrepancy. However, one has to take into account that these axons are still elongating even at 2 weeks in culture and beyond. Hence, at any time point, there is a new axonal compartment recently added, and hence, with low βII-spectrin and no organized MPS. Also, the dynamical evolution of the gaps-and-patches structure has to take into account the rate of βII-spectrin supply and transport. If supply is somehow lower than a given threshold, it is expected that there will be more gaps, given the new, more distant parts of the axons have a lower supply of βII-spectrin. To explore this formally, we are working on simulations of these multifactorial dynamic systems to better understand this, that together with key experimental observations would enhance our understanding into our model of MPS assembly in growing axons. However, findings for this project will be the subject of another manuscript.

      (4) It is surprising that Latrunculin A reduces gap formation induced by staurosporine (also seems to increase MPS correlation) while it decreases actin filament content. How can this be understood? If the idea is to block actin dynamics, have the authors tried using Jasplakinolide to stabilize the filaments?

      (R8) The results with the co-treatment with Latrunculin A and Staurosporine are indeed intriguing, and provide clear evidence that the gap-and-patch pattern arises from local assembly of the MPS, requiring newly formed actin filaments. On the other hand, the fact that F-actin within the pre-formed MPS seems unaffected is not surprising. There are many different populations of F-actin in axons (i.e. MPS rings, longitudinal filaments, actin patches, actin trails), all of which have a different rate of monomer turnover. Latrunculin A affects filaments indirectly. The target of Latrunculin A is not actin filaments, but free monomers. Monomer sequestration ultimately affects actin filaments: filaments are constantly exchanging monomers, but, devoid of free monomers, filaments get shorter and eventually disappear. The drastic decrease in global F-actin in LatA-treated axons reflects that. The fact that F-actin in the MPS is preserved shows that these filaments are stable -if they are not losing monomers in the time frame of the treatment, the filament remains unaffected. This subject is extensively covered in the 8th paragraph of the Discussion section.

      We have not used Jasplakinolide. The expected outcome will not mimic that of Latrunculin A since Jasplakinolide has a different mechanism of action (i.e. it binds -and stabilizes- the actin filament).

      (5) The authors speculate that the patches are formed by the condensation of free spectrins, which then leaves the immediate neighborhood depleted of these proteins. This is an interesting hypothesis, and exploring this in live cells using spectrin-GFP constructs will greatly strengthen the article. Will the patch-gap regions evolve into continuous MPS? If so, do these patches expand with time as new spectrin and actin are recruited and merge with neighboring patches, or can the entire patch "diffuse" and coalesce with neighboring patches, thus expanding the MPS region?

      (R9) We agree with the reviewer's interpretation. A virtue of our experimental model and our interpretations of the observations in fixed cells is that it gives rise to informative questions such as the ones posed by the reviewer. See R16 regarding possible live-imaging experiments using tagged βII-spectrin constructs.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Gazal et al. describe the presence of unique gaps and patches of BetaII-spectrin in medial sections of long human motor neuron axons. BII-spectrin, along with Alpha-spectrin, forms horizontal linkers between 180nm spaced F-actin rings in axons. These F-actin rings, along with the spectrin linkers, form membrane periodic structures (MPS) which are critical for the maintenance of the integrity, size, and function of axons. The primary goal of the authors was to address whether long motor axons, particularly those carrying familial mutations associated with the neurodegenerative disorder ALS, show defects in gaps and patches of BetaII-spectrin, ultimately leading to degradation of these neurons.

      (R10) We thank the reviewer for the detailed and accurate description of the data shown.

      Strengths:

      The experiments are well-designed, and the authors have used the right methods and cutting-edge techniques to address the questions in this manuscript. The use of human motor neurons and the use of motor neurons with different familial ALS mutations is a strength. The use of isogenic controls is a positive. The induction of gaps and patches by the kinase inhibitor staurosporine and their rescue by Latrunculin A is novel and well-executed. The use of biochemical assays to explore the role of calpains is appropriate and well-designed. The use of STED imaging to define the periodicity of MPS in the gaps and patches of spectrin is a strength.

      (R11) We thank the reviewer for the positive comments on the manuscript, the techniques used and the proposed model.

      Weaknesses:

      The primary weakness is the lack of rigorous evaluation to validate the proposed model of spectrin capture from the gaps into adjacent patches by the use of photobleaching and live imaging. Another point is the lack of investigation into how gaps and patches change in axons carrying the familial ALS mutations as they age, since 2 weeks is not a time point when neurodegeneration is expected to start.

      (R12) See R16 regarding possible live-imaging experiments using tagged βII-spectrin constructs.

      We don't discard the notion that axons carrying familial ALS mutations will show defects in MPS formation and/or stability when observed at longer culture times, or under culture conditions that promote neuronal aging (Guix et al., 2021). Thus, we continue to work with these cells, but the goal of such project lies well beyond the primary message of the present manuscript, as we discuss in the second paragraph of the Discussion section.

      Reviewer #3 (Public review):

      Summary:

      Gazal et al present convincing evidence supporting a new model of MPS formation where a gap-and-patch MPS pattern coalesces laterally to give rise to a lattice covering the entire axon shaft.

      Strengths:

      (1) This is a very interesting study that supports a change in paradigm in the model of MPS lattice formation.

      (2) Knowledge on MPS organization is mainly derived from studies using rat hippocampal neurons. In the current manuscript, Gazal et al use human IPS-derived motor neurons, a highly relevant neuron type, to further the current knowledge on MPS biology.

      (3) The quality of the images provided, specifically of those involving super-resolution, is of a high standard. This adequately supports the conclusions of the authors.

      (R13) We thank the reviewer for the positive comments on the manuscript, the techniques used and the proposed model.

      Weaknesses:

      (1) The main concern raised by the manuscript is the assumption that staudosporine-induced gap and patch formation recapitulates the physiological assembly of gaps and patches of betaII-spectrin.

      (R14) Along the project, various gaps-and-patches parameters were measured in different conditions and stainings. In all these examinations the only parameter that changed considerably was their abundance. While this suggests that the gaps-and-patches features are comparable between control and staurosporine-treated cells, we acknowledge as a general caution regarding negative data—that subtle qualitative differences cannot be entirely ruled out. We have now emphasized this possibility in the 9th paragraph of the Discussion section.

      (2) One technical challenge that limits a more compelling support of the new model of MPS formation is that fixed neurons are imaged, which precludes the observation of patch coalescence.

      (R15) See R16 regarding possible live-imaging experiments using tagged βII-spectrin constructs.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The reviewers all agree that the work would strongly benefit from live imaging to assess the maturation dynamics of the gap/patch pattern.

      (R16) Reviewers agreed that some of the conclusions of our manuscript would benefit from live imaging for validation. Various anticipated technical and biological challenges made these approaches not to be conducted for this initial study on human motor neurons. Just to mention the most important, from previous work of our labs, these cells themselves are difficult to transfect at 2 weeks in culture. Also, ectopically expression of tagged βII-spectrin escapes normal expression control and it has been noticed that ectopic expression yields to protein localization that does not necessarily reflect the endogenous distribution, or that produces cellular responses that precludes the observation of the phenomena under study. These difficulties in studying over-expressed tagged βII-spectrin have been reported in the field, with mentions that the analysed axons were those expressing “low levels of the construct” (Boyer et al., 2026; Zhong et al., 2014; Zhou et al., 2022). Taking this into account, we did not anticipate that, for the goals of the present project, live-imaging was to be included. However, given the positive comments and reception of our conclusions, we sought to try to perform this challenging and risky approach. To that end, we used a C-terminus tagged mouse βII-spectrin-GreenLantern plasmid to transfect our cells (a kind gift from Dr. Subjohit Roy, UCSD, USA). After 3 rounds of differentiating cells and trying various combinations of plasmid quantity, lipofectimine-to-DNA ratios and times of transfection (amongst other parameters), we have got an extremely low efficiency of transfection, and the few expressing neurons showed a distribution of βII-spectrin-GreenLantern that did not match our observations of immunolocalization of endogenous βII-spectrin. Taking all these into account, the present version of the manuscript will not include live-cell imaging on expressed tagged βII-spectrin. Given that reviewers found that some statements in the initial submission would have been better supported by live-imaging, we made changes in the manuscript so as to acknowledge the limitations of concluding dynamic mechanisms from fixed samples (see for example last sentences on 5th paragraph of the Discussion section). Having said so, we hope to be able, in the future, to overcome these experimental challenges and be able to establish live-imaging of βII-spectrin in neurons. For example, to avoid unregulated transgene expression, Heller and colleagues recently generated a βII- spectrin-mNeonGreen conditional knock-in (cKI) mice, consisting of a LoxP- flanked alternative final exon of endogenous βII-spectrin with a C- terminal mNeonGreen fusion that is expressed upon Cre expression (Heller et al., 2025). The implementation and further development of such approaches will be very helpful in new studies on the dynamics of βII-spectrin and the MPS as a whole. However, the scale of work needed to accomplish those approaches represent stand-alone projects.

      Reviewer #1 (Recommendations for the authors):

      In the section "The MPS is absent in beta-II spectrin gaps, the authors mention that the presence of MPS in patches suggests that the axons are not undergoing degeneration. I don't think this is a good criterion to use, despite the citations they take support from.

      (R17) We agree with the reviewer's suggestion: in virtue of the unlikely connection between the cited developmental axon degeneration process in sensory neurons and the possible axon degeneration of long term cultures of human-iPSCs-derived motor neurons studied here, we have eliminated the sentence of reference

      The authors show that degradation by proteases does not happen in their case. In this regard, they may want to discuss the recent article by Heller et al, Science 2025 (https://doi.org/10.1126/science.adn6712) and Hofmann et al, Sci. Rep., 2022 (https://doi.org/10.1038/s41598-022-18562-5)

      (R18) By western blot analysis, we did not see evident changes in proteolysis-derived fragments. However it is likely that even when finding phenotypes with protease inhibitors, protein fragments accumulation is below the sensitivity of western blots. We were expecting gross changes observable by western blot in the case proteolysis explained gap formation.

      Calpain and Caspase activity has been shown to be relevant in different aspects of MPS biology. To the works cited by the reviewer, now one has to add the very recent work by Fei and colleagues (Fei et al., 2026). We have modified part of the Discussion section to analyse our results in this broader context.

      Briefly, Hofmann and colleagues found that acute treatment with calpain inhibitors right before axotomy lead to an increase in percentage of periodic βII-spectrin (referred by authors as “periodicity”) in the regenerated axons in a 2-hour period. Interestingly, the βII-spectrin patches they describe at distal portions did not increase in number, but they increased in size. This indicates that in the particular situation of axonal regeneration calpain activity puts a brake into MPS formation within patches. This invited us to re-examine our own protease inhibition experiments, and measured patch length in this. The new results are shown in Supplementary Fig. 6 and and further analysed in the Discussion section. In summary, our changes were much less notable than the ones found in regenerating axons, but follow the same trend: protease inhibitors made patches longer.

      On the other hand, Heller and colleagues found in live-imaging studies that calpain activity contributes to the steady-state dynamics of βII-spectrin exchange in a mature MPS lattice. More recently, Fei and colleagues found that caspase or calpain inhibition does not change the steady-state organization of a mature MPS lattice when observing treated axons after fixation samples. Fei and colleagues find a relevant role for calpains whenever massive endocytosis (of any kind) is engaged experimentally. Interestingly, all these studies, including ours, examined calpains roles in MPS in different scenarios. When looked in detail, we don’t believe that these are contradictory results among them, and a complete picture of calpains (and caspases) roles in MPS assembly, growth, maintenance and remodeling will have to take into account all the above mentioned results, including ours. All these analyses are now included in the Discussion section.

      Minor comments:

      (1) "Recently, it was proposed that this continuous MPS organization arises from the coalescence of discontinuous "patches" of incomplete MPS units that originate in the distal axon and migrate proximally (Zhong et al. 2014)." Please check the citation. Should it be Hoffman et al. 2022?

      (R19) The reviewer is correct. The proper citation has now been included.

      (2) Is there an established link between ALS and spectrin? I would suggest decreasing the emphasis on this as no clear conclusions are achieved.

      (R20) As stated in the text, the study of ALS mutations is justified from two aspects: one aspect is that there are several tubulin and other cytoskeletal proteins whose mutations are linked to ALS (Castellanos-Montiel et al., 2020) and microtubules dynamics has been shown to affect the cortical skeleton (Qu et al., 2017). Second, since human motor neurons are affected in ALS, we thought that a complete characterization of the βII-spectrin cortical cytoskeleton in these cells should include ALS-related mutations. We have now included an a basic MPS description in TDP43 and SOD1 mutation (Suppl. Fig. 5).

      The aspect of ALS-related mutations only occupies two short paragraphs in the main text and some panels in Supplementary information. To follow the suggestions by the Reviewer, we have downplayed the relative relevance of these results in the text, without compromising the amount of data we show.

      (3) There is a typo in the approximate symbol used for 150 kDa in the section where calpain and caspase activity is reported.

      (R21) Typo corrected.

      (4) Please add the Latrunculin concentration used in the main text, as it makes it easier for the reader.

      (R22) Done.

      (5) In the Discussion, paragraph starting with "We further showed ...", there is a typo where Zhong et al is cited.

      (R23) Corrected.

      (6) Supplementary Figure 1B: attachment instead of 'atachment'.

      (R24) Corrected.

      (7) Include DIVs or time in the schematic. It is easier for the reader to understand.

      (R25) We have now included time references in schematics of Suppl. Fig1B.

      (8) Supplementary Figure 1C

      Unable to distinguish βII-spectrin and βIII-tubulin in the merged image. Separate figure panels will help.

      (R26) The merged images in the reconstructions are merely to better show the tracing individual axons at such low magnification. Relevant portions with only βII-spectrin channels are shown in C1 and C2. Separated individual channels are shown elsewhere across the manuscript.

      (9) Supplementary Figure 4D

      Why is there so much cleavage product for αII-spectrin across DMSO and treatment? It varied over batches as well. Doesn't this mean that αII-spectrin is going through more proteolytic cleavage? Why?

      (R27) The amount of cleavage product for αII-spectrin is not a surprise to us. For instance, although calpains and caspases can potentially process both α- and β-spectrin, in in vivo scenarios where calpain activity is triggered there are much more fragments of α-spectrin being produced (Czogalla & Sikorski, 2005). On the other hand, our staining of cleaved-αII-spectrin by the SNTF antibody by immunofluorescence (Fig4C) parallels the findings by western blot -high levels of cleaved-αII-spectrin across treatments. A similar strong staining using this antibody has been recently shown in the intact axon (Heller et al., 2025). It will be interesting in the future to address if these fragments have any biological significance beyond being mere byproducts of αII-spectrin processing.

      Reviewer #2 (Recommendations for the authors):

      Suggestions for improving the quality of the manuscript:

      (1) Live imaging in combination with FRAP assays will help define whether the capture of spectrin from gaps into patches is true. Fixed neurons only provide static information and may not reflect real-time physiological effects.

      (R28) See R16 regarding possible live-imaging experiments using tagged βII-spectrin constructs.

      (2) Could the presence of F-actin trails in axons facilitate the formation of patches? Will the use of formin/Arp2/3 inhibitors rescue the effect of staurosporine, similar to Latrunculin A?

      (R29) Very interesting suggestion. It is likely that different pools of F-actin contribute to the dynamic of MPS formation, and actin trails are definitely worth investigating in this context.

      (3) Figure 8 lacks a latrunculin A treated condition? Why is this not present?

      (R30) The quantification of that treatment was excluded for space and readability. We have now included the values of group LatA + DMSO in Fig8Cand D and rearranged the whole figure.

      (4) Does neuronal stimulation have any effect (KCl treatment) on gaps and patches?

      (R31) Very interesting suggestion. Unfortunately, we have not examined whereas neuronal stimulation affects any parameter of the gaps-and-patches structure.

      (5) Please check the manuscript for typos and reference insertion points in the text. More than a couple were noted.

      (R32) We have corrected typos.

      Reviewer #3 (Recommendations for the authors):

      This is a very interesting study that supports a change in paradigm in the model of MPS lattice formation.

      (1) One major concern is the assumption that staudosporine-induced gap and patch formation recapitulates the physiological assembly of gaps and patches of betaII-spectrin, solely based on their morphological similarity. This should be further discussed in the manuscript. Further analysis of additional cytoskeleton components, including microtubules in staurosporine-treated neurons, could also be provided.

      (R33) See R14.

      (2) In Figure 1E, betaIII-tubulin and NF-H seem to accumulate in betaII-spectrin-rich axonal enlargements. If these are patches, how do you reconcile this finding with Figure 2C-D, where NF-M and alphaII-tubulin are not specifically enriched in betaII-spectrin patches?

      (R34) We actually show that axonal enlargements and patches are structurally unrelated, in many aspects. We mention these axonal enlargements as a way to perform an exhaustive characterization of all βII-spectrin features found in these axons.

      (3) One technical challenge that limits a more compelling support of the new model of MPS formation is that fixed neurons are imaged, which precludes the observation of patch coalescence. This should be further discussed in the revised version of the manuscript.

      (R35) The limitation of the experimental approach is now further discussed (see for example last sentences on 5th paragraph of the Discussion section).

      (4) On a more general note, the title of some of the Results sub-sections could be revised to convey the findings of those sub-sections and not the Methods that were used (example: "Quantitave and Qualitative analyses of betII-spectrin distribution....").

      (R36) According to the suggestion, we have changed the title of this subsection.

      References

      Boyer, N. P., Sharma, R., Wiesner, T., Parperis, C., Delamare, A., Pelletier, F., Jullien, N., Bhatt, A. M., Parra-Rivas, L. A., Kearney, P. J., Shavarebi, F., Leterrier, C., & Roy, S. (2026). Spectrin condensates provide a nidus for assembling the axonal membrane-associated periodic skeleton. iScience, 29(1), 114454. https://doi.org/10.1016/j.isci.2025.114454

      Castellanos-Montiel, M. J., Chaineau, M., & Durcan, T. M. (2020). The Neglected Genes of ALS: Cytoskeletal Dynamics Impact Synaptic Degeneration in ALS. Frontiers in Cellular Neuroscience, 14, 594975. https://doi.org/10.3389/fncel.2020.594975

      Czogalla, A., & Sikorski, A. F. (2005). Spectrin and calpain: A “target” and a “sniper” in the pathology of neuronal cells. Cellular and Molecular Life Sciences: CMLS, 62(17), 1913–1924. https://doi.org/10.1007/s00018-005-5097-0

      Guix, F. X., Capitán, A. M., Casadomé-Perales, Á., Palomares-Pérez, I., López Del Castillo, I., Miguel, V., Goedeke, L., Martín, M. G., Lamas, S., Peinado, H., Fernández-Hernando, C., & Dotti, C. G. (2021). Increased exosome secretion in neurons aging in vitro by NPC1-mediated endosomal cholesterol buildup. Life Science Alliance, 4(8), e202101055. https://doi.org/10.26508/lsa.202101055

      Heller, E., Kurup, N., & Zhuang, X. (2025). The membrane skeleton is constitutively remodeled in neurons by calcium signaling. Science (New York, N.Y.), 389(6760), eadn6712. https://doi.org/10.1126/science.adn6712

      Qu, Y., Hahn, I., Webb, S. E. D., Pearce, S. P., & Prokop, A. (2017). Periodic actin structures in neuronal axons are required to maintain microtubules. Molecular Biology of the Cell, 28(2), 296–308. https://doi.org/10.1091/mbc.E16-10-0727

      Zhong, G., He, J., Zhou, R., Lorenzo, D., Babcock, H. P., Bennett, V., & Zhuang, X. (2014). Developmental mechanism of the periodic membrane skeleton in axons. eLife, 3, e04581. https://doi.org/10.7554/eLife.04581

      Zhou, R., Han, B., Nowak, R., Lu, Y., Heller, E., Xia, C., Chishti, A. H., Fowler, V. M., & Zhuang, X. (2022). Proteomic and functional analyses of the periodic membrane skeleton in neurons. Nature Communications, 13(1), 3196. https://doi.org/10.1038/s41467-022-30720-x

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Butler et al. explores a novel physiological role for connexin 32 (Cx32) hemichannels in Schwann cells at peripheral nerves. Building on the authors' prior work on CO<sub>2</sub> - sensitive gating of connexins, this study proposes that mitochondrial CO<sub>2</sub> production dependent on neuronal activity promotes the opening of Cx32 hemichannels in the paranode, which in turn modulates neuronal activity by reducing conduction velocity. This hypothesis is addressed using a multifaceted approach that includes immunofluorescence microscopy, dye uptake assays, calcium imaging, computational modeling, and extracellular recordings in isolated sciatic nerves.

      Among the strengths of the study are the interdisciplinary integration of imaging, in silico approaches, and functional data. Also, this study proposes a new mechanism with profound physiological relevance. Specifically, Butler et al. provide new insights into glial modulation of electrical conduction in sensory/motor myelinated nerves.

      In the current state, the study has some limitations. The evidence linking Cx32 to the observed dye uptake and conduction velocity changes relies primarily on pharmacological inhibition with carbenoxolone, which lacks specificity. The imaging data show overlapping marker signals that preclude the anatomical distinction between nodes and paranodes. FITC uptake, while convincing to test Cx32 hemichannel gating, lacks spatial-temporal information and validation of distribution and localization to viable intracellular compartments. Moreover, while the findings are intriguing, functional proof that Cx32 regulates conduction velocity through ATP release or other downstream effects remains incomplete. Further work using targeted genetic tools, live-tissue imaging, and additional controls would strengthen the mechanistic conclusions.

      Overall, the manuscript offers compelling preliminary evidence that supports a new role for Cx32 in peripheral nerve physiology and raises important questions for future investigation.

      We thank the reviewer for their comments and agree that the evidence for involvement of Cx32 is indirect. We have now used viral expression of Cx32<sup>DN</sup> in SCs to remove CO<sub>2</sub> sensitivity from the endogenous Cx32 to strengthen this link. We have reviewed our presentation of the morphology in terms of the node/paranode/juxtaparanode distribution and adjusted accordingly. We have added new data using GCaMP transduced into Schwann cells that provides the live-tissue imaging that the reviewer requests.

      Reviewer #2 (Public review):

      Summary:

      This article aims to demonstrate that local production of CO<sub>2</sub> at the axonal node opens Cx32 hemichannels in the Schwann cell paranode, and that CO<sub>2</sub> diffuses through the AQP1 channel to reach Cx32 and trigger its opening. The authors also present evidence supporting a physiological role for this regulatory mechanism. They propose that CO<sub>2</sub>-dependent Cx32 activation mediates activity-dependent Ca<sup>2+</sup> influx into the paranode, and by increasing the leak current across the myelin sheath, it contributes to a slowing of action potential conduction velocity.

      The study presents a very interesting and novel mechanism for the physiological regulation of Cx32 hemichannels. The findings are relevant to the field, and the methods and results are of good quality, with some improvements in interpretation and explanation required, and some minor experimental suggestions.

      Strengths:

      The article is solid in terms of the novelty of the findings and relevance for the physiology of myelinated axons. In addition, it is of major interest for the Connexin field because it explores a physiological way to open Cx32 hemichannels. The experiments are well elaborated, and most of them are sufficient for the main points described by the authors. The finding that nervous activity will trigger the mechanism of hemichannel opening by CO2 is probably the most relevant biological mechanism derived from this article.

      Weaknesses:

      Throughout the manuscript, the authors interpret their findings as if the described mechanism specifically occurs in the node and paranode regions. However, there is no direct evidence identifying the precise site of CO<sub>2</sub> production or the activation site of Cx32 hemichannels. Therefore, statements such as the one in the title ("activity-dependent CO<sub>2</sub> production in the axonal node opens Cx32 in the Schwann cell paranode") should be reconsidered or removed, as they may be misleading and are not essential to the interpretation of the data. In addition, the participation of aquaporin AQP1 as the main conduit for CO2 diffusion through the plasma membrane could have another interpretation.

      We thank the reviewer for their comments and agree that we do not have direct evidence for the site of CO<sub>2</sub> production or the site of activation of Cx32 hemichannels. This direct evidence is extremely difficult to obtain, and we therefore depend on indirect arguments. Mitochondria represent the major source of CO<sub>2</sub>, and their distribution will therefore indicate where CO<sub>2</sub> is likely to be produced. We agree that this is not essential to the interpretation of the data and have adjusted the text as recommended. We have added a section to the Discussion to consider this point in more detail. The reviewer alludes to a reported interaction between AQP1 and NaV1.8 as a possible alternative interpretation. We can confidently rule this out as the AQP1 blocker has no effect on the compound action potential.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Main comments:

      (1) While the imaging system used in this study is technically capable of resolving nodes and paranodes, interpretation depends critically on marker specificity and tissue orientation. In some figures, markers such as Caspr or KCNA2 appear to partially overlap with KCNQ2 or the putative axonal node, which could reflect biological proximity but may also result from incomplete spatial separation in the z-dimension or the curvature of teased fibers. Similarly, Cx32 immunoreactivity or FITC signal is occasionally seen within nodal gaps, raising questions about how accurately this data supports the author's hypothesis. Additionally, while the authors claim that AQP1 is localized in nodes, the data suggest the opposite. Clarifying these patterns using fluorescence intensity line scans or additional nodal markers such as Nav1.6 or Ankyrin G would help distinguish overlapping signals from true domain-specific localization and reinforce the spatial conclusions of the study.

      We have changed our presentation of the localisation studies. We have concentrated on colocalization of Cx32 and AQP1 (now Fig 2) and moved the other studies to supplements to this figure. While we have retained the same images of Cx32 and AQP1 localisation, we have emphasized that these are SIM images and thus higher resolution than conventional LSM images, and also from a single optical plane. We have also clarified that the colocalization studies are restricted to analysis of the node/paranode regions.

      (2) To strengthen the conclusion that Cx32 specifically mediates the observed dye uptake, additional data or an alternative approach would be valuable. One feasible, though technically demanding, strategy would be the use of AAV-mediated delivery of Cx32-targeting shRNA directly into the sciatic nerve, ideally under a Schwann cell-specific promoter. This approach could achieve localized, cell-type-specific knockdown of Cx32 within a relevant time frame. Alternatively, the authors are encouraged to consider using additional pharmacological inhibitors to exclude the contribution of other conduction pathways, such as pannexin channels. These complementary strategies would reduce the interpretive ambiguity associated with non-specific blockade.

      We agree that this is desirable and have used Cx32<sup>DN</sup> under the control of the Mpz promoter (delivered by AAV via intranerval injection). This approach has several advantages -the Cx32<sup>DN</sup> subunit coassembles with endogenous Cx32<sup>WT</sup> and the heteromeric assemblies lack CO<sub>2</sub> sensitivity (first shown in Butler & Dale, 2023; and this strategy used with Cx26 to demonstrate its role in the control of breathing van de Wiel, 2020). This is a new figure (Fig 9). We have included supplemental figures with Fig 9 to document the coassembly of Cx32<sup>DN</sup> with Cx32<sup>WT</sup> by FRET.

      These new data test a very specific hypothesis: that CO<sub>2</sub> binding to Cx32 is responsible for the CO<sub>2</sub> sensitivity of the nerve. We find by comparing transduced and non-transduced fibres in the same nerve that Cx32<sup>DN</sup> essentially abolishes activity dependent loading of FITC into the Schwann cells.

      (3) Related to FITC experiments: Assuming the hypothesis of the authors is correct and CO2 release is restricted to the node, one should expect that if the major source of CO2 is in the nodal mitochondria, the hemichannels adjacent to the node will open first, assuming the spatial-temporal diffusion of CO2. To demonstrate this point, I would strongly suggest performing tissue imaging with real-time dye uptake. This approach should capture the FITC wave starting from the Cx32 channel opening in the paranode, as expected. Visualization of uptake in fixed and sectioned tissue is not the ideal approach to detect functional hemichannel opening in intact, viable cells, and at this point, they do not demonstrate that the uptake occurs in the node. From my perspective, if real-time experiments using isolated axons are feasible, it would make this paper more solid.

      The suggested method is not practical as the FITC in solution will be fluorescent and thus obscure the entry of FITC into the paranode. We have however expressed GCaMP8 under the control of the Mpz promoter, and this is expressed at paranodes and gives a CO<sub>2</sub> and activity-dependent Ca<sup>2+</sup> signal at the paranode. This gives a real time measure of the effect of CO<sub>2</sub> on the nerve. The GCaMP8 signal is enhanced by AZ and blocked by TC AQP1-1 (see below).

      (4) In Figure 5, Supplement 1, the authors present data using GRAB-ATP to suggest that Cx31.3 hemichannels do not release ATP under CO<sub>2</sub> stimulation. However, control experiments with GRAB-ATP alone (without Cx31.3 expression) are not shown, and parallel conditions with Cx32-expressing cells are lacking. Including these controls would strengthen the manuscript. Finally, testing the permeability of Cx31.3 to FITC directly, using the same conditions as in the main experiments, would clarify whether the discrepancy reflects differences in molecular permselectivity or CO<sub>2</sub> sensitivity.

      Figure 5 supplement 1, does show GRAB<sub>ATP</sub> alone without Cx31.3 expression (in the box plot). However, we have now added raw traces for this to the figure in panel B. CO<sub>2</sub>-dependent and voltage dependent ATP release via Cx32 has been previously shown in two papers (Butler & Dale 2023, Frontiers Cell Neurosci; Lovatt et al 2025, J Biol Chem). The Cx32<sup>DN</sup> result (above) further eliminates any contribution of Cx31.3.

      (5) Suggestion: It would be valuable to explore whether the proposed mechanism is conserved across both motor and sensory neurons, as this would broaden its physiological relevance. Since the sciatic nerve contains both fiber types, selective analysis or comparative data could clarify whether hemichannel activity is differentially regulated or restricted to a specific neuronal subtype.

      This is a great idea, but well beyond the scope of this paper. In an ex vivo preparation it would be very difficult to selectively stimulate the sensory vs motor fibres.

      Suggestions to improve data presentation and other minor comments:

      (1) Reduce/reorganize the figures to make the paper straightforward. For example, (a) immunofluorescence data showing the CO2 signaling machinery could be represented in one single figure; (b) Figure 1 could include all the findings and keep it as a final figure to summarize what the authors claim.

      We thank the reviewer for these suggestions. We prefer to keep Fig 1 up front to have our hypothesis clear for the reader to assist their interpretation as they go through the paper. We have altered the balance of figure supplements and main figures that document the immunolocalisation studies to concentrate on the main areas of novelty (AQP1 and Cx32 colocalisation and CA localisation).

      (2) The following phrase in the Results section is incomplete: "There was colocalization between Cx32 and CytC in the Schwann cell paranode, and (Fig 2, mean; 95% confidence interval, M1: 0.314; 0.198, 0.431 and M2: 0.261; 0.165, 0.357)."

      We have corrected this

      Additionally, the three values for M1 and M2 should be clearly defined and contextualized. In the current state, I couldn't understand them.

      The three values are mean and lower and upper 95% confidence limit:

      M1: mean 0.314; 95% CI, 0.198 to 0.431

      We have now made this clearer in the text.

      (3) It is unclear whether the authors calculate Manders' coefficients across the whole image or selectively at the node/paranode. Clarifying this would help interpret the specificity of co-localization claims.

      The Manders’ coefficients were selectively calculated at the node/paranode and we have amended the text to clarify this.

      (4) It is possible that mislocalization of CytC and SFXN1 could reflect antibody unspecificity or post-isolation alterations in protein distribution (e.g., apoptosis or stress). The authors briefly discussed this observation, but it could be a good idea to consider the use of an additional antibody to validate mitochondria localization.

      Apoptosis or stress is unlikely as the isolated nerves were fixed immediately after isolation with little dissection prior to fixation.

      The SFXN1 antibody was validated by Fowler et al 2013, and IP-HTMS confirmed SFXN1 as an interacting partner with Cx32. In this paper they also described SFXN1 as being present at the plasma membrane, the speculation being that it was taken there by Cx32.

      We think this is probably a valid result and we have further cited the Fowler et al 2013 paper in our discussion of this point.

      (5) Figure 4: The legend states: "Arrow heads indicate the node, and arrows depict the outer myelin." However, no arrows are visible in the figure. Please check.

      Corrected.

      (6) Figure 5: Keep consistency: Include in panel N that trpa1 inhibitor is in the presence of 70mmHg PCO2, as indicated for cbx in the same panel.

      Done

      (7) Figure 5 Supplement 1: Normalization using 1 concentration of ATP could not be appropriate if the sensor-dependent signal is not linear. If possible, authors should make a concentration-response curve and fit the data using the appropriate equation.

      Over the range we are measuring ATP (low µM) GRAB<sub>ATP</sub> is approximately linear to allow a single point calibration -we documented this in Butler and Dale 2023. This is also shown in the original paper describing GRAB<sub>ATP</sub> (Wu et al 2022 Neuron). We have clarified this point in the methods by referring to these papers.

      (8) Figure 6: The increase in FITC signal could represent a basal uptake over time. Authors should clarify the magnitude/rate of the basal uptake. Another option is showing a picture of the uptake using the control frequency at a time of 10 min. Legend: It is not clear in panel C if this picture corresponds to frequency stimulation. If so, it would be beneficial to specify the time.

      Could dye loading in this Fig simply be time dependent rather than stimulation dependent? Our data show that this is not the case -the dye loading controls of Fig 5A were exposed to FITC for 10 mins at 35 mmHg PCO<sub>2</sub> -very little loading is apparent. We now explicitly make this point in the text. Our use of Cx32<sup>DN</sup> also eliminates this explanation, by demonstrating the necessity of CO<sub>2</sub> binding to Cx32 for dye loading to occur.

      As there is no panel C in this figure, we assume the referee means panel B and have added the frequency of stimulation and time duration used to achieve the loading.

      (9) Please revise the legend of Figure 7. It seems to refer to a previous version of the manuscript's figure.

      Thanks for pointing this out. We omitted giving a letter to one of the panels and we have corrected this so that legend and figure now correspond.

      (10) Figures 10 and 11. Please consider including a bright field image or indicating with an arrow where the node and/or paranode is located.

      The old Fig 11 has been omitted. The old Fig 11 is now Fig 10. Unfortunately, we cannot add a bright field image as we did not save these in this experiment.

      (11) Figure 11. The authors could consider doing this experiment in the presence of Cx32 blockers to strengthen their conclusion.

      We have decided to remove this figure as it the information it contains is shown in the new GCaMP8 figure (Fig 12).

      (12) Figure 12: Calcium signal increases in different areas beyond the ROI. Not clear that the calcium signal is restricted to the node, as shown in previous figures. Please clarify if the preparation is different.

      We agree that this is a limitation – there is a lot of out of focus light due to Fluo4 being membrane permeable and loading many fibres within the nerve (potentially both axon and Schwann cell). Importantly, this phenomenon occurs in the in-focus ROI (for which we show BF image).

      As we think this is basically a limitation of using Fluo4-AM, we have now produced better data using GCaMP8 under the Mpz promoter (new Fig 12). This expresses at the paranode and in far fewer fibres so the resolution of the recordings is better. We have added these new data into the main body of the paper and relegated the Fluo4 data as a figure supplement to Fig 12 that provides independent supporting information.

      (13) Figure 13: Please indicate the stimulation frequency. The authors could consider attaching Figure 7 Supplement 1 to this figure to make the manuscript straightforward.

      Frequency now indicated.

      With regard to the original Figure 7 supplement 1 -thanks for this suggestion. After consideration, we have split this up and attached it as figure supplements to the relevant figures (Figure 6 and Figure 8). We have added equivalent data to Fig 7 (effect of H<sub>2</sub>O<sub>2</sub>). We think this simplifies presentation for the readers.

      (14) Figure 7 Supplement 1 and Figure 8 Supplements: Please indicate trace colors in panel A of these figures. Also, correct the spelling issue in the legend of Figure 8 Supplement 1 (for panel B).

      Corrected

      (15) Statistical clarifications: The authors should specify which experimental groups were included in some statistical analysis where p-values are reported, but the information about which groups are compared is missing.

      Corrected

      Reviewer #2 (Recommendations for the authors):

      (1) Localization of CO<sub>2</sub> production and Cx32 activation

      Throughout the manuscript, the authors interpret their findings as if the described mechanism specifically occurs in the node and paranode regions. However, there is no direct evidence identifying the precise site of CO<sub>2</sub> production or the activation site of Cx32 hemichannels. Therefore, statements such as the one in the title ("activity-dependent CO<sub>2</sub> production in the axonal node opens Cx32 in the Schwann cell paranode") should be reconsidered or removed, as they may be misleading and are not essential to the interpretation of the data.

      We agree that we have not shown this -and now exercise more caution in the description of the results and discuss this point.

      (2) Figures 2 and 3 - Cx32, mitochondria, and AQP1 localization

      In Figures 2 and 3, it is difficult to clearly discern the localization of Cx32, mitochondria, and AQP1 in the nodal and paranodal regions. The addition of zoomed-in images and 3D reconstructions (or at least orthogonal views) would greatly help clarify whether these components are indeed localized to the axon or Schwann cell, and whether they are specifically enriched in nodal or paranodal domains. As currently presented, the images suggest that all components of this "triad" are broadly distributed within the cells, not restricted to, nor particularly enriched in, nodal or paranodal areas. This observation further supports the concern raised in point 1.

      We have revised our presentation of the localisation more clearly and added a section to the discussion to consider this point more fully. We now explicitly mention that these are SIM images and in a single optical plane, therefore colocalization is genuine. We have also clarified that the calculation of Manders’ coefficients was performed only at the node/paranode regions. However, we accept that these components are distributed more widely than the node/paranode.

      (3) Figure 5 - Clarify legend labels

      In the graph shown in Figure 5, the legend would benefit from more descriptive labeling of the experimental groups. For clarity, indicate that FCCP was applied alone, and that HCO30031 was co-applied with high PCO<sub>2</sub>, to simplify interpretation for the reader.

      Corrected

      (4) Additional experiment to block mitochondrial CO<sub>2</sub> production

      An experiment should be added to completely or significantly inhibit mitochondrial CO<sub>2</sub> production, for example, by combining FCCP treatment with a TCA cycle inhibitor such as fluoroacetate. This would more directly demonstrate that CO<sub>2</sub> generation is required for hemichannel opening during FCCP treatment. It is important to control for this because FCCP can increase ROS production as a result of compensatory metabolic activity (i.e., increased NADH/FADH<sub>2</sub> generation). Since Cx32 hemichannels are known to be modulated by ROS, and can also regulate mitochondrial ROS production, it is crucial to distinguish the role of CO<sub>2</sub> from that of ROS in these experiments.

      Thanks for this great comment, as it gave us the idea of linking activity-dependent (rather than FCCP-evoked) gating of Cx32 to the TCA cycle and, as the reviewer says, CO<sub>2</sub> generation more directly. As fluoroacetate is only effective at inhibiting the TCA cycle in glial cells, we used H<sub>2</sub>O<sub>2</sub> at 50 µM which is highly effective at blocking aconitase in neurons (Tretter & Adam-Vizi, 2000). This greatly reduced FITC dye loading in response to activity. We now include these data in the paper (Fig 7).

      We note that our new data with Cx32<sup>DN</sup> further establishes the link to CO<sub>2</sub> as opposed to ROS.

      Furthermore, to complement the experiments involving carbonic anhydrase (CA) manipulation, additional controls or mechanistic validation may be necessary to support the conclusions drawn.

      We think that our use of Cx32<sup>DN</sup> greatly strengthens our conclusions that CO<sub>2</sub> is the messenger from the axon that gates Cx32 in the paranode.

      (5) AQP1 and Na<sup>+</sup> channel interaction - alternative interpretation

      It has been reported that AQP1 interacts with voltage-gated Na<sup>+</sup> channels, influencing action potential generation. For example, in AQP1 knockout mice, current injection-evoked action potentials show a reduced peak inward current, suggesting impaired Nav1.8 function (Zhang et al., J. Biol. Chem., 2010; doi: 10.1074/jbc.M109.090233). This raises the possibility that the observed effects of AQP1 inhibition (e.g., with TC AQP1-1) could also result from altered Na<sup>+</sup> channel activity, not just impaired CO<sub>2</sub> transport. I suggest that this alternative interpretation be acknowledged and discussed, as the current data do not rule it out.

      While constitutive KO of AQP1 does alter action potential generation in DRGs and an interaction between AQP1 and Nav1.8 has been documented, we do not think that this is a viable alternative interpretation of our data. We have measured the CAP during all our manipulations including the use of TC AQP1-1, and its amplitude is unaltered (see Fig 8 fig supplement 1 and Fig 13D). Our data therefore shows that, in the context of our experiments, application of the AQP1 blocker, TC AQP1-1, does not alter Na<sup>+</sup> channel activity. The difference between our data and the evidence from AQP1 knock-out may arise from the nature of an acute application of an antagonist (short term effect without changing protein expression) and constitutive knock out, which is likely to have longer term effects. We have added some discussion to address this point (last few lines, Page 9).

      (6) Figures 11A and 12C - Add heat map calibration

      In Figures 11A and 12C, the changes in Ca<sup>2+</sup> signals are difficult to interpret. In some areas, color changes appear to occur outside of cellular structures. I recommend including a heat map calibration scale for both figures to facilitate the interpretation of the signal intensity and localization.

      We agree that these data are limited by the technique used, and as mentioned above we now have GCaMP8 data that has better resolution and strengthens our conclusions.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In the manuscript entitled "Flexible and high-throughput simultaneous profiling of gene expression and chromatin accessibility in single cells," Soltys and colleagues present easySHARE-seq, a method described as an improvement upon SHARE-seq for the simultaneous measurement of RNA transcripts and chromatin accessibility.

      The authors demonstrate the utility of easySHARE-seq by profiling approximately 20,000 nuclei from the murine liver, successfully annotating cell types and linking cisregulatory elements to target genes. The authors claim that easySHARE-seq supports longer read lengths potentially enabling better variant discovery or allele-specific signal assessment, though they do not provide direct evidence to support these specific claims.

      A key strength of the protocol is enhanced sequencing efficiency, achieved by shortening the Index 1 read from 99 to 17 nucleotides. This reduction does not come at a significant cost to barcode diversity, retaining approximately 3.5 million combinations. Additionally, the approach allows for the sequencing of a sub-library to assess quality prior to final barcoding and sequencing which seems quite clever.

      While the increase in RNA transcript recovery is substantial, it appears to come at a cost: there is a notable decrease in ATAC fragments per cell compared to the original SHARE-seq (and other platforms). Likely as a result, the dimensionality reduction (UMAP) shows good resolution for RNA profiles but relatively poor resolution for accessibility profiles. Furthermore, the presented data suggests potential ambient RNA contamination; specifically, the detection of Albumin in HSCs and B cells is likely an artifact of the protocol rather than a biological signal.

      Overall, the study is well-presented and represents a promising advance. However, there are significant shortcomings that should be addressed, particularly regarding "leaky" transcript recovery and reduced ATAC performance.

      Recommendations:

      (1) To provide a comprehensive view of the current field, the authors should include Scale Biosciences (Scale Bio) in their discussion of available commercial platforms.

      We added Scale Biosciences to the relevant part in the introduction.

      (2) A head-to-head comparison with the 10x Genomics Multiome platform would be of significant interest to the single-cell genomics community and would better contextualize the performance of easySHARE-seq.

      We agree that a comparison to the 10x Multiome technology would be of interest in the community. Therefore, we included such a dataset profiling murine liver nuclei in the comparison in Figure 1 E&F as well as Suppl. Fig. 1 L&M. The resulting comparison remains consistent - easySHARE-seq compares favourably to other multiomic technique in RNA-seq data quality (UMIs/cell) but not in ATAC-seq data quality (fragments/cell).

      (3) Optimizing ATAC Performance: I strongly suggest exploring methods to improve ATAC sensitivity. As the authors note, the improvement in RNA recovery may result from fewer processing steps and stronger fixation. It would be valuable to test if decreasing fixation back to 2% (as in the original SHARE-seq) recovers ATAC data quality, and to determine if the fixation level or the number of steps is the key variable in preserving transcripts.

      We thank the reviewer for this suggestion. We agree that knowing the specific step(s) impacting ATAC-seq data quality would be highly valuable. Unfortuantely, we are not in a position to perform the additional wetlab experiments. It remains an area of improvement as we develop the technique further. We can confirm, however, that our early trials showed that the extent of fixation is negatively correlated with ATAC-seq data recovery.

      (4) The authors allude to the possibility of scaling this assay using a barcoded poly(T). Explicit inclusion or demonstration of this capability would dramatically increase interest in this protocol. Perhaps ATAC could be scaled using a barcoded Tn5?

      We thank the reviewer for this suggestion. Since we cannot perform further experiments, we expanded and clarified on upscaling this assay in our Supplementary Notes and referred to them in the text.

      We also added a paragraph specifically discussing the use of barcoded Tn5 in the Supplementary Notes.

      (5) The number of HSCs and B cells expressing Albumin is problematic and suggests significant ambient RNA issues that need to be addressed or computationally corrected.

      We thank the reviewer for pointing out this potential issue. We have used ‘decontX’ to estimate and ‘de-contaminate’ our UMI counts. We have added a histogram of estimated fraction of contaminated counts per nuclei to Suppl. Fig. 1. We have used the decontaminated counts to re-generate the analysis in Fig. 2 B&C and Suppl. Fig. 2 F. This filtering step did not change the results of these analyses; in fact it strengthened the results and improved clarity. We have added the relevant information to the Methods section and codebase and discussed the results and implications in the Supplementary Notes which we briefly summarize here:

      “As reported in Suppl. Fig. 10, decontX identifies mean contaminated counts of 9.6% and median contaminated counts of 1.4%, suggesting that few cells that are heavily contaminated strongly inflate the overall estimation of contaminated counts. This could be due to 1) doublets or b) wrongly assigned cell types. The authors of decontX report contamination values of 1-4% in commercial droplet-based protocols and 11-14% in plate-based protocols, suggesting that easySHARE-seq performs better than other plate-based assays.”

      We again want to thank the reviewer for this suggestion. It has improved the manuscript.

      Reviewer #2 (Public review):

      Aims:

      The authors sought to optimize SHARE-seq, a multimodal single-cell method, to improve the simultaneous profiling of gene expression and chromatin accessibility. Their goal was to enhance barcode design for better sequencing efficiency and cost savings, while improving overall data quality. They then applied their optimized method, easySHARE-seq, to study liver sinusoidal endothelial cells (LSECs) to demonstrate its utility in examining gene regulation and spatial zonation.

      Strengths:

      The improved barcode design is an advance, increasing the proportion of sequencing reads dedicated to biological information rather than barcode identification. This modification offers practical benefits in terms of sequencing costs and read length, potentially reducing alignment errors. The method also demonstrates improved RNA detection compared to the original SHARE-seq protocol. The biological applications showcase how simultaneous measurement of both modalities enables analyses that would be practically impossible with single-modality approaches, particularly in examining how chromatin states change along developmental or spatial trajectories.

      Weaknesses:

      There is a notable reduction in chromatin accessibility detection compared to the original SHARE-seq method, likely limiting the broad use of the method. While the authors are transparent about this tradeoff, additional discussion would be helpful regarding how this affects data interpretation. Comparisons showing consistency between easySHARE-seq and SHARE-seq chromatin accessibility patterns at the single-cell level would strengthen confidence in the method.

      Overall:

      The authors achieve their aim of creating an optimized protocol with improved barcode design and enhanced RNA detection. The method represents a useful advance for specific experimental contexts where the tradeoffs are appropriate. Recommendations for the authors:

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Figure 1F appears identical to Supplementary Figure 1M. This should be corrected if this is in error.

      Fixed.

      Reviewer #2 (Recommendations for the authors):

      The following comments are intended to strengthen the work.

      (1) scATAC-seq Performance and Data Consistency

      While I appreciate the authors' transparency regarding scATAC-seq performance, the extent of underperformance warrants greater emphasis. Additionally, does the average ATAC-seq signal recapitulate previously published results? At the single-cell level, how consistent are easySHARE-seq and SHARE-seq data? I suspect that increased dropout in scATAC-seq may distort consistency between datasets. This should be explicitly discussed in terms of data interpretation.

      We thank the reviewer for this suggestion. We have cross-referenced the open chromatin regions in this study and we summarise the result at the end of the ‘benchmarking’ paragraph. We have further expanded on the limitations in our study in the ATAC-seq data given the lower data quality in the relevant part of the discussion. We should note that a direct comparison between SHARE-seq and this study is challenging due to different sample tissues.

      (2) LSEC Biological Investigations

      The biological investigations could be strengthened (though this may reflect my limited expertise with LSECs).

      (a) Enhancer analysis depth

      While the authors quantify potential enhancers through RNA-ATAC correlations within individual cells and identify genes regulated by multiple enhancers, a deeper exploration of enhancer biology would strengthen the manuscript. Potential questions include: Do genes sharing correlated enhancer activity also show correlated expression? How do enhancer number and strength relate to gene expression levels? How do RNA-ATAC correlations scale with ATAC peak height? Are stronger enhancers more tightly linked to gene expression? Perhaps the authors explored these questions without finding significant patterns, but this should be clarified.

      We thank the reviewer for this suggestions. We performed several analyses aimed at exploring enhancer biology with this dataset. We added a simple comparison for UMIs per gene between genes with at least one associated peak compared to those without in Suppl. Fig. 3I. We provide the corresponding plot for fragments per peak in Suppl. Fig. 3J. We also explored the relationship between gene expression and chromatin accessibility; here, we found that gene expression levels do not correlate with peak heights of chromatin accessibility (possibly because chromatin accessibility signals were somewhat binary). The corresponding plot has been added to Suppl. Fig. 3K. We added a small paragraph discussing these findings in the main text.

      (b) Correlation magnitude interpretation

      The reported correlation values are extremely small. Does this reflect weak biological linkages or primarily experimental noise? If experimental noise, how does variation in detection per gene influence the confidence in this type of analysis?

      We thank the reviewer for raising this potential issue. We identify a total of 40,957 significant peak-gene associations with a mean Spearman correlation of 0.1 (± 0.056; Suppl. Fig. 3E). This analytical workflow to identify these gene-peak associations was first described alongside SHARE-seq in Ma et al.. For context, they reported significant peak-gene associations to have a mean Spearman correlation of 0.026 (± 0.015; Ma et al. Table S4).

      Generally, we hypothesize that these low correlation values in this type of analysis are the results of sparseness of single-cell data, especially in chromatin accessibility. Therefore, the power to detect gene–peak associations increases with cell number (Ma et al., Fig. 3B) and the limited cell numbers in the analysis in this study likely results in an enrichment of the most strongly correlated associations among those detected. We have added a comparison of UMIs per gene for genes with and without a significant gene-peak correlation, illustrating this dynamic (Suppl. Fig. 3I). Furthermore, we have described this relationship and limitation in the relevant part of the results section.

      (c) Zonation analysis framing

      The zonation analysis is compelling, but the authors should more explicitly emphasize that defining pseudotime and examining chromatin state dynamics is only possible because both modalities are measured simultaneously. And more detail on the Monocle3 pseudotime analysis is needed, as it is unclear how this was really done.

      We expanded our description on the pseudotime analysis using Monocle in the relevant section in the Methods. Furthermore, we explicitly point out that this type of analysis relies on simultaneous measurements of both modalities at the end of the results section.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Yang et al. investigates the relationship between multi-unit activity in the locus coeruleus, putatively noradrenergic locus coeruleus, hippocampus (HP), sharp-wave ripples (SWR), and spindles using multi-site electrophysiology in freely behaving male rats. The study focuses on SWR during quiet wake and non-REM sleep, and their relation to cortical states (identified using EEG recordings in frontal areas) and LC units.

      The manuscript highlights differential modulation of LC units as a function of HP-cortical communication during wake and sleep. They establish that ripples and LC units are inversely correlated to levels of arousal: wake, i.e., higher arousal correlates with higher LC unit activity and lower ripple rates. The authors show that LC neuron activity is strongly inhibited just before SWR is detected during wake. During non-REM sleep, they distinguish "isolated" ripples from SWR coupled to spindles and show that inhibition of LC neuron activity is absent before spindle-coupled ripples but not before isolated ripples, suggesting a mechanism where noradrenaline (NA) tone is modulated by HP-cortical coupling. This result has interesting implications for the roles of noradrenaline in the modulation of sleep-dependent memory consolidation, as ripple-spindle coupling is a mechanism favoring consolidation. The authors further show that NA neuronal activity is downregulated before spindles.

      Strengths:

      In continuity with previous work from the laboratory, this work expands our understanding of the activity of neuromodulatory systems in relation to vigilance states and brain oscillations, an area of research that is timely and impactful. The manuscript presents strong results suggesting that NA tone varies differentially depending on the coupling of HP SWR with cortical spindles. The authors place their findings back in the context of identified roles of HP ripples and coupling to cortical oscillations for memory formation in a very interesting discussion. The distinction of LC neuron activity between awake, ripple-spindle coupled events and isolated ripples is an exciting result, and its relation to arousal and memory opens fascinating lines of research.

      Weaknesses:

      I regretted that the paper fell short of trying to push this line of idea a bit further, for example, by contrasting in the same rats the LC unit-HP ripple coupling during exploration of a highly familiar context (as seemingly was the case in their study) versus a novel context, which would increase arousal and trigger memory-related mechanisms. Any kind of manipulation of arousal levels and investigation of the impact on awake vs non-REM sleep LC-HP ripple coordination would considerably strengthen the scope of the study.

      We agree that conducting specific behavioral tests before electrophysiological recordings, as well as manipulating arousal during the recording session, would strengthen the study. These experiments are planned for future work, and we acknowledged this point in the discussion.

      We added the following text in the Discussion: “Conducting behavioral assays prior to electrophysiological recordings, along with spatially and temporally precise modulation of LC activity during recording sessions, will be essential for achieving a mechanistic understanding of network dynamics and its functional role for memory consolidation in future investigations.”

      The main result shows that LC units are not modulated during non-REM sleep around spindle-coupled ripples (named spRipples, 17.2% of detected ripples); they also show that LC units are modulated around ripple-coupled spindles (ripSpindles, proportion of detected spindles not specified, please add). These results seem in contradiction; this point should be addressed by the authors.

      The detection of coupled events - spindle-coupled ripples (spRipple) and ripple-coupled spindles (ripSpindle) - was performed independently, although, some overlap cannot be excluded. We found that LC suppression was generally weak around both types of coupled events. Specifically, LC suppression around spRipples and ripSpindles reached significance (exceeding the 95% confidence interval) in 4 sessions (from 3 rats) and 3 sessions (from 2 rats), respectively, out of a total of 20 sessions (from 7 rats).

      We revised the manuscript by providing additional information in the Results section and adding a Supplementary Figure 5 showing a significant correlation (Pearson r = 0.72, p = 0.0003) between the modulation index (MI) for spRipple and ripSpindle.

      Results are displayed per recording session, with 20 sessions total recorded from 7 rats (2 to 8 sessions per rat), which implies that one of the rats accounts for 40% of the dataset. Authors should provide controls and/or data displayed as average per rat to ensure that results are now skewed by the weight of that single rat in the results.

      High-quality recordings from the LC in behaving rats are technically challenging and relatively rare; therefore, we included all valid datasets in analysis. The average modulation index (MI), calculated per animal and per session, fell within a consistent range (Supplementary Figure 3) despite variability in the number of recording sessions (2–8 sessions per rat).

      In its current form, the manuscript presents a lack of methodological detail that needs to be addressed, as it clouds the understanding of the analysis and conclusions. For example, the method to account for the influence of cortical state on LC MUA is unclear, both for the exact methods (shuffling of the ripple or spindle onset times) and how this minimizes the influence of cortical states; this should be better described. If the authors wish to analyze unit modulation as a function of cortical state, could they also identify/sort based on cortical states and then look at unit modulation around ripple onset? For the first part of the paper, was an analysis performed on quiet wake, non-REM sleep, or both?

      The LC activity around rippled was modulated at multiple temporal scales. First, we observed a relatively sharp drop in the LC firing rate ~ 2 s before the ripple onset. When computing peri-ripple LC activity over a longer time window ([–12, 12] sec), we observed a rather slow decrease in the LC firing rate beginning as early as 10 s before the ripple onset (Supplementary Figure 2).

      Considering two temporal scales, we hypothesized that slow modulation of LC activity might be related to fluctuations of the global brain state. We quantified the ongoing cortical state using a synchronization index (SI), calculated as a power ratio (1–4 Hz/30–90 Hz) of the EEG within 4-s windows and computed the corresponding ripple and LC-MUA rates. Figure 3A (in the main manuscript) illustrates that a higher SI (more synchronized cortical population activity) corresponded to a lower arousal state and reduced LC tonic firing; this brain state was associated with a higher ripple activity. As shown in the new Figure 3B, the LC firing rate was negatively correlated with the SI and ripple rate. Thus, slow LC modulation was likely driven by cortical state transitions.

      To correct for the influence of the global brain state on the peri-ripple LC activity, we generated surrogate events by jittering the times of detected ripples. First, we confirmed that triggering the hippocampal LFP on the surrogate events lacked the ripple-specific frequency component (main Figure 3C) and the SI state did not differ around ripples and surrogate events (main Figure 3D). Plotting the LC activity around surrogate evens captured its state-dependent dynamics (Figure 3 or Supplementary Figure 2, orange trace). To extract state-independent peri-ripple LC modulation, we subtracted the state-related LC activity (orange trace) from the ripple-triggered LC activity (blue trace). The resulting trace yielded a corrected estimate of ripple-associated LC activity that was largely free from the confounding influence of cortical state transitions (main Figure 3E).

      In the Results subsection “LC-NE neuron spiking is suppressed around hippocampal ripples”, we reported LC modulation without accounting for the cortical state (main Figure 2). The state-dependent effects were instead examined in the subsequent Results subsection, “LC firing and ripple occurrence are state-dependent and inversely related” we report state-corrected LC modulation (main Figure 3). Finally, in the Results subsection “Peri-ripple LC modulation depends on the cortical–hippocampal interaction,” we characterized LC activity around ripples across different cortical states (quite awake and NREM sleep).

      We revised Methods and Results to provide more methodological details and a rationale for each analysis, as requested.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors studied the synchrony between ripple events in the Hippocampus, cortical spindles, and Locus Coeruleus spiking. The results in this study, together with the established literature on the relationship of hippocampal ripples with widespread thalamic and cortical waves, guided the authors to propose a role for Locus Coeruleus spiking patterns in memory consolidation. The findings provided here, i.e., correlations between LC spiking activity and Hippocampal ripples, could provide a basis for future studies probing the directional flow or the necessity of these correlations in the memory consolidation process. Hence, the paper provides enough scientific advances to highlight the elusive yet important role of Norepinephrine circuitry in the memory processes.

      Strengths:

      The authors were able to demonstrate correlations of Locus Coeruleus spikes with hippocampal ripples as well as with cortical spindles. A specific strength of the paper is in the demonstration that the spindles that activate with the ripples are comparatively different in their correlations with Locus Coeruleus than those that do not.

      Weaknesses:

      The claims regarding the roles of these specific interactions were mostly derived from the literature that these processes individually contribute to the memory process, without any evidence of these specific interactions being necessary for memory processes. There are also issues with the description of methods, validation of shuffling procedures, and unclear presentation and the interpretation of the findings, which are described in the points that follow. I believe addressing these weaknesses might improve and add to the strength of the findings.

      We believe that our responses to the Reviewer 1 and Reviewer 2, corresponding revisions of the manuscript and new figures adequately addressed all issues raised by the Reviewer 2.

      Reviewer #3 (Public review):

      Summary:

      This manuscript examines how locus coeruleus (LC) activity relates to hippocampal ripple events across behavioral states in freely moving rats. Using multi-site electrophysiological recordings, the authors report that LC activity is suppressed prior to ripple events, with the magnitude of suppression depending on the ripple subtype. Suppression is stronger during wakefulness than during NREM sleep and is least pronounced for ripples coupled to spindles.

      The study is technically competent and addresses an important question regarding how LC activity interacts with hippocampal and thalamocortical network events across vigilance states.

      Weaknesses:

      The results are interesting, but entirely observational. Also, the study in its current form would benefit from optimization of figure labeling and presentation, and more detailed result descriptions to make the findings fully interpretable. Also, it would be beneficial if the authors could formulate the narrative and central hypothesis more clearly to ease the line of reasoning across sections.

      We improved the presentation of results by incorporating additional figures and expanding the detail in the figure captions. In the main text, we clarified specific hypotheses and provided a rationale underlying each analysis.

      Comments:

      (1) Stronger evidence that recorded units represent noradrenergic LC neurons would reinforce the conclusions. While direct validation may not be possible, showing absolute firing rates (Hz) across quiet wake, active wake, NREM, and REM, and comparing them to published LC values, would help.

      We added the requested data and a Supplementary Figure 1 in the revised manuscript: “The average firing rates of LC single units were 1.70 ± 0.21 Hz during wakefulness, 0.51 ± 0.07 Hz during NREM sleep, and 0.014 ± 0.01 Hz during REM sleep (Supplementary Figure 1). Firing rates differed significantly across arousal states, with the highest activity during wakefulness, reduced activity during NREM sleep, and minimal activity during REM sleep (one-way ANOVA: F(2,38) = 39.8, p < 0.0001). This firing pattern is characteristic of LC-NE neurons and is consistent with existing literature.”

      (2) The analyses rely almost exclusively on z-scored LC firing and short baselines (~4-6 s), which limits biological interpretation. The authors should include absolute firing rates alongside normalized values for peri-ripple and peri-spindle analyses and extend pre-event windows to at least 20-30 s to assess tonic firing evolution. This would clarify whether differences across ripple subtypes arise from ceiling or floor effects in LC activity; if ripples require LC silence, the relative drop will appear larger during high-firing wake states. This limitation should be discussed and, if possible, results should be shown based on unnormalized firing rates.

      We agree with the reviewer that a longer pre-event window provides a clearer estimate of baseline LC activity. However, given that both ripples and spindles are brief oscillatory events, we tested a range of time windows and found that a 12-s interval adequately captures baseline LC activity dynamics. Accordingly, we included plots with extended pre-event windows (−12 to 12 s), as requested.

      We added in the revised manuscript absolute firing rates for well-isolated LC single units. Because the number of neurons contributing to LC multi-unit activity (LC-MUA) is unknown, we avoided averaging absolute firing rates for this signal. For LC-MUA, we implemented a normalization approach in which firing rates (50-ms bins) around ripple or spindle are scaled to a baseline period preceding the trigger event (−12 to −10 s). Importantly, unlike z-scoring, this normalization method preserves baseline differences across behavioral states. As shown in Author response image 1A and new Figure 5 in the main manuscript, baseline LC firing rates were highest prior to awake ripples and lowest prior to sleep spindles. During ripples occurring in wakefulness, LC activity did not decrease to the levels observed during sleep. In contrast, during NREM sleep, LC activity was downregulated during both ripples and spindles, although it did not reach complete silence around either oscillatory event.

      Author response image 1B illustrates a slow downward drift in the LC firing rate preceding either ripple or spindle. The slow LC dynamics likely reflected gradual transitions toward more synchronized brain state, which is optimal for ripple generation. In contrast, event-specific LC modulation had faster dynamics (Author response image 1B, highlighted interval) and was largely absent in cases where spRipples and ripSpindles were not associated with LC suppression (Author response image 1C).

      To minimize the influence of global state fluctuations and emphasize event-related dynamics, we therefore presented the main results using state-corrected and z-scored PETHs.

      Please also refer to our response to Reviewer 1 regarding the two temporal scales of LC modulation.

      Author response image 1.

      LC modulation around sleep oscillations. (A) Peri-event LC-MUA during awake and NREM sleep. LC activity and the range of peri-event LC modulation differed across behavioral states; it was overall higher preceding ripples occurring in wakefulness than in NREM sleep, and it was the lowest around sleep spindles. Despite the state-dependent differences in the firing rate, LC modulation was observed around all oscillatory events. During wakefulness, LC activity did not decrease to the levels observed during NREM sleep. During NREM sleep, LC activity was down-regulated around both ripples and spindles, and the LC firing did not completely cease around either oscillatory event. (B) Peri-event LC-MUA around isolated oscillatory events. LC activity exhibited fast peri-event dynamics (highlighted interval) superimposed on slower, state-dependent fluctuations. (C) Peri-event LC-MUA around coupled oscillatory events. Fast peri-event LC modulation was absent, while slow fluctuations were preserved around coupled oscillatory events. For all plots, LC-MUA firing rate was scaled to a pre-event baseline interval [-12 to -10 sec] to preserve baseline differences in LC activity across behavioral states. Bin size: 50 ms. isoRipple – isolated ripple, isoSpindle – isolated spindle, spRipple - spindle-coupled ripple, ripSpindle - ripple-coupled spindle.}

      (3) Because spindles often occur in clusters, the timing of ripple occurrence within these clusters could influence LC suppression. Indicate whether this structure was considered or discuss how it might affect interpretation (e.g., first vs. subsequent ripples within a spindle cluster).

      We did not consider spindle clusters and classified the event as ripple-coupled spindle if the ripple occurred between the spindle on and offset.

      (4) While the observational approach is appropriate here, causal tests (e.g., optogenetic or chemogenetic manipulation of LC around ripple events and in memory tasks) would considerably strengthen the mechanistic conclusions. At a minimum, a discussion of how such approaches could address current open questions would improve the manuscript.

      We agree that conducting causal tests would strengthen the study. We added the following text in the Discussion: “Conducting behavioral assays prior to electrophysiological recordings, along with spatially and temporally precise modulation of LC activity during recording sessions, will be essential for achieving a mechanistic understanding of network dynamics and its functional role for memory consolidation in future investigations.”

      (5) Please show how "Synchronization Index" (SI) differs quantitatively across behavioral states (wake, NREM, REM) and discuss whether it could serve as a state classifier. This would strengthen interpretations of the correlations between SI, ripple occurrence, and LC activity.

      We plotted the awake state-normalized SIs for awake and NREM sleep. Due to small number of REM sleep episodes, SI for REM sleep is not shown. The average SI during NREM sleep was significantly higher than during awake state, consistent with the well-established dominance of low-frequency (1-4 Hz) oscillatory power and reduced high-frequency (30-90 Hz) power during NREM sleep.

      Although SI could potentially serve as a behavioral state classifier, we have chosen not to address this point to maintain the focus in the discussion on new results.

      Author response image 2.

      Synchronization index differentiates behavioral states.

      (6) The current use of SI to denote a delta/gamma power ratio is unconventional, as "SI" typically refers to phase-locking metrics. Consider adopting a more standard term, such as delta/gamma power ratio. Similarly, it would be easier to follow if you use common terminology (AUC) to describe the drop in LC-MUA rather than using "MI" and "sub-MI".

      The ranges of delta and gamma bands might vary across studies; therefore, we prefer using SI, as defined here and in our previous publications (Novitskaya et al., 2016; Yang et al., 2019, 2021). We calculated the modulation index (MI) as the area under the curve of the peri-event time histogram within the 1 second preceding ripple onset. To avoid potential confusion with the AUC calculated over the entire signal window, we opted to use MI.

      (7) The logic in Figure 3 is difficult to follow. The brain state (delta/gamma ratio) appears unchanged relative to surrogate events (3C), while LC activity that is supposedly negatively correlated to delta/gamma changes markedly (3D-E). Could this discrepancy reflect the low temporal resolution (4-s windows) used to calculate delta/gamma when the changes occur on a shorter time scale?

      We appreciate the reviewer’s question. We revised the results and Figure 3 legend to clarify this point. The main Figures 3E and 3F show the 'state-corrected' peri-ripple LC activity. The purpose of generating ‘surrogate’ events was precisely to capture the component of LC activity dynamics that can be explained by cortical state fluctuations alone. As shown in Supplementary Figure 2, the orange trace represents LC activity aligned to surrogate events and, as the Reviewer noted, shows a clear decrease, yet at a slower time scale. We interpret this surrogate-aligned signal as the LC modulation attributable specifically to cortical state fluctuations. Importantly, shuffled events were associated with similar SIs (cortical state), but absent HPC LFP power increase in the ripple range (140-250 Hz), as shown in the main Figures 3C and 3D, respectively. To isolate the peri-event LC dynamics, we subtracted the state-related component (Figure 3, orange trace) from the ripple-triggered LC activity (blue trace). This correction yielded an estimate of ripple-associated LC activity that is largely independent of the confounding influence of ongoing cortical state.

      Please, see our detailed response to the Reviewer 1 about multiple time scales of LC dynamics.

      (8) There are apparent inconsistencies between Figures 4B and 4C-D. In B, it seems that the difference between the 10th and 90th percentile is mostly in higher frequencies, but in C and D, the only significant difference is in the delta band.

      We repeated this analysis, clarified inconsistency, and revised Figure 4 legend.

      (9) Because standard sleep scoring is based on EEG and EMG signals, please include an example of sleep scoring alongside the data used for state classification. It would also be relevant to include the delta/gamma power ratio in such an example plot.

      We replaced ‘standard’ with ‘previously established” sleep scoring procedure and added a Supplementary Figure 4 showing representative NREM sleep and wake episodes with corresponding EEG and SI.

      (10) Can variability in modulation index (subMI) across ripple subsets reflect differences in recording quality? Please report and compare mean LC firing rates across subsets to confirm this is not a confounding factor.

      We agree that considering recording quality and unit stability over time as potential confounding factors is important. We therefore carefully evaluated each dataset to ensure the absence of significant drift in the LC firing rate. However, we find that comparing mean LC firing rates across subsets of ripples, as suggested by the Reviewer, is insufficient to control for recording stability, as LC activity varies substantially across behavioral states. At present, we are not aware of a robust method to fully eliminate variability related to recording quality and unit stability over time.

      (11) Figure 6B: If the brown trace represents LC-MUA activity around random time points, why would there be a coinciding negative peak as relative to real sleep spindles? Or is it the subtracted trace?

      We have revised Figure 7 (original Figure 6) and its legend to improve clarity and readability.

      (12) On page 8, lines 207-209, the authors write "Importantly, neither the LC-MUA rate nor SIs differed during a 2-sec time window preceding either group of spindles". It is unclear which data they refer to, but the statement seems to contradict Figure 6E as well as the following sentence: "Across sessions, MI values exceeded 95% CI in 17/20 datasets for isoSpindles and only 3/20 for ripSpindles". This should be clarified.

      We have revised the corresponding text to improve clarity and readability.

      (13) The results in Figures 5C and 6F do not align. It seems surprising that ripple-coupled spindles show a considerably higher LC modulation than spindle-coupled ripples, as these events should overlap. Could the discrepancy be due to Z-score normalization as mentioned above? Please include a discussion of this to help the interpretation of the results.

      In the original manuscript, Figure 6F was mistakenly labelled for ripple-coupled (ripSpindles) and isolated (isoSpindles) spindles. Now it has been corrected.

      Please, also see our response to the Reviewer 1 weaknesses.

      (14) The text implies that 8 recordings came from one rat and two each from six others. This should be confirmed, and it should be explained how the recordings were balanced and analyzed across animals.

      Since high-quality recordings from LC in behaving animals are challenging and rare, we used all valid sessions. We addressed the same point in our response to the Reviewer 1 weaknesses.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Below are some suggestions for clarification/information that are needed to improve the paper's readability (and the understanding of the analysis and methods).

      (1) The authors describe a consistently negative correlation between cortical EEG synchronization index and ripple rate or LC-MUA, show an example in Figure 3A, and report a range of r values in the text with a mention of p < 0.01. The reported p-value is presumably the highest p-value for the correlations - please specify. Visualization of the results might be improved by adding example correlations (also true for later correlations in Figure 6).

      We revised the result description accordingly and included correlation plots in Figures 3 and 7.

      (2) Description of statistical testing is missing for Figure 3C (nothing in the text or the figure legend); there is also no statistics section in the methods. For Figure 4, the statistics are reported for the Friedman test but not the post-hoc tests. Exact p-value and statistics should be reported for the comparison of LC-MUA rate and SI in the 2 s preceding spindles.

      We have added the statistical results requested and revised figure legends by providing additional information. We added the Statistic Analysis section in the Methods.

      Figure 3D (original Fig.3C): “Average Synchronization Index (SI) around ripples and shuffled events. The cortical state preceding shuffled events and ripples was comparable, as confirmed by the absence of significant differences in SI (Wilcoxon signed-rank test; shuffled: Z = -0.20, p = 0.84; ripples: Z = 0.14, p = 0.88). Cortical synchrony increased following both events (shuffled: Z = -3.50, p = 0.00044; ripples: Z = -3.66, p = 0.00026). Similar cortical state dynamics surrounding shuffled events and ripples indicate that the surrogate events adequately capture the cortical state associated with ripple occurrence.

      Figure 6: Intra-ripple frequency (A) and peak amplitude (B) for different ripple types. Boxwhisker plots show the median, the 1st and 3rd quartiles, and min/max. Gray dots show data from individual rats. *** - p < 0.001 for post hoc pairwise comparisons (Wilcoxon signed-rank tests with Holm–Bonferroni correction for multiple comparisons).

      We revised the Results accordingly: “The ripple subtypes differed in the intra-ripple frequency (Friedman test, chi2 = 35.62, p < 0.0001, post hoc pairwise comparisons were performed using Wilcoxon signed-rank tests with Holm–Bonferroni correction for multiple comparisons. awRipple vs isoRipple: p = 0.00003 awRipple vs spRipple: p = 0.00004 isoRipple vs spRipple: p = 0.0002}), with awRipples being the fastest and spRipples the slowest (Figure 6A).There was no difference in the ripple peak amplitude (Friedman test, $\chi$2 = 3.7, p = 0.16; Figure 6B).”

      (3) The method description of ripple-spindle coupling detection is missing.

      We have added the description of ripple-spindle coupling detection in the Methods.

      (4) Based on Figure 6D, the authors report that ripple-coupled spindles are significantly shorter than isolated spindles. What are the measurements reported on lines 206-207, and how do they relate to the averaged spectrograms shown in Figure 6D?

      Spindle duration was calculated as the time between spindle onset and offset (as described now in the Methods and Figure 7 legend). Ripple-coupled spindle was considered if at least one ripple occurred between the spindle onset and offset. The duration of ripple-coupled and uncoupled spindles was statistically compared (the stats is reported in text). In Figure 7E, the peri-event averaged EEG spectrograms are plotted for isolated and ripple-coupled spindles, highlighting the difference in the event duration.

      (5) None of the color scales have legends (Figures 2A, B, C, Figure 3D, etc.).

      We have added the color scales on all Figures.

      (6) Description of what is represented in the box plots is missing.

      We have added the description.

      (7) Figure 4C, D, legend for the color code is missing.

      We have added color scales legends.

      (8) Figure 5A legend, assuming this should read intra-ripple frequency instead of inter-ripple.

      We corrected the typo.

      (9) Figure 5E, while LC units are not modulated before, it could still be informative to overlay the z-scored firing rate on the same graph for comparison.

      Figure 6E (original Figure 5E) shows overlay for awRipples and isoRipples.

      (10) The discussion states a 4s resolution for cortical state quantification (line 237), but the methods mention 2.5s (line 382).

      We corrected this discrepancy.

      (11) Results, p.5, line 138, Methods and materials, p.13, line 423: 30% in result text but 20% in method, please correct.

      We corrected this discrepancy.

      (12) The manuscript cites the biorxiv version of Osorio-Forero et al., but the paper has been published since then; please update.

      We updated this reference.

      (13) Results, p.2, line 70. The average duration of a session is presented in seconds. Minutes or hours would be more meaningful to the reader.

      We consider this suggestion as optional.

      (14) Figure 2C is not referenced.

      We added the reference to Figure 2C.

      (15) Reference missing line 406.

      We added the reference.

      (16) Lines 352-356: There seems to be an error in the sentence (an extra verb, or an "and" missing somewhere).

      We have corrected this sentence.

      (17) Figure 3C "synchronization".

      We corrected this typo.

      Reviewer #2 (Recommendations for the authors):

      (1) Line 94 states that "A significant peri-ripple decrease in LC-SUA"; however, which test and how many samples were used are unclear.

      We revised this text as follows: “A significant peri-ripple (± 6 s) decrease in LCSUA, detected by the firing suppression exceeding 2 SDs, was observed in 13 of 15 cases (n = 4 rats).”

      (2) Line 96 states that "we calculated the modulation onset, duration, and magnitude". Please define modulation before presenting the comparisons.

      We now illustrate the extraction of quantitative variables in Figure 2D.

      (3) Line 119 states that "we generated surrogate time series for each session by shuffling ripple onset times" which gives the impression that ripple events were shuffled throughout the sleep; however, the method section states that it was jittered within a specific time window for each event. Please clarify the matter.

      We have substantially revised this section to improve clarity and readability.

      (4) Line 120 states that "Comparisons of SI values before and after ripples and surrogate events confirmed that surrogate events preserved the cortical states in which ripples occurred". Ripple power doesn't seem to be different in pre vs post in the shuffled data (Figure 3B). If ripple timing was randomized, please clarify the observation shown in Figure 3C that the shuffled events had higher SI after than before, as also seen in the real SI data? Please also elaborate what specific groups were significantly different in before vs after bars; data, shuffle, or both?

      We have substantially revised this section to improve clarity and readability.

      (5) Line 113 and Figure 3A: Because both LC activity and HPC ripples were correlated to SI, the direct relationship between LC and HPC independent of SI (a covariate) was not clear. The authors might be able to conduct a partial correlation analysis to show this effect.

      We appreciate this suggestion and added the correlation plots in Figures 3 and 7. After careful consideration, we believe that the suggested partial correlation analysis does not contribute substantially beyond the main findings already presented.

      (6) Figure 5A: Inter-ripple frequency needs definition, not provided in the paper nor in the reference paper. The value (180 Hz) suggests a time interval of around 5 ms, which I fail to understand.

      We apologize for this typo. In Figure 6A (original Fig.5A), intra-ripple frequency is plotted. We have corrected this typo in the text and figure legend.

      (7) Figure 5D: Comparison between aw and sp ripples should also be shown. Please explain the dashed line at 10 (y-axis) a.u.

      Figure 6E (original Fig.5E) shows LC activity around awRipples and isoRipples.

      (8) Figure 5E: Legend states aw and iso ripples, but the caption says NREM sleep. Please clarify this matter.

      We have revised Figure 6 legend (original Figure 5).

      (9) Figure 6B: If the spindle time is permuted randomly, why is LC activity in the permuted data still modulated by the spindle times? Can you test the significance of the modulation index of the shuffled data?

      The LC modulation around shuffled time points was not significant. Figure 7C shows LC modulation dynamics around spindles; brown trace showing state-corrected LCMUA trace (after subtraction of LC-MUA around shuffled events).

      (10) Line 203: Is the unit in Hz (events per second) correctly calculated or shown? ~15 events per second seems arbitrarily large.

      We corrected the units for the event rate. We report the mean oscillatory frequency of spindles ~15 Hz, not events per second.

      (11) Line 207 states that "neither the LC-MUA rate nor SIs differed during a 2-sec time window preceding either group of spindles"; however, from Figure 6E, the average trace and errors around them (errors need to be stated clearly, for e.g., SEM or SD) show that they are non-overlapping and different. I suspect tests such as the rank-sum test, which test the difference in the central tendencies (as opposed to the KS test, which tests the overall trend in the distribution of the continuous data), might reveal the difference between these values.

      We compared the absolute (not normalized) LC-MUA rate and SI during 2 sec time window preceding spindle onset and did not find any statistical differences. In Figure 7F, the difference during ~ 2 sec before the spindle onset is due to the z-score normalization to their own baseline.

      We revised the Result text to improve clarity.

      (12) Line 209: Modulation seems to be greater in ripp-spindles as shown in fig 6E-F, yet, the text and the interpretation are the opposite i.e,. iso spindles had greater modulation. Hence, authors might have to provide further clarifications or analyses.

      We corrected the labelling in all plots.

      (13) Line 316: Claims of "suppression of noradrenergic system facilitating the generation of hippocampal ripples and sleep spindles by memory synchrony" are not fully supported by data, as the data seem to be correlational. Also, claims of "preserved LC activity during ripples coinciding with sleep spindles suggest a role for NE in facilitating cross-regional communication underlying memory-related information transfer" lack clarity and contradict the earlier mechanism. Both "suppression" as well as "preservation" of LC neurons are proposed to mechanistically support memory synchrony and/or consolidation in two different brain states (awake and sleep). The authors might need to clarify how both suppression as well as preservation (which I assume is not an activation or positive modulation) of LC neurons can help in memory synchrony or consolidation.

      We revised this part of discussion by making it less speculative.

      Reviewer #3 (Recommendations for the authors):

      I would recommend that the authors optimize their figure and result presentation, as the current version of the manuscript is unclear in several places, limiting the interpretation of results.

      We substantially revised the manuscript to improve the results presentation and readability.

      (1) Multiple results are described but not shown quantitatively. Please plot quantifications and statistics (mean {plus minus} error and individual values) in relevant figures. For example, the results referenced on p. 4 (l. 113-116), p. 5 (l. 129-133, 143-147), p. 6 (l. 159161), p. 7 (l. 188-190), and p. 8 (l. 203-207) should be supported by explicit data plots.

      We have revised the manuscript to ensure all results are supported by quantitative and statistical analyses. We revised figures and legends and added new plots showing individual datapoints.

      (2) Improvements in figures and descriptions are needed. Below are some examples I found:

      (a) All figures with color scales lack labeling of the color axis, i.e., measure and unit.

      We have revised the figures accordingly.

      (b) Use precise labeling of axes such as "ripple-band power" and "LC-MUA firing rate", rather than just "power" and "firing rate".

      We have revised the figures accordingly.

      (c) Figure 1: Indicate behavioral state (wake vs. sleep) in the example trace.

      We have indicated the behavioral state (quiet awake) in the figure legend.

      (d) Define "peri-ripple" windows explicitly (e.g., {plus minus}6 s or {plus minus}30 s).

      We have revised the text and figure legends accordingly.

      (e) Clarify how "modulation magnitude" is calculated (line 96).

      We now illustrate the extraction of quantitative variables in Figure 2D

      (f) Figure 2C: The white overlaid mean trace lacks Y-axis labeling.

      We have added y-axis labeling.

      (g) Figure 3A: The labeling of "amplitude" is confusing when referring to firing frequency.

      We have corrected the figure labelling.

      (h) Figure 4B: Is the X-axis time from ripple onset?

      We have corrected the figure labelling.

      (i) Figure 4C-D lacks an X-axis or color legend.

      We have added x-axis and color legend.

      (j) Figures 5-6: Include tonic firing rates and time scales.

      We have added in the main text the time scales and average firing rates for LC single units and also show it in Supplementary Figure 1. Because the number of neurons contributing to LC multi-unit activity (LC-MUA) is unknown, we avoided averaging absolute firing rates for this signal. For LC-MUA, we implemented a normalization approach in which firing rates (50-ms bins) around ripple were scaled to a baseline period preceding the trigger event (−12 to −10 s). Importantly, unlike z-scoring, this normalization method preserved baseline differences across behavioral states, as shown in new Figure 5.

      (k) Add tonic firing rate baselines where relevant.

      We have added the Supplementary Figure 1 and new Figure 5 showing the difference in the LC baseline firing rate across behavioral states.

      (3) Minor Comments to add more clarity

      (a) Clarify "spike train" selection criteria (Methods, p. 4, line 93).

      We revised the text as follows: “In six out of twenty LC-MUA recordings, we could reliably isolate spikes from a total of 15 single units (LC-SUA, n = 4 rats).”

      (b) Define "EEG transients" (p. 4, line 109) and support with data.

      We revised the text as follows: “Indeed, transient spectral changes in the prefrontal EEG coincided with the occurrence of hippocampal ripples (Figure 2B).”

      (c) You refer to Figure 3E as a histogram (p. 5, line 128), but I believe it shows an average trace.

      We have corrected this typo.

      (d) Standard sleep scoring procedures normally involve EMG measurements (p. 6, line 154).

      We have replaced ‘standard’ with “previously established”.

      (e) Explain how surrogate shuffling preserves the distribution of behavioral states.

      We revised the text as follows: “We first verified that hippocampal LFPs (140– 250 Hz) triggered on these surrogate events lacked the ripple-specific frequency component (Figure 3C), and that the SI state did not differ between real ripples and surrogate events (Figure 3D).”

      (f) You refer to inter-ripple frequency (p. 6, line 168), which suggests time between ripples. Do you mean the "intra-ripple" or simply ripple frequency?

      We have corrected this typo.

      (g) Ensure all references cited in the text (e.g., p. 12, line 406) are included in the bibliography.

      We have updated the bibliography.

      (h) On p. 10, line 304-305 authors refer to observations related to offline memory consolidation. However, the present study does not contain any behavioral memory data.

      We have revised the Discussion to make it less speculative about the role of describe LC dynamics for offline memory consolidation.

      References

      Novitskaya Y, Sara SJ, Logothetis NK, Eschenko O (2016) Ripple-triggered stimulation of the locus coeruleus during post-learning sleep disrupts ripple/spindle coupling and impairs memory consolidation. Learn Mem 23:238-248.

      Yang M, Logothetis NK, Eschenko O (2019) Occurrence of Hippocampal Ripples is Associated with Activity Suppression in the Mediodorsal Thalamic Nucleus. J Neurosci 39:434-444.

      Yang M, Logothetis NK, Eschenko O (2021) Phasic activation of the locus coeruleus attenuates the acoustic startle response by increasing cortical arousal. Sci Rep 11:1409.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This paper aims to improve the accuracy of predictions of the impact of ITN strategies by developing a method to estimate duration of ITN access and use over time on a subnational scale from cross-sectional survey data and the numbers ITNs received annually. The subnational estimates are then input into a mathematical model to predict clinical cases under different ITN distribution strategies.

      Strengths:

      The approach is novel and addresses a useful and timely topic. It makes use of available routine data, and has considered all of the relevant components of ITN distributions.

      The authors have made revisions, particularly to the methods, appendices and title - leaving the paper easier to follow, and with a clear, consistent aim. The assumptions are clearly stated.

      Weaknesses:

      The weaknesses are shared with other models of a similar complexity - it is not easy for a casual reader to fully understand the model or the implications of the assumptions which were required to be made. That routine data is used is good for availability, but data quality may be an issue in some places.

      Reviewer #2 (Public review):

      Summary:

      The authors design a custom Bayesian model to estimate the probabilities of access, use and use given access of insecticide-treated nets in six African countries, providing sub-national estimates and inferring the average duration of ITN use and access. An individual-based model was employed to simulate malaria epidemics and estimate the effectiveness of different ITN distribution strategies. The study finds that the mean probability of use or access did not reach 80% (a universal coverage formerly targeted by WHO) for any of the regions even for biennial campaigns, demonstrates that switching from triennial to biennial distribution campaigns increases population use by 7.9%, and evaluates the impact of employing more efficient ITNs on P. falciparum prevalence.

      Strengths:

      The authors developed a data-driven model that accounts for data collection imperfections and sources of uncertainty while differentiating between ITN use and access. They developed a methodology to infer the timing of mass campaign from publicly available data instead of assuming fixed dates. The probability of use given access allows determining the regions where ITN distribution is least effective. This work can help better inform future interventions by identifying regions where increasing mass campaign frequency or employing better ITNs are most effective. Finally, in addition to insights on ITN access and use for the six countries analyzed, the paper contributes with a methodological framework that can likely be extended to other countries.

      Weaknesses:

      Since the models employed are rather complex, the methodology description may be hard to follow for some readers. In addition, the models assume many hypotheses, including exponential decay of ITN use/access and narrow prior distributions. It is worth noting that, in the revised version of the manuscript, the authors justified the choice of exponential decay and narrow prior distributions, and made a significant effort to clarify the methodology and the model equations.

      Comments on revised version:

      I appreciate the improvements made to the text. The methodology description is much clearer now. I have no further suggestions.

      We thank the reviewers and editors for their constructive and insightful comments throughout the review process.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      P8 'Improving ITN use' L218 

      The numbers do not seem add up to me. "...increases across all settings of 14.5% (95% CrI:14.5, 14.6), from 41.7%% to 49.6%. Greater increases are predicted to be seen for ITN use with mean use across all settings increasing from 58.0% to 66.2%, an increase of 19.5% CrI (95% CrI:19.5, 19.6)."

      Thank you for highlighting this. We have reviewed all reported results on mean use, access and use given access. The previous text reported a mixture of absolute and relative % changes, as well as a mixture of raw mean estimates across all regions and population-weighted means across regions. In the extract above we had inadvertently mixed different metrics. Given administrative-one regions can vary notably in population between different countries, we have ensured estimates are now consistently reported as population-weighted means, so that countries with finer-scaled administrative-one regions, such as Burkina Faso, do not artificially bias a raw mean estimate across all sub-national regions. We have also reported % changes as absolute percentage-point increases throughout, rather than relative ones to improve clarity.

      Methods p18: There is notation in the text which does not seem to be explained. It is in the appendices, but the appendices should be optional extra information rather than essential for understanding. 

      We have reviewed the text in the main methods to check notation explanations. Following this, we have removed a use of subscript $i$, which is only used in the appendices to explicitly indicate region-specific parameters, and have clarified that lambda is a decay parameter.

      There are assumptions made and these are clearly explained in the text. However, how much the highlighted results rest on the assumptions was not clear, and there was little on this in the discussion. 

      For example, it might seem disappointing that changing from triennial to biennial ITN campaigns would only lead to an increase from 41.7% to 49.6%. The most important assumptions driving this could be clearer. Additionally, after reading I was not sure what the likely consequences of the assumption that ITN are used continuously were.

      We have added some additional text “to the discussion to clarify the modest predicted increase under biennial campaigns may, in part, be influenced by our assumed exponential loss function, and have highlighted that larger increases in mean use could plausibly be predicted under alternative ITN loss functions”. However, we have also commented that our mean use estimates are broadly in agreement with time series modelled estimates by Bertozzi-Villa et al. (2021) who utilised a sigmoidal/smooth-compact loss function.

      In relation to the assumption of continuous use, we have added additional text in the ‘Historical use, access and retention times’ methods section to clarify that “if ITN use were systematically higher during high-transmission rainy seasons, our assumption of continuous use may underestimate the protective impact of ITNs during these periods”. As stated at the start of that paragraph, the data available from DHS surveys was too infrequent to investigate seasonal fluctuations.

      P14 The text seems to imply that current transmission intensity is the only criterion for decisions about interventions. However, it is likely that the reasons for the current intensity, such as vectorial capacity, historical transmission and interventions should also play a role. The wording could reflect this.

      We have added additional text to clarify that current transmission intensity should not be treated as the only criterion for deprioritisation decisions:

      “However, current incidence should be considered alongside the factors that gave rise to that transmission intensity, with caution exercised when deprioritising mass campaigns in areas where historically higher transmission may currently be suppressed by high ITN access, high use given access, or other interventions.”

      Minor points 

      There are several definite numbers in the first paragraph of the Introduction - these are estimates rather than the absolute truth, but the wording does not acknowledge that there is uncertainty.

      We have made minor edits to clarify that these values are estimates rather than exact quantities. Measures of uncertainty, such as credible intervals were not always possible to source; for example, some of these are median estimates inferred from figures in Bertozzi-Villa et al. (2021).

      L634 typo - logisitic 

      Now corrected.

      L1731 typo https://https://

      Now corrected.

      L881 "access at random" - perhaps not the easiest for non-modelers

      We have re-written this to clarify “when ITNs in a household can provide access to more individuals than the number of users, access is assigned at random to non-users within each household under our framework”.

      Appendix 1, table 1: Using alpha for both age and also overdispersion on use or access is of course valid, but I found it a little confusing.

      To avoid confusion, we have added the following clarification in brackets:

      “Meanwhile, the overdispersion parameter, $\alpha_i^0$ (unrelated to the notation for ITN age), controls the variability of the probability of individual access around the mean”

      I suspect that the model was actually fitted in Stan via the R interface rstan (L589, L1151 and elsewhere).

      We have now clarified this throughout.

    1. Author response:

      Response to Reviewer #1

      Our work builds upon the foundations of what we term the “CM family”, specifically the Connectome Model (CM) introduced by Kovács et al.. This was a deliberate choice, as our objectives substantially overlap with those of works in this family. Moreover, we wished to avoid reinventing the wheel—starting instead from a solid body of work with validations we found convincing (thereby inheriting this solidity) and, importantly, addressing the same research community using a “familiar” conceptual language. We therefore wish to clarify how our contributions indeed constitute new conceptual insights into the genomic specification of neural circuitry.

      The function implemented by a neural circuit clearly depends on how information propagates between its nodes and connections; the contribution of synapses—their number and properties—cannot be neglected when understanding, manipulating, or designing such function. To the best of our understanding, in Kovács et al., the primary objects of interest are binary connectomes (presence or absence of synapses) or weighted connectomes where “in the occasion of multiple [genetic] rules contributing to the same link”, “the weight of each link correspond[s] to the number of rules involved”. In Barabási et al., a “relaxed” version of the CM directly provides weights for an artificial neural network without explicitly specifying how each weight might result from the combination of a specific number of synapses and their respective properties. The random variable formalism and the introduction of conductances that we propose precisely add this further—yet important—element of complexity and representational detail: synaptic multiplicity. This extends existing models with the hope of laying the groundwork for what could, in the distant future, become a technology capable of producing neural circuits genetically programmed to implement a defined function.

      Regarding the proposed validation, we acknowledge its limitations, but we clarify that at the time this work was conducted, to the best of our knowledge, no public datasets existed to perform validation as the reviewer envisions. We therefore did the best that was materially feasible: we assumed the biological correctness of the model (also based on the validations accompanying the models upon which ours was built) and verified, through simulation, that it could be used to obtain genetic variables of interest capable of producing neural agents able to solve a pre-specified task—even with the additional constraint of genetic rules derived from experimental data.

      Response to Reviewer #2

      We address the points raised by Reviewer #2 in the following paragraphs.

      Regarding point (1), we agree with the reviewer that considering single-gene expression features is a simplification, especially in the case of chemical synapses. However, as with the CM, our model can also be extended to account for combinatorial rules. One possibility is to add columns to the X matrix, as many as there are gene expression patterns of interest. For each new column, a function would be defined to compute the expression feature from the expression features of the genes involved in the pattern, and this function would be used to populate the values of the new columns. The O matrix would likewise be updated with the corresponding new probabilities. While such extension is possible, it is important to note that this gives rise to the problem of combinatorial explosion of genetic rules, with the consequent construction of matrices whose dimensionality becomes difficult to handle. Moreover, the biological plausibility of the model would then shift toward how these functions are defined, along with the interpretation of the values contained in the X matrix. Depending on the use case of our model, one possible solution to the combinatorial explosion problem could be to consider only expression patterns valid for synapse formation by extracting this information from available experimental data, thereby restricting the number of rules. We acknowledge that this problem remains open and will require more precise formulations and future work.

      Regarding point (2), Equation (11) can be derived from the assumption that the various synapses between two neurons behave as resistors in parallel. Accepting this, the equivalent conductance Guv, as denoted in the paper, can be expressed as the sum of all conductances between neurons u and v. Moving to the random variable formalism and having defined 𝒢 as the random variable representing the “signed conductance of a synapse randomly selected from the ones that connect neurons u and v”, the equivalent conductance (as a random variable) becomes ℬ·𝒢. Recall that ℬ is the random variable representing the number of synaptic connections between two neurons of interest. At this point, under the further assumption that the random variables ℬ and 𝒢 are independent, the expectation of the equivalent conductance can be calculated as the product of the expected values of ℬ and 𝒢. Equation (11) follows immediately from this. We acknowledge that these assumptions may not correspond to biological reality, but we consider them a reasonable starting point for addressing the problem.

      Finally, we explain the reasons why the baselines suggested by the reviewer are not included in the work. We did not train classical MLPs because the main objective of the work was not to develop new bio-inspired architectures aimed at generically improving the performance of neural networks in RL, and we deemed it an additional source of confusion to propose a comparison that would suggest this direction. The main objective of the work is instead to contribute to the modeling of synaptogenesis and to lay the groundwork for—or advance the state of knowledge of—what will be a future technology that allows us to manipulate it (synaptogenesis). A similar reasoning applies to a potential baseline in which the weight matrix is constructed from Equation (7). Again, the interest is not in verifying that conductances provide a performance advantage, but rather that they are a necessary element for a sufficient level of biological plausibility. Beyond this, the exclusive and direct use of matrix B in the simulation of synaptogenesis introduces a quantization problem as described in the Appendix.

      Response to Reviewer #3

      We believe the concerns raised by the reviewer regarding the weaknesses of the work are legitimate. We wish to emphasize that all claims made in the paper were made in good faith, with the intent to generate enthusiasm for the discipline while avoiding excess or the assertion of anything incorrect or untruthful. Given that the work is inherently interdisciplinary, we recognize that reader expectations depend on their reference community, and we clarify that our primary area of expertise is AI, and that the biological claims were therefore made from this perspective.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      While the revised manuscript includes additional methodological details and a supplementary comparison with conventional NMF, it would be great if the authors could add the point below as limitations in the manuscript or change the title and abstract accordingly, since core issues remain:

      (1) The study claims to evaluate rehabilitation outcomes without demonstrating that patients actually improved functionally

      (2) The comparison with existing methods lacks the quantitative rigor needed to establish superiority

      (3) The added value of this complex framework over much simpler alternatives has not been demonstrated

      The strength of evidence supporting the main claims remains incomplete. I would encourage the authors to consider discussing these points

      (1) including or adding a limitation section about functional outcome measures that go beyond clinical scale scores, (2) providing/discussing quantitative benchmarks showing their method outperforms alternatives on specific, predefined metrics, and (3) clarifying the clinical pathway by which these biomarkers would inform treatment decisions.

      We thank the reviewer for their thoughtful consideration of our study, and now better understand their perspective on the limitations of the study. We now see the importance of the aspects of functional recovery the reviewer has highlighted in the context of our work, as the clinical measure we focused on (i.e. FMA-UE) does not capture recovery at activities and participation-levels of the ICF model. Although the FMA-UE is a gold standard measure for assessing post-stroke recovery, it is limited in scope to gross motor functions.

      To more accurately describe the aspects of functional recovery the biomarkers in our study reflected, we have extensively revised the terminology used throughout the paper. For example, in the abstract we now include “…From these patterns, we derived new biomarkers that stratified patients by gross motor impairment severity and therapeutic responsiveness, each associated with unique physiological signatures.” and go on in the abstract to now highlight the limited scope of the evidence towards functional recovery more broadly also: “Future research should employ this framework to identify biomarkers of activities- and participation-related functional recovery.” In the rest of this paper, we also make this distinction clear, for example at the beginning of the results section: “The cohort of stroke survivors overall experienced a statistically significant increase at FMA-UE (Pre-treatment: 43.1±13.2, Post-treatment: 49.1±13.6 (t= -7.84, p<0.001)), representing a clinically important effect from rehabilitation on the gross motor functions of the upper-extremity (Page et al., 2012).” Finally, we have now added a limitations section, as the reviewer advised, where we specifically detail the scope of evidence provided in this study and how future research could build on it:

      “Limitations

      Although the FMA-UE is a gold standard measure of post-stroke treatment outcomes (Meyer et al., 1975; Page et al., 2012), it does not capture the impact of rehabilitation on patients' ability to perform activities-of-daily-living or to participate in daily life. Hence, interpretations of the identified biomarkers are currently limited to gross motor function impairment and recovery. Future research should employ this framework to quantify biomarkers that correspond to other important aspects of patients' recovery (e.g. functional independence, subjective experiences), thus offering a more complete evidence base for its clinical utility.”

      With these changes, we believe this manuscript more accurately describes the scope of the biomarkers analysed and hence no longer offers incomplete evidence towards stated claims.

      Regarding the reviewers second and third points concerning the validity and advantages of this framework against current approaches, in this study we applied a framework that builds on two previous papers (O’Reilly & Delis, 2022; O’Reilly & Delis, 2024). In both of these papers, we compared basic aspects of the framework to the current prevailing approach and most relevant comparative for this line of research in muscle synergy analysis, that is non-negative matrix factorisation (NNMF).

      To briefly outline this existing foundation of evidence, in O’Reilly & Delis, 2022 we dedicated most of the discussion section (i.e. sections 4.1 and 4.2) along with a supplementary materials document to comparisons with this approach. In section 4.1, we illustrate the continuity of this framework with what has come before in simpler methodologies such as NNMF and then went on in section 4.2 to show the novel insights and opportunities that can generated from our framework. Additionally, in the corresponding supplementary materials of that paper, we directly compared our framework with three different models from the established NNMF approach (i.e. spatial, temporal and space-time) by applying them to the same datasets, again highlighting points of congruence and additional utility with our framework. Building on this work, in O’Reilly & Delis, 2024, we also ensured that developments of this framework both align with previous research and credibly improve upon them methodologically. For example, Fig.5 and Fig.6 and associated text of that paper illustrates a direct comparison of our framework with the NNMF methodology, showing that it provides additional functional and physiological relevance and predictive capacity to the components extracted. Further, in the results of that paper we also directly compared the generalisability of the extracted components when extracted using our chosen dimensionality reduction approach vs other approaches promoted in the neurosciences more generally (e.g. non-negative Canonical-Polyadic (CP) tensor decomposition (Williams et al (2018)), showing that we extracted more robust components across participants and tasks.

      This previous work directly supports the credibility of basic aspects of the framework and its outputs compared to other established approaches. We have directed readers towards this previous research in the methods section of the current study: “Further comparisons with conventional approaches can be found in our previous work developing this framework (O’Reilly & Delis, 2022; O’Reilly & Delis, 2024).”

      Continuing, and building on the credibility of these basic aspects of the framework, as the reviewer previously suggested, we have included additional supplementary material in the current study illustrating how the biomarkers generated from our approach could not be found using conventional methods. In these supplementary analyses, we employed a much simpler but conceptually aligned pipeline involving NNMF and agglomerative clustering on the same dataset and directly compared the outputs, highlighting commonalities and where our approach improves significantly upon this established approach. The advancements we demonstrate here also address recognised limitations in the current NNMF approach for clustering activation coefficients (see Scano et al 2017), a point we now highlight directly in the revised manuscript:

      “Enhanced interpretability of extracted components and clusters.

      As our framework maps muscle interactions to a specific task parameter, we yield population-level motor components that correspond more consistently to meaningful biomechanical and physiological functions that can be interpreted across the dimensions of the specified task parameter. The proposed clustering approach also offers enhanced interpretability, addressing key limitations in the application of clustering approaches to the activation space of conventional muscle synergy analysis (e.g. different activation timings) (Scano et al., 2017).”

      Taken together, we believe the extensive comparisons made in our previous work on this framework and direct comparisons made in this study provide sufficient evidence towards its added value for the field beyond current approaches.

      References

      Ó’Reilly D, Delis I. A network information theoretic framework to characterise muscle synergies in space and time. Journal of Neural Engineering. 2022 Feb 1;19(1):016031.

      O'Reilly D, Delis I. Dissecting muscle synergies in the task space. Elife. 2024 Feb 26;12:RP87651.

      Williams et al. (2018) Unsupervised discovery of demixed, low-dimensional neural dynamics across multiple timescales through tensor component analysis. Neuron 98:1099–1115.

      There are specific, relatively minor points, that require attention

      The authors write: "we did not focus on such complementary evidence in this study." This is a weakness for a paper claiming to provide "biomarkers of therapeutic responsiveness." The FMA-UE threshold defines responders, but there's no independent validation that patients actually functioned better in daily life. Can you please clarify?

      See above for our response on this important aspect of the reviewer’s commentary.

      Maybe I missed the exact point about this, but with the added NMF plot, the authors list 'lower dimensionality' among their framework's advantages, but the basis for this claim is not clear because given that 12 network components were extracted compared to 11 "conventional" synergies. Can you please clarify, as it is not clear. You claim 'lower dimensionality' as an advantage of the proposed framework (in the Supplementary Materials), yet you extracted 12 components (5 redundant + 7 synergistic networks) compared to 11 synergies from the conventional NMF approach, which does not support a clinical / outcome advantage of this method. Please clarify.

      We agree with the reviewer that this statement is confusing given that overall, across separate decompositions for redundant and synergistic networks compared to the single decomposition using NNMF, there are more dimensions to consider in our frameworks output. For this reason, we have removed this statement from the updated manuscript.

      Reviewer #2 (Public review):

      This study presents an important analysis of how interactions between muscles can serve as biomarkers to quantify therapeutic responses in post-stroke patients. To do so, the authors employ an information-theoretical metric (co-information) to define muscle networks and perform cluster analysis.

      I thank the authors for improving the clarity of the Methods section; the newly added Figure 5 is very helpful.

      One minor suggestion is that the authors should avoid overloading the notation "m" for both the EEG measurement and the matrix of II values (Eq. 1.1), which I now realise was the source of some of my initial confusion. I suggest that the authors use separate notation for these two quantities.

      We thank the reviewer for their consideration and positive outlook on our study. In the updated manuscript, we have adjusted the notation for equation 1.1 so that it doesn’t cause confusion with earlier text.

      Recommendations for the authors:

      Reviewer #1 raised critical concerns about the method's ability to identify functional improvements resulting from rehabilitation protocols. In this regard, the study's translational impact remains limited, and the authors should address these limitations in a revised version. The Reviewing Editor and both reviewers agree that the "Strength of Evidence" of the manuscript cannot be improved without a major revision, given the above-mentioned aspects.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses

      (1) One of the main EEG results is based on the weighted phase lag index (wPLI) between oscillations in the alpha and theta bands. In my opinion, this is problematic, as wPLI measures the locking of oscillations at the same frequency. It quantifies how reliably the phase difference stays the same over time. If these oscillations have different frequencies, the phase difference cannot remain consistent. Even worse, modeling data show that even very small fluctuations in frequency between signals make wPLI artificially small (Cohen, 2015).

      In response authors stated : "Additionally, the present study referenced previous research by using the wPLI index as a measure of cross-frequency coupling strength31,64-66"

      Unfortunately, after checking those publications, we can see that in paper 31 there is no mention of "wPLI" or "PLV." In 64 and 65, the authors use wPLI, but only to measure same-frequency coherence, whereas cross-frequency coupling is computed by phase-amplitude coupling or cross-frequency coupling also known as n:m-PS. In 66, I cannot find any cross-frequency results, only cross-species analysis. This is very problematic, as it indicates that the authors included references in their rebuttal without verifying their relevance.

      31 de Vries, I. E. J., van Driel, J., Karacaoglu, M. & Olivers, C. N. L. Priority Switches in Visual Working Memory are Supported by Frontal Delta and Posterior Alpha Interactions. Cereb Cortex 28, 4090-4104, doi:10.1093/cercor/bhy223 (2018).<br /> 64 Delgado-Sallent, C. et al. Atypical, but not typical, antipsychotic drugs reduce hypersynchronized prefrontal-hippocampal circuits during psychosis-like states in mice: Contribution of 5-HT2A and 5-HT1A receptors. Cerebral Cortex 32, 870 3472-3487 (2022).

      65 Siebenhühner, F. et al. Genuine cross-frequency coupling networks in human resting-state electrophysiological recordings. PLoS Biology 18, e3000685 (2020).

      66 Zhang, F. et al. Cross-Species Investigation on Resting State Electroencephalogram. Brain Topogr 32, 808-824, doi:10.1007/s10548-019-00723-x (2019).

      We thank the reviewer for this critical methodological correction. We fully agree that the weighted phase lag index (wPLI) is designed for same-frequency phase synchronization and is not appropriate for cross-frequency coupling (CFC). In our original rebuttal, we incorrectly cited references that did not support the use of wPLI for CFC. We apologize for this error and have thoroughly revised our analysis and manuscript.

      What we have done:

      (1) Replaced wPLI with proper 1:2 cross-frequency phase synchrony (CFS).

      We now compute 1:2 CFS using the phase-locking value (PLV) between theta (4–7 Hz) and alpha (8–14 Hz) oscillations, following established methodologies (Siebenhühner et al., 2020, PLoS Biol; Palva et al., 2005, J Neurosci). Specifically, for each electrode pair we compute:

      .The factor 2 accounts for the 1:2 frequency ratio (theta:alpha = 1:2).

      (2) Updated all relevant sections – Methods (“Interregional connectivity”), Results (Figure 8, Figure 9), Discussion, and Figure legends – replacing “wPLI” with “1:2 CFS (PLV)” and providing the correct formula and citations.

      (3) Corrected the reference list to include the appropriate methodological papers (Siebenhühner et al., 2020; Palva et al., 2005) and removed irrelevant citations.

      We believe this revision fully resolves the reviewer’s concern. Notably, the empirical results remained qualitatively unchanged (PLV and wPLI gave highly consistent values due to the absence of zero‑lag artifacts in cross‑frequency coupling), so the main conclusions of the paper are unaffected.

      (2) Another result from the electrophysiology data shows that the attentional capture effect is positively correlated with the mean amplitude of alpha power. In the presented scatter plot, it seems that this result is driven by one outlier. Unfortunately, Pearson correlation is very sensitive to outliers, and the entire analysis can be driven by an extreme case. I extracted data from the plot and obtained a Pearson correlation of 0.4, similar to what the authors report. However, the Spearman correlation, which is robust against outliers, was only 0.13 (p = 0.57) indicating a non-significant relationship.

      Cohen, M. X. (2015). Effects of time lag and frequency matching on phase based connectivity. Journal of Neuroscience Methods, 250, 137-146

      We thank the reviewer for raising this important statistical issue. We have conducted a thorough robustness analysis and revised our interpretation accordingly.

      What we have done:

      (1) Removed the original scatter plot (Figure 7) to avoid overinterpretation. No replacement figure is provided; instead, all results are reported in text.

      (2) Conducted leave‑one‑out cross‑validation.

      The Pearson correlation remained positive across all 24 iterations (range: 0.183–0.497, mean r = 0.430 ± 0.055), confirming that no single participant solely drove the direction of the effect.

      (3) Reported Spearman rank correlation (r = 0.13, p = 0.57), which is more robust to univariate outliers.

      (4) Acknowledged the sensitivity – p‑values from leave‑one‑out iterations ranged from 0.0158 to 0.4025, indicating that statistical significance is not fully robust to sample composition.

      (5) Revised the text to present this as preliminary evidence rather than a definitive conclusion. Specifically, we state: “Thus, we interpret this as preliminary evidence that occipital alpha activity may be associated with the priority state within VWM, warranting replication in larger samples.” The Discussion also includes a dedicated limitation paragraph.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      (1) (Figure 1): Quantification of CSF1R-GFP<sup>+</sup> and CD11c-eYFP<sup>+</sup> cells in PDPN<sup>+</sup>LYVE-1<sup>-</sup> vs. PDPN<sup>+</sup>LYVE-1<sup>+</sup> regions. “This would demonstrate selective accumulation or retention of myeloid cells at the cribriform plate niche."

      We thank the reviewer for this important suggestion. The representative images in Figure 1(Bottom) establish the partial justification for the cell sorting and sequencing strategy in Figure 3, which relies heavily on myeloid cells in contact with PDPN. Importantly, our previous publication Hsu et al. 2022 has quantified elevated Cd11c+ cells in contact with the Cribriform lymphatic niche. Figure 1 in this context seeks to show PDPN as an additional and broader marker for the meningeal and lymphatic tissue at the brain's border. Because PDPN represents more surface area, PDPN+Lyve-1- regions would likely show more immune cell accumulation but our primary argument is simply that myeloid cells also accumulate in both PDPN regions. As a result we argue the quantification of cells in Lyve-1+ and negative regions is not necessary. We have added a sentence to the text which explains the intention of the figure.

      “Additionally, while PDPN labels the cribriform plate lymphatic vasculature, it also defines the meningeal-immune interface at the border of both the olfactory bulb and olfactory nerve bundles.”

      (2) While the PostContact-seq strategy is innovative (Figure 3), additional justification is needed to demonstrate that tissue dissociation did not artificially disrupt PDPN-myeloid contacts. The relatively small proportion of live PDPN-rich doublets (~2.5% total aggregates and ~18% PDPN+ within total aggregates) raises questions about representativeness compared with in situ observations. The authors should also more explicitly elaborate on why PostContact-seq was favored over alternative approaches such as PIC-seq.

      We acknowledge this important methodological concern. We have expanded the Methods section and added a dedicated paragraph in Results addressing the following:

      Tissue dissociation controls: Dissociation protocols were used specifically to minimize cell-cell adhesion. Unfortunately, we cannot perform parallel dissociations of naive (non-EAE) cribriform plates for scRNAseq because PDPN<sup>+</sup>-containing doublets are essentially non-existent. This is also supporting the representativeness compared to in situ observation. Doublets are significantly enriched in EAE tissue compared to naive controls, arguing against artifactual aggregate formation.

      Representativeness of ~2.5% doublets: While the absolute proportion of doublets is modest, this is consistent with in situ observations where myeloid-PDPN contacts are spatially restricted to the outer perineural and meningeal niche rather than globally distributed. We argue this is simply the enrichment of a rare interaction rather than a limitation.

      PostContact-seq vs. PIC-seq: “PIC-seq (Giladi et al., 2020) sequences intact doublets and relies on specialized deconvolution tools to parse apart data. PostContact-seq leverages the cellular contact signatures post-dissociation, making it a more accessible system. However, we now explicitly discuss this comparison in the Results and acknowledge PIC-seq as a complementary future approach in discussion.

      (3) (Figure 4B): Clarification of integration across four methods; consideration of CellChat/NicheNet: The authors stated that results regarding cell-cell interactions were integrated across four intercellular communication methodologies (Figure 4B), but this integration is not clearly described in either the Results or Method sections. This needs clarification. Moreover, the interaction analysis in Figure 4B seems to rely on TALKIEN, which does not incorporate prior ligand-receptor knowledge. Given the availability of widely used tools, such as CellChat and NicheNet, the authors may consider cross-referencing their findings.

      We have revised the Methods sections to clearly describe our strategy for the TALKIEN analysis. Importantly, TALKIEN does integrate ligand receptor libraries from four sources: CellChat, CellPhoneDB, iCellNet, and the Ramilowsky datasets to generate its figures. Interactions reported in Figure 4B are those supported by this analysis, and we updated the text accordingly.

      (4) Pseudotime trajectory analysis of CCR2<sup>+</sup> monocyte differentiation.

      "A pseudotime trajectory analysis may be valuable to test whether CCR2<sup>+</sup> monocytes preferentially differentiate into CHI3L3<sup>+</sup> macrophages, PD-1<sup>+</sup> DCs, or other subsets."

      We thank the reviewer for this insightful suggestion. We added a complete pseudotime analysis (Author response image 1). The pseudotime analysis tracks a continuous developmental trajectory starting from cDC2 and early macrophage populations (Pseudotime = 0, dark purple) and progressing through the main macrophage body toward an activated terminal state (Pseudotime = 16, yellow). Crucially, the trajectory correctly excludes non-continuous lineages such as resident microglia and lymphoid cells. This progression is functionally validated by the transient upregulation of the recruitment marker CCR2 during intermediate stages, which subsequently downregulates as cells transition into a mature phenotype.

      Author response image 1.

      (5) FACS-based validation of macrophage immunosuppressive signatures.

      "Validation using the same post-contact vs. no-contact sorting strategy would strengthen the conclusions."

      This is an excellent suggestion and will be the topic of future detailed investigation focusing on the cellular and molecular reprograming of the immunosuppressive microenvironment at the cribriform plate.

      (6) Identity of CD45IV<sup>+</sup> cells in contact with PDPN<sup>+</sup> cells (Figure 6B-C); gating strategy; tissue co-labeling. "Provide a gating strategy demonstrating that these are CD11b<sup>+</sup>CD11c<sup>+</sup> DCs... whether dying cells are PD-1<sup>+</sup>... co-labeling for PD-1, cleaved caspase-3, and CD11c-eYFP."

      A full gating strategy (now Figure S5) demonstrate sequential gating from Cells → Doublet → PDPN doublet → CD11b<sup>+</sup> CD11c <sup>+</sup> → CD45IV<sup>+</sup> (intravascular exclusion positive) within the doublet gate.

      (7) (Figures 1F-H): Morphological differences of CD11c<sup>+</sup> cells.

      We have added commentary to the Results section noting that “CD11c<sup>+</sup> cells in the olfactory bulb parenchyma display a ramified, microglia-like morphology consistent with tissue-resident or parenchymal surveillance cells, whereas those infiltrating the cribriform plate perineural niche show a rounded, non-ramified morphology more consistent with recently recruited monocyte-derived DCs or macrophages.” This morphological distinction aligns with our scRNAseq-defined population differences and supports the notion that the cribriform plate niche shapes distinct myeloid states.

      Reviewer #1 (Recommendations for the authors):

      (1) (Figure 1C): MHCII counts vs. MFI discrepancy

      Thank you for catching this. The text has been corrected to reflect that we counted number of cells in the PDPN+ region of the cribriform plate

      (2) Proximity ligation assay (PLA) for macrophage-fibroblast ligand-receptor pairs

      We appreciate this suggestion. PLA validation of all predicted pairs is beyond the scope of this revision, and we are primarily interested in interactions occurring in vivo and in situ. Future studies will investigate properties of these cells using PLA.

      (3) (Figure 2E vs 2G): Inconsistent quantification strategies; CSF1R-GFP/CD11c-eYFP validation of CHI3L3<sup>+</sup>/Arg1<sup>+</sup> cells

      Arg1 signal was more broadly expressed and it was hard to distinguish 1 cell vs 2 cells in close proximity. Which is why we elected to use %Arg1 in PDPN+ regions. Conversely CHI3L3 staining revealed more easily identifiable single cells for quantification. Nonetheless both methods achieve the purpose of outlining that these cells increase in number a the cribriform plate lymphatic regions.

      (4) (Figure 3E): Pro-inflammatory features of migratory DCs vs. suppressive interpretation.

      "Pdcd1lg2, Cd80, Cd83 are associated with T-cell activation — how does this align with an immunosuppressive niche?"*

      This is an excellent point that we now explicitly address in the Discussion. The co-expression of Pdcd1lg2 (PD-L2), Cd80, and Cd83 by migratory DCs likely reflects a tolerogenic activation state rather than a conventional immunostimulatory one. PD-L2 co-expression with costimulatory molecules has been documented in tolerogenic DCs that can engage T cells while simultaneously delivering inhibitory signals via the PD-1/PD-L2 axis (inhibiting rather than amplifying T-cell responses). Furthermore, the lower abundance of migratory DCs in post-contact samples relative to no-contact samples may reflect that cells expressing this immunological synapse machinery are preferentially undergoing programmed cell death (consistent with Figure 6 findings), leaving a post-contact population enriched for the macrophage-dominated tolerogenic signature. We now discuss this interpretation explicitly.

      (5) (Figure 5F-G): Gating strategy for PD-1<sup>+</sup> DCs — PDPN inclusion

      The gating strategy has been clarified in Figure S6 (new figure) and the Methods section. PD-1<sup>+</sup> DCs shown in Figures 5F-G were gated from the PDPN<sup>+</sup> doublet fraction specifically, paralleling the outlined scRNAseq approach. We have added PDPN as an explicit gate in the updated Figure S6A

      (6) (Figure 5H): Discrepancy between text and data — "lowest genes" in PD-1neg DCs.

      We apologize for this error. The text has been corrected: the data in Figure 5H show that chemokines, ISGs, and MHC genes are among the highest expressed in PD-1<sup>+</sup> DCs (not PD-1<sup>-</sup>), consistent with the heatmap shown. This aligns with the interpretation that PD-1<sup>+</sup> DCs, while tolerogenic, retain antigen-presentation and chemokine-signaling capacity.

      (7) Figure 6 reference errors in Results text

      Corrected throughout — all references to cell death/apoptosis data now correctly cite Figure 6.

      Reviewer #2 (Public review):

      (1) Sorted populations — in vivo interactions vs. ex vivo aggregation artifacts

      As detailed in our response to Reviewer 1 (Weakness 2), due to the non-detectable doublet frequency in non-EAE mice, we believe that PDPN<sup>+</sup> doublet enrichment is EAE-dependent. We also used cold dissociation conditions. We also note that the transcriptional signatures recovered from PDPN<sup>+</sup> doublets are not simply a mix of independently sorted PDPN<sup>+</sup> and myeloid single-cell transcriptomes, they contain unique interaction-associated gene programs (e.g., elevated Pdcd1, tolerogenic markers) not present in non-contact controls, arguing for biologically meaningful contact rather than artifactual aggregation.

      (2) PDPN as stromal vs. lymphatic endothelial cells — which is most relevant?

      We have clarified throughout the manuscript that PDPN in IHC marks at least two distinct populations at the cribriform plate: (1) PDPN<sup>+</sup>LYVE-1<sup>+</sup> lymphatic endothelial cells and (2) PDPN<sup>+</sup>LYVE-1<sup>-</sup> meningeal fibroblasts/perineural sheath cells. It is hard to dissociate which is most relevant in the present study.

      (3) Descriptive nature; lack of functional correlates; implications need further discussion.

      We appreciate this honest assessment. We agree that functional experiments (e.g., conditional deletion of DC populations at the cribriform plate, blockade of PD-1/PD-L1 axis, lymphatic ablation) will be critical for establishing causality and are ongoing in the laboratory. In this revision, we have:

      (1) Added a pseudotime analysis as a computational functional inference.

      (2) Refined the Discussion to explore functional implications, including how tolerogenic conditioning at the cribriform plate may limit cervical lymph node priming, parallels with perineural immunosuppression in cancer, and therapeutic opportunities (e.g., modulating this niche to enhance or dampen CNS autoimmunity).

      Reviewer #2 (Recommendations for the authors):

      (1) (Figure 1E): What does PDPN thickness increase represent?

      We have added clarification to the Results and Discussion. Based on our data, the increased PDPN<sup>+</sup> layer thickness during EAE most likely reflects a combination of: (1) increased PDPN expression per cell (supported by elevated MFI in flow cytometry), (2) cellular hypertrophy of existing PDPN<sup>+</sup> cells. However we cannot fully discriminate between these mechanisms with the current data and acknowledge this as a limitation.

      (2) (Figure 2A): In Figure 2A, can the authors provide a healthy control example to pair with 2A? Is the Chi3L3 expression "below" the plate...in the mucosa, associated with EAE, or the same in steady state? The images in 2D are hard to appreciate at the current size.

      Healthy (naive) control images are included in Figure 2D for direct comparison with EAE tissue, we added zoomed images of each panel to provide clearer context for the disease-associated changes in myeloid cell distribution and M2 marker expression.

      (3) What is the denominator for the quantification in 2E? Is this per unit area? If so, is it the PDPN area or the total cribriform plate region area? If the area of PDPN increases (as the authors show), then the potential area that can hold YM1+ cells also increases, so the absolute number of cells comparison isn't that fair.

      We have added this distinction to the results.

      (4) The same goes for 2G; however, in G, the quantification is "% Arg1+" ----percentage of what? The increase in Arg1 expression is striking, but it's also striking how similar the PDPN network appears between healthy and EAE in Figure 2F.

      We have added this distinction to the results And added a label of quantification to Figure 2G.

      (5) Are these increases in Arg1+ cells occurring in the meninges of EAE mice? Or is this specific to perineural areas at the cribriform plate? In a sagittal plane, are these cells clustered tightly at the cribriform plate, or do they extend outward along the ON tracts?

      These are clustered tightly in the meningeal regions and along ON tracts. We do not have any sagittal sections available for further proper analysis.

      (6) In Figure 2, some panels are labeled "merge" -what does this mean? The DAPI label within the figure is also impossible to see.

      Figure labels have been adjusted. Merge is a common label which identifies panels with all channels merged together in a series.

      (7) Figure 3: The authors sort cells that interact with PDPN+ CD31+ double-positive cells before the scRNAseq analysis. However, it's not clear from these data that the PDPN expansion observed in their histochemistry is on stromal or endothelial cells. As the authors note, PDPN "also efficiently labels meningeal layers surrounding them along the olfactory nerve layer, including fibroblasts and their associated extracellular matrix (ECM)". Can the authors more clearly explain the rationale for using CD31 in this gating strategy?

      We sorted for CD11b+CD45+ (immune), CD31+ (endothelial), PDPN+ (meningeal fibroblasts). CD31 was used to isolate myeloid cells and endothelial cells at the brain’s borders.

      (8) Also, without having to do scRNAseq, could the authors compare the interacting populations for cells stuck with PDPN+CD31neg cells? Figure 3B indicates that a good number of these PDPN+CD31neg cells are present in the sort.

      We did not isolate PDPN+CD31- cells from our sort, in our experience these are mostly fibroblasts though. Future studies will look at cells which adhere specifically to PDPN+CD31- aggregates.

      (9) The interacting cells seem to have a particular affinity for the sorted endothelial cells. However, it's not clear if these cells are simply seizing an opportunity to stick together once the cells are mechanically separated and spun down, or were together in vivo. The authors should determine how many of these cell types are maintaining an in vivo contact or simply are efficient at making new contacts ex vivo. One approach would be to take EAE tissues from CD45.1 and CD45.2 congenic animals and mechanically separate them together. Then the composition of doublets can be analyzed for the frequency of CD45.1/2 doublets or CD45.1 and CD45.2 single positive doublets....and also which cell types are contributing to these doublets. This will test how much of this interaction is driven by ex vivo stickiness or in vivo, and also give some idea about the inherent ability of these immune cells to find and engage PDPN cells.

      This is a limitation of the current study, and you have provided an excellent experiment and one we have added to discussion.

      (10) Figure 4: I'm confused about Figure 4. If I'm reading this correctly, these are the same data from Figure 3 that were sorted for CD31 positivity. If that's the case, how are there fibroblasts in these data? Does this represent an aggregation of endothelial, fibroblast, AND immune? (CD31, PDPN, and CD11c).

      Yes we suspect that endothelial, fibroblast, AND immune aggregates are highly heterogeneous. Without negative sorting/gating we are left with high number of immune cells in or sorting paradigm.

      (11) The authors comment on the relatively unclear biological significance of PD1 expression by DCs (non-T cells) and note their previous report on PD1 ligand expression in this cribriform region. Do the authors detect differential PD1 ligand expression in this current study (singlet vs aggregate)?

      We have not detected any significant difference in CD274 expression between non-interactor and interactor populations.

      (12) Are the FACS data Supplemental Figure 2 on singlet vs doublet DCs performed after Liberase treatment? The FACS plots for both doublet and singlet populations look very different in how they are rendered, with large cell numbers in the 10^-4 range for the doublet groups. Why is this?

      No liberase treatment was given in these experiments, we have updated the figure legend.

      (13) It seems like the figure labeling has gone awry. On page 9, what should be Figure 5 is being called Figure 4...and further on, Figure 5 is being used for Figure 6 ("Blood derived" data)---this makes it pretty confusing.

      This has been corrected. Thank you.

      (14) On page 10, the authors have written "Lowest genes in PD-1- DCs included chemokines CXCL9, CXCL10, IL-12b, interferon-stimulated genes (Ifit1, Ifit2 and Ifit3) and several MHC-related genes (H2-M2, H2-Eb2, H2-DMb2)". Is this correct? Based on my reading of the figure, "5H" is that not PD-1+ DCs instead of PD-1- DCs? Also, there is a typo, "Cxck10".

      Thank you for pointing this out. We have corrected.

      (13) It's not clear what the statement "...these data support that Pdcd1 expression in migratory DCs exhibits an immunosuppressive gene signature..." means. The PD-1 marker cannot "exhibit" anything by itself. Is this intended to say that migratory DCs expressing PD1 exhibit an immunosuppressive phenotype?

      Yes this is a better way to say it, it has been corrected.

      (14) Figure 6: These are really cool data about the influx of peripherally derived cells to the cribriform plate during EAE. However, it would be more meaningful to have other compartments to compare with. What is the IV+ percentage within the CNS or meninges more generally? And also, how do these CD11c+ CD11b+ aggregates differ in IV+ from "singlets"? The authors show that T cells are caught in the scRNA aggregates. Are these IV+? Can the authors provide additional discussion about the relevance of the Ghost+ data? What does this really mean? In Figure 6, Olfactory is misspelled 2x in A...and the "D" in CD45 is missing from B.

      Spelling mistakes have been corrected, thank you. Future investigations will compare IV+ recruitment and aggregations to dural and other brain regions. We suspect that some of the IV+ populations are T cells but our experiments do not allow for this distinction. We have added additional information regarding our interpretation of the Ghost+ data.

      (15) The title of the paper indicates that a suppressive myeloid network is assembled, and certainly, there is gene and protein expression data that are consistent with the presence of "suppressive" cells. However, can the authors demonstrate that this "network" is performing a suppressive function in vivo?

      This is a great point. Our IHC is highly indicative of classical M2 phenotype accumulating at meningeal regions around the olfactory bulb. One experiment we are interested in is local ablation of macrophages at the CP, to determine their role in EAE disease progression.

      (16) At the end of the discussion, the authors state, "They describe unique DC populations at the cribriform plate, one displaying pro-inflammatory and migratory features while the PDPN-associated population displayed more immunoregulatory characteristics". This seems a little bit misleading, or at least not giving the macrophages their due. A good part of the migratory DCs (as put in the figures) are associated with the Arg1+/Chi3l3+ macrophages. It's possible that suppression -if it's happening- could come from one or both cell types.

      We have removed that line and altered the discussion to more accurately reflect the results with respect to DCs and Macrophages.

      (17) In this study, the authors focus on dendritic cell and macrophage populations in the context of autoimmune disease and chronic CNS inflammation. In a recent study, the authors show an important recruitment of immune cells in the cribriform plate during a CNS infection by Mycobacterium tuberculosis. Do Arg1+/Chil3l3+ macrophage and tolerogenic DC populations still exist in this context? It would significantly strengthen the field's understanding of how the cells of the cribriform behave in different conditions if you could describe whether these cells are context-specific or is it really specific to cribriform plate tissue?

      This is an excellent suggestion and will be the focus of future investigations.

      We believe these revisions substantially strengthen the manuscript and directly address major concerns raised by both reviewers. We remain committed to the functional follow-up studies that both reviewers rightly identify as the natural next chapter of this work.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      Kaku and Flenniken investigate the mechanistic pathways through which specific viral infections alter the flight capabilities of honey bees. Building on their previous discovery that DWV impairs flight while SBV unexpectedly enhances it, the authors hypothesized that these behavioral shifts are driven by interactions with the insect's octopamine (OA) signaling pathway, which is responsible for the "fight-or-flight" neurohormonal stress response and energy mobilization. To test this, the authors experimentally infected adult honey bees with DWV or SBV and pharmacologically manipulated the OA pathway using either octopamine supplementation or epinastine (EP), an OA-receptor antagonist. They then evaluated the bees' flight performance (distance, duration, and speed) on custom flight mills and profiled their gene expression using qPCR and RNA sequencing.

      Strengths:

      A major strength of this study is the high prevalence of preexisting background DWV and SBV infections in the honey bee cohorts, which meant there were no completely "virus-free" control groups. However, the authors successfully mitigated this limitation by rigorously quantifying viral RNA copies for every individual bee via qPCR and utilizing these viral abundances as continuous variables in powerful linear mixed-effect models.

      Weaknesses:

      The primary weakness lies in the methodology used for targeted pharmacological manipulations, as well as the lack of OA quantification across different treatments. Thus, their claims are not sufficiently supported by the current data.

      We thank Reviewer #1 for these comments.

      (1) The authors utilize Epinastine to block octopamine signaling, describing it as a highly specific OA receptor antagonist. However, pharmacological inhibitors often lack absolute specificity. Epinastine might bind to other octopamine receptor subtypes present in honey bee neural and flight muscle tissues, or it could potentially cross-react with tyramine and dopamine receptors. Without further genetic validation (e.g., RNA interference targeting specific receptors), it is difficult to definitively conclude that the altered flight performance is solely due to the blockade of the specific Oβ−2R pathway.

      We thank the reviewer for this thoughtful comment and agree that pharmacological approaches have inherent limitations with respect to receptor specificity. However, among the available octopamine receptor antagonists, epinastine is considered one of the most selective compounds for insect octopamine receptors. Roeder et al. (1998) reported that epinastine exhibits affinities for octopamine receptors that are at least four orders of magnitude greater than those for other insect biogenic amine receptors, including dopamine, tyramine, histamine, and serotonin receptors.

      Honeybees encode four β-adrenergic-like receptors AmOARβ1- AmOARβ4) and one αadrenergic-like receptor (AmOARα1). Our transcriptomic analyses indicated that expression of AmOARβ2 was substantially higher than that of other octopamine receptor genes. Specifically, AmOARβ4 transcripts were not detected in our RNA-seq datasets, while AmOARβ1 and AmOARβ3 were expressed at very low levels in most samples (Supplementary Table S9; Figure S5). Although AmOARα1 transcripts were detected in some samples, expression levels were consistently lower than those of AmOARβ2. These observations support the interpretation that the physiological effects observed following epinastine treatment are primarily mediated through disruption of AmOARβ2 signaling. We agree that receptor-specific genetic approaches would provide valuable complementary evidence. RNAi-mediated knockdown of AmOARβ2 is an attractive future direction; however, RNAi efficacy in honey bees is variable and influenced by factors including transcript turnover rates. In addition, dsRNA treatments can induce sequence independent antiviral effects that could confound interpretation in studies involving viral infection (Flenniken and Andino, 2013). We have revised the manuscript to more explicitly acknowledge these limitations and to clarify the basis for our interpretation of the epinastine experiments.

      (2) As a natural neurotransmitter, insects have evolved highly efficient "cleanup" mechanisms. OA is rapidly cleared from the synaptic cleft via reuptake transporters and quickly inactivated by enzymes such as N-acetyltransferase (NAT) or Monoamine Oxidase (MAO). Consequently, an injection of OA produces only a transient "pulse" of activity. It is often a poor "tool" for inducing prolonged physiological effects compared to synthetic formamidines like Amitraz.

      We thank the reviewer for this important point regarding the pharmacokinetics of octopamine. We agree that octopamine is rapidly metabolized and cleared under physiological conditions and that exogenous administration is unlikely to precisely mimic endogenous signaling dynamics. Our goal was not to induce a prolonged pharmacological activation of octopamine signaling comparable to that produced by synthetic agonists such as amitraz, but rather to determine whether increasing octopaminergic signaling could mitigate the flight impairments associated with DWV infection. Octopamine was administered either by injection or through feeding (Lines 86-89), both of which resulted in significant improvements in flight performance in DWV-infected bees (Figure 2). The observation that two independent delivery methods produced similar outcomes supports the conclusion that enhanced octopaminergic signaling can partially rescue the DWV-associated flight phenotype. We have revised the manuscript to clarify this distinction and to acknowledge that exogenous octopamine administration likely produces transient elevations in signaling rather than sustained receptor activation.

      (3) The study relies heavily on transcriptomics and quantitative PCR to measure the mRNA expression of key synthesizing enzymes, namely tyrosine decarboxylase (tdc) and tyramine βhydroxylase (tβh), to infer the activation or suppression of the octopamine pathway. However, changes in enzyme synthesis at the RNA level are often insufficient to accurately reflect the true physiological levels of biogenic amines. To robustly prove the authors' hypothesis of a "feedback loop that regulates intracellular OA concentrations", direct quantification of actual octopamine and tyramine titers in the bees (e.g., using high-performance liquid chromatography or mass spectrometry) is necessary.

      We thank the reviewer for this comment and agree that octopamine and tyramine quantification would strengthen the mechanistic interpretation of our findings. Previous studies have successfully quantified OA in honey bees using HPLC-based approaches, including KayaZee et al. (2022, eLife), who measured OA in honey bee muscle tissue (both naturally occurring levels and levels post-treatment with 10 mM OA), and Cook et al. (2017, J. Exp. Bio) who quantified OA in pooled honey bee brain samples.

      Prior to submission, we inquired with our institutional mass spectrometry facility regarding the feasibility of measuring OA in individual honey bee samples. The expected concentrations of OA in our samples was below their limit of detection, so we did not pursue these analyses at that time.

      We are exploring the possibility of analyzing a subset of samples at external facilities that may have the sensitivity required to quantify OA and tyramine in honey bee tissues. However, our initial discussions indicate that such analyses would require substantial resources, with estimated costs of approximately $5,000–10,000 for 12–15 samples. While we acknowledge that direct measurements of OA and tyramine would provide valuable complementary evidence, the current study relies on multiple independent lines of evidence including gene expression analyses, OA supplementation experiments, and behavioral measurements that collectively support a role for octopaminergic signaling in mediating the observed effects.

      Reviewer #2 (Public review):

      Summary:

      This highly original and well-designed study provides insight into how honeybee picorna-like viruses, Deformed wing virus (DWV) and Sacbrood virus (SBV), affect flight performance, and reveals the role of the octopamine (OA) pathway in virus-honeybee interactions. The authors used a flight mill to quantify the flight performance of bees with different levels of DWV and SBV. Bees were treated with OA and/or epinastine (EP) - an OA receptor antagonist; the study also quantified virus loads and expression of two key genes involved in OA biosynthesis.

      The results showed that reduced flight performance associated with high DWV levels could be alleviated by OA administration. In contrast, increased levels of SBV had the opposite effect, leading to enhanced flight performance. This suggests distinct physiological responses to DWV and SBV infections. Administration of EP had led to a reduction of flight performance in SBVinfected bees, indicating the involvement of the OA pathway.

      The authors also quantified levels of mRNAs of enzymes involved in OA synthesis, tyrosine decarboxylase (TDC) and tyramine beta-hydroxylase (TbH), and concluded that DWV induced expression of TbH, while SBV upregulated expression of TDC. Furthermore, the study identified upregulated and downregulated genes in response to SBV, DWV and DWV in combination with OA.

      Strengths:

      The study reported opposing effects of infections of related viruses, SBV and DWV, on honeybee flight performance, and identified the central role of the octopamine (OA) signaling pathway in the effect of viruses on honeybee flights.

      These findings were achieved by using a combination of approaches, including experimental measurement of flight distance, virus infections, and introduction of OA and EP. Experimental work with honeybees is technically challenging and requires specialized expertise, which makes the results produced in this study more valuable.

      DWV and SBV are among the most important honeybee pathogens affecting honeybee health and threatening the pollination service. Therefore, an understanding of the mechanisms underlying DWV and SBV pathogenesis has the potential to develop novel approaches to mitigate the negative impact of these viruses.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      We thank Reviewer #2 for these comments

    1. Author response:

      Reviewer #1 (Public review):

      Weaknesses:

      While the breadth of techniques is impressive, the central premise of the work-the structural and functional relationship between polyQ assemblies and the Golgi apparatus-is not supported by sufficiently rigorous cell biological evidence.

      A major concern is that much of the cell biology data remains descriptive and lacks mechanistic depth. The findings are fragmented and not integrated into a coherent molecular or cellular model. Instead of building a logical progression of experiments, the study presents a collection of observations that appear disconnected and, at times, driven more by technical capability than by hypothesis-driven design.

      Critically, the key claim that polyQ HTT functionally disrupts the Golgi (Golgipathy) is not convincingly demonstrated. Many observations could be more simply explained by the polyQ HTT localization to the Golgi and known Golgi sensitivities to perturbations (e.g., starvation or Brefeldin A treatment), rather than by a specific mechanistic role of polyQ HTT.

      The manuscript also suffers from issues in organization and clarity, including imprecise descriptions and figures that are difficult to interpret.

      We thank the Reviewer for their time, valuable comments, and recognition of our technical expertise and resources. With our specialized background in pathology and super-resolution microscopy, our research heavily relies on structurally precise histological methods to address these fundamental biological questions. Furthermore, our laboratory maintains one of the largest repositories of patient-derived and healthy control fibroblasts, as well as iPSC lines, within the Huntington's disease (HD) research community. Because these patient-derived and engineered cell models express endogenous mutant HTT (mHTT) within an authentic genetic background, they provide a uniquely powerful system for decoding HD pathogenesis.

      We appreciate the Reviewer’s comment regarding hypothesis-driven design. Classically, a hypothesis-driven approach relies on well-established, highly stable experimental platforms. However, a key finding of our study is the highly fragile and volatile nature of polyQ assemblies, particularly when subjected to post-fixation and oxidative stress. Because these structures can behave unpredictably under stress, we utilized an unbiased, data-driven approach leveraging our high-resolution imaging pipeline to explore polyQ assemblies in both healthy and HD cells.

      Importantly, hypothesis-driven and data-driven methods are complementary rather than mutually exclusive. For instance, if a real-time tracking method were developed to endogenously label native HTT in living cells, it would open the door for direct hypothesis testing regarding polyQ assembly mechanics. Despite the current technical limitations of the field, our study successfully overcomes these challenges to reveal the spatial tomography and unique dynamics of polyQ assemblies directly within patient-derived cells. We fully discussed the limitations of this research in the discussion section.

      We appreciate the Reviewer’s critical assessment regarding the functional disruption of the Golgi apparatus (Golgipathy). To rigorously investigate this phenomenon, we employed a comprehensive suite of methodologies ranging from live-cell imaging to single-cell RNA sequencing. Our findings build directly upon a well-established body of literature. We previously demonstrated that mutant huntingtin (mHTT) disrupts Golgi function within the neural tubes of human cortical organoids (hCO) (Liu et al., 2024), aligning with broader neurodevelopmental defects observed in HD (Barnat et al., 2020). Furthermore, prior independent studies have confirmed that both HTT knockdown and the presence of mHTT impair Golgi-to-plasma membrane trafficking, notably in primary fibroblasts from homozygous Htt<sup>140Q/140Q</sup> knock-in mice (Brandstaetter et al., 2014); mHTT also affects post-Golgi trafficking of proteins (del Toro et al., 2006). Backed by this literary consensus and our own multi-modal data, which Reviewer 2 also noted as sufficient, we are confident that our manuscript provides a robust, multi-layered demonstration of mHTT-induced Golgipathy.

      In this study, we found that polyQ assemblies and the Golgi form a Golgi-polyQ complex mediated by ARF1 and ARFIP2. Thus, the structural coupling of polyQ assemblies with the Golgi apparatus under starvation, during the cell cycle, is rational.

      Based on the reviewer suggestion, we will completely revise the entire manuscript. Hopefully, this revision will meet the requirement of smoothness and clarity.

      Major Concerns:

      (1) Golgi localization

      The localization of polyQ HTT relies entirely on the antibody 3B5H10, which is foundational to the study. However, previous reports using the same antibody have described predominantly cytosolic localization. This discrepancy must be addressed rigorously by independent validation using alternative antibodies or tagged, exogenously expressed polyQ HTT constructs that should be shown to colocalize with 3B5H10 signals.

      Despite historical inconsistencies across existing publications (Barnat et al., 2020; Hickman et al., 2022; Shen et al., 2019; Tousley et al., 2019), we noticed that the immunostaining results of multiple HTT antibodies are consistent with our data (DiFiglia et al., 1995; Ko et al., 2001; Velier et al., 1998; Wheeler et al., 2000). Although these pioneering studies lacked modern 3D high-resolution imaging and standardized staining protocols, their reported 2D distribution patterns heavily resemble our results. For instance, transmission electron microscopy (TEM) immunolabeling originally revealed that HTT localizes along Golgi cisternae (DiFiglia et al., 1995) and formed organized and parallel fibrils (DiFiglia et al., 1997). Furthermore, immunostaining with a panel of distinct antibodies, including MV2, 3, 4, 5, 6, and 1F8, demonstrated characteristic Golgi-like distribution patterns for HTT (Ko et al., 2001). In addition, our polyQ antibody immunostaining in human fetal brain, which is reflective of polyQ assembly, is nearly identical to the staining results of Barnat et al., publication in Science (Barnat et al., 2020).

      We have carefully checked two early publications, which reported that 3B5H10 only binds expanded polyQ but does not bind a normal polyQ (non-disease causing), which displays a part of a neuron that has a cytosolic diffuse pattern of HTT in 3B5H10 staining (Legleiter et al., 2009; Miller et al., 2011). Based on our extensive experience with HTT immunohistochemistry, we hypothesize that this diffuse signal may reflect nonspecific background artifacts, often caused by high antibody concentrations, poor tissue fixation, inadequate post-incubation washing, or the presence of effete cells, or premature fragmentation of the polyQ tract prior to staining. Interestingly, Miller et al. utilized a rapid tissue-perfusion and sectioning protocol originally published in Brain Research Bulletin (Ko et al., 2001), which is optimized to preserve intact polyQ assemblies. When reviewing the original Brain Research Bulletin study (Ko et al., 2001), we noted that the immunostaining profiles for polyQ-containing HTT peptides (specifically using antibodies MW2, MW3, MW4, MW5, and 1F8) are entirely consistent with our data, yet completely diverge from the patterns reported by the Muchowski group (Legleiter et al., 2009; Miller et al., 2011) (please see Ko et al., 2001.Legleiter et al., 2009; Miller et al., 2011). Furthermore, contrary to the Muchowski group's claims, subsequent biophysical evidence by Owens et al. (2015) independently confirmed that 3B5H10 binds to both normal and expanded polyQ sequences in huntingtin exon 1 fusion proteins (Owens et al., 2015). Together, these observations strongly support the validity of our staining profiles.

      Several antibodies, including MV1, 1C2, and 3B5H10, were previously reported to recognize the expanded, pathogenic polyQ tracts of HTT (Khoshnan et al., 2002; Miller et al., 2011; Wang et al., 2008). However, emerging studies reveal that these antibodies actually bind both short and long polyQ sequences (Klein et al., 2013; Owens et al., 2015). Because a standard antibody Fab epitode typically spans only 5 to 15 amino acids (or 3 to 4 sugar residues), and normal HTT polyQ repeats range from 18 to 24, it is theoretically impossible for an antibody to exclusively target expanded polyQ while sparing normal polyQ.

      We previously noticed that 3B5H10 antibody immunostaining signals are located in the long projection of striatal neurons. As we did not notice an intact neuron in the two publications, we have no idea about the 3B5H10 antibody signals in the neuronal projections of those images.

      We investigated whether the polyQ assemblies detected by the 3B5H10 antibody contain full-length or large fragments of huntingtin (HTT). To test this, we selected two distinct HTT antibodies: EM48, which binds the first 256 amino acids (excluding the polyQ stretch), and 3E10, which targets the HDA region (amino acids 1171–1177). In patient fibroblasts, the immunostaining patterns for both EM48 and 3E10 were nearly identical to those observed with 3B5H10. These results demonstrate that the polyQ assemblies in fibroblasts are primarily composed of HTT proteins (see Author response image 1). We will include the results of EM48 and 3E10 immunostaining in the revised version.

      Author response image 1.

      (A) GFAP and EM48 antibodies staining of the astrocytes derived from HD patient and healthy sibling iPSCs showed polyQ assembly in the astrocytes derived from iPSC. (B). The spindle of polyQ assembly formed a dent on the nuclear surface of astrocytes. (C, D) Coimmunostaining of GM130 antibody with 3E10 or EM48 antibody in fibrobalsts revealed that polyQ assemblies contain great amount of HTTs. The middle and right panel are the sectional view boxed region (D) and the boxed region are a magnified part or rendering part (C, D) .

      We also check whether the transfected exogenous HTT fragment of the first exon can be recruited into polyQ assemblies. We transfected fibroblasts with three vectors of the HTT first exon containing 19, 23, and 74 CAGs, respectively. The exogenous HTT fragments of the first exon did not significantly recruit into endogenous polyQ assemblies-Golgi complexes of fibroblasts (please refer to Reviewer only figure 5). We will include this part in the revised version.

      In the cover letter, we told the editor that we have been studying this structure for over ten years. The results have been stable for over ten years.

      Furthermore, the Golgi is identified solely using GM130, a cis-Golgi and ER exit site marker. This raises ambiguity: does polyQ HTT associate with the entire Golgi or only recruit GM130? Could the observed signal correspond to a sub-Golgi compartment?

      Thank you for highlighting the precise sub-Golgi localization of GM130 as resolved by electron microscopy. We agree that transmission electron microscopy (TEM) demonstrates GM130 is restricted to the cis-Golgi network, intercisternal regions, and tubular structures, and is absent from the trans-Golgi (Nakamura et al., 1995). Given that individual Golgi cisternae measure approximately 20 nm in width, resolving cis- versus trans-Golgi sub-compartments exceeds the physical resolution limits of our microscopy system. Consequently, GM130 was utilized here as a robust, widely accepted pan-Golgi marker rather than a tool for sub-compartmental differentiation. To specifically evaluate the trans-Golgi network (TGN), we tracked Clathrin+ vesicles, which actively sort at the TGN (Klumperman, 2011). Our Clathrin staining confirms that polyQ assemblies localize to both the cis- and trans-Golgi compartments, as clearly demonstrated in the new lateral view projections provided in revised Figure 4C.

      If polyQ HTT is indeed Golgi-associated, several key observations become expected rather than novel. For example, in Figure 4I-M, sensitivity to Brefeldin A is unsurprising, as Golgi structure collapses upon such treatment; in Figure 4N-O, co-fragmentation with the Golgi is expected under Golgi-disrupting conditions.

      We agree that our data demonstrate the formation of a functionally coupled polyQ assembly–Golgi complex. Physically and structurally, the dynamics of polyQ assemblies are intrinsically linked to Golgi dynamics under distinct physiological states, including cell cycle progression and energy deprivation. This structural coupling is mediated by ADP-ribosylation factor 1 (ARF1). Specifically, the polyQ tract of HTT interacts with ARFIP2, which is one of the key effector proteins that physically bind active ARF1. Mechanistically, ARF1 is recruited to the Golgi membrane upon GDP-to-GTP exchange catalyzed by guanine nucleotide-exchange factors (GEFs). Consequently, treatment with Brefeldin A (BFA)—which inhibits ARF1 activation—effectively decouples the polyQ assemblies from both intact and fragmented Golgi structures.

      Regarding the question of novelty, we define experimental novelty based on generating entirely unprecedented, empirical data that either confirms or redefines biological expectations, rather than evaluating conceptual expectations themselves. We believe the uncovering of this real-time, stimulus-responsive coupling mechanism provides fundamentally novel insights into HTT biology.

      (2) 3D rendering

      The extensive use of 3D rendering appears unnecessary and, in some cases, misleading. The rendered images do not provide additional insight beyond conventional 2D fluorescence images. Serial 2D fluorescence sections should be more objective in representing the 3D organization.

      Thanks for pointing out the 3D rendering. While 3D rendering provides an essential spatial approximation of fluorescently labeled architectures, it offers significantly more precise structural information than conventional 2D or serial section imaging alone (Cao et al., 2023; Han et al., 2021; Hexige et al., 2015). A primary objective of our study was to evaluate these subcellular features within their intact, native three-dimensional context rather than relying solely on two-dimensional cross-sections. Crucially, without complete volumetric rendering, it is mathematically and visually challenging to accurately delineate complex morphological features, such as the nuclear gorge or true intranuclear accumulation. Consequently, 3D volumetric analysis and rendering are entirely indispensable for the accurate interpretation of the structural data presented in this study.

      In Figure 2A and Figure 5A, red line features in 3D beige polyQ HTT structures resemble unrelated biological structures, such as vasculature, which is inappropriate.

      We would like to clarify that there is no vasculature present within the referenced 3D rendering. The features the reviewer is highlighting are artifacts of the pseudo-coloring used exclusively to mask and visualize the surface tomography. In volumetric 3D rendering, pseudo-colors are assigned strictly to enhance visual clarity and contrast for the reader; they carry no intrinsic biological meaning or cellular identity. Furthermore, from a structural standpoint, the narrow red features in the rendered image are orders of magnitude smaller than true microvasculature. Functional microvessels possess a minimum diameter of 6 to 45 micrometers and exhibit a defined vascular lumen, endothelial cells, a basement membrane, pericytes, and a tunica intima. Therefore, based on both the scale of the image and established histological criteria, these features cannot biologically or structurally represent vasculature.

      There is also an inconsistency in rendering. For example, fine mesh-like structures are shown in some figures (e.g., Figure 2A, Figure 4A), whereas others appear as amorphous aggregates (e.g., Figure 5A, Figure S2B), without explanation.

      The selection of opacity and color masks in our 3D volumetric reconstructions is systematically chosen to optimize the visual clarity and spatial relationships between intersecting sub-cellular structures. For example, as shown in the fourth and fifth panels, an opaque blue mask was applied to clearly define the outer surface tomography. Conversely, in the third panel, a semi-transparent blue mask was utilized for the nucleus. This transparency is methodologically necessary because a subset of polyQ fragments is embedded within or localized directly inside the nuclear envelope; a transparent mask allows for the unambiguous visualization of these internal structures. Similarly, the inset in Figure 5A illustrates the distinct intranuclear occupancy pattern of polyQ, which also necessitates a transparent nuclear boundary. Collectively, these volumetric rendering strategies provide critical spatial and structural depth that cannot be captured by conventional 2D cross-sections or unconstructed serial imaging.

      (3) Quantification of area and volume

      The manuscript extensively quantifies the area and volume of polyQ assemblies (e.g., Figure 2B, C and Figure 3B, C, E, G, H). These measurements are not reliable. First, the structures appear filamentous and likely below the diffraction limit. Second, fluorescence signals are broadened by the point spread function (PSF), artificially inflating measured dimensions. Last, even with 3D SIM (~100 nm resolution), fine structural details remain unresolved. Thus, these quantitative measurements lack physical meaning and might not be used to support conclusions.

      We appreciate the reviewer’s thoughtful critique regarding the quantification of area and volume. Our measurements are derived from immunofluorescent signals captured via structured illumination microscopy (SIM) and confocal imaging. If the reviewer's concern is that antibody-labeled structures do not perfectly match the absolute physical dimensions of native polyQ assemblies due to the linkage error of the primary-secondary antibody complex, we agree conceptually.

      However, our imaging pipeline is optimized to minimize these discrepancies. Our SIM resolution reaches approximately 64 nm. Given that the total observed thickness of the fluorophore-labeled polyQ assemblies exceeds 200 nm, these structures reside well within the detectable range of our super-resolution system, minimizing diffraction-induced overestimation. Regarding the point spread function (PSF) and optical distortion, we emphasize that all comparative quantifications across experimental groups were conducted under identical imaging parameters and thresholds, ensuring a standardized baseline. Furthermore, our acquisition systems (Leica, Nikon, and Zeiss) utilize advanced deconvolution algorithms specifically designed to mitigate PSF-related blur. While we observed that deconvolution yielded negligible baseline improvements when using high-numerical-aperture objectives (63x) or 100x, oil immersion), it validates that our raw high-resolution scanning was already highly optimized.

      We acknowledge that an immunolabeled complex is not structurally identical to a naked, pure polyQ tract. Nonetheless, indirect immunofluorescence remains the most robust method to evaluate spatial distribution in situ. Indeed, cryo-EM studies have highlighted that native polyQ tracts are highly flexible and structurally dynamic, making them exceptionally difficult to resolve in their native state (Guo et al., 2018). Intriguingly, we observed that antibody-bound polyQ assemblies remain structurally stable for several weeks with minimal fragmentation, suggesting that antibody binding may structurally stabilize these highly flexible regions. Consequently, indirect immunolabeling provides an indispensable framework for capturing these assemblies within the cellular environment.

      (4) Interpretation of structural features (Figure 2A)

      Descriptions such as "parallel spindles" and "ring-like assemblies" are not clearly supported by the data. The terminology is ambiguous, and the claimed structures are not discernible. The use of the term "interaction" with the nuclear membrane is also inappropriate. At best, the data suggest colocalization, which itself is not convincingly demonstrated.

      Please refer to Fig. 2A (middle upper), Fig. 2F, and Fig. 4C for “parallel spindles”. Please refer to Fig. 5I, J, and Fig.S3C (right panel) for additional clear “ring-like assemblies”. Due to the unique spatial distribution of the 'ring-like assemblies', observing multiple rings within a single spindle is technically challenging. Accordingly, we have tempered our statement in the revised manuscript to accurately reflect this limitation. Furthermore, it is important to note that the visualized structures represent the fluorescent signal from secondary antibodies rather than direct imaging of the proteins themselves. Consequently, we cannot definitively confirm whether this immunostaining pattern precisely replicates the native state of polyQ assemblies within the cellular environment.  

      (5) Mitotic fragmentation (Figure 2E)

      The conclusion that polyQ assemblies fragment during mitosis lacks proper controls. It is unclear whether these cells exhibited intact "fabric-like" assemblies during interphase, or the observed structures were already fragmented prior to mitosis.

      We thought that we had displayed enough non-mitotic cells in this study (Fig. 2A, D, Fig. 4A, F). Most of the cells in this study are non-mitotic cells (G1+S+G2). Thus, we consider the control of non-mitotic cells to be redundant here.

      (6) Fixation-induced fragmentation (Figure 2F)

      The claim that fixation-induced fragmentation reflects a unique dynamic property of polyQ assemblies is likely an overinterpretation. This phenomenon may simply represent a fixation artifact. Therefore, it cannot be used as evidence for in-cellulo structural dynamics.

      By definition, a laboratory artifact refers to any unintended structural detail, distortion, or error introduced by experimental equipment or the preparation process. We contend that the observed phenomenon represents a native chemical characteristic of the HTT polyQ domain inside cells following paraformaldehyde (PFA) fixation, rather than a technical artifact. Similar structural features have been documented by other investigators in tissue samples (Ferrante et al., 1997). A classic textbook example of an artifact is the lamina lucida of the basal lamina, which is artificially generated during electron microscopy tissue processing and does not exist in living tissue. In contrast, the fragmentation of polyQ assemblies occurs naturally both in living cells subjected to stress and during post-fixation processing.

      (7) Nuclear localization claims (Figure 5A)

      The assertion that polyQ assemblies "almost completely occupy the nucleus" is not supported. The images are more consistent with perinuclear localization, typical of the Golgi region. There is no clear evidence for nucleoplasmic distribution.

      Please refer to the rendering image in the upper left inner insert of HD neurons (the blue [transparent] is the nucleus and the pink white is polyQ). The almost complete occupation of the nucleus is crystal clear in these images (rendered inner inserts, upper left). In iPSC-induced HD neurons, it is not only distributed in the nucleus but also in the cytoplasm. Based on your description, you might refer to the cytoplasmic polyQ assemblies but not the nucleus in the rendering image of the upper left (left panel). We will add a label in the revised version for clarity (white arrows for nuclear accumulation). In this manuscript, we have enough figures that clearly show the nuclear accumulation. Please also refer to Fig. 7 and Fig. S2 for additional images of nuclear accumulation.

      (8) Drug treatment and data interpretation (Figure 3D-E)

      The x-axis in Figure 3E is non-linear, which is inappropriate unless explicitly justified. Furthermore, the rationale for using Onjisaponin F is unclear. What is its known mechanism? Does it affect the Golgi organization? Without this context, observed effects may reflect Golgi perturbation rather than specific effects on polyQ assemblies.

      We appreciate the reviewer pointing out Figure 3E. In this experiment, Huntington's disease (HD) fibroblasts were cultured in a low-glucose medium for the first 72 hours, which accounts for the linear trend observed across the first four data points. Following this 72-hour period, the cells were switched to a high-glucose medium and cultured for an additional 48 hours to evaluate subsequent dynamic changes in the polyQ assemblies. To improve visual clarity, we have color-coded these distinct treatment conditions in the revised manuscript, using red to denote low-glucose treatment and green to denote high-glucose treatment.

      Regarding the choice of Onjisaponin treatment (a concern also raised by another reviewer), Onjisaponin is an active component derived from Radix Polygalae (Yuan Zhi). Previous literature indicates that Onjisaponin B enhances autophagy, accelerates the degradation of mutant α-synuclein and huntingtin in vitro, and activates the AMPK-mTOR signaling pathway (Wu et al., 2013). To optimize our experimental model, we screened multiple variants—specifically Onjisaponin B, D, and F. We determined that Onjisaponin F exhibits remarkably low cytotoxicity while maintaining a robust autophagy-enhancing capacity in both human fibroblasts and iPSC-derived neurons. Consequently, Onjisaponin F was selected for our human cell line experiments (please refer to the Reviewer only image 3). While we did not previously assess Golgi apparatus alterations under Onjisaponin F treatment, we recognize the value of this metric. We are currently evaluating changes to both the Golgi apparatus and neuronal firing rates following Onjisaponin F exposure, and this new dataset will be integrated into our revision.

      Reviewer #2 (Public review):

      […] Overall, this work reports a novel polyQ assembly, which was previously reported as a pathogenic factor, has not been reported before for HTT, is related to Golgi activities and vesicular transport, and is dismantled in HD patient cells. The intensive immunostaining and super-resolution scanning are impressive and definitely strengthened by the impact of the findings. The scRNAseq data adds another layer to the observed Golgi impairments and their suggested relationship to Golgi function. The drug testing for polyQ assemblies, especially polyQ assemblies in HD cells, is preliminary. However, the data in this study are enough to support the existence of polyQ assemblies in human cells and their specific relationships with the Golgi apparatus.

      We sincerely thank the reviewer for their time, dedication, and insightful evaluation of our manuscript. We agree that the drug screening component represents an initial phase of discovery, and we appreciate the opportunity to clarify this in our text. As the reviewer notes, executing high-throughput or exhaustive drug screenings in human brain organoids is exceptionally resource- and time-intensive due to prolonged culture requirements. We will provide more mechanistic and physiological details of these drug in the future.

      Strengths:

      In this study, the authors used the cells from a large HD family and fetal/child brain samples to decode the structure of endogenous polyQ assemblies. This part is impressive. The intensive staining and super-resolution scanning are amazing. The spatial relationships of polyQ assemblies with the Golgi apparatus and mitochondria are well illustrated.

      Weaknesses:

      Although they used healthy sibling cells as a control, an isogenic control (genetic correction of the mutant gene) is lacking. Based on the Golgipathy of mHTT, they did a drug screening. The drug testing for polyQ assemblies is preliminary. More rigorous validation, such as scRNA seq and proteomic analysis, etc., is necessary to reach a systemic conclusion.

      References

      Barnat, M., Capizzi, M., Aparicio, E., Boluda, S., Wennagel, D., Kacher, R., Kassem, R., Lenoir, S., Agasse, F., Braz, B.Y., et al. (2020). Huntington's disease alters human neurodevelopment. Science 369, 787-793.

      Brandstaetter, H., Kruppa, A.J., and Buss, F. (2014). Huntingtin is required for ER-to-Golgi transport and for secretory vesicle fusion at the plasma membrane. Dis Model Mech 7, 1335-1340.

      Cao, L., Ma, L., Zhao, J., Wang, X., Fang, X., Li, W., Qi, Y., Tang, Y., Liu, J., Peng, S., et al. (2023). An unexpected role of neutrophils in clearing apoptotic hepatocytes in vivo. Elife 12.

      del Toro, D., Canals, J.M., Gines, S., Kojima, M., Egea, G., and Alberch, J. (2006). Mutant huntingtin impairs the post-Golgi trafficking of brain-derived neurotrophic factor but not its Val66Met polymorphism. J Neurosci 26, 12748-12757.

      DiFiglia, M., Sapp, E., Chase, K., Schwarz, C., Meloni, A., Young, C., Martin, E., Vonsattel, J.P., Carraway, R., Reeves, S.A., et al. (1995). Huntingtin is a cytoplasmic protein associated with vesicles in human and rat brain neurons. Neuron 14, 1075-1081.

      DiFiglia, M., Sapp, E., Chase, K.O., Davies, S.W., Bates, G.P., Vonsattel, J.P., and Aronin, N. (1997). Aggregation of huntingtin in neuronal intranuclear inclusions and dystrophic neurites in brain. Science 277, 1990-1993.

      Ferrante, R.J., Gutekunst, C.A., Persichetti, F., McNeil, S.M., Kowall, N.W., Gusella, J.F., MacDonald, M.E., Beal, M.F., and Hersch, S.M. (1997). Heterogeneous topographic and cellular distribution of huntingtin expression in the normal human neostriatum. J Neurosci 17, 3052-3063.

      Guo, Q., Bin, H., Cheng, J., Seefelder, M., Engler, T., Pfeifer, G., Oeckl, P., Otto, M., Moser, F., Maurer, M., et al. (2018). The cryo-electron microscopy structure of huntingtin. Nature 555, 117-120.

      Han, X., Ma, L., Gu, J., Wang, D., Li, J., Lou, W., Saiyin, H., and Fu, D. (2021). Basal microvilli define the metabolic capacity and lethal phenotype of pancreatic cancer. J Pathol 253, 304-314.

      Hexige, S., Ardito-Abraham, C.M., Wu, Y., Wei, Y., Fang, Y., Han, X., Li, J., Zhou, P., Yi, Q., Maitra, A., et al. (2015). Identification of novel vascular projections with cellular trafficking abilities on the microvasculature of pancreatic ductal adenocarcinoma. J Pathol 236, 142-154.

      Hickman, R.A., Faust, P.L., Marder, K., Yamamoto, A., and Vonsattel, J.P. (2022). The distribution and density of Huntingtin inclusions across the Huntington disease neocortex: regional correlations with Huntingtin repeat expansion independent of pathologic grade. Acta Neuropathol Commun 10, 55.

      Khoshnan, A., Ko, J., and Patterson, P.H. (2002). Effects of intracellular expression of anti-huntingtin antibodies of various specificities on mutant huntingtin aggregation and toxicity. Proc Natl Acad Sci U S A 99, 1002-1007.

      Klein, F.A., Zeder-Lutz, G., Cousido-Siah, A., Mitschler, A., Katz, A., Eberling, P., Mandel, J.L., Podjarny, A., and Trottier, Y. (2013). Linear and extended: a common polyglutamine conformation recognized by the three antibodies MW1, 1C2 and 3B5H10. Hum Mol Genet 22, 4215-4223.

      Klumperman, J. (2011). Architecture of the mammalian Golgi. Cold Spring Harb Perspect Biol 3.

      Ko, J., Ou, S., and Patterson, P.H. (2001). New anti-huntingtin monoclonal antibodies: implications for huntingtin conformation and its binding proteins. Brain Res Bull 56, 319-329.

      Legleiter, J., Lotz, G.P., Miller, J., Ko, J., Ng, C., Williams, G.L., Finkbeiner, S., Patterson, P.H., and Muchowski, P.J. (2009). Monoclonal antibodies recognize distinct conformational epitopes formed by polyglutamine in a mutant huntingtin fragment. J Biol Chem 284, 21647-21658.

      Liu, Y., Chen, X., Ma, Y., Song, C., Ma, J., Chen, C., Su, J., Ma, L., and Saiyin, H. (2024). Endogenous mutant Huntingtin alters the corticogenesis via lowering Golgi recruiting ARF1 in cortical organoid. Mol Psychiatry.

      Miller, J., Arrasate, M., Brooks, E., Libeu, C.P., Legleiter, J., Hatters, D., Curtis, J., Cheung, K., Krishnan, P., Mitra, S., et al. (2011). Identifying polyglutamine protein species in situ that best predict neurodegeneration. Nat Chem Biol 7, 925-934.

      Nakamura, N., Rabouille, C., Watson, R., Nilsson, T., Hui, N., Slusarewicz, P., Kreis, T.E., and Warren, G. (1995). Characterization of a cis-Golgi matrix protein, GM130. J Cell Biol 131, 1715-1726.

      Owens, G.E., New, D.M., West, A.P., and Bjorkman, P.J. (2015). Anti-PolyQ Antibodies Recognize a Short PolyQ Stretch in Both Normal and Mutant Huntingtin Exon 1. Journal of Molecular Biology 427, 2507-2519.

      Paulson, H.L., Bonini, N.M., and Roth, K.A. (2000). Polyglutamine disease and neuronal cell death. Proc Natl Acad Sci U S A 97, 12957-12958.

      Shen, M., Wang, F., Li, M., Sah, N., Stockton, M.E., Tidei, J.J., Gao, Y., Korabelnikov, T., Kannan, S., Vevea, J.D., et al. (2019). Reduced mitochondrial fusion and Huntingtin levels contribute to impaired dendritic maturation and behavioral deficits in Fmr1-mutant mice. Nat Neurosci 22, 386-400.

      Tousley, A., Iuliano, M., Weisman, E., Sapp, E., Richardson, H., Vodicka, P., Alexander, J., Aronin, N., DiFiglia, M., and Kegel-Gleason, K.B. (2019). Huntingtin associates with the actin cytoskeleton and alpha-actinin isoforms to influence stimulus dependent morphology changes. PLoS One 14, e0212337.

      Velier, J., Kim, M., Schwarz, C., Kim, T.W., Sapp, E., Chase, K., Aronin, N., and DiFiglia, M. (1998). Wild-type and mutant huntingtins function in vesicle trafficking in the secretory and endocytic pathways. Exp Neurol 152, 34-40.

      Wang, C.E., Tydlacka, S., Orr, A.L., Yang, S.H., Graham, R.K., Hayden, M.R., Li, S., Chan, A.W., and Li, X.J. (2008). Accumulation of N-terminal mutant huntingtin in mouse and monkey models implicated as a pathogenic mechanism in Huntington's disease. Hum Mol Genet 17, 2738-2751.

      Wheeler, V.C., White, J.K., Gutekunst, C.A., Vrbanac, V., Weaver, M., Li, X.J., Li, S.H., Yi, H., Vonsattel, J.P., Gusella, J.F., et al. (2000). Long glutamine tracts cause nuclear localization of a novel form of huntingtin in medium spiny striatal neurons in HdhQ92 and HdhQ111 knock-in mice. Hum Mol Genet 9, 503-513.

      Wu, A.G., Wong, V.K., Xu, S.W., Chan, W.K., Ng, C.I., Liu, L., and Law, B.Y. (2013). Onjisaponin B derived from Radix Polygalae enhances autophagy and accelerates the degradation of mutant alpha-synuclein and huntingtin in PC-12 cells. Int J Mol Sci 14, 22618-22641.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study builds on previous work from the same authors to present a conceptually distinct workflow for cryo-EM reconstruction that uses 2D template matching to enable highresolution structure determination of small (sub-50 kDa) protein targets. The paper describes how density for small-molecule ligands bound to such targets can be reconstructed without these ligands being present in the template. However, the evidence described for the claim that this technique “significantly” improves the alignment of the reconstruction of small complexes is incomplete. The authors could better evaluate the effects of model bias on the reconstructed densities.

      We have addressed both concerns. Regarding the claim that 2DTM “significantly” improves alignment, the most direct evidence is the controlled comparison in Fig. 3: using the same particle stack and the same reconstruction software (RELION), 2DTM-derived orientations yield a 3.1 Å reconstruction whereas RELION auto-refinement of the same particles yields 3.7 Å. Because the orientations are the only variable, this comparison directly demonstrates that 2DTM produces more accurate alignments.

      We further evaluated RELION auto-refinement with initial low-pass filters of 3, 5, 10, and 15 Å (Fig. 3c); the final resolution remained between 3.7 and 4.0 Å across all conditions, indicating that the achievable resolution difference reflects a fundamental distinction between the two approaches. 2DTM directly leverages high-resolution signal in the template during alignment, which is particularly advantageous for small particles.

      To assess whether this improvement extends beyond the ligand pocket, we constructed a composite omit map (Fig. 5) assembled from 36 reconstructions, each generated using a template with a different subset of residues deleted. The composite shows that density can be recovered at distributed locations across the kinase, including peripheral and surface-exposed regions further away from the alignment center. Recovery varies across sites, with some regions exhibiting weaker or fragmented density, consistent with local differences in structural heterogeneity and residual alignment error. Together, these results indicate that the orientation estimates support global density recovery rather than being confined to the ligand-binding region.

      Regarding model bias, we have strengthened both the quantitative and visual analyses. Specifically, we have (i) updated the template-bias metric Ω in Fig. 4, (ii) added grouped occupancy refinement showing that omitted residues 222–227 refine to 0.55–0.80 (mean 0.72), ATP to 0.61, and Mn to 0.28, while template-included control residues 150–155 remain near 1.0 (0.88–1.00; mean 0.96), and (iii) completed the composite omit map described above. Together, these results provide consistent evidence that densities corresponding to omitted regions are not driven by the template and can be recovered from the data, while template-included regions show some, albeit limited evidence of overfitting, as expected.

      Reviewer #1 (Public review):

      Summary:

      This paper describes an application of the high-resolution cryo-EM 2D template matching technique to sub-50kDa complexes. The paper describes how density for ligands can be reconstructed without having to process cryo-EM data through the conventional single particle analysis pipelines.

      Strengths:

      This paper contributes additional data (alongside other papers by the same authors) to convey the message that high-resolution 2D template matching is a powerful alternative for cryo-EM structure determination. The described application to ligand density reconstruction, without the need for extensive refinements, will be of interest to the pharmaceutical industry, where often multiple structures of the same protein in complex with different ligands are solved as part of their drug development pipelines. Improved insights into which particles contribute to the best ligand density are also highly valuable and transferable to other applications of the same technique.

      Weaknesses:

      Although the convenient visualisation of small molecules bound to protein targets of a known structure would be relevant for the pharmaceutical industry, the evidence described for the claim that this technique “significantly” improves alignment of reconstruction of small complexes is incomplete. The authors are encouraged to better evaluate the effects of model bias on the reconstructed densities in a revised paper.

      We thank the reviewer for these constructive comments. We have updated the template-bias metric Ω in Fig. 4 and added two further quantitative controls: grouped occupancy refinement of omitted residues and a composite omit map spanning the entire protein. Full details are provided in our responses to Comments 1 and 2 below.

      Reviewer #1 (Recommendations for the authors):

      Main Comments

      (1) For the 1ATP structure: Q-scores for deleted residues/ligands are worse than the Q-scores for residues in the template. This means that the reconstructed map must suffer from template bias. Another indication of this bias is that the density for the ATP (and the omitted residues) appears to be weaker than the density for the residues in the template (although this is not easy to assess from the figures). The authors should perform additional experiments to quantify this bias.

      (a) One option could be to do what the X-ray crystallographers call an OMIT map, and omit allresidues, a few at a time, from the template in multiple 2DTM runs. They could then assemble a density map from all the omitted residues together and measure the resolution of the omit map against the known template by FSC.

      (b) Another insightful experiment would be to take the various 2DTM reconstructed maps describedin the paper and perform a refinement of the atom occupancies of all residues in the structure. Residues included in the template should refine to values close to 1. In the absence of bias, the occupancies of the omitted residues should be 1 too; if the reconstructed map were completely biased, those occupancies would refine to 0. Therefore, the refined occupancies of omitted residues could perhaps serve as a measure for the amount of bias in the reconstructed map.

      We thank the reviewer for these detailed and constructive suggestions. We agree that the lower Q-scores for omitted regions indicate weaker density and that template bias exists at residues that are included in the template. To quantify this more directly, we corrected the template-bias metrics at the omitted region (mask from the full–omit template difference) in Fig. 4.

      Following the reviewer’s suggestion, we performed Phenix real-space grouped occupancy refinement against the omit reconstruction using the docked full model. The results are shown in Table. S2. We refined occupancies for the omitted residues (chain E 222–227), ATP, Mn, and template-included control residues (chain E 150–155), while excluding waters. The omitted residues refined to occupancies of 0.55–0.80 (mean 0.72), ATP to 0.61, and Mn to 0.28, whereas the control residues remained near 1.0 (0.88–1.00; mean 0.96). These results indicate substantial recovery of density in the omitted regions, but also some degree of bias.

      The substantially lower refined occupancy of Mn<sup>2+</sup> may reflect genuine partial occupancy in the dataset. While compact features can be especially sensitive to residual alignment error, we cannot conclude from the present analysis that alignment effects alone account for the weak Mn<sup>2+</sup> density.

      Finally, we have constructed a composite omit map to assess density recovery across the protein. We generated 36 omit templates, each deleting ∼10 non-overlapping residues scattered across the structure (including peripheral and surface-exposed regions). For each template, an independent 2DTM search and reconstruction was performed. Local density patches were extracted within 3 Å of the omitted atoms (with neighboring residues excluded as described in Methods) and assembled into a composite map (Fig. 5). The composite map shows that density can be recovered at distributed locations across the protein and is not restricted to the central binding pocket. Recovery is variable across sites, with some regions exhibiting weaker or fragmented density, consistent with local differences in signal-to-noise, structural heterogeneity, and residual alignment error.

      (2) The claim that 2DTM leads to “Improved” reconstruction (title) and “alignment and reconstruction [...] can be significantly improved” (abstract) is not supported by the data presented in the paper. The smallest single particle structure to resolutions sufficient for de novo atomic modelling is currently the ACA2 complex, with an ordered mass of less than 40 kDa, which was reconstructed using Blush regularisation in RELION. This paper should be referenced, and statements about single particle analysis (SPA) not working for sub-50 kDa complexes should be toned down. In general, I would say that 2DTM and SPA are not competing techniques, and the paper would be better if it focused on the intrinsic advantages of 2DTM (like ease-of-use for screening of pharmaceutical compounds) and useful findings described that make 2DTM better, e.g., excluding thick ice.

      We thank the reviewer for this important perspective and have added the Blush regularization reference Kimanius et al. (2024) to the revised manuscript, noting that the 40 kDa Aca2–RNA complex was reconstructed to 2.5 Å resolution using this approach (at L451). Furthermore, Blush regularization could be applied to reconstructions derived from 2DTM-based particle stacks, and a combination of both approaches may yield further improvements.

      We agree that 2DTM and SPA are complementary rather than competing techniques and have revised the manuscript to reflect this. We have also toned down claims in the abstract, which now states that 2DTM “reconstructed a previously intractable ∼43 kDa kinase complex and improved the density of its ligand-binding site” rather than making broad claims about SPA limitations. In the discussion, we now describe 2DTM as broadening possibilities for structural studies of targets “that have remained difficult to reconstruct” rather than implying they are impossible by SPA.

      Regarding the intrinsic advantages of 2DTM: beyond ligand screening, the composite omit map (Fig. 5, described in Comment 1) demonstrates that 2DTM-derived orientations support density recovery throughout the entire protein, including peripheral and surface-exposed residues, using roughly an order of magnitude fewer particles than conventional SPA workflows.

      (3) Given the uncertainties about the amount of template bias in the reconstructed 2DTM densities, I have trouble interpreting the predictions in Table 1. Where would the 1ATP structure lie in Figure 8? How much bias would there be in a 2DTM reconstruction at SNR n = SNR s? Could the authors perform tests on simulated data to confirm these predictions? At the point of SNR n = SNR s, how would a 2DTM reconstruction look, and what would refined occupancies for deleted residues be?

      (This may reflect a misunderstanding on my part, but I don’t really see how the SNR n = SNR s is completely dependent on the number of orientations searched (through Equation 1). In Figure 8, is the full search in a 4k x 4k micrograph, or inside a particle box? And what are the relevant search ranges? Perhaps as a consequence of this misunderstanding, I do not understand how one would decide on the amount of noise in the simulated data for these tests.)

      We thank the reviewer for this important question and agree that this point needed clearer explanation. In our framework, is the expected alignment-noise level from maximizing many cross correlations, where N<sub>s</sub> is the total number of sampled hypotheses in the 5D search (in-plane angle, out of-plane angles, and x, y shifts), not only the number of orientations. Thus, the relevant search is the per-particle alignment search window (full or constrained), not a full 4k×4k micrograph area.

      At SNR<sub>n</sub> = SNR<sub>s</sub>, the true-match and noise-maxima levels are at a threshold; one could imagine if SNR<sub>s</sub> is only slightly larger than SNR<sub>n</sub>, the correct pose is favored on average, so with sufficiently large particle numbers real omitted-region density should accumulate, but with residual pose errors that attenuate high-frequency amplitudes (effectively a large positive B-factor). In that regime, sharpening (negative-B correction) can improve visibility once signal is accumulated. Therefore, we expect partial recovery rather than fully unbiased recovery at this threshold, with omitted-region occupancies remaining between 0 and 1 and below template-included controls (consistent with our measured values), and improving as SNR<sub>s</sub> − SNR<sub>n</sub> and particle number increase. Simulations at this exact threshold would require a very large particle number to achieve sufficient statistics, and we leave this to future work. We have added this clarification to the Supporting Information.

      (4) The strong (> 5 sigma!!) and ubiquitous difference densities in Figure 9A imply that the authors have a serious problem with their forward model, which could explain some of the effects of model bias discussed above. I recommend they investigate these differences in detail. It would be good to see negative and positive densities in different colours to understand these differences better. The text speaks about incomplete capture of the solvent background, but the difference densities appear to be of much higher spatial frequencies than those typical for background/solvent effects (e.g., 15-20A). It may thus also be helpful to analyse these differences in Fourier space.

      We thank the reviewer for this important point. In our previous analysis, we did not incorporate an appropriate protein mask when generating the difference map, which contributed to widespread residual densities. We have now regenerated the map using the program diffmap.exe (https: //grigoriefflab.umassmed.edu/diffmap) with a protein soft mask and moved it to the Supplementary Information (Fig. Figure 1—figure supplement 4, contour SD = 20). With this controlled setup, the strongest coherent residual densities localize to the omitted ATP pocket and residues 222–227, consistent with recovery of omitted features. We have revised the figure/text accordingly and clarified that remaining diffuse residuals are likely due to forward-model mismatch (including solvent/background representation). We also added to the manuscript that improved template generation may be achieved by incorporating recent methods that learn environment-aware scattering factors directly from experimental cryo-EM maps.

      Other Comments

      (1) P.1: Alongside reference 2, a reference to the 1.2 Å apoferritin structure from the Stark group should be included.

      We have added the reference at L30.

      (2) P.2: “commond line tool”

      We have corrected the typo.

      (3) P.2-3: Robust reconstruction of the ATP binding pocket: Auto-refinements in RELION without alignments do not exist, and corresponding statements need to be removed from the manuscript. If one wants to skip alignments, then there is no refinement left to be done. In that case, one should just perform a reconstruction of the 2 halves (e.g., using relion reconstruct) and then run a standard RELION postprocessing.

      We agree with the reviewer and have revised the manuscript accordingly. Technically, RELION’s relion refine with the --skip align flag runs an iterative loop that re-estimates the per-particle noise model (spectral noise σ<sup>2</sup>) and computes the gold-standard FSC between half-maps, but it does not modify the particle orientations or translations. As the reviewer correctly points out, this is effectively a 3D reconstruction followed by postprocessing, not a refinement. We have updated the text to replace “skip-alignment auto-refinement” with “3D reconstruction without angular refinement” to accurately reflect what was performed.

      (4) P.3: What are “first-quadrant p-values” and “three-quadrant p-values”?

      We apologize for the ambiguity and now define these terms explicitly in the revised text (with citation to the p-value paper). After transforming z-score and SNR to probit coordinates, “first-quadrant” (1Q) p-values use only candidate points with both coordinates > 0 (i.e., both probit-zscore and probitSNR are positive). “Three-quadrant” (3Q) p-values include candidates where at least one coordinate is > 0 (equivalently, all points except the quadrant where both are < 0).

      (5) P.5: In Equation (2), it is unclear what Q means from the main text. Would it be better to leave Equation (2) for the Appendix, and only show Equation (3) in the main text?

      Thank you for this suggestion. We kept Equation (2) in the main text to preserve the continuity of the derivation, but we now define Q(k,N<sub>i</sub>) explicitly at first use as the normalized exposure-weighting transfer function (following Grant 2015). The detailed derivation and assumptions remain in the Supporting Information.

      (6) P.6: “Remaining gaps”: this section considers differences between 200 keV and 300 keV electron beam energies. The main practical effect for cryo-EM data sets is that the current detectors are designed for detecting 300 keV electrons, and their DQE is thus a lot worse at 200 keV. The entire paper doesn’t mention detectors. Perhaps because they are assumed to be perfect, but it is still far from the case.

      Also, why were defocus searches not performed if the thickness of micrographs was up to 1500 A?

      The conclusion of this section states “Considering all these factors...”, but it then claims standard single particle analysis still remains an outstanding challenge. This concluding statement makes no sense, as this whole section was about 2DTM.

      Thank you for this comment. We agree and have revised the text to make these points explicit. First, we now state clearly that detector response (DQE) is generally more favorable at 300 keV than at 200 keV, which contributes to the experimental–theoretical gap. Second, we clarify why we did not perform a defocus search in 2DTM: after CTF/thickness filtering, the retained micrographs are predominantly in the thin-ice regime, so expected defocus spread is smaller, while adding a defocus dimension substantially increases computational cost. We also tested downstream refinement (including CTF/beam-tilt related refinement in cisTEM) and did not observe measurable improvement for this dataset (data not included in the manuscript). Finally, we revised the concluding sentence in this subsection to refer specifically to 2DTM-based alignment limits rather than standard SPA, so the section scope is now consistent.

      (7) P.7: Data-driven refinement of AlphaFold3 models: it might be worth pointing out that removing residues a few at a time from AF3 models and checking their reconstructed density by 2DTM would come at a considerable computational cost.

      We agree. We have demonstrated residue-level omission validation using the X-ray template via a composite omit map (Fig. 5), confirming that the approach is feasible. We have updated the Discussion to reflect this: extending the composite omit approach to AlphaFold3-based templates remains computationally expensive — each omission design requires an independent 2DTM search and downstream reconstruction — and we present this as an important direction for future work.

      (8) Figure 1: What is “full FSC” and what is “particle FSC”?

      Thank you for pointing this out. We have clarified the terminology in the figure legend and text using cisTEM and Frealign definitions (Grant et al., 2018). What was previously labeled “Full FSC” is now referred to as the uncorrected FSC (FSC<sub>uncor</sub>), computed within a generous mask. “Particle FSC” denotes the solvent-corrected FSC, obtained from FSC<sub>uncor</sub> using the mask-volume correction factor f as described in the cisTEM/Frealign framework (Grant et al., 2018).

      (9) Figure 3: Why were particles in class 5 discarded? The 2DTM approaches described in this paper are all about carefully selecting good particles, yet now the authors use standard 3D classification to throw away another 156 particles. This seems to be an arbitrary choice. How different would the results have been if these had been included in the reconstruction? Alternatively, did these few particles have any 2DTM metrics that would justify their exclusion?

      We thank the reviewer for raising this point. Class 5 contained only 156 particles (∼2% of the dataset). While the 2DTM p-value and SNR metrics provide principled criteria for particle selection, they are not perfect, and a small number of suboptimal particles may still pass these filters. To address the reviewer’s concern, we repeated the reconstruction including all five classes. The resulting map achieved a resolution of 3.7 Å, identical to the reconstruction without class 5, confirming that including these particles does not affect the results. We have clarified this point in the manuscript.

      (10) Figure 4C: What are the negative sample thicknesses here? Why use an inset?

      The negative sample thickness values are artifacts of the CTF-based thickness estimation algorithm in ctffind5. This algorithm fits oscillations in the 1-D power spectrum arising from the interaction between the CTF and the specimen’s finite thickness (a sinc-modulated envelope). When the ice is very thin or the power spectrum is noisy, the optimizer can converge to a physically meaningless negative value. Of the 2,488 total micrographs across both sessions (after CTF score filtering, 2,314 retained), 136 (∼5.9%) returned negative thickness estimates. We have revised Figure 1—figure supplement 1c (previously Figure 4c) to show only the physically meaningful positive thickness values without the inset, which gives a clearer view of the unimodal distribution peaked near 350–400 Å.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Zhang et al describe a method for cryo-EM reconstruction of small (sub50kDa) complexes using 2D template matching. This presents an alternative, complementary path for high-resolution structure determination when there is a prior atomic model for alignment. Importantly, regions of the atomic model can be deleted to avoid bias in reconstructing the structure of these regions, serving as an important mechanism of validation.

      The manuscript focuses its analysis on a recently published dataset of the 40kDa kinase complex deposited to EMPIAR. The original processing workflow produced a medium resolution structure of the kinase (GSFSC ∼4.3 Å, though features of the map indicate ∼6-7 Å resolution); at this resolution, the binding pocket and ligand were not resolved in the original published map. With 2DTM, the authors produce a much higher resolution structure, showing clear density for the ATP binding pocket and the bound ATP molecule. With careful curation of the particle images using statistically derived 2DTM p-values, a high-resolution 2DTM structure was reconstructed from just 8k particles (2.6 Å non-gold standard FSC; ligand Q-score of 0.6), in contrast to the 74k particles from the original publication. This aligns with recent trends that fewer, higher-quality particles can produce a higher-quality structure. The authors perform a detailed analysis of some of the design choices of the method (e.g., p-value cutoff for particle filtering; how large a region of the template to delete).

      Overall, the workflow is a conceptually elegant alternative to the traditional bottom-up reconstruction pipeline. The authors demonstrate that the p-values from 2DTM correlations provide a principled way to filter/curate which particle images to extract, and the results are impressive. There are only a few minor recommendations that I could make for improvement.

      We appreciate the positive assessment. In response to the bias-related concerns raised elsewhere, we have: (i) updated the template-bias metric Ω reported in Fig. 4, (ii) added grouped occupancy refinement showing that omitted residues 222–227 refine to a mean occupancy of 0.72 while template-included control residues remain near 1.0, and (iii) assembled a composite omit map (Fig. 5) from 36 partial-deletion reconstructions spanning the entire protein. These additions are described in the revised Results and in the rebuttal below.

      Reviewer #2 (Recommendations for the authors):

      (1) On page 3, “Finally, by comparing Figure 2a and b, we observed that deleting IP20 strongly reduced signal at several residues.” Looking at Figure 2a and 2b, it was unclear which residues they were referring to.

      We have revised the text to explicitly list the affected residues. In the updated Figure 2, we now label the omitted residues with the lowest backbone Q-scores in the structural views (column 2) and include per-residue backbone Q-score plots (column 4), making the comparison between panels (a) and (b) quantitative. For example, when IP20 is additionally deleted (Fig. 2b), residues Phe54, Gly55, Lys72, Glu127, Glu170, and Asp184 all fall below a backbone Q-score of 0.5, compared with only Ser53 and Glu127 in the within-3 Å deletion alone (Fig. 2a).

      (2) Figure 1a. Both the published density map and the text “Template” are gray, but the 2DTM template density map is yellow.

      Thank you for catching this inconsistency. We have updated Figure 1a so that the 2DTM template density is now rendered in gray, consistent with the X-ray crystal structure (PDB) coloring. The published single-particle map is shown in wheat and the 2DTM reconstruction in blue, providing a clear three-way color distinction.

      (3) Figure 1b. I would recommend the x-axis label of “spatial frequency” instead of “resolution” (which is overloaded). Furthermore, the fact that this is not a GSFSC should be clearly labeled in the figure to prevent confusion with a standard GSFSC.

      We agree with both suggestions. The x-axis has been relabeled “Spatial Frequency (1/Å)” in the revised figure. We have also added a note in the figure caption stating that these FSC curves are not gold-standard FSCs, as the reconstruction uses orientations determined by template matching rather than independent half-set refinement.

      (4) Figure 2: The usage of the negative sign in the labels “-3 Å”, “-5 Å” to indicate within a given radius is a bit confusing. “Within 3 Å”, perhaps?

      Thank you for this suggestion. We have changed the labels in Figure 2 from “−3 Å” and “−5.5 Å” to “Within 3 Å” and “Within 5.5 Å.” We have also added a fourth column to Figure 2 showing per-residue backbone Q-scores for each deletion experiment, with omitted residues distinguished by color and marker shape. The residues with the lowest backbone Q-scores among the omitted set are circled in red and correspond to the labeled residues in the structural views.

      (5) Figure 4c: Why does the sample thickness histogram go to negative values (-20,000 A)?

      As noted in our response to Reviewer 1, the negative thickness values are artifacts of the ctffind5 thickness estimation, which fits a sinc-modulated envelope to the 1-D power spectrum. For micrographs with very thin ice or noisy power spectra, the fit can converge to unphysical negative values. These account for ∼5.9% of micrographs. We have revised Figure 1—figure supplement 1 (originally Fig. 4c) to display only positive thickness values, removing the inset and providing a clearer histogram.

      (6) Figured 4d: Should the label be “(Before Filtering)” instead of After?

      Yes, thank you for catching this. The original Figure 4d was mislabeled—it showed particle counts before filtering but was titled “After Filtering.” We have corrected the labels: Figure 1—figure supplement 1d (originally Fig. 4d) now reads “Before Filtering” and Figure 1—figure supplement 1e (originally Fig. 4e) reads “After Filtering.”

      (7) Supplementary Note 1: Please provide units for d, p, D, and k max in equation S4 and the preceding text.

      We have added units to the text preceding Eq. S4: d = 1/k<sub>max</sub> is the high-resolution alignment limit (Å), k<sub>max</sub> is the maximum spatial frequency (Å <sup>−1</sup>), p = d/2 is the ideal pixel size (Å/pixel), and D is the particle diameter (Å).

      (8) What does the map-model FSC look like with the template as the model vs. the AF3 structure as the model?

      We have computed the map–model FSC for both the X-ray crystallographic template (PDB 1ATP) and the AlphaFold3-predicted template against their respective 2DTM reconstructions (Fig. Figure 6—figure supplement 1). Both curves cross the FSC = 0.143 threshold at ∼2.3 Å. We note that the map–model FSC in this context should be interpreted with caution, because the vast majority of the structure lies outside the omitted region and is present in the template, so template bias in those regions will dominate the map–model FSC and obscure differences in the small omitted region.

      Reviewer #3 (Public review):

      Summary:

      Due to the low SNR of cryo-EM micrographs necessitated by radiation damage, determining the structure of proteins smaller than 50 kDa is exceedingly challenging, such that only a handful have been solved to date. This work aims to improve the reconstruction of small proteins in single-particle cryo-EM by using high-resolution 2D template matching, an algorithm previously used to locate and align macromolecules in situ, to align and reconstruct small proteins. This approach uses an existing macromolecular structure, either experimentally determined or predicted by AlphaFold, to simulate a noise-free 3D reference and generates whitened projections, crucially including high-spatial-frequency information, to align particles by the orientation with maximal cross-correlation. They demonstrate the success of this approach by generating a 3D reconstruction from an existing dataset of a 41.3 kDa protein kinase that had previously evaded attempts at high-resolution structure determination. To alleviate concerns that this is purely from template bias, they demonstrate clear density at two regions that were not present in the template: 6 residues in an alpha helix and an ATP in the ligand binding pocket. The latter is particularly important for its implications in determining structures of ligand-bound proteins for drug discovery. Additionally, the authors provide an update to the classic calculation in Henderson 1995 to predict the minimum molecular mass of a protein that can be solved by single-particle cryo-EM.

      Strengths:

      I am in no doubt that this technique can be used to gain valuable insights into the structures of small proteins, and this is an important advancement for the field. The ability to determine the structure of ligands in a binding site is particularly important, and this paper provides a method of doing that which outperforms traditional single-particle cryo-EM processing workflows.

      The claim that using high-spatial frequency information is essential for aligning small proteins is a valuable insight. A recent pre-print published at a similar time to this manuscript used high-resolution information in standard ab-initio reconstruction to generate a high-resolution reconstruction from the same dataset, supporting the claims made in the manuscript.

      The theoretical section outlined in the appendix is also theoretically sound. It uses the same logic as Henderson, but applies more up-to-date knowledge, such as incorporating dose-weighting and altering the cross-correlation-based noise estimation. This update is valuable for understanding factors preventing us from reaching the theoretical limit.

      Weaknesses:

      Given that this technique creates template bias, only parts of the reconstruction not in the template can be trusted, unlike standard single-particle processing, where the independent half-maps from separate, ab initio templates are used to generate a 3D reconstruction. Although, in principle, one could perform the search many times such that every residue has been omitted in at least one search, this will be extremely computationally intensive and was not demonstrated in this manuscript. It is therefore currently only realistically applicable when only a small portion of the sub-50 kDa protein is of interest.

      The applicability of this technique to more than a single target was also not demonstrated, and there are concerns that it may not work effectively in many cases. The authors note in the results that “the ATP density was consistently recovered more robustly than nearby residues” and speculate that this may be because misalignments disproportionately blur peripheral residues. Since the region of interest in a structure is not necessarily in the center, this may need further investigation. The implications of this statement may also be unclear to the reader. For example, can this issue be minimized by having the region of interest centered in the simulated volume?

      In Figure 3, the authors demonstrate that it is not solely improved particle filtering and a noise-free reference that improves alignment, but that the high spatial frequency information is important. This information is very valuable since it can be applied to other, more standard methods. However, this key figure is not as clear or convincing as it could be. The FSC curves are possibly misleading, since the reduced resolution could be explained by reduced template bias when auto-refining with a map initially low-pass filtered to 10 A. Moreover, although the helix reconstruction does look slightly better using the 2DTM angles, the improvement in density for ATP in the binding pocket is not clear. A qualitative argument only clear in one out of two cases is not as convincing as a quantitative metric across more examples.

      We address these concerns in three ways: (i) we quantify template bias using Phenix real-space grouped occupancy refinement: omitted residues 222–227 refine to occupancies of 0.55–0.80 (mean 0.72) and ATP to 0.61, while template-included control residues 150–155 remain near 1.0 (mean 0.96), confirming that recovered density is genuine rather than a template artifact; (ii) we have now completed a composite omit-map experiment (Fig. 5), in which 36 partial-deletion templates, each omitting ∼10 non-overlapping residues, were used to perform independent 2DTM searches and reconstructions; local density patches from all 36 reconstructions were assembled into a composite map showing density recovery at distributed locations across the protein, including peripheral and surface-exposed regions, although recovery is variable across sites; and (iii) we have expanded the discussion to clarify that, while the primary scope of this work is omitted-region validation for the ligand-binding site, the composite omit-map result demonstrates that the approach generalizes beyond the central pocket.

      Reviewer #3 (Recommendations for the authors):

      In addition to the comments on the public review, I have some more specific suggestions that could improve the manuscript.

      (1) Another recent pre-print posted on BioRxiv shortly before this manuscript (Kim et al. Highresolution ab initio reconstruction enables cryo-EM structure determination of small particles) determined a high-resolution structure of the same protein from the same dataset, as well as determining the structures of other small proteins. Since both manuscripts rely on high-spatial frequency information, I think that the paper strengthens the claims in this manuscript and should be cited.

      We thank the reviewer for this suggestion. We agree that the recent preprint by Kim et al. strengthens the relevance of high-spatial-frequency information for small-particle cryo-EM reconstruction. We have now added this work to the revised manuscript and included a brief discussion comparing its ab initio strategy with our 2DTM-based approach.

      (2) The claim in the abstract that “we were able to reconstruct previously intractable targets under 50 kDa and improve the density of the ligand-binding sites in the reconstructions” should be altered to make it clear that this is only a single previously intractable target.

      We agree. The revised abstract now reads “. . . we reconstructed a previously intractable ∼43 kDa kinase complex and improved the density of its ligand-binding site” making clear that a single target is demonstrated in this work.

      (3) Q-scores in the manuscript were sometimes used to quantify the improvement in map to model fit for the ATP binding pocket, but never for the 6 residues of the alpha helix. They were also not reported in every case for the ATP-binding pocket. This could lead a reader to think it is only being reported when the Q-score matches the expectation. For transparency, I would suggest either using Q-scores in every comparison or in no cases and simply relying on the qualitative result.

      We agree with the reviewer. In the revised manuscript, we now report Q-scores consistently for both ATP and residues 222–227 across all conditions: individual residue Q-scores for the omitted residues 222–227 in Fig. 1 are reported in the main text and figure caption; per-residue backbone Q-score plots for all deletion experiments in Fig. 2 are shown as the fourth column of each panel; Fig. 3 (RELION reconstruction) does not include Q-scores as the focus is on orientation accuracy rather than map-model fit; and average Q-scores for all four particle selection conditions in Fig. 4 are listed in Figure 4—source data 1.

      (4) The sigma values used for viewing the maps should also be stated in several figures, particularly Figure 3 and Figure 6.

      We have added contour levels (σ) to the captions of Fig. 3 and Fig. 4 (originally Fig. 6) in the revised manuscript.

      (5) I have a slight concern about how well this method applies away from the region centered in the alignment. If parts on the periphery of the structure are removed, do these also reconstruct? Is it required that the omitted region be centered in the simulation of the 3D volume for each alignment? If so, this should be clearly stated.

      2DTM determines particle orientations by matching the full projected template to the image, so alignment is driven by the global structure rather than a localized region. As a result, the recovered orientations define the reconstruction throughout the entire particle, not only near the center. The omitted region does not need to be centered in the template volume. Any region of the protein can be omitted and its density evaluated after reconstruction.

      To directly test whether peripheral regions are recovered in the same manner as central ones, we performed a composite omit-map experiment. We generated 36 omit templates, each deleting ∼10 non-overlapping residues distributed across the entire protein, including peripheral and surface-exposed regions. For each template, an independent 2DTM search and reconstruction was performed. Local density patches corresponding to the omitted regions were then extracted and assembled into a composite map (Fig. 5). The resulting map shows density at distributed locations across the protein, indicating that density recovery is not restricted to regions near the alignment center and that peripheral regions can be reconstructed under the same alignment framework, although the quality of recovery varies across sites.

      (6) I was confused by the difference between the FSCs in Figure 1 and Figure 3. I understand Figure 1 is from cisTEM and Figure 3 from RELION, but I expected the unmasked FSC and full FSC to be similar. Do the authors have any insights into why there is such a large difference? I would also consider removing the FSCs in Figure 3, since the reduced resolution may only be due to reduced template bias, meaning including this may be misleading.

      Thank you for raising this point. The apparent discrepancy arises from multiple differences between the two figures: different FSC definitions, different half-maps (reconstructed with different software and slightly different particle sets), and different masks.

      In cisTEM (Fig. 1), two FSC curves are reported: the uncorrected FSC (FSC<sub>uncor</sub>), measured within a spherical mask, and the “Particle FSC”, which applies an analytical solvent-fraction correction (Grant et al., 2018) to account for solvent dilution within the mask. The Particle FSC crossed the 0.143 threshold at ∼2.6 Å, whereas FSC<sub>uncor</sub> crossed at ∼3.0 Å. In Fig. 3, RELION postprocess applied phase-randomization correction with a soft mask, yielding ∼3.1 Å. However, the Fig. 3 FSC was computed on different half-maps (RELION skip-alignment reconstruction of 7,197 particles after 3D classification) with a different mask.

      To directly compare the two packages, we computed the FSC on the same cisTEM half-maps using both methods (Figure 3—figure supplement 1). The cisTEM Particle FSC (spherical mask + solvent correction) gave ∼2.6 Å, while RELION image handler with a tight 3D protein mask gave ∼2.7 Å. These two approaches converge to a similar resolution through different mechanisms: cisTEM compensates for a generous spherical mask using the solvent-fraction correction, while RELION uses a tight mask that excludes most solvent directly. This confirms that when the same half-maps are used, the two packages give consistent results and the apparent discrepancy between Figs. 1 and 3 is primarily due to differences in the reconstruction and particle set, not the FSC calculation.

      We agree with the reviewer that the FSC values in Figure 3 should be interpreted with caution. In this case, the particle orientations are not independently refined but are instead inherited from the 2DTM alignment, so the two half-maps are not strictly independent. We have added clarifying language in the revised manuscript to make this point explicit (Fig. 1 caption).

      (7) I would also like to see how RELION auto-refinement performs with different low-pass filtering. This could strengthen the argument that high-resolution information is necessary from the start to successfully align small particles.

      We thank the constructive suggestion from the reviewer. We performed RELION auto-refinement on the same 7,197-particle stack using different initial low-pass filter resolutions (--ini high) of 3, 5, 10, and 15 Å. The resulting post-processed resolutions were:

      Author response table 1.

      The results show that varying the initial low-pass filter has minimal effect on the final resolution. This is expected because RELION uses a gold-standard, maximum-likelihood framework in which the resolution used for alignment is determined iteratively from the data via a probability distribution, rather than being fixed by the initial reference. After the first iteration, the reference is updated from the data, and higher-resolution information is incorporated only to the extent supported by the definition of the current reconstruction. Consequently, differences in the initial low-pass filter have limited impact on the final refinement outcome.

      This behavior contrasts with 2DTM, where alignment is performed by direct cross-correlation against a fixed template. In this case, high-resolution features in the template contribute directly to the scoring function and can improve alignment accuracy.

      To directly test the importance of high-resolution information for 2DTM alignment, we performed an additional experiment in which 2DTM was run on bin4x images (2.234 Å/pixel), and the detected particle coordinates were used to extract particles from the corresponding bin2x images (1.117 Å/pixel) for reconstruction. Despite using the same bin2x images for reconstruction, the bin4x-aligned particles yielded a map in which ATP density was lost and backbone density for residues 222–227 was visibly degraded compared to the bin2x-aligned reconstruction (Fig. Figure 1—figure supplement 3). This demonstrates that access to high-spatial-frequency information during template matching is critical for accurate alignment of small particles.

      (8) The caption in Figure 3 should be more descriptive about what is being shown in each panel.

      We have substantially expanded the Figure 3 caption. It now describes each panel explicitly: (a) 3D classification results with particle counts, percentages, and per-class resolutions; (b) side-by-side comparison of reconstructions using 2DTM orientations versus RELION auto-refine, including full maps, zoomed binding-pocket views with the atomic model overlaid, orientation distributions, and FSC curves with reported resolutions; and (c) a table of RELION auto-refinement resolution as a function of the initial low-pass filter setting. We also added a new panel (c) showing that including all five classes yields the same 3.7 Å resolution, addressing the concern about Class 5 exclusion.

      (9) Figures 4 and 5 may be better suited as supplementary figures.

      We agree. Figures 4 and 5 have been moved to the Supplementary Information in the revised manuscript.

      (10) In Figure 4c, it is difficult to understand why the thickness distribution plot goes negative, especially to such a high magnitude as 1.5 microns.

      We agree this was confusing. The negative values are fitting artifacts from ctffind5’s thickness estimation, which fits a sinc-modulated envelope to the power spectrum. When the ice is very thin or the spectrum is noisy, the optimizer can converge to unphysical negative values (affecting ∼5.9% of micrographs). We have revised Figure 1—figure supplement 1c (previously Figure 4c) to show only positive thickness values, which now clearly displays the unimodal distribution peaked at 350–400 Å.

      (11) In Figure 5d, the micrograph looks a lot like a cross-grating grid used for calibration instead of crystalline ice or a fractured film.

      We agree. We have updated the caption for Figure 1—figure supplement 2d (originally Figure 5d) to read “Cross-grating calibration grid”

      (12) Figure 6 was very surprising to me if I am interpreting it correctly. It is not stated in the caption what omega is, but I am assuming it is a measurement of template bias. It is very surprising that the template bias drops when using more particles by reducing the p-value from 8.0 to 7.0. This goes against what I understood from Lucas et al. 2023, so I am curious as to why this is the case.

      We thank the reviewer for this question and apologize for the unclear presentation. We have revised Fig. 4 (previously Figure 6) and its caption to define Ω explicitly and updated the Ω values. We also identified that the mask used in the original computation was too loose; the revised mask is now constrained to the omitted region only (ATP, Mn<sup>2+</sup>, and residues 222–227), derived from the difference between the full and omit templates and shown in Figure 4—figure supplement 1. Ω is adapted from the template-bias metric introduced in (Lucas et al., 2023) and measures how much of the density in the omitted region is attributable to using the full template rather than the omit template. Specifically, for each particle selection condition we reconstruct two maps using orientations and particles derived from independent 2DTM searches with the full and omit templates (V<sub>full</sub> and V<sub>omit</sub>, respectively). Ω is the fractional reduction in density within the omission mask: . In the revised Fig. 4, Ω increases from 46% (p-value = 8.0) to 48% (p-value = 7.0), consistent with the expectation that including more, lower-quality particles increases the relative contribution of the template to the reconstruction. The Ω values are 48% for the SNR = 7.5 and 53% for the tilt conditions.

      (13) It would be useful if the in-house Python script used to calculate template bias could be made publicly available.

      We agree. The template-bias calculation (measure-template-bias) is now included in the publicly available Python package at https://github.com/kekexinz/2DTM_postprocess_tool, and can also be accessed in the official cisTEM repository at https://github.com/timothygrant80/cisTEM. The package also contains the extract-particles and filter-particles tools described in the Methods section.

      (14) The p-value used is said to be a three-quadrant p-value instead of a one-quadrant p-value. Although I assume this is simply replacing an ‘and’ statement with an ‘or’ statement, the exact difference could be made clearer to the reader.

      We have now defined these terms explicitly in the revised Methods. After probit transformation of z-score and SNR, the first-quadrant (1Q) p-value requires both values to be > 0 (logical AND), whereas the three-quadrant (3Q) p-value requires at least one to be > 0 (logical OR). The 3Q criterion is therefore looser, retaining more candidates—which is beneficial for small targets that may score well on one metric but not both.

      (15) I was, perhaps naively, surprised that z-scores could not be used. It was my understanding that by removing the rotationally invariant component from the cross-correlation, the z-score would down-weight low-resolution information compared to the cross-correlation. Given that the manuscript suggests low-resolution alignment can cause getting stuck in local minima, this is surprising to me. The authors note it led to the rejection of most particles; were there simply too many false positives when a lower threshold was used?

      The reviewer is correct that subtracting the angular mean removes the rotationally invariant component of the cross-correlation. However, the resulting z-score primarily measures how strongly a specific orientation stands out relative to other orientations. In other words, it reflects the orientation discriminability (closely related to Fisher information) rather than the absolute correlation strength. For small particles the cross correlation often varies only weakly across orientations, so CC<sub>max</sub>− CC<sub>avg</sub> remains small even when the absolute correlation is significant. As a result, using the z-score alone as a selection criterion led to the rejection of many true particles.

      Theoretical Section Improvements

      (a) The discussion on beam-induced motion could be improved by separating it into initial motion (e.g., cryo-crinkling, buckling) that can be eliminated through grid design, and pseudo-Brownian motion, which cannot. Pseudo-Brownian motion will become much more significant for small proteins (based on reference 5, for a 10 kDa protein, this would be a MSD of ∼0.1 A 2/e−/A 2, or a B-factor of over 2 A 2/e−/A 2), and Bayesian Polishing is unlikely to correct this perfectly, given that it imposes a smoothness of motion between nearby particles. The impact of not correcting for this could be quantified more explicitly.

      We thank the reviewer for this helpful suggestion. As noted, pseudo-Brownian motion of particles within irradiated ice introduces stochastic displacements that accumulate with dose and are expected to be more significant for small particles. Based on the analysis in (Mcmullan et al., 2015), and scaling with particle size, this effect can be aproximated as a dose-dependent mean-squared displacement (MSD) of ∼0.1 Å<sup>2</sup> per (e<sup>−</sup>/Å<sup>2</sup>) for a ∼10 kDa particle. Over a typical total exposure of 40–60 e<sup>−</sup>/Å<sup>2</sup>, this corresponds to an accumulated RMS displacement of ∼2–2.5 Å, sufficient to attenuate high-resolution signal.

      In practice, such motion acts as an additional high-frequency attenuation in Fourier space, analogous to an envelope function, reducing the coherent signal available for template matching. While Bayesian polishing can partially correct beam-induced motion, it assumes spatially smooth trajectories between nearby particles and therefore may not fully compensate for stochastic, particle-specific motion.

      Within the theoretical framework presented here, this effect can be interpreted as an additional frequency-dependent damping of the signal (B-factor). Its primary consequence would be to reduce the effective signal-to-noise ratio at high spatial frequencies and therefore shift the detectable molecular-weight limit somewhat upward, without altering the structure of the derivation. We have added text in the manuscript to clarify this point and to indicate the expected magnitude of this effect.

      (b) The inclusion of inelastic scattering assumes an energy filter is being used, and this should be clearly stated.

      We have added this clarification in the inelastic scattering paragraph of the Supplementary Information.

      (c) The reasons for not including other factors, such as DQE and the temporal and spatial coherence envelope functions, could be stated.

      We have added a note in the dose-weighting section clarifying that these instrument-dependent attenuation factors were not explicitly included, and that they could be incorporated as additional frequency-dependent weighting terms without changing the structure of the derivation.

      (d) The flexibility and heterogeneity in protein structures, especially at high spatial frequencies, must also be a reason for a gap from experiment to theory, but this is not clearly stated.

      We agree. We have added a statement in the “Remaining gaps” section noting that structural flexibility and conformational heterogeneity act as an additional envelope that attenuates high-resolution signal relative to the rigid-particle model assumed in our derivation.

      Additional Minor Comments

      (15) It is noted in the discussion that 2DTM-based single-particle alignment simplifies the processing pipeline. Although true, I think stating the computation time would be useful for the reader.

      We have added computation times to the Discussion. For a typical single-particle dataset of ∼2,000 micrographs (5k × 4k pixels), a 2DTM search without defocus refinement completes in approximately one day on 64 NVIDIA A6000 GPUs. Once particles are located with their orientations and positions, a single 3D reconstruction is sufficient without further refinement, eliminating the iterative 2D classification, ab initio modeling, 3D classification and refinement steps of a conventional pipeline.

      (16) There are some formatting issues with e−/A 2, sometimes losing the minus sign.

      Thank you for catching this. We have corrected all instances to consistently use e<sup>−</sup>/Å<sup>2</sup> throughout the manuscript.

    1. Author response:

      Reviewer #1 (Public review):

      The study by He and colleagues aims to investigate the molecular mechanisms driving key cell potency transitions, particularly the naïve-to-primed pluripotency transition. The authors explore the relationship between cell polarity and stemness using stem cell models combined with a comprehensive panel of experiments, including pharmacological inhibition and co-culture/conditioned medium rescue approaches. Overall, the study provides interesting observations and contributes to the understanding of the molecular mechanisms dynamically regulating stem cell differentiation.

      However, several conceptual and interpretational aspects could be strengthened:

      (1) First, the Introduction would benefit from being more focused on what is currently known regarding cell polarity during early embryogenesis and pluripotent stem cell transitions, rather than emphasizing later neurogenesis events. Such reorientation would better match the main topic of the manuscript and improve the conceptual coherence of the study.

      We thank the reviewer for this constructive suggestion. We fully agree that the Introduction should be more tightly focused on the current understanding of cell polarity during early embryogenesis and pluripotent stem cell transitions, rather than on later neurogenesis events.

      Accordingly, we will revise the Introduction in the following ways:

      (1) Reduce the discussion on later neurogenesis and move some of those details to the Discussion section where they more appropriate.

      (2) Expand the background on early embryonic development and pluripotent stem cell transitions by citing key recent and classical references, including but not limited to: cell polarity establishment in the preimplantation embryo, apical–basal polarity during lineage specification, polarity remodeling in naïve-to-primed pluripotent stem cell transition, the role of PAR complex in early mouse development.

      (3) Refocus the Introduction to clearly state: what is known about polarity in early embryogenesis and pluripotent states, what remains unknown, and how our study addresses that gap.

      (2) Similarly, Figure 6, where the authors attempt to provide clinical relevance through neural organoid formation experiments, feels somewhat disconnected from the central theme of the naïve-to-primed transition. Although this section is interesting on its own, there is already extensive literature describing polarization and morphogenetic events occurring much earlier during pluripotent state transitions. Therefore, the developmental relevance of the neural differentiation phenotypes could be better contextualized in relation to earlier morphogenetic events associated with pluripotency progression.

      We thank the reviewer for this insightful comment. We agree that the neural organoid experiments in Figure 6 are somewhat disconnected from the central theme of the naïve-to-primed transition, and that extensive literature already exists on polarization events occurring earlier during pluripotent state transitions.

      In the revised manuscript, we will better contextualize these findings by explicitly discussing how the neural differentiation phenotypes relate to the earlier morphogenetic events associated with pluripotency progression, rather than presenting them as a standalone observation. We will also incorporate relevant references to bridge this gap and strengthen the developmental relevance of our neural organoid data.

      (3) The manuscript contains a substantial amount of experimental work; however, several results would benefit from deeper discussion. For example, in Figure 1, what is the rationale behind ZO1 downregulation being observed specifically in primed PAR knockout cells but not under naïve culture conditions? In addition, in Figure 3, the authors perform co-culture and conditioned medium experiments between wild-type and knockout cells. While the authors focus on the secreted protein fraction that rescues the phenotype, they also mention that other fractions display rescuing activity. Could the authors briefly discuss what additional components may contribute to this rescue effect? For example, could other molecules within these fractions also converge on AKT signaling regulation?

      We thank the reviewer for recognizing the substantial experimental work in our manuscript and for providing these thoughtful suggestions to improve the depth of our discussion. We agree that deeper discussion of several key results will strengthen the manuscript. In the revised version, we will address the specific points as follows:

      (1) Regarding ZO1 expression in Figure 1:

      Our primary focus is actually on ZO1 localization rather than its total expression level. In our experiments, RNA-seq and immunofluorescence analysis revealed that the total expression level of ZO1 does not change significantly in PAR knockout cells. However, ZO1 localization is markedly altered in PAR knockout primed cells. Specifically, in wild-type primed cells, ZO1 is predominantly localized at the cell membrane, whereas this specific membrane accumulation is not observed in PAR knockout primed cells. Furthermore, this phenomenon is observed specifically under primed state and does not occur under naïve culture conditions. This is likely due to the differential requirement for PAR complex components in maintaining tight junction integrity during distinct pluripotency stages.

      (2) Regarding the rescue activity of other fractions in Figure 3:

      In our experiments, we found that beyond the secreted protein fraction, the WT CM-Exosome fraction exhibited limited rescue efficacy, particularly during the later stages of NPT. Based on our literature review, we suggest that these exosomal components may still contribute to the observed rescue effect, potentially through the delivery of functional proteins, miRNAs, or other signaling modulators that converge on AKT signaling regulation. This discussion will provide a more comprehensive understanding of the paracrine communication between wild-type and knockout cells, while acknowledging the limited contribution of exosomes relative to the secreted protein fraction.

      (4) Importantly, transitions in cell potency are frequently associated with coordinated morphogenetic changes. For example, during mouse embryogenesis, naïve pluripotent inner cell mass cells progressively polarize into a rosette-like structure with apical domain specification before lumen formation and epithelialization during progression toward the primed epiblast state. This developmental context could help strengthen the biological interpretation of the study.

      We sincerely thank the reviewer for providing this valuable developmental context. The example of naïve pluripotent inner cell mass cells progressively polarizing into rosette-like structures with apical domain specification before lumen formation and epithelialization during progression toward the primed epiblast state is highly insightful and directly relevant to our study.

      In the revised manuscript, in the Introduction section, we will incorporate this developmental perspective to strengthen the biological interpretation of our findings. Specifically, we will place greater emphasis on the role of Par complex-mediated cell polarity in coordinating both pluripotency transitions and morphogenetic changes during early embryogenesis. We believe this contextualization will significantly improve the framing of our study and better connect our in vitro observations to in vivo developmental processes.

      (5) There are also several claims throughout the manuscript that appear to be overinterpreted or insufficiently quantified. For example, in Figure 1, the authors state that CDH1 expression is uniform; however, this is difficult to appreciate from the images shown, and quantitative analysis would be necessary to support this conclusion.

      We thank the reviewer for this important comment. We agree that the claim that "CDH1 expression is uniform" in Figure 1 is overinterpreted based on the images shown, and we apologize for the lack of quantitative support.

      Upon re-examination, we realize that our focus should be on CDH1 localization rather than its expression level or uniformity. In the updated manuscript, we will rephrase the statement about uniformity and instead present appropriate quantitative analysis (e.g., RNA-seq or fluorescence quantification across multiple cells) to better support our conclusions regarding CDH1 distribution. We will also adjust our data presentation to more clearly reflect the localization changes we observe.

      (6) Another example appears in Figure 2, where the authors claim that "heatmap analysis revealed that transcriptomic profiles of PAR knockout cells progressively diverged from wild type from day 3 onwards". This conclusion is not fully supported by the presented data for two reasons: (1) transcriptomic divergence is more appropriately assessed through principal component analysis, clustering, or distance-based methods rather than by visual inspection of a heatmap alone; and (2) although some genes displayed in panel E begin to show genotype-associated differences from day 3, the overall transcriptomic structure shown in the PCA and heatmap remains primarily dominated by temporal progression rather than genotype.

      We thank the reviewer for this careful and constructive critique. We apologize for the imprecise claim regarding the heatmap analysis in Figure 2. We agree that 1) transcriptomic divergence should be assessed by PCA, clustering, or distance-based methods rather than by visual inspection of a heatmap alone, and 2) the overall transcriptomic structure shown in PCA and heatmap remains primarily dominated by temporal progression rather than genotype.

      In fact, our main point in this figure was to show that differentially expressed genes (DEGs) between PAR KO and WT become more numerous and more pronounced from day 3 onwards, and the supporting data for this claim are presented in Supplemental Figure 2 A–B. The number of DEGs between PAR knockout and wild-type cells is 480 at day 1, 523 at day 3, 1088 at day 4, and 1893 at day 6. Furthermore, we focused on specific genes within particular signaling pathways, and their expression levels began to show significant differences between PAR knockout and wild-type cells from day 3 onwards.

      We realize that our original wording was misleading. In the revised manuscript, we will rephrase our conclusion to more accurately reflect what the data actually show, focusing on the timing and extent of differential gene expression rather than suggesting a global divergence of transcriptomic profiles.

      (7) In this context, it remains unclear whether PAR knockout cells truly retain a more naïve pluripotent transcriptomic identity. To support this claim, the authors should compare the knockout transcriptome directly against a naïve pluripotent population. The phenotype observed in the knockout cells may instead represent an incomplete or aberrant primed transition rather than maintenance of naïve pluripotency itself. Intermediate morphogenetic states, such as rosette-like epithelial stages, could also explain the observed phenotype.

      We apologize for the confusion caused by our imprecise wording. We realize that our original manuscript may have inadvertently suggested that Par knockout cells retain a naïve pluripotent transcriptomic identity, which was not our intended claim.

      To clarify, Par knockout naïve cells lose their naïve identity and differentiate toward a primed state during the NPT process described in this manuscript. Unlike wild-type primed cells, PAR-knockout primed cells exhibit altered morphology: they cannot establish or maintain the typical flat morphology, and possess distinct expression profile. In terms of naïve identity, key naïve markers (e.g., Esrrb or Oct4) are downregulated to comparable levels in both wild-type and Par knockout primed cells. Although the two cell types differ in their overall expression profiles, several core primed markers (e.g., Fgf5 or T) show normal expression in both groups. Collectively, these results indicate that Par knockout naïve cells do lose their naïve identity and undergo differentiation toward a primed state during NPT, even though the final primed states of the two cell populations are distinct.

      In the revised manuscript, we will:

      (1) Revisit and revise our wording to avoid any misinterpretation that Par knockout cells retain a naïve identity.

      (2) Directly compare the transcriptome of Par knockout cells against a true naïve pluripotent population (e.g., naïve ESCs) to further support our conclusion that the knockout cells are not maintaining naïve pluripotency, but rather exhibit an aberrant primed state with morphological abnormalities.

      (3) Discuss the possibility that the observed phenotype may represent an intermediate morphogenetic state (e.g., rosette-like epithelial stages) rather than genuine naïve pluripotency maintenance.

      (8) Strengthening this aspect of the study would substantially improve its developmental and in vivo relevance, which currently appears somewhat limited. In particular, it would be interesting to determine whether this mechanism operates during embryogenesis itself. The authors could consider relatively simple but informative experiments, such as perturbing PAR signaling or Furin activity during embryo culture.

      We thank the reviewer for this constructive and forward-looking suggestion. We agree that the current manuscript focuses primarily on in vitro cellular mechanisms, and we have not sufficiently explored the developmental and in vivo relevance of our findings. We acknowledge that this aspect of the study is currently somewhat limited.

      In the revised manuscript, we will:

      (1) Explicitly acknowledge this limitation in the Discussion section.

      (2) Incorporate more background on early embryogenesis, particularly regarding pluripotency transitions and morphogenetic changes during early development, to better contextualize our in vitro observations.

      (3) We will attempt to use embryo-like models to investigate whether the PAR complex–Furin–Lefty–FAK signaling axis also operates during embryogenesis itself. As the reviewer suggested, simple but informative experiments—such as perturbing PAR signaling or Furin activity during embryo culture—would be valuable next steps to determine the in vivo relevance of our proposed mechanism. We will include these as important future perspectives.

      (9) Along the same lines, some statements in the manuscript appear overly speculative. For example, the statement that "these findings may reveal a developmental compensation mechanism during embryogenesis, whereby normal cells rescue defective cells or increase their own proportion" extends well beyond the experimental evidence presented. Such claims invoke concepts related to cell competition, abnormal cell recognition, or developmental quality control mechanisms in vivo, none of which are directly demonstrated in this study. The authors are encouraged either to substantially tone down these statements or move them to the Discussion as speculative possibilities.

      We thank the reviewer for this important critique. We agree that our original statement—"these findings may reveal a developmental compensation mechanism during embryogenesis, whereby normal cells rescue defective cells or increase their own proportion"—is overly speculative and extends beyond the experimental evidence presented in our study. We also acknowledge that it was inappropriate to directly extrapolate from in vitro cellular mechanisms to in vivo developmental rules without proper justification.

      In the revised manuscript, we will:

      (1) Substantially tone down this claim from the Results section.

      (2) Move this speculation to the Discussion section, where we will explicitly present it as a speculative possibility rather than a conclusion supported by our data. We will also clearly state that concepts such as cell competition, abnormal cell recognition, or developmental quality control mechanisms remain to be tested in future studies.

      (10) Another important conceptual point concerns the relationship between PAR complex regulation and Lefty signaling. If this mechanism indeed reflects a physiological or homeostatic process operating during embryogenesis, what would be the developmental rationale for the PAR complex regulation of Lefty? Lefty is well known for its role during gastrulation and anterior epiblast patterning. It would therefore be interesting if the authors could further discuss potential links between these developmental contexts.

      We thank the reviewer for raising this important conceptual point. In our manuscript, we have indeed demonstrated that the PAR complex regulates Lefty signaling under the conditions of this study, and we are aware from the literature that Lefty signaling plays a critical role during early embryogenesis, particularly in gastrulation and anterior epiblast patterning.

      However, we admit that we have not deeply considered the potential pathways and developmental rationale for PAR complex-mediated regulation of Lefty in the context of embryogenesis. This is an important gap in our current discussion.

      In the revised manuscript, we will:

      (1) Review and incorporate relevant literature to better understand and discuss the potential links between PAR complex regulation and Lefty signaling during early embryonic development, including possible connections to gastrulation and anterior patterning.

      (2) Offer speculative but informed perspectives on the developmental rationale for such regulation, while clearly distinguishing between what our data directly show and what remains to be explored in future studies.

      Minor points:

      (1) The authors state that PAR knockout cells do not exhibit major differences in self-renewal capacity; however, they simultaneously claim that these cells remain in a more naïve-like state. This interpretation requires clarification, as naïve pluripotent cells are typically associated with increased clonogenicity, enhanced self-renewal, and expression of markers such as alkaline phosphatase and SSEA1 compared to primed cells. The relationship between the observed phenotype and the proposed "naïve-like" state should therefore be discussed more carefully.

      We thank the reviewer for this comment, which addresses a similar concern as Point 7 mentioned above. Consistently, we do not claim that PAR knockout cells remain in a more "naïve-like" state. Our actual conclusion is that PAR knockout naïve cells undergo differentiation toward the primed state during NPT. However, due to loss of cell polarity, PAR knockout primed cells fail to establish and maintain the typical flat morphology and instead form dome-shaped colonies. Importantly, these dome-shaped colonies do not retain the characteristics of the naïve state, such as increased clonogenicity, enhanced self-renewal, or expression of alkaline phosphatase and SSEA1.

      In the revised manuscript, we will:

      (1) Revise our wording to avoid any misinterpretation that PAR knockout primed cells maintain a naïve-like identity.

      (2) Explicitly clarify that the observed dome-shaped morphology represents an aberrant primed state rather than a naïve or naïve-like state.

      (3) Discuss more carefully the relationship between the observed phenotype and the absence of typical naïve state features.

      (2) The authors generated several independent knockout clones, but appear to use only one clone for downstream analyses after observing similar morphogenetic phenotypes. Is this sufficient to account for potential clonal heterogeneity? Would the use of pooled clones provide a more robust experimental system?

      We thank the reviewer for raising this important concern regarding clonal heterogeneity. We agree with the reviewer that our current approach using only one representative knockout clone for downstream mechanistic analyses after confirming similar morphogenetic phenotypes across multiple independent clones is not sufficient to fully exclude potential clonal heterogeneity.

      To address this issue, we will perform additional experiments in the revised study. Specifically, we will use another independent knockout clone (ParKO6) to repeat the key mechanistic analyses. The following experiments will be carried out:

      (1) ParKO6 and wild-type ESCs will be subjected to NPT. During the NPT process, cells will be treated with an AKT inhibitor (MK2206), a FAK inhibitor (PF562271), or WT CM. We will observe whether the morphological defects of ParKO6 cells are rescued, and RT-qPCR will be performed to characterize the molecular features of ParKO6 cells under these conditions.

      (2) After treatment with the AKT inhibitor (MK2206), FAK inhibitor (PF562271), or WT CM, immunofluorescence (IF) will be used to detect p-FAK levels in ParKO6 cells.

      (3) Following the same treatments, Western blotting (WB) will be performed to detect FURIN and LEFTY protein levels in ParKO6 cells.

      These additional experiments will allow us to confirm that the observed results are not due to clone-specific artifacts from the originally used clone.

      (3) The rescue experiments using pathway inhibitors are interesting; however, the interpretation again relies primarily on colony morphology. Readers may question whether these experiments truly represent rescue of the naïve-to-primed transition itself without additional transcriptomic or molecular characterization.

      We thank the reviewer for this important comment. We apologize for the lack of clarity in our original manuscript, which may have led to the misunderstanding that our interpretation of the rescue experiments relied solely on colony morphology.

      In fact, we did perform molecular characterization on a subset of cells rescued by pathway inhibitors, and these data are presented in Supplemental Figure 2 D–E. We realize that our description of these results was insufficiently clear, and we failed to properly highlight this molecular evidence in the main text.

      In the revised manuscript, we will revise our wording to clearly state that the rescue effects are supported not only by morphological observations but also by molecular characterization.

      (4) In Figure 4, the manuscript could be strengthened by integrating transcriptomic analyses from pharmacological treatments with the secreted-factor and co-culture datasets.

      We thank the reviewer for this constructive suggestion.

      In our current manuscript (Figure 4), we have indeed performed an integrated transcriptomic analysis comparing pharmacological treatment and secreted-factor treatment, and we demonstrated that both treatments converge on the FAK signaling.

      Regarding the co-culture dataset, we did not include it in the integrated analysis presented in Figure 4. This is because, based on our data in Figure 3, we concluded that the rescue effect observed in co-culture is primarily mediated through secreted factors. Therefore, the secreted-factor transcriptomic data already capture the key signaling pathways responsible for the co-culture rescue effect.

      We will clarify this rationale explicitly in the revised manuscript to avoid any confusion.

      (5) The authors could better clarify the context of Furin downregulation in the knockout cells. Is this a direct consequence of altered transcriptional regulation by the PAR complex, or could it instead represent a secondary consequence of impaired progression through the primed pluripotent transition?

      We thank the reviewer for this important mechanistic question.

      Based on our experimental data, we conclude that Furin downregulation in PAR knockout cells is a direct consequence of altered transcriptional regulation by the PAR complex, rather than a secondary consequence of impaired progression through the primed pluripotent transition. Our evidence is as follows:

      (1) Transcriptomic analysis revealed that PAR knockout leads to a significant reduction in Furin RNA levels.

      (2) Western blot analysis confirmed that PAR knockout also results in a significant reduction of FURIN protein levels.

      (3) Importantly, treatment with an AKT inhibitor (upstream of the proposed pathway) significantly upregulated both Furin RNA and protein levels in PAR knockout cells. In contrast, treatment with a FAK inhibitor or WT CM (downstream) did not significantly alter Furin expression.

      These data collectively indicate that Furin downregulation is directly linked to PAR complex-mediated transcriptional regulation, rather than being an indirect consequence of defective primed state transition. We will clarify this rationale in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      The study demonstrated that Par, but not other polarity genes, Crumbs or Scrib, regulates cell polarity during PSC transition to primed state as well as neural tube formation.

      Strengths:

      The use of KO convinces the role of Par in NPT. Scrib and Crumbs KO data are informative to the field. The conditioned medium experiment is informative. They suggested the potential secreted factors over 50kDa are responsible for maintaining the polarity of NPT in Par KO.

      Weaknesses:

      (1) Most importantly, how Par is important for PSC maintenance and differentiation is not clear. The data provided are dome shape formation, endoderm lineage tendency, and neural tube formation reduction. The manuscript lacks a core message of the physiological importance of Par. Is Par critical of PSC maintenance? Is Par critical for neural system development?

      We thank the reviewer for this critical comment, which helps us better articulate the core message of our study.

      In our manuscript, we have provided clear evidence regarding the role of the PAR complex in pluripotent stem cell (PSC) maintenance and differentiation:

      (1) Regarding PSC maintenance:

      The PAR complex is not critical for PSC maintenance under self-renewing conditions. Specifically, PAR knockout does not significantly affect the expression levels of pluripotency genes (Figure 1 B–C and Supplemental Figure 1 C). Moreover, PAR knockout PSCs can be continuously cultured for at least 30 passages without notable changes in cell morphology or proliferation capacity (Figure 1 D–F). These findings are consistent with previous literature, which demonstrates that the core function of the PAR complex is to establish and maintain cell polarity, rather than directly regulating the transcriptional network of pluripotency genes.

      (2) Regarding PSC differentiation:

      The PAR complex is important for proper differentiation. PAR knockout leads to multiple differentiation defects, including: Failure to establish normal cell morphology during (NPT) (Figure 1 G–K). Impaired formation of proper three-germ-layer structures during embryoid body (EB) and teratoma differentiation (Figure 5 F–G). In particular, the type and quantity of ectodermal tissues are significantly reduced. Consistent with our findings, previous literature has reported that PAR complex deficiency leads to neural developmental defects in mouse embryos, resulting in mid-gestation embryonic lethality.

      (3) Regarding neural system development:

      The PAR complex is critical for neural development. During neural stem cell (NSC) differentiation, PAR knockout cells exhibit a significantly reduced efficiency of Nestin-positive cells and fail to form the classical rosette structures (Supplemental Figure 5 B–C). During neural tube organoid induction, PAR knockout cells show significantly impaired lumen formation and spontaneous elongation efficiency. Moreover, during subsequent maturation, PAR knockout cells fail to differentiate into neurons, leading to a marked reduction in neural tube organoid maturation efficiency (Figure 6 B–E).

      These findings are consistent with previous literature showing that in zebrafish embryonic development, mislocalization of the PAR complex leads to neural tube abnormalities while PAR complex deficiency results in severe hydrocephalus; in mouse embryonic development, PAR complex deficiency causes neural developmental defects leading to embryonic lethality; and disruption of the PAR complex impairs the formation of apical tight junctions in the neuroepithelium and subsequent neuroepithelial tissue polarization, resulting in neural tube closure defects in humans.

      In the revised manuscript, we will incorporate classical literature to discuss the essential roles of the PAR complex in early embryonic development, thereby providing a broader developmental context for our findings.

      (2) Secondly, AKT-FURIN-...... axis still lacks supportive data. Various inhibitors were used to rescue the Par KO. But the link between each component in the axis is missing and rather superficial.

      We thank the reviewer for this critical comment. We acknowledge that the proposed AKT–FURIN–LEFTY–ECM-integrin–FAK signaling axis has certain limitations, particularly that the connection between LEFTY and ECM-integrin lacks direct experimental support. Therefore, in the revised manuscript, we will de-emphasize the role of ECM and integrin and revise the signaling axis to AKT–FURIN–LEFTY–FAK.

      We believe the current data and previous publications support this revised signaling axis well. Accordingly, we have summarized the relevant information as follows. In addition, we plan to perform additional experiments to further support the new signaling axis, which are also included in the following text.

      (1) AKT-FAK

      We found that Par KO cells exhibit defects during NPT, and these defects can be rescued by AKT inhibitor (MK2206), FAK inhibitor (PF562271), and WT CM. Through transcriptomic analysis, we found that both AKT inhibitor and WT CM share similar expression profiles with WT and converge on FAK signaling. Notably, through Western blotting analysis, we found that Par KO led to upregulated p-AKT levels, which were effectively suppressed by MK2206 treatment, but WT CM did not decrease p-AKT levels. In contrast, through immunofluorescence analysis, we found that FAK signaling was hyperphosphorylated in Par knockout primed cells compared to WT primed cells, and MK2206, WT CM, and PF562271 all effectively reduced p-FAK levels. Given that both MK2206 and WT CM attenuated the elevated p-FAK, we propose that all three treatments restore the flat monolayer morphology by regulating FAK signaling homeostasis, with WT CM acting downstream of AKT signaling. The relevant data are presented in Figure 1G-I, Figure 2F, Figure S2C, Figure 3H-I, Figure 4A-D, and Figure S4A.

      (2) AKT-LEFTY

      Through integrated proteomic and transcriptomic analysis, we identified a set of functional proteins. Overexpression screening revealed that LEFTY exhibited the most significant rescue effect in Par KO cells during NPT. Proteomic analysis revealed that the protein levels of LEFTY were significantly higher in WT CM compared to KO CM, suggesting that WT cells modulate FAK signaling via secretion of LEFTY proteins. It is therefore reasonable to infer that MK2206 rescues the defects in Par KO primed cells through upregulation of LEFTY expression. Western blotting analysis confirmed this, showing that MK2206 significantly increased LEFTY protein levels in Par KO primed cells. The relevant data are presented in Figure 4E, Figure 4H and Figure S4C-D.

      (3) LEFTY-FAK

      Proteomic analysis indicated that WT CM treatment supplied extracellular LEFTY to Par KO ESCs, thereby rescuing the phenotypic defects of Par KO primed cells, and significantly reduced p-FAK levels in these cells. Concordantly, LEFTY overexpression also reduced p-FAK in Par KO primed cells. These results are consistent with the reported role of LEFTY in suppressing FAK signaling (Alowayed et al., 2016). The relevant data are presented in Figure 4D, Figure 4F, and Figure S4D-E.

      (4) AKT-FURIN

      LEFTY proprotein requires FURIN-mediated cleavage for secretion and function (Dubois et al., 2001). Through transcriptomic analysis, we found that Par KO downregulated Furin mRNA expression, while MK2206 treatment restored its expression levels. Through Western blotting analysis, we found that MK2206 increased FURIN protein abundance and cleaved LEFTY levels. The relevant data are presented in Figure 4G-H.

      (5) FURIN-LEFTY

      To validate the role of FURIN in LEFTY maturation, we treated WT cells with BOS318, a highly specific and potent inhibitor of FURIN that irreversibly binds to the protease by mimicking its natural substrate (Ivachtchenko et al., 2024). BOS318 induced WT primed cells to adopt a dome-shaped morphology resembling Par KO primed cells, confirming that inhibition of FURIN prevents LEFTY secretion and function, leading to defective primed cell morphology. The relevant data are presented in Figure 4I-J. To further strengthen the role of FURIN in regulating LEFTY, we will treat wild-type cells with BOS318 and examine the expression changes of LEFTY.

      (6) ECM/integrin

      Integrated analysis of both transcriptomic and proteomic data revealed that Par KO leads to significant enrichment of pathways associated with ECM and integrin (Figures 2D, 3F, 3K, and S4B). Notably, both MK2206 and WT CM treatment co-upregulated the ECM-receptor interaction pathway (Figure 4C). The FAK signaling pathway serves as a central node that integrates upstream inputs from both PKC and AKT pathways while transducing extracellular cues derived from ECM-integrin interactions into intracellular signaling cascades (Sakthivel et al., 2025). We therefore propose that secreted LEFTY acts as an extracellular signal that activates specific ECM receptors and modulates integrin complexes, thereby regulating FAK phosphorylation and maintaining normal cell adhesion and morphology. However, this speculation still lacks direct experimental evidence. We will endeavor to perform additional experiments to support this proposed connection in the future. Nevertheless, we have decided to de-emphasize the role of ECM and integrin in the AKT–FURIN–LEFTY–FAK signaling axis in the current manuscript.

      References

      Alowayed, N., Salker, M. S., Zeng, N., Singh, Y., & Lang, F. (2016). LEFTY2 Controls Migration of Human Endometrial Cancer Cells via Focal Adhesion Kinase Activity (FAK) and miRNA-200a. Cellular Physiology and Biochemistry, 39(3), 815-826. https://doi.org/10.1159/000447792

      Dubois, C. M., Blanchette, F., Laprise, M.-H., Leduc, R., Grondin, F., & Seidah, N. G. (2001). Evidence that Furin Is an Authentic Transforming Growth Factor-β1-Converting Enzyme. The American Journal of Pathology, 158(1), 305-316. https://doi.org/10.1016/s0002-9440(10)63970-3

      Ivachtchenko, A. V., Khvat, A. V., & Shkil, D. O. (2024). Development and Prospects of Furin Inhibitors for Therapeutic Applications. International Journal of Molecular Sciences, 25(17). https://doi.org/10.3390/ijms25179199

      Sakthivel, K., Kotowska, A., Fan, Z., Portner, E. J., Merry, C., Nordenfelt, P., Simonsen, A. C., Wright, A. J., & Swaminathan, V. S. (2025). Integrin‐Piezo1 Axis Drives ECM Remodeling and Invasion of 3D Breast Epithelium. Advanced Science. https://doi.org/10.1002/advs.202509932

    1. Author response:

      The following is the authors’ response to the original reviews.

      In response to the reviewers’ comments, we have made revisions to the manuscript. Specifically, we have:

      (1) Increased the sample size in the whole-brain imaging and demixed principal component analysis (dPCA) analyses presented in Figures 1 and 3, strengthening the statistical support for our conclusions;

      (2) Revised the presentation of Figure 3B to clarify that the displayed dPC1 traces were scaled for visualization purposes only (dPC1 / max(dPC1)), rather than normalized for quantitative comparison across animals;

      (3) Expanded the main text and supplementary figures to provide more intuitive explanations and geometric illustrations of dPCA and hyperbolic space analysis, and clarified the interpretation of correlation matrices and principal-angle analyses to improve readability;

      (4) Substantially expanded the sections on Bayesian multidimensional scaling and hyperbolic embedding, including additional methodological details and validation analyses to strengthen the computational framework and its interpretation;

      (5) Expanded the Discussion to incorporate recent studies and discuss potential mechanisms underlying DRN 5-HT-mediated motor suppression.

      We believe that these revisions have substantially strengthened the manuscript and addressed the major concerns raised during peer review.

      Reviewer #1 (Public review):

      The wide-ranging serotonergic projections emerging from the Dorsal Raphe nucleus (DRN) are suggestive of a central role in regulating brain-wide activity and behavioural states. DRN activity has been associated with diverse functions, ranging from mood, motivation and pain regulation to sleep and cognitive flexibility. Its far-reaching connectivity made it challenging to assess the brain-wide effect of its activation, especially during behaviour.

      The present study by Qi et al. addresses these challenges by combining state-of-the-art tracking microscopy with the whole-brain accessibility of the larval zebrafish model. To investigate the effect of DRN activation, the authors leveraged the Tg(tph2:ChrimsonR) line to optogenetically activate tph2-positive neurons in the DRN, while monitoring changes in brain-wide activity, locomotion and auditory-stimuli evoked responses.

      Optogenetic activation had a suppressing effect on locomotion, which the authors distinguished from inducing sleep by the maintenance of posture and its sleep disturbing effect of nighttime stimulations. Further, the authors report a distinct effect of DRN activation on motor-related, but not auditoryrelated neuronal subspaces, identified by demixed principal component analysis.

      In addition, rather than affecting all motor-correlated neurons similarly, tph2+ DRN-mediated suppression focused on neurons encoding high-amplitude or turning motion.

      In summary, the work of Qi et al. provides solid evidence for a predominant role of the DRN in wake-state motor suppression by aptly combining the vast data-acquisition possibilities of the larval zebrafish model with computational methods to extract relevant information.

      The brain-wide scope of the analysis is a key strength, reducing bias, confirming the involvement of known motor and auditory regions, and providing a valuable dataset for future analyses.

      While the results well support the conclusion of the authors, certain biological and technical aspects demand discussion.

      We thank you for the positive and thoughtful evaluation of our work. We also appreciate your constructive comments on the biological and technical aspects of the study. We have carefully considered these concerns and addressed them point-by-point below, with corresponding revisions to the manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) Further samples required:

      Figure 1D relies on n=3 with lots of variability; the author should add more Ns to illustrate their point (typically 10-15 fish used per study to show reliability across fish).

      Figure 3 also relies only on 5 fish in each condition; the authors should increase to 10-15 to show variability.

      Thank you for this valuable suggestion. To address this concern, we have increased the sample size in the revised manuscript. Specifically, the number of animals in Figure 1D has been increased from n = 3 to n = 5, and additional statistical analyses have been included to strengthen the quantitative support for our conclusions. Note that the error bars are plotted as standard deviation (SD), which may make the variability appear larger. In Figure 3, the number of animals was also increased from n = 5 to n = 8.

      In addition, our findings are consistent with previous work showing a strong association between elevated dorsal raphe nucleus (DRN) activity and reduced locomotion in zebrafish [1, 2, 3]. Importantly, across animals, the variance explained by the dPCA components and the rapid modulation of whole-brain state remain highly consistent, supporting the robustness and reproducibility of our observations.

      Given this increased sample size together with consistency across animals and convergence with prior studies, we believe the current dataset provides sufficient statistical and biological support for our conclusions.

      (2) Further steps to be added to the analysis to fully support the claim:

      It appears that the individual brains are registered and individually clustered into areas by combining highly-correlated nearby neurons.

      dPCA is then computed for individual brains. Evidence for our interpretation of individual dPCA spaces:

      (1) Figure 3A depicts separate dPCs for different fish.

      (2) Line 488–489 describes normalization of the value range of dPCs to compare across fish, which implies separate dPCs.

      While the authors normalize the projections onto the principal components, the dPCA spaces remain individual, as does the meaning of their components. It is thus questionable how to conclude from data across fish in a rigorous manner.

      Instead, we recommend that the authors build voxels for each individual’s brain and calculate dPCA across all brains, not individual ones, so that components could become truly comparable across the brains of given individuals.

      We thank the reviewer for this important comment. We would like to clarify that our analysis does not aim to construct a shared dPCA space across animals or to quantitatively compare dPC scores between individuals. In this analysis, dPCA was performed separately for each fish to capture the dominant low-dimensional population dynamics within each individual brain.

      The purpose of Figure 2 is to demonstrate that DRN activation induces a rapid and robust transition in whole-brain activity, rather than to define a common population subspace across animals.

      We also attempted to register and pool data across animals for a joint analysis, as suggested by the reviewer. However, our dataset includes zebrafish at slightly different developmental stages (6–12 dpf). Although the behavioral effects of DRN activation (including motor suppression and global brain-state modulation) were robust across this age range, developmental differences introduced substantial anatomical variability in brain size and morphology, which reduced registration accuracy and made voxel-wise correspondence across animals unreliable.

      We realize that our previous description of “normalization” may have caused confusion. To clarify, the dPC1 traces shown in Figure 2 were only scaled for visualization by dividing each fish’s projection by its maximum value (dPC1 / max(dPC1)), so that trajectories from different fish could be displayed on the same axis. This scaling does not alter the underlying dPCA space, does not constitute normalization for cross-animal comparison, and was not used for any quantitative analysis.

      Importantly, despite being computed independently for each fish, we observed a consistent temporal pattern across animals: DRN activation was reliably accompanied by a rapid transition captured by dPC1 in each individual fish. We have revised the Methods and corresponding text in the manuscript to make this distinction explicit and avoid ambiguity.

      Reviewer #2 (Public review):

      Summary:

      The authors examine the effects of activating the dorsal raphe nucleus serotonergic system using a combination of calcium imaging and optogenetics in freely moving larval zebrafish. Their findings show that optogenetic stimulation induces a state of behavioral quiescence.

      They further investigate whether this state corresponds to sleep or reduced motor activity. Analyses of posture and sleep-related paradigms indicate that serotonergic activation primarily suppresses motor output rather than promoting sleep. Notably, this suppression appears to be bout type-dependent, with stronger effects on neurons associated with larger tail amplitudes and turning angles.

      In addition, auditory stimulation experiments reveal no significant impact of serotonin on sound encoding.

      We thank the reviewer for the careful and thoughtful summary of our work.

      Strengths:

      The study combines advanced experimental techniques with state-of-the-art analytical methods, enabling precise and compelling insights into the role of serotonergic modulation. The experiments and analyses are well aligned with the questions being addressed, and the results appear robust and reliable.

      Moreover, the implementation of experiments that combine calcium imaging and optogenetics in freely moving animals is technically challenging and appears well justified in the context of the research questions.

      We thank you for the positive assessment of our work and for recognizing the technical and analytical strengths of our experimental approach.

      We address the reviewer’s specific comments in detail below.

      Weaknesses:

      While the analytical techniques employed are sophisticated and appear to be appropriately applied, their presentation makes the manuscript difficult to follow. Although the explanations are provided in the Methods section, including more guidance in the main text, such as how to interpret each analytical approach and what outcomes would be expected under different scenarios, would help readers who are less familiar with these techniques.

      Providing this context would better guide the reader in navigating the figures, broaden the accessibility of the work, and ultimately increase its impact.

      We thank you for this important suggestion. To improve clarity and accessibility, we have revised the main text to provide more intuitive explanations of both demixed principal component analysis (dPCA) and hyperbolic space analysis, with additional emphasis on how to interpret their outputs and what different outcomes imply biologically.

      Additionally, we have included new supplementary figures (Figure S2 and Figure S6) with geometric illustrations and simplified examples to provide a more visual and conceptual understanding of these methods. We hope these revisions make the analytical framework easier to follow and improve the accessibility and impact of the manuscript.

      While the authors discuss different quiescent states mediated by serotonin reported in previous studies, their interpretation is limited to stating that “a common feature shared by these distinct behavioral states is a pronounced reduction in movement,” and consequently proposing that activation of dorsal raphe nucleus is not sufficient to specify a particular behavioral state, but rather plays a primary role in driving motor suppression.

      In my view, a more thorough attempt to determine whether the observed state corresponds to any of the previously described forms of quiescence, or represents a subset or variant of them, would strengthen the manuscript. This would help better integrate the findings with the existing literature.

      For example, given that the authors have access to whole-brain activity data, it would be valuable to examine and discuss whether there are shared patterns of activation with previously reported quiescent states.

      Thank you for the insightful suggestion. To address this, we compared our whole-brain activity patterns with key neural signatures reported in previously characterized zebrafish quiescent states.

      A recent study reported that exposure to conspecific alarm substance (CAS) induces a quiescent but vigilant state associated with elevated DRN 5-HT activity and low-frequency synchronized forebrain activity [3]. In our dataset, although DRN 5-HT activation similarly induced robust locomotor suppression, we did not detect comparable low-frequency synchronized forebrain dynamics during the stimulation period. These results suggest that while DRN 5-HT activation is sufficient to induce motor suppression, it does not recapitulate the full neural signature of CAS-induced vigilant quiescence. We have incorporated this comparison and its interpretation into the Discussion section of the revised manuscript.

      Following the termination of optogenetic stimulation, we observed a gradual recovery of locomotory speed, consistent with the behavior in an earlier study [3], although our recovery was much faster. Interestingly, whole brain imaging also revealed a transient increase in forebrain activity. This elevated forebrain activity gradually returned to baseline as locomotor activity recovered. In accordance with the reviewer’s suggestion, we propose that these forebrain dynamics represent a common motif that facilitates the transition out of the DRN-induced quiescent state (Author response image 1.).

      The manuscript largely avoids discussing the mechanisms underlying the observed motor suppression. For instance, is this effect driven directly by serotonin release onto target neurons? Is it mediated by glial activity, as suggested in other studies? Are additional neuromodulatory systems being recruited?

      While addressing these questions may require substantial further work, potentially beyond the scope of the present study, the availability of whole-brain data provides an opportunity to at least explore or

      Author response image 1.

      Forebrain activity increases following termination of DRN optogenetic stimulation. (A) Following the termination of optogenetic stimulation of DRN 5-HT neurons, locomotor speed in Tg(tph2:ChrimsonR) zebrafish gradually recovered and returned to control levels. (B) Neural activity in forebrain regions showed a transient increase immediately after stimulation offset and gradually returned to baseline as locomotor activity recovered. discuss these possibilities. In particular, it would be interesting to examine the recruitment of regions not directly stimulated but known to be associated with other neuromodulatory systems or promoting glial activation (e.g., the locus coeruleus).

      We thank you for this important suggestion. In the revised Discussion, we now frame our findings in relation to several candidate mechanisms.

      Our results are most consistent with a direct neuromodulatory action of serotonin on downstream motor-related circuits. This is supported by the known projection patterns of DRN 5-HT neurons [4], which target midbrain and hindbrain regions involved in motor control, as well as by prior serotonin imaging studies showing elevated 5-HT levels in hindbrain regions during low-motor states, where inhibitory HTR1-family receptors are enriched [5]. In addition, recent voltage imaging studies have shown that DRN serotonergic neurons are embedded within a broader motor-state-dependent circuit, in which they are dynamically regulated by local GABAergic inputs [6]. We have incorporated a discussion of these potential mechanisms into the revised Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) Lines 91-97 page 2.

      “dPCA separates neural population activity into components tied to specific experimental variables, allowing us to isolate DRN-dependent changes (Methods). Components associated with DRN activation explained significantly more variance in Tg(tph2:ChrimsonR) zebrafish than in controls (Fig. 3A), indicating a strong serotonergic impact on brain-wide neural activity. The small stimulation-related variance in controls likely reflected visual responses to laser.”

      Directly stimulated neurons are not included, as stated in the Methods, but I think it would be better to mention this explicitly in the main text.

      We thank you for this helpful suggestion. We agree that explicitly stating this point in the main text improves clarity. In our analysis, neurons directly stimulated by the laser were excluded (as described in the Methods) to ensure that the identified components reflect whole brain responses rather than direct optogenetic activation. We have now added a clarifying sentence in the Results section to make this explicit.

      (2) Lines 113 - 115 page 3.

      “To examine how DRN 5-HT neuron activation affects sensorimotor processing (Fig. 4C), we next recorded whole-brain neural activity in head-fixed, tail-free larvae embedded in agarose to capture transient calcium signals with minimal motion artifacts.”

      Lines 117-119 page 3.

      “Because head-fixed larvae rarely enter natural sleep, we applied 1 mM mepyramine, a sleep-promoting antihistamine, to induce a sleep-like state (41), which markedly changed auditory responses (Fig. 4E, Fig. S2C)”

      Why not perform these experiments in freely moving fish instead? To what extent do movements in freely moving animals affect segmentation? Is it actually problematic to apply dPCA in that case? You used it in the previous section.

      We thank the reviewer for raising this important point. In principle, freely moving preparations would provide a more natural behavioral context. However, reliable application of dPCA requires stable neuron identification and accurate trial alignment across time, both of which are substantially compromised in freely moving larvae due to motion-induced imaging noise and segmentation errors.

      In our hands, whole-brain calcium imaging in freely moving fish introduces significant variability in segmentation and signal extraction, which in turn leads to unstable and noisy low-dimensional decompositions, preventing robust estimation of task-related components. By contrast, the head-fixed preparation enables consistent neuron tracking and precise alignment to sensory stimuli, which are critical for dPCA.

      We have now clarified in the manuscript that all dPCA analyses were performed on head-fixed animals.

      (3) Line 117 page 3.

      Why do you use cosine similarity? Are the results different when using other metrics?

      I can see the matrix, but what exactly are you looking for in it to support the claim ”DRN activation preserved the structure of the auditory population code”? I think explaining some of these concepts more clearly, or at least providing expectations or interpretations for the different metrics and analyses, would make the manuscript easier to follow.

      We thank you for this question. Cosine similarity is widely used to quantify similarity between population activity patterns because it captures relative activity across neurons while ignoring overall gain.

      In our analysis, each trial is a population activity vector, and the cosine similarity matrix encodes pairwise relationships between these vectors. We assess preservation of the auditory population code by testing whether this similarity structure (i.e., the geometry of population responses) remains consistent across conditions. We have expanded the text to clarify how these matrices are constructed and interpreted.

      In addition, we computed alternative similarity measures based on Pearson correlation, which is equivalent to the cosine similarity of two vectors after they have been centered (subtracting the mean of each vector) (Author response image 2A). We further quantified pairwise trial distances using the Euclidean chord distance on the unit hypersphere, defined as

      D<sub>ij</sub> = √2(1−C<sub>ij</sub>), where C<sub>ij</sub> is Pearson correlation; smaller distances indicate higher similarity (Author response image 2B). Both alternative measures yielded qualitatively consistent results, showing that DRN 5-HT neuron activation preserves the similarity structure across trials.

      (4) Figure 4D.

      If “significant alignment between DRN activation and motor-related neural subspaces, with the sound related subspace being nearly orthogonal” is correct, shouldn’t there be some visible overlap between blue and red, and little to no overlap with yellow? This is not easy to see. Perhaps plotting all three in a single panel would help.

      We thank you for this helpful suggestion. We would like to clarify that the “alignment” we refer to is defined in terms of the angle between neural subspaces, rather than the spatial overlap of neurons. In other words, significant alignment indicates that the corresponding population activity patterns occupy similar directions in a high-dimensional activity space.

      As a result, even statistically significant aligned subspaces (see further exposition below) do not necessarily involve overlapping sets of neurons with large PC weights. This distinction is important because subspace geometry is defined at the population level and cannot be directly inferred from spatial overlap in low-dimensional visualizations. In addition, the visualization shown in Fig. 4D highlights only brain regions containing neurons with relatively high weights for illustrative purposes.

      We also note that the current visualization is based on a maximum intensity projection of a 3D volume, which can create the appearance of overlap in two dimensions even when the underlying neurons are spatially segregated in three dimensions. To provide a clearer spatial reference, we have re-plotted the three subspaces in a three-dimensional representation.

      (5) Figure 4F.

      Do the arrows represent the values for each combination? This is not clear to me. Perhaps it could be clarified in the paragraph. Most of the values, including those being compared, are around 87 plus minus 2 degrees, i.e., mostly orthogonal. Does this imply no overlap between patterns (again, this is hard to see in Figure 4D)? The values are different from the null model but still close to orthogonal. The phrase “significant alignment between DRN activation and motor-related neural subspaces” could be interpreted as strong alignment, but the values do not seem to support that, do they?

      Author response image 2.

      Alternative similarity measures reveal preserved trial-to-trial similarity structure. (A) Trial-by-trial similarity matrix quantified using Pearson correlation. Higher correlation indicates greater similarity between trials (B) Pairwise trial distances quantified using the Euclidean chord distance on the unit hypersphere (D<sub>ij</sub> = √2(1−C<sub>ij</sub>)), where smaller distances indicate greater similarity between trials.

      Author response image 3.

      Three-dimensional visualization of DRN activation-, motor-, and sound-related subspaces. Threedimensional rendering of the high-weight neurons in the DRN 5-HT activation, motor-related, and sound-related subspaces. Colors are consistent with Figure 4D.

      We thank the reviewer for this important clarification.

      We agree that the phrase “alignment” could be interpreted as implying strong spatial overlap in the anatomical space, which is not what we intend to convey. In our analysis, “alignment” refers to a statistically significant deviation from a null distribution.

      In high-dimensional spaces, random vectors are expected to be nearly orthogonal, with angles tightly concentrated around 90°. To demonstrate this phenomenon, we conducted simulations using random vectors over a range of dimensionalities (100–10,000 dimensions) and observed that the expected angle distribution over 1000 trials becomes progressively more concentrated around 90° as the dimensionality increases (Author response image 4). Therefore, even modest deviations from 90° reflect a systematic bias and indicate structured overlap beyond chance. So, “significantly aligned” means the motor–DRN angle is significantly less than the random baseline, and “significantly orthogonal” for sound–DRN means the angle is significantly closer to 90° than the random baseline. We will revise the text to clarify this point and avoid potential misinterpretation.

      Regarding Figure 4D, we agree that the meaning of the arrows was not sufficiently clear. The arrows represent the mean angle, computed across all fish, between the DRN 5-HT activation subspace and the motor-related subspace (left), and between the DRN 5-HT activation subspace and the sound-related subspace (right). We will update the figure legend to explicitly define these elements.

      Author response image 4.

      Random vectors become increasingly orthogonal in high-dimensional spaces. Simulated distributions of pairwise angles between random vectors across different dimensionalities (100–10,000 dimensions; 1000 repetitions per dimensionality). As dimensionality increases, the angle distribution becomes increasingly concentrated around 90°.

      (6) Lines 125 - 126 page 5.

      “After detecting bouts, we computed each bout’s direction and amplitude and classified them into 12 types.”

      It would be interesting to see how the distribution of bouts looks in the direction-amplitude space, in order to better visualize the 12 bout types (perhaps using different colors). It might also be useful to include examples of the 12 bout types in the supplementary material.

      We thank you for this helpful suggestion. To better visualize the distribution of bouts and the definition of the 12 bout types, we have added a new supplementary figure showing the distribution of all bouts in the direction–amplitude space, with each bout color-coded according to its assigned category, consistent with the scheme used in the main text.

      We further quantified the frequency of each bout type across the dataset, which comprises 1,493 bouts from 7 animals. Among these, 4 animals exhibited all 12 bout types and were therefore included in subsequent regression analyses that require complete coverage of all categories.

      In addition, we have included examples of representative bout types in the supplementary material. These additions improve the clarity and interpretability of the bout classification scheme.

      (7) Lines 131 - 133 page 5.

      “Some neurons exhibited activity related to all bout types with similar amplitudes, yielding low coefficient variability, whereas others responded selectively to specific bout types - typically those with larger tail amplitudes and turning angles - exhibiting higher variability in regression coefficients (Fig. 5B).”

      I would appreciate some quantification of “typically.”

      We thank you for this suggestion. Fig. 5B (bottom) shows a neuron with large variability in regression coefficients across bout types, quantified by the coefficient of variation (CV). Bout types with large amplitudes and turning angles (e.g., type 12) have larger regression coefficients than others. We will remove “typically” from the text.

      (8) Lines 546 - 547 page 15.

      “Fish whose baseline tail movements were insufficient to cover all 12 bout types were excluded from further analysis.”

      It would be useful to report the number or proportion of animals that did not exhibit all 12 bout types. Which types of bouts are less frequently observed?

      Thank you for this helpful suggestion. In the full dataset (n = 7 fish), 4 animals exhibited all 12 bout types. We have now added a supplementary figure showing the occurrence probability of each bout type across all animals.

      (9) Line 147 page 5.

      Honestly, the Bayesian multi-dimensional scaling is difficult to follow, and it is not clear what new insight it provides. I assume that ”hyperbolic geometry indicates complex hierarchical organization” is the main point, but its meaning in this context is not sufficiently explained. This paragraph would benefit from being rewritten for clarity or potentially removed if it does not contribute essential information.

      We appreciate your insightful comments. In response, we have substantially expanded the section on Bayesian multidimensional scaling. First, we now provide an intuitive exposition (see Figure S6) of hyperbolic geometry and multidimensional scaling, clarifying why this framework constitutes a powerful approach for uncovering the geometric and functional organization of neuronal populations. Second, we show that multidimensional scaling in a curved hyperbolic space more accurately captures the correlation structure among neurons than embeddings in a flat Euclidean space. Third, and most notably, we find that the inferred curvature of the hyperbolic embedding space tightly scales with the degree of quiescence: fish in which dorsal raphe nucleus (DRN) stimulation nearly abolished locomotor activity exhibit the largest curvatures (new Figure 5F). Collectively, these computational analysis indicate that the curvature of the embedding space serves as a quantitative signature of the quiescent state.

      References

      (1) J. C. Marques, M. Li, D. Schaak, D. N. Robson, J. M. Li, Internal state dynamics shape brainwide activity and foraging behaviour. Nature 577, 239–243 (2020).

      (2) V. Choudhary, C. R. Heller, S. Aimon, L. de Sardenberg Schmid, D. N. Robson, J. M. Li, Neural and behavioral organization of rapid eye movement sleep in zebrafish. bioRxiv pp. 2023–08 (2023).

      (3) Y. Zhao, C.-X. Huang, Y. Gu, Y. Zhao, W. Ren, Y. Wang, J. Chen, N. N. Guan, J. Song, Serotonergic modulation of vigilance states in zebrafish and mice. Nature Communications 15, 2596 (2024).

      (4) Z. Song, C.-X. Huang, H. Zhang, C. Ye, N. Guan, J. Song, Integrated single-cell atlases unveil the operation principles of whole-brain 5-ht neuronal subsystems. Science Advances 11, eadv8128 (2025).

      (5) R. Haruvi, R. Barbara, I. Shainer, A. Rosenberg, L. Moshe, D. Malamud, J. Toledano, D. Braun, H. Baier, T. Kawashima, Global and compartmentalized serotonergic control of sensorimotor integration underlying motor adaptation. BioRxiv pp. 2024–09 (2024).

      (6) T. Kawashima, Z. Wei, R. Haruvi, I. Shainer, S. Narayan, H. Baier, M. B. Ahrens, Voltage imaging reveals circuit computations in the raphe underlying serotonin-mediated motor vigor learning. Neuron (2025).

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      The authors addressed all my concerns.

      We sincerely appreciate your recognition of our efforts to address the reviewers' suggestions and improve the manuscript.

      Reviewer #2 (Public review):

      (1) All the treatment arms (A-control, MgIG-25 mg/kg, MgIG-50 mg/kg) showed significant body weight loss compared to the untreated controls (Supplemental Figure 1A), but the body weight significantly increased in the treatment arms (A-control and MgIG-50 mg/kg) compared to the untreated controls (Figure 1E). Why?

      We appreciate the reviewer’s careful observation regarding the apparent discrepancy between Supplemental Figure 1A and Figure 1E. We apologize for any confusion caused by the presentation of these data.

      We would like to clarify that Supplemental Figure 1A and Figure 1E represent two different parameters. Supplemental Figure 1A shows absolute body weight, whereas Figure 1E presents the liver-to-body weight ratio (LW/BW), as indicated in the revised figure legend.

      In the NIAAA alcohol-fed model, chronic ethanol exposure typically results in reduced body weight gain or relative body weight loss compared with normal diet-fed control mice, which is consistent with the findings shown in Supplemental Figure 1A. In the preliminary dose-finding study, all alcohol-fed groups (EtOH groups, MgIG 25 mg/kg, and MgIG 50 mg/kg) exhibited lower absolute body weight compared with the untreated control group, which is a common feature of ethanol-induced liver injury models.

      By contrast, Figure 1E reflects changes in the LW/BW ratio rather than total body weight. Ethanol feeding induces hepatomegaly and hepatic steatosis, thereby increasing the LW/BW ratio. Although the LW/BW ratio in the MgIG-treated group remained higher than that in the untreated control group, MgIG treatment significantly reduced the ethanol-induced increase in LW/BW ratio compared with the EtOH group, consistent with its hepatoprotective effects and reduced hepatic lipid accumulation. We hope this clarification could well answer this concern. Thank you very much!

      (2) Mice with MgIG (25 mg/kg) showed the lowest body weight, compared to either A-control or MgIG (50 mg/kg) treatment. According to the authors' explanation, the MgIG (25 mg/kg) caused bodyweight loss are attributed to inter-individual variability, differences in metabolic adaptation, or sample size-related variation. Did these differences happen in MgIG (25 mg/kg) only? or in all other groups? The mouse group assignment should be randomized; however, a large variation in bodyweight was seen in MgIG (25 mg/kg) group. It is not convincing for the author to select MgIG (50 mg/kg) group for subsequent animal experiments, because of a large variation in MgIG (25 mg/kg) group, and because that MgIG (50 mg/kg) group demonstrated more consistent and stable improvements across multiple parameters. The author should reanalyze and compare all the raw data between MgIG (50 mg/kg) group and MgIG (25 mg/kg) group, and address the issues being pointed out and justify rationale for the animal group assignment.

      We appreciate the reviewer’s careful evaluation regarding the variability observed in the MgIG (25 mg/kg) group and the rationale for dose selection.

      Supplemental Figure 1A presents data from our preliminary dose-finding study (n=5 per group, independent cohort), in which all alcohol-fed groups showed expected body weight loss relative to the normal-diet control, as is typical in the NIAAA model. The 25 mg/kg group exhibited numerically greater variability (likely due to inter-individual metabolic differences and small sample size), but no statistically significant difference was observed among the three alcohol-fed groups (A-control, 25 mg/kg, and 50 mg/kg) in final body weight (one-way ANOVA with post-hoc test).

      Mice were randomized by initial body weight and age prior to diet feeding. To address the reviewer’s concern, we have now included Supplementary Table Body weight-raw data with individual animal body weight data (raw values, mean ± SD) for both the dose-finding and main experiments, together with statistical comparisons. We selected 50 mg/kg for all subsequent experiments because it provided more consistent and statistically significant improvements across multiple key parameters (ALT, AST, TG, TC, NAS score, Oil Red O staining, and LW/BW ratio) compared with 25 mg/kg. The 25 mg/kg group showed greater variability in several indices, which is why it was not chosen for mechanistic studies.

      To further clarify this point, we have added detailed descriptions of the randomization procedure and dose-selection rationale in the revised Methods section. Please refer to Page 5, line 106-108 and Page 10, line 276-277. In addition, we will provide the original data on mouse body weight changes, together with the corresponding statistical analyses, in the supplementary materials to further enhance transparency and facilitate reference.

      (3) The author's response did not answer my question. If the authors believe it could be experimental constraints associated with the MgIG formulation, then it is questionable for this MgIG formulation used in all other associated experiments. The experiments, at least those the MgIG formulation associated experiments, need to be repeated.

      We sincerely appreciate the reviewer’s concern regarding the potential impact of the MgIG formulation on the reliability of the associated experiments.

      As clarified in our previous response, the commercially available MgIG preparation used in this study is a clinically approved injectable formulation (5 mg/mL). During the preliminary in vitro dose-ranging experiments, achieving the highest testing concentration (1.0 mg/mL) required the addition of a relatively larger volume of stock solution, which slightly reduced the effective culture medium volume and may have contributed to minor effects on cell status. Consistently, CCK-8 and LDH assays showed a slight reduction in cell viability only at the highest concentration tested.

      Importantly, this phenomenon was observed exclusively in the 1.0 mg/mL group. All subsequent functional and mechanistic experiments were performed using the optimized non-toxic concentration (0.25 mg/mL), at which MgIG consistently and significantly improved IL-6, Acc1, Scd1, and other relevant parameters in a dose-dependent manner (P < 0.05), without detectable cytotoxicity.

      In addition, vehicle controls with volume-matched conditions were included for the high-concentration (1 mg/mL) condition to exclude potential confounding effects caused by solvent volume differences. The protective effects observed at 0.25 mg/mL were highly reproducible and were further supported by multiple independent lines of evidence, including RNA-seq analysis, enzyme activity assays, and knockdown/overexpression experiments, all of which demonstrated consistent mechanistic trends.

      Therefore, we believe that the current data obtained using the optimized concentration remain reliable and interpretable, and that the formulation-related issue observed at the highest concentration does not affect the validity of the main conclusions. Nevertheless, to further address the reviewer’s concern, we are willing to provide additional replicate data for the 1.0 mg/mL cell viability/toxicity assays, as well as repeat qPCR analyses under volume-matched vehicle control conditions in the Supplementary File . Please refer to Supplementary Figure 2E.

      (4) The author explained the relative expression was normalized to GAPDH (fold change), but they did not answer my question. My question is for Figure 5B. in Figure 5B (left, Hsd11b1-KD), scramble control showed over 100 (unit), however, in Figure 5B (right, Hsd11b1-OE), scramble control showed only 0.5-1 (unit). The data seemed that authors used same scramble control for both KD and OE? If yes, they should provide more details of the KD and OE experiments and explain why this happened. If they used plasmid for OE control, they also need to clarify it. In addition, qPCR is not a good assay to show the success of KD or OE, Western blotting should be done as convincing data to show the success of KD or OE.

      We apologize that our previous response did not fully clarify the details of Figure 5B. The left panel of Figure 5B shows the Hsd11b1 knockdown experiment using Hsd11b1 siRNA with scramble siRNA as the corresponding control, whereas the right panel shows the Idi1 overexpression experiment using the Idi1 expression plasmid with empty vector as the corresponding control. These are two independent experiments with separate control groups, rather than a shared scramble control. We recognize that the labeling and figure presentation may have caused confusion, we have revised the legend for Figures 3B, 3C and Figures 5B, 5C as suggested.

      For both experiments, relative mRNA expression levels were normalized to GAPDH and analyzed independently using the 2<sup>^−ΔΔCt</sup> method relative to their respective controls. Therefore, the numerical values shown in the two panels are not directly comparable. The apparent difference in baseline expression levels reflects independent normalization and the intrinsic expression characteristics of different genes, rather than the use of the same control group or any data inconsistency.

      We have confirmed that transfection efficiencies were consistent with expectations and did not significantly affect cell viability.

      We also agree with the reviewer that protein-level validation would provide stronger evidence for the success of knockdown and overexpression. Accordingly, we have performed Western blot analyses for Hsd11b1 knockdown and Idi1 overexpression and will include these data in the revised manuscript to complement the qPCR results (Please refer to revised Supplementary Figure 3C and 4D).

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      Tropical single-island endemic bird populations are particularly vulnerable to climate change. The authors investigate genetic evidence of how such species dealt with climate changes in the past as a possible predictor for how they will respond to change in the future, which could provide an important example for the fields of conservation genetics and island biogeography. The authors' integration of genomics and habitat modeling is commendable, but we find that the support for their conclusions is incomplete: at times, the results presented appear to contradict each other, the authors do not fully account for key variables, and the limited taxonomic scope may cause problematic biases for the conclusion.

      We thank the editors for supporting the premise of this study and highlighting the importance of the study approach. Based on the lacuna identified by the editors and the reviewers, we have modified the manuscript and details of the same are given below. We believe that these revisions have now substantially improved the flow and scope of the manuscript and have addressed the concerns raised by the reviewers.

      Reviewer #1 (Public review):

      Summary:

      The authors combine PSMC and habitat modeling to try to connect habitat change during the Last Glacial Period to changes in Ne.

      Strengths:

      Observing how tropical single-island endemic bird species responded to habitat change in the past may help inform conservation interventions for these particularly vulnerable species. The combination of genomics and habitat modeling is a good idea - this sort of interdisciplinary thinking is what is needed to tackle these complex questions. Additionally, the use of PSMC makes it possible to perform this analysis on poorly-studied species with only a single genome available.

      Room for Improvement:

      Why coalescent Ne is a better predictor of extinction risk than current genomic diversity, or current Ne, isn't explicitly explained. PSMC in particular has many caveats, and some are not acknowledged or adequately addressed by the authors. For example, the authors note that population structure is a confounding factor with PSMC, but that it is not a problem in this instance. They do not provide compelling evidence for why this would be the case, they simply state that the species studied are all single-island endemics. However, single-island endemic species are not necessarily panmictic; this is even less likely to be true for species studied here that inhabit a large geographic area (ie, Australian species). Differing PSMC parameters may also impact results: the differences between passerines and non-passerines were one of their main results, but they do not provide any analysis to show that this difference was not driven by the different mutation rates used for the two groups.

      Parameters for many steps are not described, and choices that are described (such as the PSMC parameters) are not always fully explained. It is unclear why all data was mapped to the autosomes rather than removing reads that map to the sex chromosomes first. Using all the data, the reads belonging to the sex chromosomes could potentially map to other areas of the genome. It does not seem like a mapping quality filter was used, so these potential spurious alignments would not have been removed prior to analysis.

      There are points where the results are described in ways that appear to potentially differ from the supplementary figures. The authors state that even for species where PSMC results differed between models, "trends of Ne increase or decrease from the LIG to LGM were robust across all three PSMC models considered." The figures in the supplement for Pachycephala philippinensis, Rhynochetos jubatus, and Zosterops hypoxanthus appear to potentially contradict this statement, but it is difficult to tell, as the time period observed is not clearly marked on the graphs. How this robustness of trends was determined is not explained, leaving the precision of the analysis unclear.

      Table 1 also includes some information that contradicts what is in the Supplementary Tables, leading to a lack of clarity. Centropus unirufus, Chaetorhynchus papuensis, and Cnemophilus loriae are not included in Supplementary Table 4. Table 1 says Eulacestoma nigropectus, Paradisaea rubra, and Parotia lawesii did not undergo PSMC analysis, but Supplementary Table 4 says PSMC and modeling trends matched for these species. Table 1 says Rhagologus leucostigma underwent both PSMC and climate modeling, but Supplementary Table 4 says "NA" as if it was missing one of these analyses.

      Additionally, some of the results appear to contradict each other. For example, they show that there is no impact of habitat change in larger-bodied species, but also that larger-bodied species saw a decrease in Ne during the LGP. In another example, they state that when a species saw an increase in habitat during the LGP, they also had an increase in Ne. However, they also state that this was not the case for non-passerines.

      Ecosystems are highly complex; there may also be other variables influencing past demographic change other than those explored here. Results should be interpreted with caution.

      We thank the reviewer for their comments, which has helped us in improving the scope of the manuscript while also removing errors in the supporting information. We have improved the section of the manuscript which addressed the drawbacks of PSMC in our revised version. Details and rational for parameter choice are now included in the revised manuscript.

      We performed additional PSMC analyses for a subset of the samples (n = 5), wherein the scaffolds mapping to the sex chromosome were removed only after mapping the reads. We compared the new approach suggested by the reviewer to our original approach and no differences in the PSMC pattern were observed, highlighting the robustness of the results (Supplementary Information Fig. S3).

      Additionally, we have included multiple box-plot and tables in the revised manuscript that helps with interpreting the changes in effective population size. The details of the revisions are presented below in the “Recommendations for the authors” section. We believe that these changes have improved the scope of the manuscript and removed any redundancies and conflicts.

      Reviewer #2 (Public review):

      Summary and strengths:

      In this manuscript, Karjee and colleagues used coalescent-based effective population size reconstruction (PSMC) from single genomes to understand past population trends in island birds and related this to life history traits and glacial patterns. This concept is fairly new, as there are still relatively few multiple PSMC synthesis studies. I also thought that the focus on island endemics was unique and adds value to this paper. I enjoyed seeing a paper focused on South East Asia and think that this could help contribute to our knowledge of the important biodiversity within this region.

      Major weaknesses:

      My biggest concern with this paper is that the analyses are limited to 20-30 species, and significant taxonomic bias is present (there are multiple species of passerine but only 1-2 representatives of other groups). While this is not an issue alone, many of the life history traits or geographical traits are conflated with phylogenetic diversity (e.g., there are no large-bodied passerines). Thus, it is my opinion that the impact of these drivers of past population size is conflated and cannot be disentangled with the current data. The authors themselves state that the core hypothesis surrounding Ne and habitat availability is not supported by their entire dataset (only seen in Passerines). This was not clear enough in the abstract, and conclusions cannot be drawn here as the impact of taxonomy cannot be separated from data richness, traits, etc. The PSMC analysis was done according to the most recent recommendations, and this part of the manuscript is fairly robust. However, in several places, it is incorrectly stated that the PSMC measures or can infer genetic diversity; PSMC only infers past effective population size. It cannot measure genetic diversity in the past. I cannot review the habitat reconstruction modelling as I am a conservation genomics specialist.

      Appraisal:

      I am not convinced about the findings within the paper. I do not think that the results are sufficiently supported at this time, largely due to the conflation of taxonomy with other variables. As this type of comparison is new, I do think that there is a chance for reasonable impact on the field of genomics and island biogeography if the manuscript's constraints are addressed. I do not see scope for impact on conservation at this time and find the conclusions in the abstract regarding conservation relevance to be unfounded.

      We thank the reviewer for highlighting the unique and robust analytical approaches we have taken in this study. We agree with the reviewer that our sample size currently is small. However, we do observe a robust correlation between habitat fluctuation and change in effective population size. Further, the study also highlights the predicament of tropical island endemics, which are currently understudied and future studies are necessary to safeguard the biodiversity. We have highlighted this while also addressing the concerns in the revised version of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Overall:

      This starts with a great premise - looking at how tropical single-island endemic bird species dealt with climate changes in the past may be a predictor of how they will respond to change in the future. Since these species are at high risk of extinction in the face of climate change, tailored approaches to conservation are a good idea. While the premise is solid, I have some questions and recommendations. At times while reading, I did feel a bit confused, which may be due to the fact that this isn't my exact area of expertise. However, if I'm confused, that means a reader from a general audience is also likely to be confused. Some results appear to be conflicting, some claims about data seem possibly inaccurate, and some major limitations are not acknowledged or fully addressed.

      Below I've noted areas that I feel could benefit from revisions. That being said, I liked the integration of habitat modeling and genomics! These sorts of multifaceted approaches are necessary when it comes to unraveling the complex dynamics involved in ecology and evolution.

      Crucial Issues to Address:

      (1) Line 75: With the lower sea levels and habitat change, you say animals can disperse across barriers of land and sea. When it comes to these single-island endemics, were they always confined to a single island? Is there no possibility of introgression with ancient populations of birds on other islands during these periods?

      We thank the reviewers for identifying the potential artifact in effective population size estimates that may occur due to hybridization/introgression. Most of our species belong to small and oligotypic families as has been addressed in the discussion section already, making them likely to be newly arisen lineages rather than refugial ones. There is scant information available in the literature on where the species in our dataset originated from, and further species-specific studies are required to identify signatures of hybridization/introgression. However, we have included this caveat in the revised version of the manuscript (line numbers: 73-78 and 303–305).

      (2) Lines 149-151 "However, in these species as well, trends of Ne increase or decrease from the LIG to LGM were robust across all three PSMC models considered." Please double-check this claim. Some of your figures in the supplement appear to contradict this. In particular, Pachycephala philippinensis, Rhynochetos jubatus, and Zosterops hypoxanthus appear to differ a bit in the time frame described, but it is difficult to tell-I would recommend adding some shading on the graphs to indicate the time period observed. If there was a way you determined this that is more precise than eyeballing the figures like I did, this should also be explained.

      We thank the reviewer for this comment and have reworded the sentence by cross verifying with the PSMC graphs. In addition, we have calculated the precise values of effective population size at the Last Interglacial (LIG) and Last Glacial Maximum (LGM) for each species using custom scripts and used these to evaluate whether the change in Ne during the Last Glacial Period (LGP) was significantly different for the three PSMC settings used. A table depicting these effective population size changes from LIG to LGM are also included in the revised version of the manuscript (Supplementary table S4; line numbers: 145­-156 and 345-357).

      (3) Lines 280-292: Issues with PSMC that are not acknowledged here are my largest concern. The situation being investigated does not necessarily meet all the assumptions PSMC makes (ie, neutral evolution and panmixia), which should be explained in this section. I'll point out the two issues I think should be acknowledged and addressed: First, selection is a confounding factor with PSMC, which is not mentioned here. While that's likely not an issue due to the size of the genome, this is still something that should be stated and explained. Second, the following statement is what I take the most issue with: "Population structure is thus a confounding factor. However, this is unlikely to be a problem given that all our species are single-island endemics". This needs justification. You state that in the past, islands could be connected (see my first comment regarding line 75), so it seems unlikely that 1) migration between past populations on other islands never happened, and 2) there is no population structure *on* the island.

      We thank the reviewer and have modified the PSMC caveats section of the revised version of the manuscript (line numbers: 289-307).

      (4) Line 310: Mapping all the data to the autosomes seems inappropriate to me. The sex chromosome reads could potentially map to other areas of the genome. Unless this information was accidentally left out of the methods section, it doesn't seem like any mapping quality filter was used, so spurious alignments aren't being removed. To remove sex chromosome data, I would instead align data to the whole genome, remove all reads that map to the sex chromosomes, and then map the remaining reads to the autosomes.

      As mentioned earlier, for a subset of the species (n =5), we directly mapped raw reads files onto the genome and then called SNPs on only autosomal regions using the SAMtools mpileup-bcftools pipeline, after which we performed PSMC as above (Supplementary Information Fig. S3). We did not observe and significant difference between the two approaches. Further, only high-quality mapped reads were used for SNP calling as mentioned in the previous version of the manuscript (line numbers: 338-343; Supplementary Information Fig. S3).

      (4) Table 1 includes some information that contradicts what is in the Supplementary Tables: Centropus unirufus, Chaetorhynchus papuensis and Cnemophilus loriae are not included in Supplementary Table 4. Table 1 says Eulacestoma nigropectus, Paradisaea rubra, and Parotia lawesii did not undergo PSMC analysis, but Supplementary Table 4 says PSMC and modeling trends matched for these species. "Pseudorectes ferrugineus" and "Rhynochetos jubatus" are spelled differently in Supplementary Table 4. Table 1 says Rhagologus leucostigma underwent both PSMC and climate modeling, but Supplementary Table 4 says "NA" as if it was missing one of these analyses.

      We thank the reviewer for identifying the errors and we have corrected for these in the revised version of the manuscript. Please see the detailed changes for these comments outlined below

      Centropus unirufus, Chaetorhynchus papuensis and Cnemophilus loriae are not included in Supplementary Table S4 (Now Supplementary table S2): we have added these species to the revised table S2.

      Table 1 says Eulacestoma nigropectus, Paradisaea rubra, and Parotia lawesii did not undergo PSMC analysis, but Supplementary Table 4 says PSMC and modeling trends matched for these species: The genomes for these samples were obtained from museums and exhibited high error rates. Hence, we excluded these samples from further analysis. However, the supplementary table S2 was not updated, and we have corrected this error in the revised version of the manuscript.

      "Pseudorectes ferrugineus" and "Rhynochetos jubatus" are spelled differently in Supplementary Table 4 (Now table S2): we have corrected the typographical error in the revised manuscript.

      Table 1 says Rhagologus leucostigma underwent both PSMC and climate modeling, but Supplementary Table 4 (Now table S2) says "NA" as if it was missing one of these analyses: This was a typographical error, and we have updated it to “mismatch”.

      Major Issues to Address:

      (1) Lines 97-99: "Information on tropical, single-island endemics' demographic responses to past climate change can inform conservation efforts, owing to the genomic signatures that predispose a species to extinction". This needs more explanation. For example, why couldn't we just look at these genomic signatures instead of recreating demographic responses? I'm not sure I fully understand what you mean here.

      We thank the reviewer for this comment and have modified the introduction to highlight the importance of demographic history in predicting species extinction. Comparison of genomic diversity and demographic history of over 200 mammalian genomes, highlights the importance of demographic history in predicting species endangerment and extinction risk (Wilder et al., 2023) (line numbers: 99-104).

      (2) Line 181-182: Whether or not a species was a passerine was an important predictor of Ne only in combination with the change in habitat from LIG to LGM". This is a major finding, but "respond positively to habitat change" (line 183) is a bit ambiguous. Were they responding to habitat expansion? Habitat contraction? Increase in rainfall? What is the change happening? Not all habitat changes are equal.

      We thank the reviewer for this comment and have modified this section for clarity in the revised results and discussion section of the manuscript. We observed a positive correlation between effective population size and availability of suitable habitat. Further, we observed precipitation of the warmest quarter to be the largest contributing bioclimatic variable for all but one Caribbean species (line numbers: 172-­191; 196-211).

      (3) Line 184-185: "The interaction between habitat change and body mass (β = 10.05, 95% CI: [-0.3, 24.41) suggests that there is no impact of habitat change in larger species." Doesn't this contradict the earlier finding of larger-bodied species seeing a decrease in Ne? Or do you mean the decrease in Ne was not due to habitat change?

      We have edited this section for clarity. With the inclusion of additional species, we observed a significant positive relationship between body size and effective population size (line number: 191-193).

      (4) Lines 206-207: "Our results also reveal that both passerine and non-passerine island endemics have entered the Holocene with low genetic diversity." How does this align with the statement that passerines responded positively to habitat change?

      The observation that passerines respond positively to habitat change is based on a systematic analysis of the last glacial period. However, a close look at the entire species’ demographic history reveals the often the Ne is at the lowest following the LGM, and coinciding with the advent of Holocene, the current interglacial. We have therefore modified the sentence in the revised version of the manuscript (line numbers: 213-214).

      (5) Line 215: If we already know flightless birds and endemics are particularly prone to extinction, what is the benefit of this study? Be clear about how your method can be used in a way that is better than what people are already doing. It would be good to explicitly explain why coalescent Ne is a better predictor of extinction risk than current genomic diversity or current Ne.

      We thank the reviewer for this comment and have modified this section in the revised version of the manuscript (line numbers: 221-224).

      (6) Line 259-261: "Habitat change in the LGP was positively associated with Ne fluctuations (Figure 3, β = 9.45), that is, species which showed an increase in habitat in the LGP also showed a concurrent increase in Ne." Is this true in all instances? I thought you found it had no effect for some, or did I misunderstand?

      We thank the reviewers for pointing this out. Species which showed an increase in habitat in the LGP did not always show a concurrent increase in Ne. Our results instead reflect an overall trend and this is clarified in the revised version of the manuscript (line numbers: 268-269).

      Lines 328-330: Could the different mutation rates used for passerines and non-passerines be driving the differences found between the two groups?

      The difference in the mutation rate is low and using the passerine specific mutation rate for non-passerines only shifts the PSMC graph slightly. As our analysis is considering the change in Ne across the LGP, this shift is minimal and does not affect the overall results.

      How are you connecting the demographic changes to species traits? I'm a bit confused about that, so I think some further explanation would be beneficial.

      We have modified the discussion to highlight the role of species traits in shaping the species response to habitat modification and ultimately the change in effective population size. We have included this in the revised version of the manuscript (line numbers: 437­-439).

      Minor Issues to Address:

      (1) Lines 165-168: "Habitat change was poorly associated with change in Ne for the 20 species for which both PSMC and ENM analyses were possible (Cramer's V = 0.15). However, passerine species only showed a strong association (Cramer's V = 0.96), while non-passerines showed a weak negative association (Cramer's V = -0.15)." This is phrased in a way that is a bit confusing. I'd consider rephrasing for clarity.

      We have modified this section in the revised version of the manuscript (line numbers: 167­-170).

      (2) Line 177: The confidence interval says "16.27, -2.61". I think it's supposed to be -16.27?

      We have corrected the typographical error in the revised version of the manuscript.

      (3) Line 185-187: "Finally, the random intercept for Country (sd (Intercept)) showed a marginal positive influence (β = 0.85, 95% CI: [0.04, 2.24])". What does this mean? This needs further explanation.

      We modified this sentence in the revised version of the manuscript (line number: 189-191).

      (4) Line 204: landbridge is misspelled as "landbride".

      We have fixed the typographical error.

      (5) Line 310: What were your Trimmomatic parameters?

      We have included the parameters used for Trimmomatic in the revised version of the manuscript (line numbers: 324-326).

      (6) Line 311: What were your bwa parameters?

      We used default parameters for bwa alignment and this is included in the revised version of the manuscript (line numbers: 328-329).

      (7) Line 322-324: Why did you choose those specific parameters for PSMC? Splitting up the first time window makes sense (as shown in Hilgers 2025), but why did you choose t=5, r=1, and 84 atomic time intervals? Did you choose these parameters independently, or did you decide to use them because they were used by Nadachowska-Brzyska et al? Either way, that information is important to state.

      The parameter selection followed the suggestions based on both Hilgers et al. 2025 and Nadachowska-Brzyska et al. 2016. The information is included in the revised version of the manuscript (line numbers: 345-350).

      (8) Lines 325-326: What did you use for bootstrapping? If not Psmcfa, why?

      We have used “splitfa” to generate files for bootstrap analysis and have included this information in the revised version of the manuscript (line numbers: 350-351).

      (9) Lines 350-354: Please explain the reasoning behind using the different resolution and worldclim for Amazona guildingii.

      Based on the reviewer’s comment, we have re-run the habitat model with the same resolution for Amazona guildingii and include this in the revised version of the manuscript.

      (10) Line 412-413: "For the response variable i.e., the change in Ne, a Bernoulli distribution with a logit link because it is a binary response variable." I think this sentence might be missing some words.

      We have fixed the typographical error in the revised version of the manuscript (line numbers: 444-445).

      (11) Figure 1 is difficult to read, especially the top left panel. I would consider presenting this differently.

      We have supplemented Figure 1 with boxplots of effective population size values estimated during the Last Interglacial and the Last Glacial Maximum which should aid in clarity.

      Reviewer #2 (Recommendations for the authors):

      The authors state that they intentionally chose to remove several avian species that would be suitable for this analysis, because they were subject to larger studies elsewhere. This seems like an unnecessary constraint, and it is my opinion that the authors need to add this data in. I am not aware of what species were excluded, but I hope this will increase the non-passerine proportion of their dataset to help them robustly address their questions. An alternative solution would be for the authors to only include passerines, but this will come at the expense of statistical power with the current dataset and so would also require an increase in sample size. Overall, I recommend including more non-passerine species with traits similar to your passerine species.

      This was a typographical error from the previous versions of the manuscript arising from the fact that we excluded museum species from our samples. We have modified this sentence in the revised version of the manuscript as well as included one new species (Melanocharis versteri) in our study panel (line number: 311-314).

      It was not clear how or if PSMC bootstrapping was included in the comparisons across species, i.e. how did you include bootstrapping when you turned PSMC into a response variable within your statistical analysis? Failing to account for it would introduce measurement error into the data, and I would suggest that the authors explore how to incorporate this.

      We thank the reviewer for this comment and have calculated the precise values of effective population size during the LIG and the LGM using custom scripts to generate boxplots. These boxplots were used to investigate if effective population size values were significantly different during the LGP for all three PSMC parameter settings. Non-significant results were treated as “no change” in effective population size for further statistical analyses. The bootstrap values were used for this analysis, in addition to circumventing the issue of selection on the genome.

      I would also like to see a greater discussion on what aspects of the PSMC curve were used for comparisons and the limitations therein. These cross-species comparisons are still relatively new, and I think they will add value to this paper.

      In our study, the change in Ne from LIG to LGM is considered. We have elaborated this in the revised version of the manuscript. Addition analysis, depicting the changes in Ne as box plots were also included to help understand the fluctuations in Ne.

      Lines 164-168, which refer to your core hypothesis, are really unclear. What was actually found here? Please rephrase.

      We have rephrased the sentence for clarity in the revised version of the manuscript (line numbers: 169­-172).

      PSMC measures effective population size, not genetic diversity. Please change throughout.

      Based on the reviewer’s comment we have changed this in the revised version of the manuscript.

      I was surprised to see some references to conservation within the abstract of the paper. It is important that this is also included in the discussion so that the authors ensure their logic is accessible to managers. It would also be good to discuss the risks of using PSMC to inform conservation from just one genome, as I see these being quite high.

      We thank the reviewer for this comment and have included both pros and cons of using PSMC in the revised version of the manuscript (line numbers: 229-237).

      As this paper is based on public reference genomes, it is best practice that the original notes or reference genome papers are cited to acknowledge the data holders.

      We thank the reviewer for this comment and have included a supplementary table (Supplementary Table S7) acknowledging all the data holders.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary of goals:

      The authors' stated goal (line 226) was to compare gene expression levels for gut hormones between males and females. As female flies contain more fat than males, they also sought to identify hormones that control this sex difference. Finally, they attempted to place their findings in the broader context of what is already known about established underlying mechanisms.

      Strengths:

      (1) The core research question of this work is interesting. The authors provide a reasonable hypothesis (neuro/entero-peptides may be involved) and well-designed experiments to address it.

      (2) Some of the data are compelling, especially positive results that clearly implicate enteropeptides in sex-biased fat contents (Figures 1 and 3).

      We thank the Reviewer for this overall positive assessment of our work.

      Weaknesses:

      (1) The greatest weakness of this work is that it falls short of providing a clear mechanism for the regulation of sex-biased fat content by AstC and Tk. By and large, feminization of neurons or enteroendocrine cells with UAS-traF did not increase fat in males (Figure 2). The authors mention that ecdysone, juvenile hormone or Sex-lethal may instead play a role (lines 258-270), but this is speculative, making this study incomplete.

      Figure 2 shows pan-neuronal or EE-specific expression of the female-specific Tra isoform (UAS-traF) did not explain sex differences in mRNA levels of EE cell-derived factors (we did not test body fat in this figure). We therefore agree that we did not pinpoint the upstream regulator of this difference, and suggest in our revised manuscript that identifying this regulator(s) will be an important future direction of our work.

      “Another important task for future studies will be to elucidate how sex differences in neuropeptide expression are established. The first step in understanding these mechanisms will be to determine which factors specify the sex bias in neuropeptide mRNA levels. Because our data shows that sex determination gene tra does not regulate the sex bias in neuropeptide expression in either the brain or the gut, the role of other factors that influence sexual identity and sexual differentiation must be assessed. One strong candidate is the steroid hormone ecdysone, as virgin females have higher ecdysone titers than males. Ecdysone plays a role in regulating sexual differentiation and development, and contributes to male-female differences in multiple aspects of intestinal physiology (e.g., intestinal stem cell proliferation) and brain development. Another candidate is juvenile hormone, which has been shown to regulate sexual maturation in Drosophila and other insects. While it remains unclear whether juvenile hormone titers differ between virgin males and females, juvenile hormone regulates many aspects of gut physiology in mated females (e.g., intestinal lipid accumulation, ISC proliferation) and influences brain development. Other than hormones, it is possible that sex determination gene Sex-lethal plays a role in regulating the sex difference in mRNA levels of EE cell-derived hormones, as tra-independent effects of Sex-lethal have been described in the brain.”

      (2) Related to the above point, the cellular mechanisms by which AstC and Tk regulate fat content in males and females are only partially characterized. For example, knockdown of TkR99D in insulin-producing neurons (Figure 4E) but not pan-neuronally (Figure 4B) increases fat in males, but Tk itself only shows a tendency (Figure 3B). In females, the situation is even less clear: again, Tk only shows a tendency (Figure 3B), and pan-neuronal, but not IPC-specific knockdown of TkR99D decreases fat.

      We thank the Reviewer for raising this point. In terms of general data interpretation, unless the ‘experimental genotype’ (e.g., cell type-specific gain/loss of a gene) shows a significant difference in gene expression or body fat (e.g., lower body fat/gene expression) from both control genotypes (UAS control, GAL4 control), the cell type-specific manipulation of a gene is not considered to have a biologically meaningful effect as it does not differ in phenotype from the parental strains.

      To ensure reader clarity on this issue we added the following text:

      “For these data, cell type-specific Tra overexpression was considered to have a significant effect on EE cell-expressed hormones only if the experimental genotype (e.g., tissue-GAL4>UAS-tra<sup>F</sup>) significantly differed from both parental strains (e.g., tissue-GAL4>+ and +>UAS-tra<sup>F</sup>) with the same direction of effect.”

      “For all fat storage data, cell type-specific RNAi was considered to have a significant effect on fat storage only if the experimental genotype (e.g., tissue-GAL4>UAS-RNAi) significantly differed from both parental strains (e.g., tissue-GAL4>+ and +>UAS-RNAi) with the same direction of effect.”

      Thus, in Figure 3B our data shows that gut-specific loss of Tk caused a trend toward decreased body fat in females ((p<sup>GAL4</sup>=0.1109 and p<sup>UAS</sup>=0.0118) with no effect in males (p<sup>GAL4</sup><0.0001 and p<sup>UAS</sup>=0.5704).

      In Figure 4B our data shows that pan-neuronal loss of TkR99D caused a significant decrease in female body fat ((p<sup>GAL4</sup><0.0001 and p<sup>UAS</sup><0.0001) with no effect in males ((p<sup>GAL4</sup>>0.9999 and p<sup>UAS</sup>>0.9999).

      In Figure 4E our data shows that IPC-specific loss of TkR99D caused a significant increase in male body fat ((p<sup>GAL4</sup><0.0001 and p<sup>UAS</sup>=0.0003) with no effect in females ((p<sup>GAL4</sup>=0.0321 and p<sup>UAS</sup>=0.0724).

      To summarize our findings for the reader, in our revised manuscript we added text to the Results section:

      “This suggests a role for gut-derived AstC and a potential role for gut-derived Tk in regulating female body fat, whereas gut-derived AstC or Tk do not play a role in regulating male body fat.”

      “These findings are interesting for several reasons. For example, in males, loss of EE cell-derived Tk and loss of TkR99D across neurons had no effect on fat storage, in contrast to the greater fat storage observed with IPC-specific TkR99D loss. This suggests that Tk derived from outside of the gut, and likely in the head, regulates fat storage via effects on TkR99D in the IPC. Future experiments will be needed to test this model, and to determine how Tk affects IPC biology. Further studies will also be needed to understand why IPC but not pan-neuronal loss of TkR99D causes an effect on body fat. Possible explanations include greater knockdown in the IPC using Dilp2-GAL4 or that Tk mediates opposing effects on body fat via effects on additional neuron groups with pan-neuronal TkR99D loss. In females, more work will be needed to identify the neurons upon which Tk acts to regulate body fat, and to test the relative contributions of EE cell- and brain-derived Tk in regulating body fat.”

      (3) The text sometimes misrepresents or contradicts the Results shown in the figures. UAS-traF expression in neurons or enteroendocrine cells did sometimes alter fat contents (Figure 2H, S), but the authors report that sex differences were unaffected (lines 164-166). On the other hand, although knockdown of Tk in enteroendocrine cells caused no significant effect (Figure 3B), the authors report this as a trend towards reduction (lines 182-183). This biased representation raises concerns about the interpretation of the data and the authors' conclusions.

      In Figure 2 we show the effects of UAS-traF expression in either EE cells or in neurons on mRNA levels of EE cell-derived factors (not body fat). Figure 2H shows the effect of UAS-traF in EE cells on Tk mRNA levels in the head, and Figure 2S shows the effect of pan-neuronal UAS-traF on NPF mRNA levels in the head.

      We thank the Reviewer for pointing out we should comment on the significant findings in 2H and 2S even though the direction of effect does not contribute to the sex difference in mRNA levels. In our revised manuscript we added the following text to this effect:

      “However, we note that Tra expression in EE cells further augments the male bias in head Tk mRNA levels (Figure 2H), whereas Tra expression in female neurons paradoxically decreases NPF mRNA levels in the head (Figure 2S).”

      (4) The authors find that in males, neuropeptide expression in the head is higher (Figure 1F-J). This may also play an important role in maintaining lower levels of fat in males, but this finding is not explored in the manuscript.

      We thank the Reviewer for pointing this out.

      In response to an earlier comment, one of the phrases we added to the revised manuscript was to acknowledge that the increased body fat we observed due to IPC-specific loss of TkR99D in males was likely mediated by Tk in the head, as there was no significant effect of loss of EE cell-derived Tk on body fat in males.

      “These findings are interesting for several reasons. For example, in males, loss of EE cell-derived Tk and loss of TkR99D across neurons had no effect on fat storage, in contrast to the greater fat storage observed with IPC-specific TkR99D loss. This suggests that Tk derived from outside of the gut, and likely in the head, regulates fat storage via effects on TkR99D in the IPC. Future experiments will be needed to test this model, and to determine how Tk affects IPC biology.”

      Appraisal of goal achievement & conclusions:

      The authors were successful in identifying hormones that show sex bias in their expression and also control the male vs. female difference in fat content. However, elucidation of the relevant cellular pathways is incomplete. Additionally, some of their conclusions are not supported by the data (see Weaknesses, point 3).

      Impact:

      It is difficult to evaluate the impact of this study. This is in great part because the authors do not attempt to systematically place their findings about AstC/Tk in the broader context of their previous studies, which investigated the same phenomenon (Wat et al., 2021, eLife and Biswas et al., 2025, Cell Reports). As the underlying mechanisms are complex and likely redundant, it is necessary to generate a visual model to explain the pathways which regulate fat content in males and females.

      We agree with the Reviewer that sex differences in fat storage are complex. We were also surprised that our findings regarding EE cell-derived hormones did not contribute to sex differences in the Akh- and insulin-producing cells. This suggests the regulation of sex differences in body fat is highly complex and involves many different factors. In our revised manuscript, we added text to this effect, and a graphical abstract to synthesize our past and new findings together into a single model.

      “Interestingly, these effects were not mediated by the IPC or APC, cells that we have previously shown contribute to the sex difference in fat storage. Taken together, our data provide additional insight into the highly complex mechanism(s) by which unmated female flies achieve higher fat storage than male flies (Fig. 5).”

      Reviewer #2 (Public review):

      Summary:

      This manuscript by Biswas and Rideout investigates sex differences in the expression and function of hormones derived from Drosophila enteroendocrine cells (EE). The authors report that while whole-body and head expression of several EE hormones (AstA, AstC, Tk, NPF, Dh31) is male-biased, gut-specific expression of AstC, Tk, and NPF is female-biased. Intriguingly, this sex-specific effect is not dependent on Tra - a surprising and important result. The authors then used an RNAi-based approach to demonstrate that gut-derived AstC and Tk promote fat storage specifically in females. Similar effects are observed when their receptors are knocked down in neurons. In addition, the authors were able to demonstrate that while Tk promotes female body fat via the insulin-producing cells. Together, these findings suggest that EE cell-derived hormones contribute to sex-specific fat storage regulation.

      We thank the Reviewer for their positive assessment of our paper.

      Strengths:

      Overall, I find the paper quite interesting. While the findings are brief, they reveal novel aspects of the sex-specific lipid storage program that I believe are important. As noted by the authors in the discussion, there are many open questions, including how these neuronal effects translate into systemic sex-specific regulation of lipid storage. Regardless, I find the results to be convincing - this paper will serve as the launching point of many future studies.

      Weaknesses:

      My main criticisms are focused on two points:

      (1) If the sex specific differences are eliminated by tra overexpression, what else might be responsible? As the authors note, the differences in 20E titers might be responsible. I would encourage the authors to simply feed adult flies with food containing 20E and determine if this alters sex-specific 20E expression.

      We agree that there are many candidates (e.g., ecdysone, juvenile hormone) that might contribute to sex differences in mRNA levels of EE cell-derived hormones. We suggest this is an important future direction of our work.

      “Another important task for future studies will be to elucidate how sex differences in neuropeptide expression are established. The first step in understanding these mechanisms will be to determine which factors specify the sex bias in neuropeptide mRNA levels. Because our data shows that sex determination gene tra does not regulate the sex bias in neuropeptide expression in either the brain or the gut, the role of other factors that influence sexual identity and sexual differentiation must be assessed. One strong candidate is the steroid hormone ecdysone, as virgin females have higher ecdysone titers than males. Ecdysone plays a role in regulating sexual differentiation and development, and contributes to male-female differences in multiple aspects of intestinal physiology (e.g., intestinal stem cell proliferation) and brain development. Another candidate is juvenile hormone, which has been shown to regulate sexual maturation in Drosophila and other insects. While it remains unclear whether juvenile hormone titers differ between virgin males and females, juvenile hormone regulates many aspects of gut physiology in mated females (e.g., intestinal lipid accumulation, ISC proliferation) and influences brain development. Other than hormones, it is possible that sex determination gene Sex-lethal plays a role in regulating the sex difference in mRNA levels of EE cell-derived hormones, as tra-independent effects of Sex-lethal have been described in the brain.”

      (2) I'm quite intrigued by the discovery that Tra does not eliminate the sex-specific differences. There are quite a few recent studies demonstrating that fruitless influences sex-specific neuronal function - here to I would encourage the authors to examine whether this aspect of the sex-determination pathway is involved in the lipid accumulation phenotype.

      We thank the Reviewer for raising this point. Transcripts derived from the fruitless-P1 promoter, which is largely responsible for the production of male-specific Fru<sup>M</sup> proteins in the CNS, are spliced by Tra. Therefore, while we cannot definitively rule out a role for fruitless, it is less likely given that the Tra expression in males (which would eliminate Fru<sup>M</sup> proteins in males) did not have a significant effect. In the revised manuscript, we added text to clarify this important point.

      “Future studies will also need to test additional members of the sex determination pathway. While sex differences in expression of EE cell-derived hormones does not involve tra, and is therefore unlikely to involve known tra targets such as fruitless, without further experiments we cannot fully rule out these additional sex determination pathway members.”

      Reviewer #1 (Recommendations for the authors):

      (1) The authors should explain why they focused on AstA, AstC, Tk, NPF and Dh31 but not Bursicon, CCHamides 1 and 2, and sNPF, especially since the latter four are also important entero-peptides.

      We thank the Reviewer for raising this point. In our revised manuscript we clarify that evaluating sex differences in all EE cell-derived hormones will be an important future direction of our work.

      “In particular, we focused on hormones known to influence whole-body fat metabolism, though an important future direction of this work will be to assess sex differences in all EE cell-expressed hormones.”

      (2) The authors initially compare peptide gene expression in males vs. females (Figure 1), but all subsequent comparisons (Figures 2-4) are experimental group vs. controls. It is necessary to directly compare males vs. females for these experiments as well, since the sex-biased difference is the focus of the paper. This may also help with variable performance of controls for some experiments (e.g. Figure 2), which makes interpreting these data difficult.

      We thank the Reviewer for making this point. In terms of general data interpretation, as with our response to an earlier point, unless the ‘experimental genotype’ (e.g., cell type-specific gain/loss of a gene) shows a significant difference in gene expression or body fat (e.g., lower body fat) from both control genotypes (UAS control, GAL4 control), the cell type-specific manipulation of a gene is not considered to have a biologically meaningful effect as it does not differ in phenotype from the parental strains.

      To ensure reader clarity on this issue we added the following text to the Results section:

      “For all fat storage data, cell type-specific RNAi was considered to have a significant effect on fat storage only if the experimental genotype (e.g., tissue-GAL4>UAS-RNAi) significantly differed from both parental strains (e.g., tissue-GAL4>+ and +>UAS-RNAi) with the same direction of effect.”

      In terms of comparing the sexes, all of our analyses used a two-way ANOVA and tested for a sex:genotype interaction. This allowed us to test whether males and females showed a statistically distinct response to the different genetic manipulations. To ensure clarity for readers, we include p-values for all the sex:genotype interactions in figure legends.

      (3) The organization of Figure 1 is unintuitive because the authors change the order of peptides in the last row of panels (Figure 1 K-O). The authors should keep the same order, so that every column corresponds to the same peptide, to make the figure easier for readers to follow.

      We thank the Reviewer for pointing out that we should make every row the same order of EE cell-derived peptides. We made this change in our revised manuscript.

      (4) The authors should explain why mRNA levels in whole-body samples are so highly skewed towards males (sometimes approaching 3-fold expression), whereas in the constituting tissues (head, guts), the differences are much milder and also in opposite directions. How do the big differences in favor of males in Figure 1A-E come about? Does the inclusion of the VNC skew expression levels so much?

      We thank the Reviewer for suggesting we clarify several points around the anatomical focus of sex differences in mRNA levels of EE cell-derived hormones. In our revised manuscript we explain that while male-biased mRNA levels in heads suggest that sex-biased expression in whole bodies may be attributed to expression in heads, that other tissues may contribute to the male-biased expression. We further state this is an interesting area for future investigation.

      “For most peptides, the male bias was due to a higher mRNA level in the head and not the fat body (Figure S1A-E); however, TkR99D mRNA levels were higher in male fat bodies with no difference in head mRNA levels (Figure S1C). We therefore cannot rule out a contribution of additional anatomical sites to the male bias in expression of EE cell-expressed hormones, which is an interesting area for future investigation.”

      (5) The authors use voila-GAL4 as a driver for enteroendocrine cells, but this line is also expressed in sensory cells. The authors should at least mention the expression pattern of this line at first mention (line 165).

      We thank the Reviewer for raising this point, we added text to this effect in the revised manuscript:

      “We found that sex differences in mRNA levels of AstA, AstC, Tk, NPF, and Dh31 were unaffected when we used either voila-GAL4 (Figure 2A-2J) which expresses in EE and sensory cells, or elav-GAL4 (Figure 2K-2T) which expresses in neurons and neuropeptide-producing cells, to drive Tra expression in these cells.”

      (6) Figure legends for Figures 2, 3 and 4 should be simplified and condensed to more concisely describe the panels. There is a lot of redundant repetition, which can easily be avoided by organizing the panels into groups (for example, in Figure 2, A-E should get a single legend entry rather than separate ones).

      We thank the Reviewer for this suggestion, we shortened our legends in the revised manuscript.

      (7) The authors refer to triglyceride contents as 'fat storage', but triglycerides can also be carried through the hemolymph via lipoproteins. The authors should use a more factual expression like 'total triglycerides'.

      We thank the Reviewer for this comment. Circulating lipoproteins in Drosophila carry primarily diacylglycerol, phosphatidylethanolamine, and sterol, with only a small fraction of triacylglycerol (PMID 22844248). Nevertheless, to ensure we are clear we added text in the Methods section to clarify that “fat storage” refers to whole-body triacylglycerol.

      “Triglyceride is the main form of stored fat in the body, with very little in the circulation. We therefore refer to whole-body triglyceride levels as ‘fat storage’ or ‘body fat’.”

      (8) The authors should justify their use of unmated flies for their experiments (line 324) and comment if they expect similar findings and mechanisms in mated flies, especially since nutritional and energy demands are greater in mated females.

      We added text to the methods to justify our use of unmated females to uncover the genetic mechanisms that contribute to sex differences in body fat.

      “We used unmated flies to identify genetic factors that regulate the sex difference in body fat; mated females were not used to avoid mating-induced changes in physiology mediated by additional factors (e.g., Sex-peptide) and behavioral changes due to altered food preferences.”

      (9) Are there any additional AstC and/or Tk receptors that could also play a role? The authors should comment on why they focused on AstC-R2 and TkR99D alone.

      We thank the Reviewer for this interesting point. We added text in our revised manuscript to acknowledge that we tested the primary known receptors for AstC and Tk, other receptors may contribute to their effects.

      “We therefore predicted that loss of AstC-R2 and TkR99D in these cells would reproduce the reduced fat storage we observed in females with loss of EE cell-derived AstC and Tk, though we cannot fully rule out effects of Tk and AstC mediated by other receptors as we did not test these additional receptors.”

      (10) The authors cite Song et al. 2014 to justify using R57C10-GAL80 to restrict expression patterns to the gut (lines 177-179), but upon checking that paper,r I could not find that Song et al. used this approach. Please scrutinize this and remove the reference if it is incorrect.

      We thank the Reviewer for pointing out that Song et al. did not specify how they achieved gut-specific Tk-GAL4; we removed this reference.

      Reviewer #2 (Recommendations for the authors):

      (1) Line 70 - the statement "In males, body fat is maintained..." seems too generic. I would suggest a small edit - "In males, body fat levels are maintained...".

      This is a good suggestion, thank you, we made the appropriate adjustment.

      “In males, body fat levels are maintained by higher expression and activity of two catabolic pathways that promote fat breakdown.”

      (2) Lines 78-81 - These statements suggest an either/or scenario, but I assume this is more a function of balance and equilibrium, where females have more ISS signaling that maintains elevated fat, while bmm pushes homeostasis in males toward catabolism. The authors should include more nuanced statements.

      We thank the Reviewer for this suggestion. In our revised manuscript we adjusted the text as follows:

      “Together, these studies have defined a model of the sex difference in fat storage in which females maintain higher levels of fat storage in part due to a higher relative activity level for anabolic pathway IIS, whereas males have lower fat storage due to higher relative activity of catabolic effectors such as bmm and Akh.”

      (3) Please provide all RRID numbers for the listed BDSC strains - the RRID numbers can be found at the bottom of the BDSC page for each strain.

      We thank the Reviewer for this suggestion, we added the RRID to the Methods.

      (4) Please cite the most recent FlyBase manuscript published in Genetics. Ideally, a statement under the fly husbandry section noting that Flybase was used as a resource throughout the study.

      Thank you for this suggestion, we made the requested change to properly acknowledge this critical community resource.

      “We acknowledge FlyBase as an essential resource providing genetic, genomic, and functional data and tools that supported this study.”

    1. Author response:

      The following is the authors’ response to the previous reviews

      We are pleased that Reviewer 3 appreciated our findings and found the temporal lag between the expression of TFF1 and TFF3 during signaling particularly interesting. The reviewer also advised us not to overemphasize that this lag arises from phase separation of ERα at the TFF1 locus, as the use of 1,6-hexanediol alone is not sufficient to conclusively establish whether ERα condensates undergo liquid–liquid phase separation.

      We agree with this assessment and have revised the manuscript accordingly. Specifically, we have modified the title to remove reference to phase separation and have updated the text throughout the manuscript to avoid claiming that the observed condensates are a result of phase separation.

      The revised title is:

      Ligand-dependent Enhancer Activation Indirectly Modulates Non-target Promoters in a Chromatin Domain.”

      With these changes, we are proceeding with the version of record using revised version of the manuscript.

      Thank you for your continued support.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      The authors investigate how infestation of rice plants by the small brown planthopper (Laodelphax striatellus), an important pest in rice cultivation, alters host plant carbohydrate metabolism and how these changes affect insect physiology and fitness. They show that planthopper infestation leads to a density-dependent increase in glucose levels in rice plants, which the authors suggest results from a redistribution of carbohydrates from roots to shoots. Elevated glucose levels in plants are reflected by increased glucose contents in the insects themselves, an effect that is particularly pronounced in gravid females and associated with enhanced fecundity.

      In addition, the authors demonstrate that increased glucose availability enhances tolerance of the small brown planthopper to the neonicotinoid insecticide imidacloprid. These findings suggest that insect-mediated changes in plant carbohydrate allocation may benefit insect fitness in multiple ways, including increased reproductive output and enhanced tolerance to insecticides, both of which are relevant for understanding insect population dynamics in agroecosystems.

      Beyond these physiological observations, the authors aim to elucidate the underlying molecular mechanisms. They propose that glucose functions not only as a nutritional resource but also as a signaling molecule. Specifically, they show that increased glucose availability is associated with activation of the Target of Rapamycin (TOR) pathway, a conserved nutrient-sensing signaling pathway regulating growth and metabolism across eukaryotes. Activation of TOR signaling is linked to increased juvenile hormone levels, which in turn stimulate vitellogenesis and likely contribute to increased fecundity. Furthermore, elevated juvenile hormone levels are associated with increased expression of glutathione S-transferases, suggesting a mechanism contributing to enhanced detoxification capacity. Independent of this pathway, increased glucose availability also leads to higher expression of glutamate-cysteine ligase, the rate-limiting enzyme in glutathione synthesis. Together, these mechanisms provide a non-exclusive explanation for the observed increase in imidacloprid tolerance and form the basis of the authors' proposed mechanistic framework linking glucose availability to reproduction and detoxification.

      We appreciate the reviewer for the thoughtful and positive summary of our work. We greatly appreciate the careful reading and the constructive recognition of our key findings, including the density‑dependent increase in glucose levels in rice plants, the resulting enhancement of planthopper fecundity, and the link between glucose availability and imidacloprid tolerance.

      We are also grateful that the reviewer highlighted our proposed mechanistic model, in which glucose acts as a signaling molecule to activate the TOR pathway, leading to increased juvenile hormone levels, enhanced vitellogenesis, and upregulation of detoxification-related enzymes such as glutathione S‑transferases and glutamate‑cysteine ligase.

      We have carefully addressed all other comments from the previous public reviews in the point‑by‑point response below.

      Strengths:

      A major strength of the manuscript is its substantial mechanistic depth and the extensive use of complementary experimental approaches that converge on a coherent mechanistic interpretation. The authors combine plant manipulations, dietary supplementation, injection assays, RNAi-mediated gene silencing, pharmacological inhibition, and rescue experiments to systematically test the role of glucose as a signaling molecule linking plant-derived nutrition to insect reproduction and insecticide tolerance. Results obtained from independent experimental strategies are highly consistent, and the different datasets collectively support the central conclusions of the study.

      The role of glucose is supported by multiple lines of evidence demonstrating that increased glucose availability, whether induced by prior planthopper feeding, dietary supplementation, or direct injection, consistently results in elevated glucose levels in insects, increased oviposition, and enhanced expression of vitellogenesis-related genes (LsVg and LsVgR). The specificity of this effect is further strengthened by experiments using alternative carbohydrates that release glucose upon enzymatic cleavage, as well as inhibitor and rescue experiments, supporting the interpretation that glucose acts beyond a purely nutritional role.

      The authors further establish a mechanistic link between glucose availability, TOR signaling, juvenile hormone regulation, and vitellogenesis. Activation of TOR signaling by glucose, demonstrated at the level of protein phosphorylation, together with RNAi knockdown and pharmacological inhibition, allows causal placement of TOR upstream of juvenile hormone signaling. Consistent reductions in juvenile hormone titers, vitellogenesis-related gene expression, and oviposition following TOR inhibition, as well as rescue of reproductive output by juvenile hormone analog treatment, provide strong functional support for a glucose-TOR-juvenile hormone axis regulating fecundity. The absence of additive effects following combined knockdown of TOR and juvenile hormone synthesis components further supports the interpretation that these factors act within the same signaling cascade.

      Similarly, the authors provide a detailed mechanistic analysis of glucose-mediated effects on imidacloprid tolerance. Functional assays demonstrate that glutathione S-transferases contribute to detoxification in this species and that increased glucose availability enhances GST activity, glutathione synthesis, and overall glutathione levels. Transcriptomic analyses and targeted RNAi experiments further identify specific GSTs contributing to insecticide tolerance and indicate that glucose enhances detoxification through both TOR-dependent and TOR-independent mechanisms. The combined knockdown experiments, which produce additive effects on mortality, provide particularly strong support for the involvement of multiple interacting glucose-dependent pathways.

      We appreciate the reviewer for the highly positive and thorough recognition of our work's strengths, including the mechanistic depth, convergent experimental approaches, and the proposed glucose–TOR–JH signaling cascade.

      Weaknesses:

      While I am impressed by the mechanistic depth of the study and the clarity with which the authors dissect the underlying physiological pathways, I am less convinced by the current conceptual framing of the phenomenon as a sophisticated adaptive strategy "co-opted" by the small brown planthopper. The data convincingly demonstrate that glucose availability activates conserved nutrient-sensing and endocrine pathways, including TOR signaling and juvenile hormone regulation, which in turn affect reproduction and detoxification capacity. However, these pathways are deeply conserved and likely operate in many insects in response to nutritional status. As such, the results may reflect a general physiological response to elevated carbohydrate availability rather than a species-specific, evolved strategy. Relatedly, herbivory-induced changes in plant carbohydrate allocation appear to be relatively common across plant-insect systems, and it would be helpful to discuss how specific (or general) the observed phenomenon is likely to be.

      In particular, I encourage the authors to more clearly distinguish between (i) a conserved nutrient-responsive signaling cascade and (ii) an adaptive mechanism that evolved specifically under selection imposed by insecticide exposure. The presented data strongly support the former interpretation, whereas evidence for the latter is less clear. The increased tolerance to imidacloprid appears to arise as a consequence of enhanced metabolic and detoxification capacity under elevated glucose conditions, rather than as a trait shaped directly by insecticide-driven selection. Framing this phenomenon as an adaptation to insecticide stress may therefore overextend the conclusions that can be drawn from the data. A more cautious discussion acknowledging that glucose-mediated activation of conserved metabolic and endocrine pathways may incidentally enhance insecticide tolerance, without necessarily having evolved under insecticide selection, would strengthen the conceptual clarity of the manuscript.

      We fully agree with the concerns raised regarding the evolutionary framing, conceptual definitions. We have thoroughly revised the manuscript to avoid overstatements about adaptive evolution, distinguish between conserved nutrient-responsive pathways and species-specific adaptations, supplement key definitions and literature, and address the study limitations and future directions in Discussion.

      While I am impressed by the mechanistic depth of the study and the clarity with which the authors dissect the underlying physiological pathways, I am less convinced by the current conceptual framing of the phenomenon as a sophisticated adaptive strategy "co-opted" by the small brown planthopper.

      We appreciate this comment. We replaced “how herbivorous insects exploit host nutritional signals for adaptation” with “how herbivorous insects respond to host nutritional signals to modulate their fitness traits”.

      Additionally, we uniformly revised overstated terms such as exploit, co-opt, and adaptive strategy throughout the manuscript to utilize, and nutrient-responsive mechanism, respectively, clarifying that our findings reflect a conserved physiological response of insects to host nutritional signals rather than specialized adaptive evolution under insecticide stress, thus avoiding overstatement of evolutionary adaptation.

      The specific revisions are as follows:

      (1) “exploit” was revised to “utilize”;

      (2) “manipulation” was revised to “change”;

      (3) “manipulated resource is exploited” was revised to “nutritional change is utilized”;

      (4) The first sentence of the Discussion section “Our study reveals a sophisticated adaptive strategy whereby SBPH actively manipulates host plant carbohydrate metabolism to simultaneously augment its reproductive capacity and insecticide tolerance.” was revised to: “Our study reveals a conserved nutrient-responsive mechanism whereby SBPH infestation elicits a carbohydrate metabolism shift in rice, and the insect subsequently utilizes host-derived glucose to simultaneously augment its reproductive capacity; concurrently, this glucose-mediated pathways enhances insecticide tolerance.”; t)

      (5) The second sentence of the Discussion section “we identify host-derived glucose as a central resource co-opted by SBPH and delineate two interconnected molecular cascades through which it exerts dual fitness benefits” was revised to: “we identify host-derived glucose as a central signaling molecule that modulates two interconnected molecular cascades exerting dual fitness benefits”.

      The data convincingly demonstrate that glucose availability activates conserved nutrient-sensing and endocrine pathways, including TOR signaling and juvenile hormone regulation, which in turn affect reproduction and detoxification capacity. However, these pathways are deeply conserved and likely operate in many insects in response to nutritional status. As such, the results may reflect a general physiological response to elevated carbohydrate availability rather than a species-specific, evolved strategy. Relatedly, herbivory-induced changes in plant carbohydrate allocation appear to be relatively common across plant-insect systems, and it would be helpful to discuss how specific (or general) the observed phenomenon is likely to be.

      Thank you for your comments and insights; we fully agree with your perspective. Accordingly, we have made the following revisions in the Abstract, Introduction, and Discussion of our manuscript:

      (1) The sentence “Our findings establish host-derived glucose as a central signaling molecule that SBPH exploits to simultaneously optimize reproduction and insecticide resistance.” has been modified to “Our findings establish host-derived glucose as a central signaling molecule that SBPH utilizes to modulate conserved pathways for simultaneous optimization of reproduction and insecticide resistance.”.

      (2) We added the following citation in the Introduction: “and sugar-promoted TOR activation has also been reported in Drosophila [29]”.

      (3) We revised the sentence “However, direct evidence for glucose-mediated TOR activation in insects and its functional connection to JH signaling and reproduction is lacking” by specifying “insects” as “hemipteran insects”.

      (4) In the Discussion, we revised “its sensitivity to glucose has remained elusive” to “sugar-promoted TOR activation has been reported in Drosophila [29], and our study extends this conserved regulatory mechanism to hemipteran insects”.

      (5) We added the phrase “This nutrient-responsive cascade might be conserved across insect species” at the end of the fourth paragraph of the Discussion.

      (6) Additionally, we added the following statement in the Discussion: “Notably, studies have shown that brown planthopper (Nilaparvata lugens) infestation can reshape sugar distribution in rice by altering the expression of rice sugar transporters, yet the mechanism through which planthoppers regulate these transporters remains unresolved [9]”.

      These revisions align with our data and support the reviewer’s view.

      In particular, I encourage the authors to more clearly distinguish between (i) a conserved nutrient-responsive signaling cascade and (ii) an adaptive mechanism that evolved specifically under selection imposed by insecticide exposure. The presented data strongly support the former interpretation, whereas evidence for the latter is less clear. The increased tolerance to imidacloprid appears to arise as a consequence of enhanced metabolic and detoxification capacity under elevated glucose conditions, rather than as a trait shaped directly by insecticide-driven selection. Framing this phenomenon as an adaptation to insecticide stress may therefore overextend the conclusions that can be drawn from the data. A more cautious discussion acknowledging that glucose-mediated activation of conserved metabolic and endocrine pathways may incidentally enhance insecticide tolerance, without necessarily having evolved under insecticide selection, would strengthen the conceptual clarity of the manuscript.

      We appreciate the professional comments and fully agree with your perspective. Accordingly, we have made the following revisions in the Discussion section:

      (1) The sentence “Our study reveals a conserved nutrient-responsive mechanism whereby SBPH infestation elicits a carbohydrate metabolism shift in rice, and the insect subsequently utilizes host-derived glucose to simultaneously augment its reproductive capacity and insecticide tolerance.” has been revised to:

      “Our study reveals a conserved nutrient-responsive mechanism whereby SBPH infestation elicits a carbohydrate metabolism shift in rice, and the insect subsequently utilizes host-derived glucose to augment its reproductive capacity; concurrently, this glucose-mediated activation of conserved metabolic pathways incidentally enhances insecticide tolerance.”

      (2) Original sentence: “The insect then exploits this manipulated nutritional landscape, deriving dual benefits of increased fecundity and enhanced insecticide tolerance.”

      Revised to:

      “The insect then exploits this manipulated nutritional landscape to increase fecundity, and the concurrent activation of conserved pathways by glucose incidentally enhances insecticide tolerance.”

      To remove the phrase “dual benefits,” which could imply that tolerance is an actively obtained adaptive advantage.

      (3) Original sentence:

      “Parallel to fecundity enhancement, SBPH utilizes host glucose to bolster its tolerance to the insecticide imidacloprid by supporting a novel dual-pathway model for GST activation, entailing both metabolic fueling and transcriptional regulation.”

      Revised to:

      “Parallel to fecundity enhancement, the activation of conserved metabolic and endocrine pathways by host-derived glucose incidentally bolsters SBPH tolerance to the insecticide imidacloprid, which is mediated by a novel dual-pathway model for GST activation involving both metabolic fueling and transcriptional regulation.”

      To clarify that the enhancement of insecticide tolerance is an incidental consequence of pathway activation, not a direct utilization strategy.

      Reviewer #1 (Recommendations for the authors):

      (1) Line 26 (Abstract): "how herbivorous insects exploit host nutritional signals for adaptation remains unclear." I am not sure that what is described here constitutes exploitation of a signal for adaptation. The authors convincingly unravel mechanisms by which insects benefit from elevated glucose, but the wording implies an evolved adaptation to insecticide pressure. Given that herbivore effects on nutrient allocation are likely widespread, I would recommend more cautious phrasing and clearer separation between physiological mechanisms and evolutionary interpretations.

      We appreciate this comment. We replaced “how herbivorous insects exploit host nutritional signals for adaptation” with “how herbivorous insects respond to host nutritional signals to modulate their fitness traits”.

      Additionally, we uniformly revised overstated terms such as exploit, co-opt, and adaptive strategy throughout the manuscript to utilize, and nutrient-responsive mechanism, respectively, clarifying that our findings reflect a conserved physiological response of insects to host nutritional signals rather than specialized adaptive evolution under insecticide stress, thus avoiding overstatement of evolutionary adaptation.

      The specific revisions are as follows:

      (1) “exploit” was revised to “utilize”;

      (2) “manipulation” was revised to “change”;

      (3) “manipulated resource is exploited” was revised to “nutritional change is utilized”;

      (4) The first sentence of the Discussion section “Our study reveals a sophisticated adaptive strategy whereby SBPH actively manipulates host plant carbohydrate metabolism to simultaneously augment its reproductive capacity and insecticide tolerance.” was revised to: “Our study reveals a conserved nutrient-responsive mechanism whereby SBPH infestation elicits a carbohydrate metabolism shift in rice, and the insect subsequently utilizes host-derived glucose to simultaneously augment its reproductive capacity; concurrently, this glucose-mediated pathways enhances insecticide tolerance.”;

      (5) The second sentence of the Discussion section “we identify host-derived glucose as a central resource co-opted by SBPH and delineate two interconnected molecular cascades through which it exerts dual fitness benefits” was revised to: “we identify host-derived glucose as a central signaling molecule that modulates two interconnected molecular cascades exerting dual fitness benefits”;

      (6) The phrase “This nutrient-responsive cascade might be conserved across insect species” was added at the end of the fourth paragraph in the Discussion section;

      (2) Line 37 (Abstract): To improve readability, please define "LsGST" on first use

      We appreciate this comment and have added taxonomic definitions for LsGSTe1 and LsGSTo1 at their first appearance in the Abstract: “LsGSTe1 (SBPH epsilon class GST) and LsGSTo1 (SBPH omega class GST)”.

      (3) Lines 38-39 (Abstract): The repeated framing as "signal exploitation" may not be fully justified, since glucose is simultaneously a key energetic resource that could plausibly fuel parts of the observed response. Clarifying what is meant by "signal" versus "resource" in this context would improve conceptual clarity.

      We appreciate this comment. Following your suggestion, we revised the description related to “signal exploitation”, and the sentence “Our findings establish host-derived glucose as a central signaling molecule that SBPH exploits to simultaneously optimize reproduction and insecticide resistance.” has been modified to “Our findings establish host-derived glucose as a central signaling molecule that SBPH utilizes to modulate conserved pathways for simultaneous optimization of reproduction and insecticide resistance.”.

      In addition, we have emphasized the signaling role of glucose in both the Results and Discussion sections. Through mannitol osmotic control treatments, hydrolase inhibition assays, and rescue experiments, we excluded the possibility that glucose acts merely as an energy source and confirmed its signaling function in regulating the JH pathway via TOR phosphorylation. These experiments clearly distinguish its signaling role from its nutritional/energetic role.

      (4) Lines 39-41 (Abstract): The phrase "nutrient-based control strategies" is difficult to interpret without at least a brief example or explanation. A short clarification would help readers understand the applied implications.

      We fully agree with and appreciate this comment. We added a concrete example in the Abstract: “, such as disrupting insect nutrient-sensing pathways or modulating host carbohydrate metabolism”.

      We also added a new section “The identification of the glucose‑TOR‑JH axis as a key regulator of SBPH fecundity and insecticide tolerance provides novel strategies for eco-friendly, nutrient-based pest control. Firstly, varieties that limit SBPH-induced glucose redistribution would reduce reproduction and insecticide tolerance without yield loss. Secondly, small-molecule inhibitors targeting TOR phosphorylation or JH synthesis can serve as biopesticides or synergists to improve insecticide efficacy, as they would suppress the glucose-mediated incidental enhancement of insecticide tolerance. Finally, optimized fertilization and irrigation can reduce shoot glucose accumulation and suppress SBPH outbreaks. These strategies offer sustainable alternatives to traditional insecticides and help mitigate insecticide resistance of SBPH.” in the Discussion, detailing three practical strategies (rice breeding, small-molecule inhibitor, agronomic management) to clarify the applied meaning of nutrient-based control strategies.

      (5) Line 68 (Introduction): The statement that glucose is "the dominant transportable carbon source" in plants seems inaccurate; sucrose is generally considered the main transport sugar. Consider revising.

      We appreciate this professional comment. According to the literature, glucose can be transported in plants but is not the primary sugar involved; sucrose is the main transported form. We have therefore removed the word “dominant”, and the revised description is consistent with current knowledge.

      (6) Line 82 (Introduction): The claim that "direct evidence for glucose-mediated TOR activation in insects...is lacking" may not be correct. For example, Kim & Neufeld (2015) report sugar-promoted TOR activation in Drosophila (Nat. Commun., doi: 10.1038/ncomms7846). This may also relate to statements later in the manuscript (e.g., around line 506).

      We apologize for this oversight during our initial literature review and sincerely appreciate this professional comment. We have added the relevant citation in the Introduction: “and sugar-promoted TOR activation has also been reported in Drosophila [29]”. (of the revised manuscript)

      Furthermore, we revised the sentence “However, direct evidence for glucose-mediated TOR activation in insects and its functional connection to JH signaling and reproduction is lacking” by specifying “insects” as “hemipteran insects”. (of the revised manuscript)

      In addition, we modified the corresponding statement in the Discussion section: the phrase “its sensitivity to glucose has remained elusive” was revised to “sugar-promoted TOR activation has been reported in Drosophila [29], and our study extends this conserved regulatory mechanism to hemipteran insects”. (of the revised manuscript)

      These revisions could clarify that our innovative contribution lies in extending this conserved mechanism from Drosophila to hemipteran insects, rather than reporting the first discovery of glucose-induced TOR activation. Accordingly, we have adjusted the reference numbering for all subsequent citations in the manuscript.

      (7) Line 176 (and elsewhere): Mannitol is used as an osmotic control; it would be helpful to briefly explain why osmolarity is expected to be a relevant confound in these assays and how osmotic effects might otherwise influence the measured outcomes.

      We greatly appreciate this valuable comment. We have added the following paragraph to the Discussion section: “Given that osmotic pressure, a key determinant of plant cell turgor pressure, can disrupt insect homeostasis and impair fitness when insects ingest hyperosmotic plant sap [47,48], we rigorously excluded confounding effects of rice osmotic pressure in this study.”.

      Two relevant references [47, 48] have been cited to support this statement, and we have adjusted the reference numbering for all subsequent citations in the manuscript.

      (8) Line 464 ff.: The statement that co-option of plant defenses by insects is an "emerging paradigm" seems overstated; classic examples such as sequestration of plant toxins have been known for decades. A more nuanced phrasing may be appropriate.

      We appreciate this comment and agree with your perspective. We have revised “emerging paradigm” to “classic paradigm” for greater objectivity.

      (9) Lines 491-492: This passage is somewhat confusing with respect to framing: here, elevated glucose is described as a plant stress response, whereas elsewhere (including title/abstract) it is presented as manipulation by the insect. Clarifying whether the authors view elevated glucose primarily as a plant response that insects benefit from, versus an actively induced manipulation, would improve consistency.

      We greatly appreciate your professional comment. We have revised the relevant statement from: “Given that elevated sugar levels might enhance plant stress resistance [46,47], our study reveals an intriguing ecological paradox: the plant's potential attempt to mount a stress response via glucose accumulation is effectively co-opted by the insect to enhance its own fitness and resilience.”

      To: “Notably, studies have shown that brown planthopper (Nilaparvata lugens) infestation can reshape sugar distribution in rice by altering the expression of rice sugar transporters, yet the mechanism through which planthoppers regulate these transporters remains unresolved [9]. Given that elevated sugar levels might enhance plant stress resistance [49,50], our study reveals an intriguing ecological paradox that SBPH infestation likely manipulates glucose distribution via unidentified pathways to boost its own fitness and resilience.”

      (10) Discussion (general): In addition to the demonstrated glutathione/GST mechanisms, elevated glucose could plausibly support detoxification in other ways (e.g., providing a substrate for conjugation in phase II metabolism). It may be worth briefly acknowledging such additional routes, even if not tested here.

      We appreciate this comment. We added the following text in the GST pathway section of the Discussion: “Beyond the GCL-GSH-GST and TOR-JH-GST pathways characterized in this study, elevated glucose may also enhance insecticide detoxification through additional routes (e.g., providing carbon skeletons for phase II xenobiotic conjugation reactions or fueling energy-dependent detoxification processes in insect midgut and fat body), which warrant further experimental verification.”, objectively acknowledging other potential pathways and listing them as future research directions.

      Reviewer #2 (Public review):

      Summary:

      Zhang and colleagues investigate the molecular mechanisms by which the small brown planthopper (SBPH, Laodelphax striatellus) manipulates host rice carbohydrate metabolism to enhance its own fitness. Using a combination of molecular, pharmacological, and biochemical approaches, they demonstrate that SBPH infestation induces systemic glucose reallocation in rice, as evidenced by the upregulation of glucose levels in aerial tissues and a simultaneous reduction in root glucose levels. Notably, host-derived glucose acts as a central signaling molecule, driving two key adaptive traits: enhanced fecundity via the glucose-TOR-JH-Vg signaling cascade, and increased imidacloprid tolerance through synergistic metabolic (GCL-GSH) and regulatory (TOR-JH-GST) pathways targeting GST activity. These findings uncover a sophisticated resource-manipulation strategy in SBPH and identify nutrient-sensing and detoxification pathways as potential targets for pest control.

      Strengths:

      (1) The study addresses a gap in plant-insect coevolution research by identifying glucose as a dual-function signaling molecule that coordinates SBPH reproduction and insecticide tolerance, providing valuable insights into how herbivores exploit host nutritional signals.

      (2) The experimental design is well structured and multifaceted, integrating RNAi, RT-qPCR, Western blotting, pharmacological inhibition, and biochemical assays. The use of appropriate controls (e.g., osmotic controls with mannitol and hydrolase-inhibitor rescue experiments) strengthens the causal interpretation of the results.

      (3) The mechanistic framework is clear and well-supported. The authors delineate two interconnected molecular cascades (glucose-TOR-JH-Vg for fecundity and GCL-GSH/TOR-JH-GST for tolerance) with hierarchical validation (e.g., rescue experiments with JHA), ensuring the reliability of conclusions.

      We thank the reviewer for recognizing the novelty of the scientific question, rigor of the experimental design, and clarity of the mechanistic framework in our study. We fully agree with the limitations raised regarding the generality of the findings, identification of upstream signals, range of insecticides tested, and translational applications for pest management. We have supplemented the manuscript with discussions of our study limitations and future research directions, added a section on the application of our findings in pest control, and provided key future directions such as identification of upstream signal identification, validation of generality and expansion of insecticide testing.

      Weaknesses:

      (1) The study focuses exclusively on SBPH without validating whether the observed phenomena and mechanisms are conserved in closely related planthopper species (e.g., brown planthopper Nilaparvata lugens). This limitation restricts the generalizability of the findings to other economically important rice pests.

      We appreciate this valuable comment. We have added a subsection titled “Limitations and Future Research Directions” in the Discussion section, explicitly stating that this study focuses exclusively on SBPH and the broader generality of the mechanism remains to be verified. Among the future directions outlined, “verifying the conservation of the glucose‑TOR‑JH axis in other economically important rice planthoppers” is designated as the first key research priority, and cross‑species validation experiments are planned accordingly.

      (2) The specific upstream signals that trigger glucose reallocation in rice (e.g., SBPH salivary effectors or oviposition-associated factors) are not identified. Although this represents a complex and independent research direction, the absence of such information limits the depth and completeness of the mechanistic framework and leaves open questions regarding the initiation of host metabolic manipulation.

      We greatly appreciate this insightful comment. We have incorporated this issue as a key future research direction in the Discussion section. Specifically, we added the following statement: “Notably, the upstream signals (e.g., specific salivary effectors secreted by SBPH or oviposition-associated plant response factors) that trigger glucose reallocation in rice remain uncharacterized and represent a key direction for future in-depth research.”

      In addition, we have added a subsection titled “Limitations and Future Research Directions” in the Discussion, which includes the third point: “(3) Identifying the specific SBPH salivary effectors and plant signaling pathways that trigger glucose reallocation in rice, to complete the mechanistic framework of host metabolic changes manipulated by herbivores.”

      (3) Insecticide tolerance assays are limited to imidacloprid. Extending these analyses to one or two additional commonly used insecticides (e.g., thiamethoxam) would help determine whether the glucose-mediated detoxification pathway is specific to imidacloprid or reflects a broader resistance mechanism, thereby strengthening conclusions regarding the generality of the GST activation cascade.

      We greatly appreciate this comment. Related discussion was added in the subsection titled “Limitations and Future Research Directions” as following:

      Expanding assays to other commonly used rice insecticides (e.g., thiamethoxam, pymetrozine, triflumezopyrim) to validate whether the glucose-mediated detoxification pathway confers broad-spectrum tolerance.

      (4) Given the study's potential implications for pest management, the manuscript would benefit from a brief discussion of possible practical applications, such as manipulating rice glucose metabolism through breeding strategies or developing small-molecule inhibitors targeting the TOR-JH axis. Including such perspectives would enhance the translational relevance of the work by linking mechanistic insights to real-world pest control strategies.<br />

      We greatly appreciate your professional comment. We have added a standalone paragraph in the Discussion section to discuss the novel strategies for SBPH control provided by this study, as follows: “The identification of the glucose‑TOR‑JH axis as a key regulator of SBPH fecundity and insecticide tolerance provides novel strategies for eco-friendly, nutrient-based pest control. Firstly, varieties that limit SBPH-induced glucose redistribution would reduce reproduction and insecticide tolerance without yield loss. Secondly, small-molecule inhibitors targeting TOR phosphorylation or JH synthesis can serve as biopesticides or synergists to improve insecticide efficacy, as they would suppress the glucose-mediated incidental enhancement of insecticide tolerance. Finally, optimized fertilization and irrigation can reduce shoot glucose accumulation and suppress SBPH outbreaks. These strategies offer sustainable alternatives to traditional insecticides and help mitigate insecticide resistance of SBPH.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors propose that HSV-1 infection degrades the class I histone deacetylases HDAC1 and HDAC2. The MDM2 E3 ubiquitin ligase from the DNA damage response pathway is responsible for ubiquitinating these HDACs that are subsequently degraded via proteasomes. The authors hypothesize that HDAC degradation will cause hyperacetylation of viral chromatin and enable viral gene transcription.

      Strengths:

      The ubiquitination of HDAC1 & HDAC2 by Mdm2 and the mapping studies are clear.

      Weaknesses:

      (1) Degradation of HDACs is observed late, at least 12-24 h post-infection (1 PFU/cell). Viral genes have been transcribed by that point, and the virus has replicated its genome. The kinetics do not match the proposed model.

      We sincerely thank the reviewers for their insightful and constructive feedback. The original low‑MOI condition introduced asynchronous infection and obscured early events. We repeated the time course at high MOI (MOI = 5) in HeLa cells. Under these synchronized conditions, HDAC1/2 degradation is detectable by 2 hpi and pronounced by 4‑6 hpi—preceding viral DNA replication (~3‑4 h) and coinciding with true late gene expression (ICP4). These data (Author response image 1) show that HDAC1/2 depletion is an early, virus‑directed event, not a late consequence.

      Author response image 1.

      (2) The authors need to connect these findings with their story. As of now, these findings are correlative. For example, what is the impact of MDM2 depletion on viral gene expression and progeny virus production? Leptomycin B is not specific to the HDAC cytoplasmic translocation, and its effect on the infection could be due to its effect on ICP27.

      We generated stable MDM2 knockdown HeLa cells. MDM2 depletion reduced progeny virus titers at 24 hpi and suppressed ICP0, ICP8, and gB expression at both RNA and protein levels (Figure 4G‑J). HDAC1/2 degradation was abolished (Figure 4B). Thus MDM2‑dependent HDAC1/2 proteolysis is essential for the lytic transcriptional cascade.

      To bypass LMB’s broad CRM1 inhibition, we constructed an HDAC1 mutant lacking the nuclear export signal (HDAC1‑ΔNES) that remains nuclear during infection. In HDAC1/2 double‑knockdown cells, re‑expression of HDAC1‑ΔNES failed to rescue viral replication compared to wild‑type HDAC1, and overexpression of HDAC1‑ΔNES in control cells more strongly inhibited progeny yield (Figure 5K). This genetic approach confirms that HDAC1 nuclear export is specifically required for its proviral function, independent of ICP27.

      Author response image 2.

      (3) The time point when the inhibitors were added to the cultures has not been stated in any experiment. If inhibitors were added with the virus, viral gene expression would be blocked.

      We sincerely thank the reviewers for identifying this critical oversight. Unless otherwise specified, All inhibitors were added at 1 hpi (post‑adsorption) unless otherwise noted. Time‑of‑addition experiments for MG‑132, LMB, and berzosertib are now included (Author response image 3, Figure 2K ,5A).

      Author response image 3.

      (4) The authors need to present late gene expression data in all the experiments where drugs have been used.

      In all drug-treated experimental conditions, we have performed comprehensive qRT-PCR analyses—quantifying mRNA levels of at least one immediate-early gene (ICP0), one early gene (ICP8), and one late gene (gB or gC)—to establish a temporally resolved viral gene expression profile across the entire replication cycle. This systematic assessment allows us to rigorously determine whether the observed inhibitory effects are global or selectively restricted to specific kinetic classes of viral genes. Consistent with this design, both LMB and berzosertib significantly suppressed the mRNA expression of ICP0, ICP8, and gB in HSV-1–infected cells (Author response image 4), indicating that their antiviral activity likely stems from interference with an upstream regulatory node common to the transcriptional activation of immediate-early, early, and late viral genes.

      Author response image 4.

      (5) Figure 1A, ICP4 is not detected up to 12 hours post-infection of HeLa cells with 1 PFU/cell. This cannot be true.

      This observation stems from technical artifacts in the original Western blot images—primarily insufficient signal intensity and suboptimal dynamic range. Therefore, we re-conducted the time-course experiment, systematically collecting samples at each time point and moderately increasing the sample loading volume while ensuring protein integrity. Optimized Western blot analysis confirmed robust ICP4 protein expression beginning at 12 hours post-infection, with progressive accumulation over time. Consequently, we have replaced all ICP4-related Western blot panels in Figure 1A, Figure 1E, and Figure 2H with newly acquired, rigorously exposure-calibrated images exhibiting high signal-to-noise ratios and unambiguous temporal resolution.

      (6) Leptomycin B blocks nuclear/cytoplasmic shuttling of ICP27 that brings viral mRNAs to the cytoplasm to be translated. So, the effect of LMB is not specific to the HDACs.

      This is a critical point. As outlined in our response to Reviewer 2, we will address this issue through three complementary experimental approaches: (1) generation of an HDAC1 nuclear export signal (NES)–deficient mutant to genetically abrogate its nuclear export (Figure 5K); (2) rigorous nuclear-cytoplasmic fractionation coupled with immunoblotting to quantitatively assess HDAC1 subcellular distribution and site-specific ubiquitination. Importantly, our fractionation data confirm that HDAC1 ubiquitination is predominantly cytoplasmic (Figure 5F). Collectively, these experiments will rigorously distinguish direct HDAC1 modulation from indirect, LMB-mediated off-target effects—thereby ensuring the mechanistic specificity and interpretability of our conclusions.

      (7) The key experiment is to use the degradation-resistant form of HDAC1 to evaluate its impact on viral gene transcription.

      Based on their comments, we systematically evaluated the effects of overexpression of wild-type (WT) HDAC1 and its ubiquitination site mutant K74R (anti-degradation form) on the transcription of HSV-1 viral genes and the yield of progeny viruses. Specifically, we measured the mRNA levels of representative immediate-early genes (ICP0), early genes (ICP8), and late genes (gB), and simultaneously determined the viral titer (Figure 3L). The results showed that compared with the empty vector control, overexpression of HDAC1 WT significantly inhibited the transcription of viral genes at all stages and reduced the yield of progeny viruses; while overexpression of HDAC1 K74R exhibited a stronger inhibitory effect - its inhibition of viral gene transcription and viral replication was significantly higher than that of HDAC1 WT (Figure 3M). This result provides key functional validation for the core mechanism that "HSV-1 promotes its own replication by targeting the degradation of HDAC1/2 to relieve the epigenetic inhibition of its genome".

      (8) In the experiment where Mdm2 was depleted, the authors need to demonstrate the effect on the infection. ICP4 expression is not enough. How about growth curves? After Mdm2 depletion, ICP4 expression increases, which may contradict the authors' findings. An analysis of alpha and gamma gene expression is important.

      We sincerely apologize to the reviewers for the error in Figure 4B, which arose from an oversight during experimental execution and data validation. We have rigorously repeated the experiment and confirmed that MDM2 knockdown robustly suppresses ICP4 protein expression—directly contradicting the erroneous upregulation depicted in the original figure. To fully characterize the functional consequences of MDM2 depletion on HSV-1 replication, we performed a multi-step viral growth assay and quantified mRNA levels of canonical viral genes by quantitative RT–PCR: consistent with the corrected ICP4 data, MDM2 knockdown significantly impaired progeny virus production and concurrently reduced transcript abundance of the immediate-early gene ICP0 and the early gene ICP8 (Figure 4G, 4H, 4I and 4J). We are profoundly grateful to the reviewers for identifying this critical discrepancy and for affording us the opportunity to provide a thorough correction and mechanistic clarification. In response, we have revised Figure 4B, updated all related text and figure legends, and conducted a comprehensive cross-check of all data, figures, and textual content across the manuscript.

      (9) Why did the authors analyze a liver HSV-1 infection and not a more relevant skin infection?

      We sincerely apologize for the lack of sufficient detail in our prior response, which may have inadvertently increased the reviewers’ evaluation burden. Prior to in vivo experimentation, we conducted a systematic tissue tropism profiling of HSV-1–infected mice, quantifying viral protein expression across multiple organs by Western blot (WB) (Author response image 5). This unbiased, multi-modal assessment demonstrated that the liver exhibited both the highest viral protein abundance and the greatest viral genomic load, coupled with the most pronounced and histologically reproducible pathology—including dense inflammatory infiltration, hepatocyte vacuolar degeneration, and sharply demarcated foci of necrosis. Importantly, HSV-1 infection induced a robust and coordinated downregulation of HDAC1 and HDAC2 protein levels specifically in the liver; this effect was neither as pronounced nor as consistent in other tissues examined. In contrast, although the skin serves as the natural portal of entry for HSV-1, it displayed consistently low and highly heterogeneous viral protein expression in this systemic model—rendering it unsuitable for rigorous virological or immunological quantification. Accordingly, grounded in these empirical findings and aligned with established practices in models of disseminated herpesvirus infection [1]—where the liver is routinely prioritized as the primary site of pathogenesis and immune interrogation—we designated the liver as the principal organ for in-depth analysis of viral replication dynamics and host innate and adaptive immune responses.

      Author response image 5.

      Reviewer #1 (Recommendations for the authors):

      (1) It is an HDAC class and not cluss. All figures need correction.

      We sincerely apologize to the reviewers for this oversight and confirm that the error has been duly corrected in the revised manuscript.

      (2) The authors need to quantify cells with H2AX in the nucleus in HSV-1 and PRV infections. Some blots from the analysis of the ATM signaling are not of great quality.

      As recommended, we performed quantitative immunofluorescence analysis to assess nuclear γ-H2AX foci formation in infected cells. Consistent with activation of the DNA damage response, γ-H2AX levels increased markedly in a time- and dose-dependent manner following HSV-1 or PRV infection. In addition, the immunoblot images for ATM signaling pathway components (Fig. 2H and 2I) have been replaced with higher-resolution.

      Reviewer #2 (Public review):

      Summary:

      The authors discovered that HDAC1/2 are degraded in HSV-1 and PRV infections. They attempted to establish a new mechanism by which HDAC1/2 are translocated to the cytoplasm to be degraded in HSV-1 infection, and the degradation causes changes in histone acetylation to affect the DDR pathway.

      Strength:

      (1) Interesting findings of HDAC1/2 degradation during HSV-1 and PRV infection, and it may impact more than the virology field.

      (2) Significant work to identify the ubiquitin site in HDAC1/2 and K63 linkage.

      We sincerely thank you for your positive assessment of this work and for your thoughtful, constructive feedback. Below, we provide point-by-point responses to each of your comments.

      Weaknesses:

      (1) Insufficient evidence to support the mechanism described by the authors.

      (2) Expansion of the conclusion to alphaherpesvirus without studying the intended mechanism in PRV infection.

      Overall, there may be a correlation between HDAC1/2 level, ATM/ATR phosphorylation, and HDAC1 translocation during the HSV-1 infection. However, core evidence supporting the mechanism that a) HDAC1 export causes its degradation, b) degradation of HDAC1 causes histone acetylation changes and DRR activation has not been sufficiently demonstrated.

      In direct response to the central concern raised—that “the core experimental evidence supporting the proposed mechanistic model remains insufficient”—we have performed two complementary sets of rigorous validation experiments. Specifically, we addressed the two key mechanistic steps: (a) HDAC1 nuclear export is required for its ubiquitin–proteasome-dependent degradation; and (b) HDAC1 degradation drives histone hyperacetylation and consequent activation of the DNA damage response (DDR). Our new data robustly substantiate both causal links.

      To establish causality between HDAC1 nuclear export and degradation, we employed a dual experimental approach: (i) generation of an HDAC1 nuclear export signal (NES) loss-of-function mutant (HDAC1-ΔNES), which specifically abrogates CRM1-mediated nuclear export without affecting protein stability or catalytic activity; and (ii) high-fidelity subcellular fractionation coupled with ubiquitin pull-down and quantitative immunoblotting, enabling precise quantification of HDAC1 distribution and site-specific ubiquitination across nuclear and cytoplasmic compartments. Consistent with our model, wild-type HDAC1 underwent pronounced cytoplasmic accumulation and polyubiquitination following HSV-1 infection, demonstrating that nuclear export is both necessary and sufficient for HDAC1 degradation.

      To determine whether HDAC1 degradation functionally triggers downstream DDR activation, we generated stable HDAC1/2-knockdown HeLa cell lines using validated siRNA constructs. Loss of HDAC1/2 led to significant increases in H3 and H4 acetylation levels and robust induction of canonical DDR markers—including γ-H2AX foci formation, ATM phosphorylation, and ATR phosphorylation—phenocopying the effects observed during HSV-1 infection. These gain-of-function data confirm that HDAC1/2 depletion alone is sufficient to recapitulate the epigenetic and DDR phenotypes, thereby solidifying the mechanistic hierarchy: HDAC1 export → degradation → histone hyperacetylation → DDR activation.

      (2) Expansion of the conclusion to alphaherpesvirus without studying the intended mechanism in PRV infection.

      Our prior work demonstrates that both porcine pseudorabies virus (PRV) and herpes simplex virus type 1 (HSV-1) elicit highly concordant phenotypic outcomes—including marked depletion of HDAC1/2 proteins, elevated acetylation of histones H3 (K9/K27/K56) and H4 (K8/K12), and robust activation of the DDR pathway—as evidenced by parallel assays across both viral systems. While mechanistic dissection was primarily pursued in the HSV-1 model—due to its well-established tractability for biochemical and genetic interrogation—PRV and HSV-1 are evolutionarily closely related α-herpesviruses sharing extensive conservation in genome organization, replication machinery, and key immune-modulatory effectors. Critically, all core phenotypes described herein were independently validated in PRV-infected cells (Figure 1A/B,1E/F,2A/B), thereby providing direct experimental support for generalizing the findings to the α-herpesvirus genus. Accordingly, the title’s scope is both empirically justified and scientifically precise.

      Reviewer #2 (Recommendations for the authors):

      Major issues:

      (1) Line 26: "we uncover a novel mechanism by which alphaherpesviruses exploit the DDR pathway". A mechanism has not been clearly described. The authors showed DNA damage, phosphorylation of DDR components, and viral inhibition by berzosertib in Figure 2. The authors seem to imply that DDR activation is the result of HDAC1/2 degradation, but causation has not been established. Do nondegradable HDAC1/2 identified in Figure 3 affect the ATM/ATR pathways?

      We acknowledge that causal inference required further experimental substantiation. To address this, we first established stable HDAC1/2-knockdown HeLa cell lines using validated siRNA constructs. Loss of HDAC1/2 resulted in marked elevation of histone acetylation marks—including H3K9ac, H3K27ac, H4K8ac, and H4K12ac—and robust induction of canonical DDR markers, specifically γ-H2AX foci formation, ATM phosphorylation, and ATR phosphorylation (Figure 1H and 2J). These phenotypes closely recapitulated those induced by HSV-1 infection, supporting a gain-of-function relationship. Critically, these data demonstrate that HDAC1/2 depletion alone is sufficient to drive both histone hyperacetylation and DDR activation—thereby reinforcing the proposed mechanistic cascade: HDAC1 nuclear export → proteasomal degradation → histone hyperacetylation → DDR pathway engagement. Second, to directly test whether HDAC1 degradation is functionally required for DDR activation during infection, we compared the effects of ectopically expressing wild-type HDAC1 WT versus the degradation-resistant mutant HDAC1 K74R in HSV-1-infected cells. Consistent with our model, HDAC1 WT expression partially attenuated both DDR activation and viral replication, whereas HDAC1 K74R exerted significantly stronger suppression of both endpoints—indicating that blocking HDAC1 degradation potently restrains the virus-induced DDR response and impairs viral fitness (Figure 3K and 3M). Collectively, these complementary loss- and gain-of-function experiments provide convergent evidence for a causal role of HDAC1 degradation in orchestrating the DDR during α-herpesvirus infection.

      (2) Line 30: "Strikingly, viral infection promoted nuclear export of HDAC1/2, followed by MDM2-mediated K63-linked polyubiquitination and proteasomal degradation in the cytoplasm". In Figure 5A, strong staining of HDAC1 is present in the nucleus, while a small fraction seems to be detected in the cytoplasm at 24 h infection with the vehicle treatment. Compared to vehicle treated mock infection, the HSV-1-infected cell has more HDAC1 staining, not less. The authors need to explain the contradictory results before reaching such a conclusion.

      With regard to the apparent discrepancy in HDAC1 subcellular localization depicted in Figure 5A, we have rigorously re-evaluated the immunofluorescence data. All samples were reimaged under strictly identical acquisition parameters—including exposure time, laser power, detector gain, and objective magnification—to eliminate technical variability. Quantitative analysis was performed on ≥100 randomly selected, non-overlapping cells per condition, with nuclear and cytoplasmic fluorescence intensities measured independently and normalized to yield the nuclear-to-cytoplasmic (N/C) ratio—a robust, internally controlled metric of HDAC1 redistribution. Complementing this, biochemical validation was carried out via subcellular fractionation followed by quantitative Western blotting, which confirmed a significant decrease in nuclear HDAC1 and concomitant accumulation in the cytoplasmic fraction at 24 h post-HSV-1 infection (p < 0.001 vs. vehicle-treated mock control). These findings fully corroborate our original model (Figure 5B and 5C). The elevated nuclear signal previously observed in Figure 5A arose from localized contrast enhancement applied during image processing—an artifact unrelated to biological abundance—and has now been replaced in the revised figure with raw, unprocessed images. Full details of imaging protocols, quantification methods, and statistical analyses are provided in the updated figure legend and Methods section. We sincerely apologize for any confusion this may have caused.

      (3) In Figure 5, LMB blocks the export of HDAC1 and has negative effects on HSV-1. LMB nonspecifically blocks the nuclear export of many factors. HSV-1 sensitivity to LMB has been investigated (PMID: 23740995). Is ICP27 involved in HDAC1 translocation? To clarify the causation, can the authors separate the nuclear and cytoplasmic fractions to detect the HDAC1/2 ubiquitination status? Does the K74R mutant prevent viral replication?

      This is a critical point. As outlined in our response to Reviewer 2, we will address this issue through three complementary experimental approaches: (1) generation of an HDAC1 nuclear export signal (NES)–deficient mutant to genetically abrogate its nuclear export (Author response image 2, Figure 5K); (2) rigorous nuclear-cytoplasmic fractionation coupled with immunoblotting to quantitatively assess HDAC1 subcellular distribution and site-specific ubiquitination. Importantly, our fractionation data confirm that HDAC1 ubiquitination is predominantly cytoplasmic (Figure 5F). Collectively, these experiments will rigorously distinguish direct HDAC1 modulation from indirect, LMB-mediated off-target effects—thereby ensuring the mechanistic specificity and interpretability of our conclusions.

      Furthermore, to functionally validate the physiological relevance of HDAC1 degradation in HSV-1 replication, we systematically evaluated the impact of HDAC1 WT versus its ubiquitination-resistant K74R mutant on viral gene expression and progeny production. Using qRT–PCR, we quantified mRNA levels of representative immediate-early (ICP0), early (ICP8), and late (gB) viral genes (Figure 3L); parallel plaque assays measured infectious virus yield. Strikingly, HDAC1 K74R overexpression conferred significantly stronger suppression of viral transcription across all kinetic classes and reduced progeny titers to a greater extent than HDAC1 WT (Figure 3M). These gain-of-function data provide compelling functional evidence supporting the central model that “HSV-1 promotes its own replication by inducing proteasomal degradation of HDAC1/2 to alleviate epigenetic repression of its genome.”

      (4) Figure 1E: The elevation of H4K8 and H4K12 seems to correlate to the disappearance of HDAC1/2 in the time points given. However, changes in H3K9, H3K27, and H3K56 occur between 0 and 6 hours, when HDAC1/2 levels show minimum changes. How are the experiments repeated? Can the bands be quantitated to better reflect a correlation?

      In direct response to the concerns raised, we have rigorously refined our experimental approach and expanded the dataset to strengthen mechanistic interpretation:First, to address the temporal heterogeneity inherent in low-multiplicity infections (MOI = l)—a condition that can obscure early virus–host regulatory dynamics—we performed synchronized time-course experiments in HeLa cells at a high multiplicity of infection (MOI = 5). Under these optimized conditions, HDAC1 and HDAC2 degradation is robustly detectable by 2 hours post-infection and reaches near-complete loss by 4–6 hours. Critically, this degradation kinetics precedes the onset of viral DNA replication (initiated at ~3–4 h) and coincides with the expression of true late viral proteins (e.g., ICP4), confirming that HDAC1/2 clearance is an active, early viral strategy—not a passive consequence of late-stage infection. These revised data position HDAC1/2 depletion as a causal, upstream regulator of the immediate-early-to-late transcriptional switch (Author response image 1).

      Second, to enable rigorous quantitative correlation, we conducted densitometric analysis of all histone acetylation marks (H3K9ac, H3K27ac, H3K56ac, H4K8ac, H4K12ac) and HDAC1/2 protein levels across three independent biological replicates. Quantified values are presented as mean ± SD beneath each corresponding blot panel in Figure 1, and kinetic profiles are visualized using normalized line graphs. Strikingly, the rapid hyperacetylation of H3K9, H3K27, and H3K56 during 0–6 hours exhibits strong temporal concordance with HDAC1 depletion—supporting a direct functional link between HDAC1 loss and locus-specific histone hyperacetylation on the viral genome. see Author response image 1

      Minor points:

      (1) Materials and Methods: How are mouse live tissues harvested, maintained, and infected?

      We sincerely apologize for the omission of methodological details in the original manuscript. In response to your insightful and constructive comments, we have comprehensively revised the "Methods" section (lines 92 to 102), adding detailed steps for in vitro processing of mouse tissues, including precise time points for sample collection after infection, and strictly defined in vitro infection conditions for herpes simplex virus type 1. These additions have significantly enhanced the reproducibility of the experiments, the rigor of the analysis, and the transparency of the techniques. We are deeply grateful for your thorough, meticulous, and highly valuable review comments, which have greatly improved the scientific quality of our work.

      (2) Figure 1A: class, not cluss.

      We sincerely apologize to the reviewers for this oversight and confirm that the error has been duly corrected in the revised manuscript.

      (3) Figure 2A, 2B: need a control to indicate which cells are infected.

      Regarding Figure 2A and 2B, we wish to clarify that the primary antibody against total H2AX was a mouse monoclonal antibody, whereas the anti-γ-H2AX antibody was a rabbit polyclonal antibody. Due to species incompatibility in multiplex immunofluorescence staining, simultaneous detection of γ-H2AX and viral proteins (e.g., HSV-1 ICP0 or PRV gB) using conventional two-color labeling was not feasible in those initial experiments. To rigorously address this concern, we performed additional, carefully controlled validation experiments: we conducted parallel immunofluorescence assays using the same rabbit anti-γ-H2AX antibody together with mouse monoclonal antibodies against HSV-1 ICP0 and PRV gB—employing appropriate species-matched secondary antibodies and stringent controls. As shown in the newly included data (Author response image 6), γ-H2AX foci intensity and nuclear signal intensity increased progressively in a time- and infection-dose-dependent manner, correlating robustly with viral antigen expression. These results provide direct, orthogonal support for our original conclusion that HSV-1 and PRV infection induce DNA damage signaling in host cells.

      Author response image 6.

      (4) Figure 4D, 4E; Why so small?

      We sincerely apologize for the confusion arising from the original figure layout. Figure 4D and 4E have now been revised to ensure accurate labeling, consistent scale bars, proper orientation, and full alignment with the corresponding descriptions in the text and legend.

      (5) Does the level of endogenous MDM2 change under the experimental conditions of Figure 4?

      As demonstrated in Figure 4C—representing an endogenous co-immunoprecipitation assay—the protein level of endogenous MDM2 is markedly increased following viral infection. This result is consistently observed across biological replicates and is quantified in the accompanying immunoblot analysis (Figure 4C, lower panel), confirming robust upregulation of MDM2 expression under the experimental conditions.

      (6) Line 359: "Our results are consistent across multiple cell types, including HeLa, 3D4/21, and murine liver". This statement is misleading. Only HeLa cells were used in Figures 3-5, which attempted to explain the mechanisms.

      We sincerely thank the reviewers for their careful reading and for identifying the inaccurate statements in the manuscript. All such statements have been revised, and the entire text has been systematically reviewed to ensure consistency, accuracy, and clarity across all sections.

      Reviewer #3 (Public review):

      The authors state that infection of cells by the alphaherpesviruses HSV-1 or PRV leads to a proteosome-dependent reduction in levels of HDAC1 and HDAC2 and that this leads to chromatin hyperacetylation, a DNA damage response, and greater replication of these viruses. Previously, other authors reported no change in levels of HDAC1 and HDAC2 after HSV-1 infection of human cells, but this paper is neither cited nor commented on in this new submission. The experiments are poorly designed. For instance, most of the time points analysed are way beyond the time needed for HSV-1 replication and are therefore not biologically relevant. The infections are done with a dose of virus that does not ensure that all cells are infected synchronously, but rather infection spreads from cell to cell with multiple rounds of replication. Some essential controls are missing. Additionally, this reviewer feels that the data presented do not support the conclusions drawn. Currently, links are not established between a reduction in HDAC1/ 2 and other phenomena such as hyperacetylation of histones, a DDR, and altered virus replication. The paper does not identify which HSV or PRV protein(s) induce reduction in HDACs, nor how the HDACs mediate antiviral activity; what are the HSV-1 or PRV protein targets? Lastly, the paper is not well prepared, and it does not adequately refer to prior literature.

      We sincerely thank the reviewers for their thoughtful, constructive, and highly valuable feedback. We deeply regret the shortcomings in our original submission—including incomplete literature coverage, insufficient mechanistic clarification, and gaps in experimental rigor—and fully acknowledge that these limitations affected the clarity and impact of our work. In response, we have comprehensively revised the manuscript: (i) expanded the literature review to incorporate key prior studies; (ii) added new experimental data—including time-resolved HDAC1/2 degradation assays, MDM2 knockdown/rescue experiments, and viral mutant analyses—to robustly substantiate the proposed mechanism; and (iii) rewritten the Results and Discussion sections to present a more precise, logically coherent, and evidence-based narrative. We are profoundly grateful for the reviewers’ time, expertise, and guidance, which have significantly strengthened this study.

      Reviewer #3 (Recommendations for the authors):

      Major points

      (1) Failure to cite prior literature, incorrect in-text citations, and mistakes in the bibliography.

      (a) The authors do not refer to highly relevant prior literature. For instance, a proteomic study of HSV-1-infected human cells showed that HDAC1 and HDAC2 were stable during high MOI. (Soh et al., Cell Rep, 2020. 33, 108235). This paper must be cited, and the difference between the findings of these authors and the current submission must be addressed.

      We sincerely thank the reviewers for bringing to our attention the study by Soh et al. (Cell Reports, 2020, 33: 108235). After a thorough review, we found that the paper titled "Temporal Proteomic Analysis of Herpes Simplex Virus 1 Infection Reveals Cell-Surface Remodeling via pUL56-Mediated GOPC Degradation" does not report any changes in the protein abundance of HDAC1 or HDAC2 in its full text and supplementary data. We did not detect any significant differential expression or degradation of HDAC1/2 in the main figures, supplementary figures, quantitative proteomic data tables (Supplementary Tables S1–S3), or through a full-text keyword search of the original literature.

      Furthermore, the other study that the reviewers might have in mind (Zhang et al., Cell Reports, 2019, 27: 1425–1438, DOI: 10.1016/j.celrep.2019.04.042) is about vaccinia virus (VACV) rather than HSV-1. It reports the degradation of HDAC5 and a transient downregulation of HDAC1 in the later stage of infection (see Figure 6E), but this downregulation did not reach statistical significance and was restored at subsequent time points. The study explicitly states that its findings do not apply to the HSV-1 infection system.

      Therefore, the study by Soh et al. (2020) does not provide experimental evidence that HDAC1/2 remain stable under high MOI HSV-1 infection. We have added this clarification in the revised manuscript and will more rigorously distinguish the specificity of HDAC regulation in different herpesviruses and poxviruses in the discussion section to avoid cross-reference confusion. We are grateful to the reviewers for their insightful questions, which prompted us to conduct a systematic review of the relevant literature.

      (b) In several instances, citations given in the text are not relevant to the statement made. For example, consider line 64 reference 12, line 67 reference 14, and line 76 reference 22. The sentence preceding reference 12 is about the control of cellular gene expression by modulation of chromatin: the title of reference 12 is "Functional interaction between class II histone deacetylases and ICP0 of herpes simplex virus type 1". The sentence preceding reference 14 is about type IV HDACs (HDAC11), but the title of reference 14 is "Seneca Valley virus 3C protease cleaves HDAC4 to antagonize type I interferon signaling". HDAC4 is a type II HDAC. The sentence preceding reference 22 is HDAC1 facilitates STAT1 phosphorylation and enhances interferon-stimulated gene (ISG) activation, thereby restricting influenza A virus replication. The title of reference 22 is "Positive role of promyelocytic leukemia protein in type I interferon response and its regulation by human cytomegalovirus". I have not examined every citation, so there may be other examples of this. A thorough check of every statement and associated reference is needed.

      We sincerely apologize for the oversight in verifying and updating the accuracy of the cited references. All citations have now been thoroughly reviewed and corrected to ensure full alignment between each statement and its supporting source. We are deeply grateful to the reviewers for their careful scrutiny and constructive feedback, which greatly strengthened the rigor and reliability of our manuscript.

      (c) The reference list is a mess. There are some references in which the given name of the authors is written, and the family name is abbreviated (incorrect), whereas in others the family name is written and the given name(s) are abbreviated (correct). I suspect this reflects the fact that in Mandarin, the family name is given first and the given names thereafter, whereas in English it is the other way round. But the inconsistency is careless, and modern reference management programs, such as EndNote, should eliminate these errors.

      We apologize for the errors in the original reference list and confirm that it has now been comprehensively revised: all entries have been uniformly reformatted in EndNote using the target journal’s official citation style; author names have been standardized to surname followed by initials (e.g., “Smith J”) in strict adherence to indexing and bibliographic standards; and all instances of underlined text, typographical inconsistencies, and grammatically incomplete sentences have been systematically identified and corrected.

      Overall, the failure to cite relevant literature, the incorrect citations, and the incorrectly prepared bibliography are indicative of an unacceptable level of care in the preparation of this paper. As another example, consider lines 94-99. Why is the text underlined? And the last sentence is incomplete and does not make sense.

      The underlining in lines 94–99 was originally intended to highlight the experimental treatment protocols applied to the mice; however, we acknowledge that this formatting choice was inappropriate for a formal manuscript and could impair readability and professionalism. We have therefore removed all underlining in this section and revised the text to clearly and explicitly describe the mouse treatment procedures in complete, grammatically correct sentences.

      (2) Virus infections have been done at 1 pfu/cell. This is a strange choice because the Poisson distribution shows that not all cells will be infected, and so after a first round of replication, the virus will spread sequentially from an infected cell to an uninfected cell. The fact that the level of HSV-1 protein gB is still increasing from 36-48 h pi (Fig. 1B) indicates that infection was very likely much less than 1 pfu/cell. Ditto for PRV (Figure 1). All infections must be redone at high moi (5-10 pfu/cell) so that the contribution of HDAC1 / 2 or the influence of specific pharmacological agents on the replication of virus in a single cycle can be determined. This is important because soluble factors released from infected cells can influence subsequent replication in the other cells. The release of these factors, or their influence on the uninfected cells, might be affected by the knockdown of HDACs or the addition of drugs tested. These concerns are largely eliminated by a high MOI (5-10 pfu/cell) so that all cells are infected synchronously.

      In direct response to the concerns raised, we have rigorously refined our experimental approach and expanded the dataset to strengthen mechanistic interpretation:First, to address the temporal heterogeneity inherent in low-multiplicity infections (MOI = 1)—a condition that can obscure early virus–host regulatory dynamics—we performed synchronized time-course experiments in HeLa cells at a high multiplicity of infection (MOI = 5). Under these optimized conditions, HDAC1 and HDAC2 degradation is robustly detectable by 2 hours post-infection and reaches near-complete loss by 4–6 hours (see Author response image 1). Critically, this degradation kinetics precedes the onset of viral DNA replication (initiated at ~3–4 h) and coincides with the expression of true late viral proteins (e.g., ICP4), confirming that HDAC1/2 clearance is an active, early viral strategy—not a passive consequence of late-stage infection. These revised data position HDAC1/2 depletion as a causal, upstream regulator of the immediate-early-to-late transcriptional switch.

      (3) A one-step growth curve for HSV-1 in human cells is about 12 h. So most of the time points measured (e.g., 24, 3,6 and 48 h pi) are not biologically relevant.

      In response, we have repeated the HSV-1 one-step growth curve experiment under rigorously controlled high-MOI conditions (2 PFU/cell) and extended the kinetic sampling to include precise, biologically informative time points: 0, 1, 2, 4, 6, 8, 12, and 24 hours post-infection—thereby capturing the complete early-to-late replication cascade while excluding late-phase secondary spread (Figure 2L,4J and 5J).

      Figure 2L,4J and 5J

      (4) Figure 2I and Figure 5F show that the titer of HSV-1 obtained after infection of cultured cells reaches ~10e10 pfu/cell. This is extraordinarily high in comparison to a large body of HSV literature. Usually, the titer would be between 10e7 and 10e8 pfu/cell. The authors should explain what feature of their cell culture system enables production of infectious virus up to at least 100-fold greater than that of other investigators. The data shown in Figures 2I and 5F for "vehicle" look identical. If this is the same experiment, this should be stated, and it would be much better to show different data sets.

      First, we clarify that while the “vehicle” control data in Figure 2I and Figure 5F appear visually similar, they derive from independent biological replicates conducted on separate days under identical experimental conditions—not from the same assay (Author response image 7). Second, regarding the elevated viral titers (~10^10 PFU/cell) observed in our assays relative to typical literature values (10^7–10^8 PFU/cell), we confirm that the virus stock used is HSV-1 strain F (generously provided by Dr. Chun-Fu Zheng, University of Calgary, Canada), with a validated starting TCID<sub>50</sub> of ~10^7/mL. All infections were performed in standard growth medium (DMEM + 10% FBS), ruling out medium-related artifacts. Critically, our initial time-course design extended beyond the single-cycle window—leading to secondary spread and cumulative amplification. To address this, we rigorously re-optimized the assay: using a high MOI of 2 PFU/cell and sampling precisely at 0, 1, 2, 4, 6, 8, 12, and 24 hours post-infection, we consistently recapitulated the kinetics and magnitude of viral production(Figure 2L, 4J and 5J). We sincerely apologize for the oversight in our original experimental design and thank the reviewers for prompting this essential refinement.

      Author response image 7.

      (5) Throughout the manuscript, there is little consideration given to the timing of the reduction in HDAC1/2 seen by immunoblotting, or the other changes such as hyperacetylation and DDR activation, in relation to the replication kinetics of the virus. These changes must occur early after infection to be able to influence virus replication. If they affect virus replication, what is the mechanism? At which stage during virus infection are they acting? What are the virus targets in HSV-1 or PRV-infected cells?

      First, we acknowledge the reviewers’ important point that the temporal relationship between HDAC1/2 depletion (as detected by immunoblotting), concomitant histone hyperacetylation, DDR activation, and HSV-1 replication kinetics was not explicitly addressed in the original manuscript. As these host modifications must occur early post-infection to mechanistically influence viral replication, we have now performed time-resolved immunoblotting following high-MOI HSV-1 infection (MOI =5) and confirmed that HDAC1/2 protein levels begin to decline within 2–4 hours post-infection—well before the onset of robust viral DNA synthesis (typically detectable after 4–6 hpi) (see Author response image 1). Second, consistent with our prior work [2] and independent reports [3], the DDR can activate the cGAS–STING pathway, leading to upregulation of type I interferons and proinflammatory cytokines—established antiviral effectors. Notably, our previous study demonstrated that pharmacological or genetic inhibition of BRD4 induces DDR-dependent cGAS–STING activation and potently suppresses PRV replication [4]. In contrast, α-herpesviruses—including HSV-1 and PRV—actively subvert this antiviral axis by targeting HDAC1 and HDAC2 for MDM2-mediated K63-linked polyubiquitination and proteasomal degradation. This targeted depletion promotes histone hyperacetylation, chromatin decompaction, and a transcriptionally permissive environment that facilitates efficient viral gene expression and replication. Finally, regarding the reviewers’ question about the specific viral determinant responsible for HDAC1/2 degradation, we fully agree that identifying the viral effector(s) is critical. Our ongoing studies are focused on systematically evaluating HSV-1 structural and non-structural proteins—including the E3 ubiquitin ligase activity of ICP0, the tegument protein VP16, and the viral kinase US3—to determine which factor(s) directly mediate MDM2 recruitment and HDAC1/2 ubiquitination. These experiments are underway and will be reported in future work.

      (6) The manuscript does not demonstrate that a reduction in HDAC1 or HDAC2 is responsible for the changes in hyperacetylation. The virus induces many changes in the cell; others could also directly affect hyperacetylation. What about the activity of acetylases?

      To rigorously establish causality between HDAC1/2 depletion and histone hyperacetylation, we performed loss-of-function experiments using siRNA-mediated knockdown of HDAC1 and HDAC2 in uninfected cells. Immunoblotting analysis revealed a significant increase in acetylation levels of histone H3 (at lysines K9, K27, and K56) and histone H4 (at K8 and K12) upon HDAC1/2 depletion (Figure 1H)—mimicking the hyperacetylation pattern observed during HSV-1 infection. Importantly, no corresponding increase in histone acetyltransferase (HAT) activity was detected in HDAC1/2-knockdown cells, as assessed by in vitro HAT assays using nuclear extracts and confirmed by unchanged expression levels of major HATs (p300). These data demonstrate that HDAC1/2 loss alone is sufficient to drive global histone hyperacetylation, independent of alterations in acetyltransferase activity—and thus support a direct mechanistic link between viral-induced HDAC1/2 degradation and the observed epigenetic changes.

      (7) The study needs to make knockout cell lines, lacking HDAC1 or HDAC2 or both HDACs, and then test virus replication after high MOI in these cells. If a difference is seen, the missing HDAC should then be reintroduced into the knockout cell line under an inducible promoter, and the replication of the virus checked in these cells with or without induction of the HDAC. Furthermore, HDACs have many interacting partners; so to prove that it is the histone deacetylase activity of the HDAC that is causing a change in replication, cell lines that inducibly express each HDAC (derived from the corresponding knockout cell line) with the key residues needed for catalytic activity mutated, should be constructed, and the replication of the virus tested. As an example, HDAC4 is antiviral - but does this is independent of the histone deacetylase activity (Lu et al., PNAS 2019).

      To effectively address this comment, we successfully constructed HDAC1 single knockdown, HDAC2 single knockdown, and HDAC1/HDAC2 double gene co-knockdown cells using siRNA technology mediated by transfection reagents. We first confirmed that co-knockdown of HDAC1/2 robustly enhances histone H3/H4 acetylation and activates the DNA damage response (DDR), as evidenced by increased phosphorylation of H2AX (γH2AX), ATM, and ATR—consistent with the epigenetic and DDR phenotypes observed during HSV-1 infection (Figure 1H and 2J). To further clarify the functional significance of HDAC1 degradation and its subcellular localization during HSV-1 infection, we reconstituted HDAC1 in HDAC1 stably knockdown cells with: (i) wild-type HDAC1 (HDAC1-WT), (ii) a degradation-resistant mutant with ubiquitination site mutations (HDAC1-K74R), and (iii) a nuclear retention mutant lacking the nuclear export signal (HDAC1-ΔNES). The analysis of viral replication kinetics revealed that HDAC1 knockdown significantly increased the yield of progeny HSV-1; however, the reconstitution of HDAC1-WT completely reversed this phenotype, restoring the viral titer to the level of the unknockdown control. Crucially, both HDAC1-K74R and HDAC1-ΔNES exhibited stronger anti-HSV-1 replication activity than HDAC1-WT (Figure 5K), indicating that blocking the ubiquitin-dependent degradation of HDAC1 or forcing its retention in the nucleus can more effectively inhibit viral proliferation. In conclusion, these genetic data strongly support the model that HSV-1 actively promotes the nuclear export and K48-linked ubiquitination-mediated proteasomal degradation of HDAC1 to relieve its transcriptional repression on viral gene expression, thereby optimizing its replication environment (Figure 5K). These findings demonstrate that HSV-1 exploits HDAC1 degradation—and its subsequent cytoplasmic translocation—as a proviral strategy, and that preserving nuclear HDAC1 activity is intrinsically restrictive to viral replication. Regarding the reviewers’ critical point on enzymatic specificity, we agree that definitive attribution to HDAC catalytic activity requires catalytically dead mutants (e.g., HDAC1-H141A/Y303F) expressed in isogenic HDAC1/2 knockout backgrounds under tightly regulated inducible systems. While such comprehensive genetic rescue experiments are technically demanding and beyond the scope of the current study, they represent a key focus of our ongoing work. Specifically, we are now systematically evaluating: (i) how HSV-1–mediated HDAC1 degradation mechanistically elevates histone acetylation at viral and host genomic loci; and (ii) the basis for functional divergence among HDAC family members—including differential expression, subcellular partitioning, interacting partners, and substrate selectivity—in regulating herpesviral replication. We deeply appreciate the reviewers’ insightful guidance, which has significantly strengthened the mechanistic rigor and conceptual framework of this study.

      (8) Figure 1C & D. The RT-qPCR data do not include analysis of a housekeeping gene against which the levels of mRNA for HDAC1 /2 can be compared. This is an essential missing control. The fact that the mRNA for gB is still increasing also confirms that the virus is still spreading, so the initial infection was most unlikely to have been at 1 pfu/cell.

      We confirm that all RT-qPCR data presented in Figure 1C and 1D were normalized to the endogenous control β-actin, and we have now explicitly stated this in both the Methods section and the corresponding figure legend. Regarding the observation that gB mRNA levels continue to rise over time, we agree that this reflects ongoing viral gene expression and progeny production—consistent with productive HSV-1 infection. However, this does not contradict our use of MOI = 1. In lytic herpesvirus infections, a single infectious particle initiates a cascade of gene expression, DNA replication, and assembly of new virions; therefore, increasing gB transcript levels across the time course are expected and reflect successful progression through the viral life cycle—not incomplete or suboptimal infection. Critically, Figures 1C and 1D were designed specifically to assess whether HSV-1 infection alters HDAC1/2 transcriptional regulation. The stable, MOI-independent expression of HDAC1/2 mRNA—despite progressive gB accumulation—demonstrates that the observed reduction in HDAC1/2 protein (shown in Figure 1A–B) is not due to transcriptional repression but rather results from post-translational mechanisms, such as virus-induced proteasomal degradation. This distinction strengthens our central conclusion: HSV-1 modulates host epigenetic machinery primarily via targeted protein destabilization, not transcriptional silencing.

      (9) Figure 1G. The authors should explain how intranasal infection with HSV-1 leads to infection of the liver 5 days later. HSV-1 is neurotropic, not hepatotropic. Which types of liver cells are infected? Are they the same cells as the cells in which there are changes in acetylation? No evidence is presented to show that infection and acetylation changes are linked.

      We sincerely apologize for the lack of sufficient detail in our prior response, which may have inadvertently increased the reviewers’ evaluation burden. Prior to in vivo experimentation, we conducted a systematic tissue tropism profiling of HSV-1–infected mice, quantifying viral protein expression across multiple organs by WB (Author response image 8). This unbiased, multi-modal assessment demonstrated that the liver exhibited both the highest viral protein abundance and the greatest viral genomic load, coupled with the most pronounced and histologically reproducible pathology—including dense inflammatory infiltration, hepatocyte vacuolar degeneration, and sharply demarcated foci of necrosis. Importantly, HSV-1 infection induced a robust and coordinated downregulation of HDAC1 and HDAC2 protein levels specifically in the liver; this effect was neither as pronounced nor as consistent in other tissues examined. In contrast, although the skin serves as the natural portal of entry for HSV-1, it displayed consistently low and highly heterogeneous viral protein expression in this systemic model—rendering it unsuitable for rigorous virological or immunological quantification. Accordingly, grounded in these empirical findings and aligned with established practices in models of disseminated herpesvirus infection [1]—where the liver is routinely prioritized as the primary site of pathogenesis and immune interrogation—we designated the liver as the principal organ for in-depth analysis of viral replication dynamics and host innate and adaptive immune responses.

      Author response image 8.

      (10) Figure 1G. The 3 replicates show large variations from one sample to another at the same time point. Consider H3K9, H4K12, H3K56, for instance. A statistical analysis is needed to determine if these changes are significant, but this was not included.

      We sincerely appreciate the valuable suggestion from the reviewers. To address this, we performed quantitative densitometric analysis on all Western blot images presented in the manuscript. The band intensities were normalized to each sample as a reference, and the relative protein levels were further normalized to the value of the control (time zero or untreated) sample, which was set to 1.0. The resulting normalized quantification values are now displayed directly beneath each blot lane in Figures 1–5.

      (11) Lines 233-5 and 254-8. These summary statements are not supported by the data presented. The authors have not established that these phenomena are linked.

      To directly address this point, we successfully constructed HDAC1/HDAC2 double gene co-knockdown cells using siRNA technology mediated by transfection reagents. We found that dual knockdown of HDAC1 and HDAC2 robustly enhanced histone H3/H4 acetylation and activated the DDR, as indicated by markedly increased phosphorylation of γH2AX, ATM, and ATR—phenotypes that closely recapitulate those observed during HSV-1infection (Figures 1H and 2J). These findings collectively establish a mechanistic link whereby HSV-1–mediated degradation of HDAC1/HDAC2 promotes histone H3/H4 acetylation and consequent DDR activation.

      (12) Line 246. Comet assay. A description of what is being measured is needed here. To virologists, comet assays often look at plaque morphology.

      We acknowledge the omission in our original description and have added the requisite clarification at line 258.

      (13) Figure 2I. Are these changes due to off-target effects of the drug? Cell viability assays are needed. And the authors should include an infection by a different virus that is not affected. Finally, the addition of the drug to cells lacking the target protein should be included - does this still influence virus replication?

      We sincerely thank the reviewers for their insightful suggestions. First, we assessed cell viability across all experimentally applied concentrations of Berzosertib using the CCK-8 assay and confirmed that none compromised cellular metabolic activity (Figure. 2K and 5A). Second, to control for virus-specific effects, we employed vesicular stomatitis virus (VSV) — a pathogen whose replication is independent of DDR activation — as a mechanistically distinct comparator. Consistent with this, Berzosertib treatment exerted no significant effect on VSV-GFP replication kinetics (Author response image 9), thereby excluding off-target contributions to the observed phenotypes.

      Author response image 9.

      (14) Figure 3A. MG-132 is clearly antiviral - look at the big reduction in ICP4 - so the changes in HDAC1/2 levels +/- drug likely simply reflect different degrees of infection. The 24 and 36 h pi timepoints are too late to be biologically relevant.

      While it is well established that MG132 exhibits broad-spectrum antiviral activity—including inhibition of Classical swine fever virus (CSFV) [5], Hepatitis B virus (HBV) [6], and, to a lesser extent, Hepatitis C virus (HCV) and Hepatitis E virus (HEV) replication in vitro—its primary and most rigorously validated application in mechanistic virology and cell biology remains the pharmacological inhibition of the 26S proteasome. By specifically blocking the chymotrypsin-like activity of the proteasome core particle, MG132 induces rapid accumulation of polyubiquitinated substrates, thereby enabling researchers to: (i) determine whether a given protein is degraded via the ubiquitin–proteasome system (UPS); (ii) assess the kinetics of its turnover; and (iii) distinguish UPS-mediated degradation from alternative pathways such as autophagy–lysosomal degradation [7]. Critically, MG132 is not used here as an antiviral agent per se, but as a precise biochemical tool to interrogate the degradation mechanism of HDAC1/2 during HSV-1 infection.

      (15) Figure 4. Since HDAC1 is degraded during infection, and the ligase responsible is claimed to be MDM2, HDAC1-MDM2 interaction would lead to the degradation of HDAC1/2, so less HDAC would be seen interacting with MDM2. To address this, the interaction analysis should be done in the presence of a proteosomal inhibitor such as MG132.

      We sincerely thank the reviewer for this insightful suggestion. To clarify: during HSV-1 infection, HDAC1 undergoes proteasomal degradation, a process that is strictly dependent on virus-induced ubiquitination—specifically, K48-linked polyubiquitination—which targets HDAC1 for recognition and binding by the E3 ubiquitin ligase MDM2. This ubiquitin-dependent interaction is a prerequisite for subsequent HDAC1 degradation via the 26S proteasome. While proteasome inhibition (e.g., with MG132) stabilizes HDAC1 protein levels and is routinely applied during co-immunoprecipitation (co-IP) sample preparation to prevent artifactual degradation, it concurrently dampens the physiological ubiquitination signal required for efficient MDM2–HDAC1 engagement. Consequently, although HDAC1 abundance increases under MG132 treatment, the functional ubiquitin-mediated interaction between HDAC1 and MDM2 is attenuated—not enhanced—making MG132-treated conditions suboptimal for interrogating the physiologically relevant E3–substrate interaction. Critically, our co-IP experiments—performed in the absence of MG132 but with careful attention to rapid lysis, cold buffers, and protease inhibitors—still robustly detect HDAC1–MDM2 association despite ongoing degradation, thereby providing direct biochemical evidence that HDAC1 engages MDM2 in a ubiquitin-dependent manner during active infection.

      (16) Figure 5. Lines 325-7. HSV replication is nuclear but requires export of mRNA and nascent capsids from the nucleus. So if any of these virus processes are influenced by leptomycin B, of course, the virus titer will be reduced. The link claimed is not proven。

      To better enhance the relevance of the article and further clarify the functional significance of HDAC1 degradation and its subcellular localization during HSV-1 infection, we reintroduced: (i) wild-type HDAC1 (HDAC1-WT), (ii) a degradation-resistant mutant with ubiquitination site mutations (HDAC1-K74R), and (iii) a nuclear retention mutant lacking the nuclear export signal (NES) (HDAC1-ΔNES) into HDAC1 stably knockdown cells. The analysis of viral replication kinetics revealed that HDAC1 knockdown significantly increased the yield of progeny viruses of HSV-1; however, the reintroduction of HDAC1-WT completely reversed this phenotype, restoring the viral titer to the level of the unknockdown control. Crucially, both HDAC1-K74R and HDAC1-ΔNES exhibited stronger anti-HSV-1 replication activity than HDAC1-WT (Figure 5K), indicating that blocking the ubiquitin-dependent degradation of HDAC1 or forcing its retention in the nucleus can more effectively inhibit viral proliferation. In summary, these genetic evidences strongly support the mechanism model that HSV-1 actively promotes the nuclear export and K48-linked ubiquitination-mediated proteasomal degradation of HDAC1 to relieve its transcriptional inhibition on viral gene expression, thereby optimizing its own replication environment.

      References:

      (1) B. Stefano et al., Two Fatal Cases of Acute Liver Failure Due to HSV-1 Infection in COVID-19 Patients Following Immunomodulatory Therapies. Clin Infect Dis 73, (2020).

      (2) L. Guo-Li et al., Inhibition of PARP1 Dampens Pseudorabies Virus Infection through DNA Damage-Induced Antiviral Innate Immunity. J Virol 95, (2021).

      (3) L. Tuo, C. Zhijian J, The cGAS-cGAMP-STING pathway connects DNA damage to inflammation, senescence, and cancer. J Exp Med 215, (2018).

      (4) W. Jiang et al., BRD4 inhibition exerts anti-viral activity through DNA damage-dependent innate immune responses. PLoS Pathog 16, (2020).

      (5) C. Yuming et al., MG132 Attenuates the Replication of Classical Swine Fever Virus in vitro. Front Microbiol 11, (2020).

      (6) W. Yi, L. Xiao-Liang, Y. Yong-Sheng, T. Zheng-Hao, Z. Guo-Qing, Inhibition of hepatitis B virus production in vitro by proteasome inhibitor MG132. Hepatogastroenterology 60, (2013).

      (7) H. Tianhua et al., Lipid peroxidation triggered by the degradation of xCT contributes to gasdermin D-mediated pyroptosis in COPD. Redox Biol 77, (2024).

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the Editor and Reviewers for their careful evaluation of our manuscript and for the constructive feedback. We agree with eLife’s overall assessment that, while profiling terminating ribosomes provides important insights into termination dynamics, additional clarification of the underlying mechanisms was needed. In response, we have focused our revision on three major conceptual points:

      (1) We have moderated our interpretation regarding the contribution of putative mRNA:rRNA interactions to sequence-specific termination pausing and clarified the limitations of the current evidence.

      (2) We have refined and clarified our model for the role of Rps26 in regulating translation termination.

      (3) We have expanded and strengthened the discussion of tissue-specific termination pausing, including its potential implications and current uncertainties.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors use high-resolution ribosome profiling (Ezra-seq) and eRF1 pulldown-based ribosome profiling (eRF1-seq) developed in their lab to identify a GA rich sequence motif located upstream of the stop codon responsible for translation termination pausing. They then perform a massively parallel assay with randomly generated sequences to further characterize this motif. Using mouse tissues, they show that termination pausing signatures can be tissue-specific. They use a series of published ribosome structures and 18S rRNA mutants, and eS26 knockdown experiments to propose that the GA rich sequence interacts with the 3′-end of the 18S rRNA.

      Strengths:

      (1) Robust ribosome profiling data and clear analyses clarify the subtle behavior of terminating ribosomes near the stop codon.

      (2) Novel termination or "false termination" sites revealed by eRF1-seq in the 5′-UTR, 3′-UTR, and CDS highlight a previously underappreciated facet of translation dynamics.

      Weakness:

      (1) Modest effects seen in ABCE1 knockdown do not seem to add up to the level of regulation. The authors state "ABCE1 regulates terminating ribosomes independent of the sequence context" on pg 9, and "ABCE1 modulates termination pausing independent of the mRNA sequence context" in the figure caption for Figure S4. Given the modest effect of the knockdown, such phrasing is most likely not supported. Further clarification of "ABCE1 plays a generic role in translation termination" is necessary.

      We acknowledge that the modest effects observed for ABCE1 are likely influenced by incomplete knockdown in HEK293 cells. Importantly, the increased ribosome density occurred at all stop codons rather than in a sequence-dependent manner, supporting the conclusion that ABCE1 functions broadly in termination rather than acting in a sequence-specific context. We have revised the manuscript to clarify this point and to temper our interpretation accordingly.

      (2) The authors propose that the GA rich sequence element upstream of the stop codon on the mRNA could potentially base pair with the 3′-end of the 18S rRNA. In the PDBs the authors reference in their paper and also in 3JAG, 3JAH, 3JAI (structures of terminating ribosomes with the stop codon in the A-site and eRF1), the mRNA exiting the ribosome and the 3′-end of the 18S rRNA are about 25-30 A apart. In addition, a segment of eS26 is wedged in between these two RNA segments. This reviewer noted this arrangement in a random sampling of 5 other PDBs of mammalian and human ribosome 80S structures. How do the authors anticipate the base pairing they have proposed to occur in light of these steric hindrances? RpsS26 is known to be released by Tsr2 in yeast during very specific stresses. Is it their expectation that termination pausing in human/mammalian cells happens during stressful conditions only?

      We agree that structural rearrangements in the absence of Rps26 remain speculative. In the revised manuscript, we have removed overly definitive language and clarified that, while Rps26 dissociation has been reported under stress conditions, its stoichiometry is unlikely to be exclusively stress-dependent. We now present this aspect as a working model supported by indirect evidence rather than a demonstrated structural mechanism.

      (3) The authors say, "It is thus likely that mRNA undergoes post-decoding scanning by 18S rRNA." (pg. 10). It is unclear what the authors mean by "scanning." Do they mean that the mRNA gets scanned in a manner similar to scanning during initiation? There is no evidence presented to support that particular conclusion.

      We appreciate the comment regarding the term “18S rRNA scanning.” We recognize that this wording may have been misleading and have revised the relevant text to more accurately describe post-decoding mRNA–rRNA interactions without implying an active scanning mechanism.

      (4) Role of termination pausing in the testis is highly speculative. The authors state: "It is thus conceivable that the wide range of ribosome density at stop codons in testis facilitates functional division of ribosome occupancy beyond the coding region." It is unclear what type of functional division they are referring to.

      We agree that the functional significance of testis-specific termination dynamics remains unclear. As multiple reviewers raised this concern, we have substantially expanded the discussion of tissue-specific termination pausing, explicitly outlining current limitations and framing this as an important direction for future investigation.

      Reviewer #2 (Public review):

      Summary:

      This paper presents results interpreted to indicate that sequences upstream of stop codons capable of base-pairing with the 3' end of 18S rRNA prolong the dwell time of 80S ribosomes at stop codons in a manner impeded by Rps26 in the 40S subunit exit channel, which leads to the proper completion of termination and ribosome recycling and prevents spurious translation of 3'UTR sequences by one or more unconventional mechanisms.

      Strengths:

      The standard 80S and selective eRF1 80S ribosome profiling data obtained using EZRA-Seq are of high quality, allowing the authors to detect an enrichment for purine-rich sequences upstream of stop codons at sites where termination is relatively slow and ribosomal complexes are paused with eRF1 still engaged in the A site.

      Weaknesses:

      There are many weaknesses in the experimental design, interpretation of results, and description of assay design and assumptions, the data obtained, and the interpretation of results, all of which detract from the scientific quality and significance of this work. In fact, a large proportion of paragraphs in the text and figure panels present some difficulty either in understanding how the experiment or data analysis was conducted or what the authors wish to conclude from the results, or that stem from an overinterpretation of findings or failure to consider other equally likely explanations.

      We appreciate the reviewer’s thoughtful evaluation and constructive suggestions. We recognize that our original description of the MPRA and reporter assay results may have lacked sufficient clarity, particularly regarding the sequence motifs associated with termination pausing. In the revised manuscript, we have carefully rewritten these sections to clarify the experimental design, data interpretation, and relationship between sequence context and termination dynamics. We believe these revisions address the reviewer’s concerns and improve the overall clarity of the manuscript.

      Reviewer #3 (Public review):

      Summary:

      This study from Jia et al carried out a variety of analyses of terminating ribosomes, including the development of eRF1-seq to map termination sites, identification of a GA-rich motif that promotes ribosome pausing, characterization of tissue-specific termination dynamics, and elucidation of the regulatory roles of 18S rRNA and RPS26. Overall, the study is thoughtfully designed, and its biological conclusions are well supported by complementary experiments. The tools and datasets generated provide valuable resources for researchers investigating the mechanisms of RNA translation.

      Strengths:

      (1) The study introduces eRF1-seq, a novel approach for mapping translation termination sites, providing a methodological advance for studying ribosome termination.

      (2) Through integrative bioinformatic analyses and complementary MPRA experiments, the authors demonstrate that GA-rich motifs promote ribosome pausing at termination sites and reveal possible regulatory roles of 18S rRNA in this process.

      (3) The study characterizes tissue-specific ribosome termination dynamics, showing that the testis exhibits stronger ribosome pausing at stop codons compared to other tissues. Follow-up experiments suggest that RPS26 may contribute to this tissue specificity.

      Weaknesses:

      The biological significance of ribosome pausing regulation at translation termination sites or of translational readthrough, for example, across different tissue types, remains unclear. Nevertheless, this question lies beyond the primary scope of the current study.

      We thank the reviewer for the positive assessment of our work. We agree that tissue-specific differences in termination pausing were insufficiently described in the original submission. In response, and in light of similar concerns from other reviewers, we have expanded the relevant sections in the main text and Discussion. We now more clearly articulate both the biological context and the current limitations, identifying tissue-specific regulation of termination as an open question and future research direction.

      Reviewer #4 (Public review):

      Summary:

      This manuscript by Qian and colleagues utilizes ribosome profiling, and reporter assays to dissect translation termination. Unfortunately, the data do not support the conclusions of the paper, controls are missing and several assays are not well validated and do not reproduce previous findings from others.

      Specific comments:

      Translation termination has been studied in several organisms including mammalian cells and yeast. In those cases what is analyzed is not the peak height at the stop codon, but rather the difference in the ribosome density before and after the stop. Thus, analyzing peak height is not validated. I understand that this is relevant only for the ribosome profiling experiments (and Ezra-seq) not the RF1 profiling. But much of the data was acquired that way.

      Moreover, the data do not reproduce previous findings and no effort is made to connect them to previous data. Previous data has shown that stop codon efficacy varies. This is not reproduced (S1C). Similarly, an effect from the +1 residue is not reproduced. The data isn't even stratified by different stop codons as previous work has shown that different surrounding residues have different effects in the context of different stop codons. Thus, none of the sequencing data is validated or trusted and does not reproduce previous findings.

      The GA-rich sequence identified by Ezra-Seq and RF1 seq is not the same and it differs from previous sequences (Wangen &Green).

      The authors claim that the majority of Rf1 peaks is at stop codons, but that is not true. It is only about 30% of the peaks. Also, not all mRNAs have peaks at the stop codons. That is at best problematic. Finally, there are mRNAs that are known to "suffer" from NMD, what do these look like in the Ezra-Seq and RF1-Seq? How about mRNAs that have programmed frameshifts? This raises questions on the validity of the eRF1 data.

      Figure 4: First, instead of M/P ratio, one should analyze M/M+P, to normalize out differences in the loading and effects from collisions, which are guaranteed to occur here, but not considered or analyzed. Second, the data are analyzed as if what matters are codons in the P and E site (and beyond, where there are definitely NOT recognized codons). While there is evidence for some interactions, one would think that an additional analysis based on sequence would be helpful. Also, the supplemental data indicates that very rarely are there reciprocal changes (as should be the case), and as seen for stop codons.

      Regarding the HiBit reporter assay: The two sequecnes clearly have effects on translation without considering stop codon context (Figure 4C), which need to be taken into account. Also, the effect from the sequences varies in the context of the assay in 4C and 4D (2-fold vs .5 fold), further questioning the assay. Moreover, the authors claim that re-initiation cannot account for Hibit levels, but that is clearly incorrect. The western in Figure 4E does not reproduce the data in 4D. While Hibit goes up (as in 4D, the putative GFP-fusion goes down. Finally, while the second reading frame should be more efficient is not explained and further argues for an artifact. Previous work (and work herein) suggests that read-through occurs equally in each reading frame. No controls for these assays are presented: e.g. stimulation by antibiotics, ABCE1 depletion, etc.

      Figure 5 has similar problems. I don't understand how the Figure in 5A is made, but when you overlay the cited structures on Rps26, the molecules are identical. I guess the authors used some fantasy to build non-existing sequences differently into the structure. There is no basis for that. In panel C and the same in Figure 7, the number of analyzed mRNAs varies. This could influence the outcome and the EXACT same set of mRNAs should be analyzed. But the main problem here is that the authors need to analyze readthrough and not peak height as detailed above. Essential controls are missing that show what fraction of the 18S rRNA is mutated. Previous work has shown that 2 nt truncated 18S rRNA is actively degraded. It is hard to believe how 15% of altered ribosomes can abolish 100% of the effect from the C-rich sequences. Important validation is missing: the authors should analyze rRNA sequences in their ribo-seq dataset to demonstrate that they have the mutated rRNAs, and that these enrich and de-enrich as predicted.

      In Figure 5-7 the authors develop a model that the sequence selectivity arises from base pairing between 18S rRNA and the mRNA. If so, then they should really stratify the data by number of WC pairs that can be formed. And only WC pairs, as GU pairs have a totally different geometry that will likely be discriminated against in this context. Also, the mutation is in a part of the helix that has no effect (Figure S3G). Thus, the data within the manuscript are inconsistent.

      Figure 6 does not agree with published data (Li et al., Nature 2022). Previous work did not show testis-depletion of Rps26 in purified ribosomes. This is the critical difference as the authors here did not purify ribosomes. Also, another Rps is an essential control, even if purified ribosomes are used. The validity of this dataset is thus questionable . Depletion from polysomes is hard to believe, as overall there is less signal in the polysomes.

      Figure 7 has similar problems as figure 5. Different pools of mRNAs are analyzed; peak height is not validated. Overexpression of Rps26 is not shown, as only Myc is shown, not Rps26. Beyond that, increased occupancy in ribosomes needs to be shown for the effect to come from ribosomes. Given how sick the cells are it is most likely that all effects are secondary and arise from whatever else is going on in the overexpression or depletion of Rps26. No controls are presented to show specific effects from Rps26.

      The authors need to check Rli1/ABCE levels in their cells. Their data have features that are indicative of low ABCE1 levels. These include a very small effect from ABCE1 depletion. These could be responsible for some of the effects they observe.

      We appreciate the reviewer’s engagement with our study and the opportunity to clarify several points.

      With respect to perceived inconsistencies with prior literature, we emphasize that our findings do not contradict established principles of translation termination. Rather, enabled by the development of eRF1-seq, we provide higher-resolution insight into termination dynamics that extends existing models. We have revised the manuscript to better contextualize our findings within prior studies and to avoid overstating novelty where continuity exists.

      Regarding the analysis of ribosome profiling data, we note that peak height and read density are widely used metrics for inferring ribosome dwell time and pausing. Nevertheless, we recognize that our original presentation may not have sufficiently explained this analytical framework. In the revised manuscript, we have clarified the rationale and interpretation of peak-based analyses, particularly in Figures 5 and 7 involving 18S rRNA mutants and Rps26 perturbation.

      Finally, we appreciate the reviewer’s comments concerning base pairing. We have carefully revised both the Results and Discussion sections to present mRNA–rRNA interactions as a supported but not definitively proven mechanistic model, clearly distinguishing experimental evidence from inference.

      We are grateful for the reviewers’ thoughtful feedback. We believe the revisions have strengthened the manuscript by clarifying interpretations, moderating mechanistic claims, and expanding discussion of tissue-specific regulation, while preserving the central contributions of the study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Some minor typos are present in the main text and methods section.

      We thank the Reviewer’s attention to detail in reviewing our manuscript. We have now thoroughly revised the main text and methods section.

      (2) S1I is missing or unlabelled.

      We are glad to have this opportunity to fix this mistake. Both S1I and S5D have now been added to the revised figures.

      (3) Could the authors clarify in the main text whether crosslinking was a step in the eRF1-seq protocol? Pg 5: "Without crosslinking, ribosomal proteins were minimally pulled down by the eRF1 antibody, confirming the transient nature of eRF1 binding."

      Yes, crosslinking is needed for eRF1-seq. We tried no-crosslinking but very little was pulled down, as stated in the sentence in Page 5.

      (4) Are termination events in the 5′-UTR or the CDS, as seen in the eRF1-seq data, also influenced by the GA-rich sequence? If the data is disaggregated into those two buckets, can you still pull out the motif?

      Yes, stop codons in 5’UTR and CDS share the same feature. However, the number of 5’UTR stop codons captured by eRF1-seq are too few to generate reliable sequence motif analysis.

      (5) Could the authors please clarify what peaks/fractions they are using as the monosome in Figure 4A? From the manner in which the red boxes are drawn on the sucrose gradient profile traces, it seems that the 40S, 60S, 80S monosome and half of the disome peak are included in the monosome fraction.

      The red box shown in Figure 4A is a bit misleading. For the massive paralleled reporter assay, we selected ribosome fractions based on the sucrose gradient tracing corresponding to monosome and polysomes, respectively. However, the fraction accuracy is not absolute as the fraction tube corresponding to monosome could contain traces of subunits as well as disomes. In practice, 40S and 60S are less concerned than disome, but the primary component is 80S ribosome.

      (6) On page 13, please cite references for Normal mode analysis.

      Normal Mode Analysis (NMA) using the Anisotropic Network Model (ANM) is a computationally efficient method for predicting large-scale, functional, and directional protein motions near equilibrium. We have followed the Reviewer’s suggestion by citing a review paper in the field of structural biology (Bahar, I. et al. 2005).

      Reviewer #2 (Recommendations for the authors):

      (1) The authors interpret the height of RPF peaks at stop codons in their Ribo-Seq data as an indication of pausing by ribosomes during termination, resulting from slow or inefficient decoding of the stop codon and peptide release; although it could equally result from slow recycling of the 60S subunit by ABCE1 following peptide release. Arguing against the latter possibility, they show later in the study that shRNA knockdown of ABCE1 has little effect on the stop codon RPF peaks; however, because the ABCE1 depletion does not elicit collisions near the stop codon in the manner observed in other studies, it appears that the ABCE1 depletion was insufficient to impair recycling substantially. The authors also don't attempt to support their interpretation by showing that depletion of eRF1 increases the stop codon peaks and produces collisions just upstream of the stop codon. They never specify with any precision whether it is stop codon recognition by eRF1, peptide hydrolysis, or recycling of the 60S subunit from the post-termination complex that is delayed, which is very unsatisfying.

      We agree with the Reviewer that the RPF density at stop codons only reflects the dwell time of terminating ribosomes. In fact, it is not possible to dissect molecular details from Ribo-seq data sets, same as interpreting other pausing events. Regarding ABCE1, we did observe the increased termination peak in cells with ABCE1 knockdown (Figure S4C). The lack of collisions is perhaps due to incomplete depletion of ABCE1. Notably, ABCE1 depletion selectively increased ribosome density at the –15 nt position, whereas the forward-shifted –12 nt peak was largely unaffected (Figure S4D). These results suggest that ABCE1 primarily facilitates late-stage termination or ribosome splitting, and its absence delays pre-termination progression. Nevertheless, the main focus of the study is to decipher the sequence context of termination pausing, which seems to be irrelevant to ABCE1. We thank the Reviewer for understanding.

      (2) They found enrichment for a GA-rich motif in the mRNAs with the largest stop codon peaks, which they attribute to its effect in slowing down some aspect of termination or ribosome recycling to increase the dwell time of the terminating ribosomes. They found no motif, however, in mRNAs containing the smallest RPF peaks at stop codon peaks, which presumably terminate more rapidly; even though they conclude later in the study from their massively parallel reporter assays (MPRA) that "C-richness" in the 9 nt 5' of stop codons enables rapid termination. The mRNAs with high pause scores at the stop codon that are enriched for the GA motif also show lower RPFs in 3'UTRs compared to the low pause score mRNAs, which they interpret to mean that long-lived termination complexes produce more efficient peptide termination and ribosome recycling, while short-lived complexes fail to be recycled and continue translation into the 3'UTR. However, because the 3'UTR reads are in all three frames, this could not occur simply by stop codon readthrough but would also require a frameshift upstream or at the stop codon itself to prevent termination and continued translation into the 3'UTR; and it could also arise from unconventional reinitiation by unrecycled post-termination complexes, which has been seen by others on inhibition of 60S recycling. The authors' interpretation is too simplistic.

      We thank the Reviewer’s summary about the sequence features controlling ribosome dwell time at stop codons uncovered by eRF1-seq. We are fully aware of the complex scenarios about 3’UTR translation, however, unconventional reinitiation cannot explain the results of the reporter assay shown in Figure 4D. Unlike frameshifting that generates prolonged products with mixed C-termini, reinitiation is associated with a new start. In Figure 4D, we observed products with C-terminal fused HiBiT, which cannot be explained by reinitiation. We thank the Reviewer for understanding.

      (3) They obtain support for the role of a GA-motif in pausing at stop codons from their selective ribosome profiling of eRF1-bound 80S ribosomes present at stop codons, finding a related GA-motif enriched at stop codons with high occupancies of eRF1-bound RPFs. However, once again, there is no C-rich motif enriched upstream of stop codons with low eRF1-bound RPF occupancies, at odds with later claims for such a motif. They ultimately propose that the GA motif pauses terminating ribosomes by base-pairing with the 3' end of 18S rRNA in the ribosome mRNA exit channel, principally utilizing two UU residues at the penultimate bases in the 18S rRNA that presumably base-pair with either A or G residues in the GA motif.

      The Reviewer might be confused by the results from Ribo-seq and massively paralleled reporter assay (MPRA). Ribo-seq data sets are limited to endogenous sequences that were shaped during evolution. In contrast, MPRA uses completely randomized sequences that offer unbiased analysis of sequence elements. The lack of C-rich motif in eRF1-seq data sets is due to the under-representation of such sequence elements in human genome. Perhaps this sequence bias is beneficial for termination fidelity by minimizing 3’UTR translation. We have further clarified this point in the revised manuscript.

      (4) They claim to have obtained independent confirmation of this last idea from their massively parallel reporter analysis (MPRA), in which sequences upstream of the stop codon of a uORF were randomized to determine those that appear to prevent translation downstream of the uORF and thereby place the mRNA in the monosome fraction versus those that allow downstream translation by any mechanism including leaky scanning of the uORF start codon, stop codon readthrough, or reinitiation (the assay doesn't distinguish between these mechanisms) and place the mRNA in the polysome fraction. In actuality, their results showed that the presence of only GGG triplets at any location in the 9 nt substantially prevents downstream translation, whereas only CCG and CCC proline codons enable downstream translation by one or more mechanisms. In view of their final model, it's very difficult to understand why GGG at any position would be able to base-pair with the U-U residues in the 18S rRNA when the stop codon is in the A site, and also why the many other triplets with two G's, two A's or an A and G base-all consistent with the GA-rich motif identified earlier-would not act similarly. Similarly, it's also puzzling that CCG and CCC can exert their effects at multiple positions upstream of the stop codon, and why the 7 other codons with two C's do not act similarly. Thus, it's unconvincing that a specific C-rich motif (which they refer to repeatedly but never identify) or even C-richness upstream of the stop codon confers elevated downstream translation. It's also important to note that the MPRA does not report on pausing at stop codons explicitly, only on whether ribosomes can be found downstream of the uORF stop codon, and assigning this outcome to the presence or absence of pausing during termination requires an ad hoc assumption that the authors have not identified as such.

      The Reviewer brought up excellent points in this comment regarding the MPRA result. Indeed, MPRA does not report ribosome pausing events as pointed out by the Reviewer. Additionally, MPRA is not designed to distinguish mechanisms underlying translational readthrough. As we mentioned above, both MPRA and Ribo-seq bear different experimental features that partly explain the similar, but not identical sequence motif uncovered by two assays. The prominent GGG motif identified by MPRA is intriguing, reminiscent of our prior study focusing on translation initiation (Jia et al. NSMB 2020). We propose that G-rich sequences upstream of stop codons form G-quadruplexes that block ribosome movement, resulting in monosome enrichment. Supporting this notion, the GGG motif was not identified by eRF1-seq, echoing the importance of using complementary experimental procedures in drawing conclusions.

      (5) They claim to confirm their conclusions from the profiling and MPRA data by measuring translation of the HiBiT sequence inserted downstream of the stop codon of the uORF in two reporters in which the upstream 9 nt contain either a single C-rich sequence or a single G-rich sequence. It's unclear how or why these two particular sequences were chosen. The G-rich sequence does not conform closely to either of the GA-motifs captured in the sequence LOGOs of Figures 1-2, and as noted above, there was no C-rich motif ever identified in these analyses. Thus, it's unclear whether the different effects of these two sequences are representative of sequences that pause or do not pause terminating ribosomes that they identified by the genome-wide analyses. In addition, given that the exact position of the GG or CC sequences relative to the stop codon doesn't seem to matter based on the MPRA data, it is actually possible to find the same number of base pairs with the 3' end of 18S rRNA for both of the two GA-rich and C-rich sequences analyzed in these reporter assays by sampling different registers of pairing between the mRNA and 18S rRNA. What is needed instead is be a systematic analysis using both the polysome:monosome assay, and the HiBiT translation assay of sequences that can pair perfectly with the 18S rRNA or contain increasing numbers of mismatches predicted to destabilize the putative helix that would be formed, and to determine whether the stability of the helices thus formed is highly correlated with the presence of the reporter mRNA in monosomes and with low HiBiT translation.

      We appreciate the Reviewer’s effort to improve our manuscript. The sequences inserted into the reporters were chosen based on several considerations. First, we chose the GA-motif rather than the G-rich sequences because the former represents physiological sequence element uncovered by eRF1-seq. As mentioned above, the G-rich sequences could form G-quadruplex artifacts. Second, the C-rich sequences were uncovered by both eRF1-seq (Figure 2D) and MPRA (Figure 4b). Third, only sequences top ranked were selected for the reporter assay. For the positional effects of inserted sequence elements, it is important to note that the proposed mRNA:rRNA interaction is not static because of the continuous mRNA movement along the channel. Instead of using sequences with perfect pairing, we have conducted experiments by placing the C-rich sequences at different positions of the insert. As shown in Figure S3H, the position relative to the stop codon does not seem to matter. In the revised manuscript, we have rephrased several sentences in the main text to avoid confusion.

      (6) They attempt to support their model by overexpressing a mutant 18S rRNA with mutations of the penultimate U-U residues to G-G, and present evidence that this decreases the stop codon RPF peaks on mRNAs rich in GA sequences upstream of the stop codons, and has the opposite effect on mRNAs that are C-rich; however, they never indicate the criteria used to assign mRNAs to these two bins, and whether it is based on the GA-rich motifs/LOGOs identified by genome-wide analysis or on the few triplets turned up by the MPRA. Clearly, it would be far better to conduct the same analysis of motif enrichment for high and low pause scores that produced the motif in Figure 1C and determine if the motif for high pausing switches from the GA-rich motif for WT 18S rRNA to a C-rich motif for the mutant, and vice versa for the low pause score mRNAs. It should also be noted that the C-rich sequence used in the reporter can form only 2 base pairs with the mutant 18S rRNA when the mRNA's C-C dinucleotide base pairs with the new G-G dinucleotide in rRNA, but it can actually form 4 base pairs with the WT 18S rRNA sequence in a different pairing register, undermining their interpretation of these data. Note also that there was no analysis done to determine what proportion of 40S subunits actually contain the mutant 18S rRNA, which is expected to be only a minor fraction under the best circumstances, and cannot simply be taken for granted, requiring a direct analysis of the sequences of the 3' ends of 18S rRNA in the cells expressing the mutant 18S.

      The Reviewer’s comment on 18S rRNA mutants are insightful. Given the low percentage of ribosomes incorporated with the rRNA mutants, it is not feasible to conduct motif analysis based on ribosome pausing at stop codons. As shown in Figure 5C, stop codon peaks are still evident after 18S mutant transfection albeit less prominent than the wild type. Notably, introducing 18S rRNA mutants into cells is not an easy task, and we have followed closely the protocol published previously (Burman and Mauro. NAR 2012) to obtain meaningful data. We believe (and hope the Reviewer will concur) that the experiment using the 18S rRNA mutants offers critical evidence in support of the mechanism.

      (7) They attempt to implicate Rps26 in the pausing by depleting or overexpressing (OE) the protein and comparing pausing at stop codons between the same two ill-defined GA-rich and C-rich bins of mRNAs mentioned above and by assaying the HiBit reporters. Again, they haven't determined whether the amount of Rps26 in mature 40S subunits is reduced or elevated compared to WT cells, and their interpretation of the OE data actually depends on the occurrence of 40S subunits lacking Rps26 in unstressed WT cells, which seems improbable and requires direct confirmation. Also, they haven't quantified the 80S peaks at the stop codons relative to the CDS reads immediately 5' of the stop codons, which varies with Rps26 OE versus the WT control, and doing so might well contradict their conclusion. Moreover, the C-rich and GA-rich HiBiT reporters behave identically rather than oppositely in response to Rps26 OE, which the authors fail to acknowledge or comment on.

      The Reviewer might be confused by the role of Rps26 partly due to the lack of clarity in our original description of the results. In yeast, Rps26 can dissociate from fully assembled 80S ribosomes under stress (Yang, et al. Sci Adv 2022). Therefore, although quantifying the Rps26 in mature 40S subunits is informative, it does not infer the composition of 80S ribosomes in cells with Rps26 depletion or overexpression. As pointed out by the Reviewer, we also noticed that, in cells with Rps26 depletion or overexpression, mRNAs with C-rich sequences showed no difference of ribosome density at stop codons. This is quite expected because C-rich sequences have minimal interaction with the 3’ end of rRNA. As a result, Rps26 depletion or overexpression is not supposed to affect ribosome dwell time at stop codons with upstream C-rich sequences. In contrast, only stop codons preceded with GA-rich sequences are influenced by Rps26 heterogeneity. In the revised manuscript, we have clarified this confusion in the main text.

      Additional specific comments

      (8) In the Summary statement: "We identify a sequence motif upstream of the stop codon that contributes to termination pausing, which was confirmed by massively paralleled": This is unjustified, as the MPRA showed only that a GGG triplet inserted anywhere in 9 nt 5'of the stop codon reduces ribosomes from traversing a stop codon either by blocking leaky scanning or reinitiation after an upstream uORF, and it is unclear why the position of this triplet does not matter nor why other GA-rich sequences capable of base pairing with the 3' end of 18S rRNA were not identified in the MPRA.

      As mentioned above, eRF1-seq and MPRA assays are complementary with advantages and disadvantages. Nevertheless, the Reviewer’s comments are well-taken and we have rephrased the Abstract of the revised manuscript.

      (9) A supplementary figure explaining EZRA-Seq would be very helpful.

      Since EZRA-seq methodology has been published (Mao, et al. NSMB 2023), we think a citation makes more sense. We thank the Reviewer for understanding.

      (10) The bottom plots/histograms of Figure 1A are very unclear. What is the y-axis of the bottom histogram, and relative to what elongating ribosomes have been analyzed?

      We apologize for the confusion in the histograms of Figure 1A. We stratified all mappable reads into footprints of initiating, elongating, and terminating ribosomes. Like many Ribo-seq results, the majority of footprints are of 29 nt length. If all three ribosome groups are of the same conformation, they are expected to have the same size distribution of the footprint length with the same bar height. It is true for initiating ribosomes (left) but not terminating ribosomes (right). We have now rephrased the figure legend in the revised manuscript.

      (11) Page 5: "A close inspection of stop codon footprints revealed an additional peak at -12 nt, which becomes more prominent when the reads are shorter (Figure 1B)." No explanation is offered for this finding. Do forward-shifted termination complexes have an empty A site owing to dissociation of eRF1? If so, they would be undetectable in eRF1-Seq data.

      Previous toe-printing assays have shown that eRF1 induces a forward movement of terminating ribosomes, shifting the leading edge from +13 nt to +15 nt (Pisarev, et al. Cell 2007). Moreover, single-molecule analyses have identified distinct pre- and post-termination phases catalyzed by eRF1 (Lawson, et al. Science 2023). Together, these observations suggest that the two 5’ end peaks correspond to pre- and post-terminating ribosome states, with the latter likely adopting a rotated conformation. We have revised the relevant paragraph in the main text.

      (12) Page 5: ". It is possible that the two distinct 5' end peaks represent pre- and post-terminating ribosomes, with the latter assuming the rotated conformation. We could not rule out the possibility that these terminating ribosomes have the stop codons at the P-site prior to disassembly." The logic here is difficult to follow.

      We have revised the relevant paragraph in the main text.

      (13) Figure 1C: provide coordinates relative to the stop codon on this motif.

      The motif analysis is position-independent and there is no coordinate on the logo plot.

      (14) Page 6: "This was not due to biased downstream sequences as the +4 nucleotide minimally affected the 3'UTR translation (Figure S1C)." The logic here is unclear.

      We have rephrased this sentence to “This effect could not be explained by downstream sequence bias, as the identity of the +4 nt had minimal impact on 3’UTR translation (Figure S1C).”

      (15) Page 6: "Like Ribo-seq, we also observed a forward shifting of post-terminating ribosomes from eRF1-seq (Figure 2C). " But by definition, they will have eRF1 in the A site, so why are they 26nt vs 29nt?

      Like many Ribo-seq results, the majority of footprints are of 29 nt length. However, ribosome populations with smaller footprint sizes are of physiological meanings, likely due to conformation changes.

      (16) Page 6 "In agreement with the Ribo-seq data sets, eRF1-seq revealed that not all the mRNAs exhibited eRF1 peaks at the annotated stop codons (Figure 2B), echoing the wide range of termination pausing." It should be determined whether eRF1 occupancy is correlated with 80S occupancy at stop codons in the standard Ribo-Seq. And if not, why?

      As shown in Figure 2B, there is a strong correlation between eRF1-seq and Ribo-seq in terms of termination pausing. However, the pausing index will be different between these two data sets due to distinct normalization. We thank the Reviewer for understanding.

      (17) Figure 2D: The plot on the left doesn't specify how far upstream the triplets can be from the stop codon. Is the LOGO significantly more similar to that shown in Fig. 1C than expected by chance alone?

      In Figure 2D, the codon frequency analysis is position independent. Similarly, the sequence logo in Figure 1C and Figure 2D is also position independent.

      (18) Page 7: ". Notably, three different stop codons show similar pausing features and sequence motifs (Figure S1G and S1I)." The figure citations here are incorrect.

      We apologize for the missing Figure S1I, which was also pointed out by Reviewer #1. We have now updated Figure S1 in the revised manuscript.

      (19) Page 7: The term "false termination" is a poor descriptor if termination doesn't occur.

      We have followed the Reviewer’s suggestion by replacing “false termination” with “failed termination”.

      (20) Page 8: "Consistent with previous reports 27, mutating the stop codon UAG abolished the reinitiation event that drives out-of-frame HiBiT translation (Figure 3E)." How is HiBit assayed? No details are given in the legend. This result doesn't confirm any of the actual eIF1 peaks upstream of stop codons, just that REI can occur at some level 5' of stop codons; and the eRF1 peak at the HiBit stop codon would be 3' of the peak at the main stop codon.

      HiBiT assay is a standard reporter like luciferase and Promega offers a detection kit, as described in the methods section. The result shown in Figure 3E is to confirm stop codon-associated reinitiation, which suggests that ribosomes migrated from the stop codon could contain eRF1 before reaching a start codon for reinitiation. We have revised this paragraph to avoid confusion.

      (21) Figure 4A: Unclear what position 0 to 6 in the bottom heat map corresponds to in the inserted 9 nt sequences. Are these codon positions vs. nucleotide positions? The legend lacks explanatory information.

      Figure 4A shows nucleotide positions (x axis) grouped by 3nt to reflect codon information (y axis). For the inserted 9nt random sequences, the last two nucleotides cannot be used because of the fixed nucleotides downstream of the insert. The same analysis has been reported in our prior study (Jia, et al. NSMB 2020).

      (22) Page 8: "For instance, codons enriched in frame 2 belong to NUA and NUG, another indication of in-frame stop codons (Figure S3B, bottom panel). " Need more or better explanation here.

      We have rephrased this sentence in the main text. “Codons enriched in alternative reading frames were also informative; for example, codons enriched in frame 2 predominantly belong to NUA and NUG, consistent with frameshifted presentations of in-frame stop codons (Figure S3B, bottom panel).”

      (23) "This is likely due to the faster turnover of these mRNAs because of 3'UTR translation". Need more or better explanation here.

      MPRA in Figure S3C showed that mRNA variants containing C-rich downstream sequence were depleted from both monosome and polysome fractions. Since 3’UTR translation is well-established to induce mRNA decay, it is possible that these sequences are under-represented due to mRNA turnover. We have added more explanations in this paragraph in the revised manuscript.

      (24) " Figure 4B: The logic and assumptions of this assay are not explained. How do ribosomes traverse the uORF, by leaky scanning or by stop codon read-through that is impeded by a ribosome stalled at the uORF stop codon? Presumably, it can't be read through as the uORF is out of frame and translation would likely terminate quickly.

      The rationale of Figure 4B is very similar to Figure 4A, except for the presence of the stop codon UAG. Under efficient termination, a monosome enrichment is expected, which could be promoted by termination pausing or structural hinderance by G-rich sequences. In contrast, stop codon readthrough or reinitiation would lead to polysome enrichment. We have thoroughly revised this paragraph in the main text.

      (25) Figure 4B results: It's unclear why M/P ratios are so low in Figure 4B vs Figure 4A as all constructs in 4B contain a stop codon and should have the high M/P ratios seen for the constructs in panel (A) with stop codons inserted. It's also unclear why the high M/P ratio should be so limited to GGG triplets vs. other triplets that conform to the GA-rich motifs identified above, and also why this triplet would not function at codon position 6. Similarly, it's unclear why only CCG and CCC and not CCU and CCA have an effect, and why only 3 of 9 codons with 2 or more C's have the effect, all suggesting that specific sequences and not just C-rich sequences are promoting read-through. Yet, no C-rich motif was discernible in the profiling experiments above.

      We appreciate the Reviewer’s careful reading of our manuscript. In profiling experiments shown in Figure 2, we did observe C-rich codons albeit with variations. Possible reasons include sequence differences between human genome and randomized sequence combinations. In addressing the Reviewer's question 23, we have thoroughly revised this paragraph in the main text.

      (26) Page 9: "These results are in line with the sequence specificity in termination pausing revealed by Ribo-seq and eRF1-seq." This is unjustified as the results in 4B are restricted to only GGG triplets rather than numerous triplets that equally conform to the AAGAAGA motif defined above.

      We apologize for the overstatement in this sentence. In addressing the Reviewer's question 23 and 24, we have thoroughly revised this paragraph in the main text.

      (27) Page 9: "This result is congruent with the MPRA assay, suggesting that the C-rich coding sequence preceding the stop codon not only reduces termination pausing, but also promotes downstream translation." This is unjustified as the single C-rich sequence chosen for the analysis in Figure 4C is not representative of the two C-rich triplets identified in Figure 4B, showing strong evidence of read-through.

      In Figure 4C, both C-rich and GA-rich sequences were chosen from shared elements between eRF1-seq and MPRA as they represent physiological sequences associated with termination pausing. The reporter assay is crucial in linking the lack of termination pausing with 3’UTR translation. We thank the Reviewer for understanding.

      (28) The analyses in Figures 4C-D suffer from a lack of the no-stop codon controls to allow the standard quantification of read-through as a percentage of continuous translation in the zero frame in the absence of a stop codon.

      The Reviewer might have missed the no-stop codon control in Figure 4C, which contains reporters with (bottom) and without (top) UAG stop codon. In Figure 4D, it is not feasible to include no-stop codon control for frameshifting reporters as the HiBiT value will be out-of-chart several orders of magnitude.

      (29) Page 10: "Therefore, the C-rich coding sequence triggers ribosome sliding at the stop codon, resulting in 3'UTR translation in all three reading frames." Sliding is an imprecise term. It is presumably a stop codon readthrough accompanied by frameshifting.

      We agree with the Reviewer’s suggestion and have replaced the word of “sliding” with “readthrough”.

      (30) Page 10: The citation to Figure S3H is incorrect, as there is no panel H.

      We are glad to have this opportunity to fix this error. We have now added panel H into the Figure S3 in the revised manuscript.

      (31) Page 10: "When the ribosome occupancy in the CDS was normalized, loss of ABCE1 led to a modest increase of stop codon peaks (Figure S4C)". Is this increase reproducible in replicates and statistically significant, as it seems very slight?

      The increased ribosome peak at stop codons in cells lacking ABCE1 is not significant, partly due to incomplete depletion of ABCE1 as shown in Figure S4A. Since ABCE1 is not the focus of this study, we did not attempt to knock out ABCE1, which could cause cellular toxicity.

      (32) Page 11: "Notably, the elevated ribosome density occurred at all stop codons, an indication of global effects." Where are the data substantiating this claim?

      We apologize for the confusion here. In the revised manuscript, we have deleted this sentence from the main text.

      (33) Page 11: "A closer look revealed that silencing ABCE1 increased the ribosome density at the -15 nt position". This claim is not convincing in the 29 nt read data, where it should be observed.

      We agree with the Reviewer that the increased ribosome density at the -15 nt position is more evident for shorter footprints. We have revised the sentence in the main text.

      (34) Page 11: "Since the 3' end of 18S rRNA contains a highly conserved U-rich sequence (GAUCAUUA), the GA- rich sequence element of mRNA could follow U:A and U:G base pairing near the exit site" (Figure 5A and S5A). By contrast, the C-rich sequence motif on mRNA would escape the 18S rRNA checkpoint, resulting in faster mRNA passthrough." This seems simplistic, as there would also be three G-A or A-G mispairings with 18S rRNA at other positions of the (G/A)AAGAAGA motif. Also unclear what the C-rich motif actually is, making it impossible to determine how many pairings it could make with the 18S rRNA sequence.

      Unlike base pairing on RNA structures, the putative rRNA:mRNA interaction is dynamic because of the continuous movement of mRNA along the ribosome channel. In fact, perfect base pairing might not be instrumental. Therefore, the difference between GA-rich and C-rich sequences is reflected in the accumulated effect. As mentioned above, the C-rich sequences are derived from both eRF1-seq and MPRA.

      (35) Figure S5B: Showing this sequence is misleading. While not described, it is presumably the DNA sequence of the plasmid, not the rRNA sequence, as there is 100% of the mutant sequence. They need to sequence the 3' end of rRNA isolated from ribosomes to confirm the presence of mutant ribosomes at appreciable levels.

      The Reviewer is correct that the sequences shown in Figure S5B are from the plasmids. To avoid such confusion, we have removed the sequences in the updated Figure S5B.

      (36) Page 12: "When mRNAs are stratified based on the sequence motif upstream of stop codons, we found that overexpression of the 18S mutant reduced the differential termination pausing between GA-rich and C-rich sequences (Figure 5C)". It is not explained what GA-rich or C-richness means precisely. Moreover, the same kind of analysis done in Figure 1C should have been conducted here to determine the LOGOs for high and low pausing for WT vs mutant 18S rRNA.

      We understand why the Reviewer repeatedly ask about the GA-rich and C-rich sequences, partly due to the lack of clarity in our original description of the analysis. The GA-rich transcripts were defined as those have the upstream 15-nt sequence with G or A nucleotides more than 65% (9 nt); whereas C-rich transcripts were defined as those with C more than 40% (6 nt). We have now updated the methods section in the revised manuscript.

      (37) Page 12: "Notably, the 3' end sequence of 18S rRNA is highly conserved (Figure S5D)". There is no Figure S5D in the figures.

      We are glad to have this opportunity to fix this error. We have now added panel D and E into Figure S5 in the revised manuscript.

      (38) Page 13: "Further supporting the sequence specificity of termination pausing, testis mRNAs with prominent stop codon peaks are enriched with GA-sequences upstream of the stop codon (Figure S6C). The same group of mRNAs, however, barely exhibit termination pausing in liver." Again, motif analysis of high and low pausing should have been done here.

      The motif analysis in mouse tissue samples is less informative because GA-rich sequences will be over-represented in testis, whereas the same group will be under-represented in liver. We had to select the shared mRNAs for comparative analysis. We thank the Reviewer for understanding.

      (39) Page 13: "While liver exhibited a similar distribution of Rps26 and RACK1 in polysome fractions, testis showed an evident depletion of Rps26 in polysome (Figure 6C). Notably, a substantial amount of Rps26 is present in the ribosome-free fraction of testis." They failed to normalize Rps26 levels in polysomes for bulk polysome levels, as indicated by the A260 tracings to determine if polysomes are depleted of Rps26, or rather, there is less polysomal Rps26 simply because polysomes are less abundant.

      We agree with the Reviewer’s notion regarding different polysome traces between testis and liver. Because the polysome volume is difficult to normalize, we used RACK1, a constitutive component of ribosome, to quantify the amount of polysome.

      (40) Page 14: "Indeed, normal mode analysis (NMA) by anisotropic network models suggests that, in the absence of Rps26, both the -3 to -9 extension of the mRNA and the 3' end of 18S rRNA can twist and approximate to each other with improved mutual parity (Figure 7B)." It is unclear what this means.

      Normal Mode Analysis (NMA) by Anisotropic Network Model (ANM) is a coarse-grained computational method used to study biomolecular dynamics by modeling proteins as a network of nodes connected by springs. Unlike the Gaussian Network Model (GNM), ANM calculates the full 3D directional preference of motion, enabling characterization of conformational changes, domain movements, and flexibility in large macromolecules. We have added a citation (Bahar, I. et al. 2005) in the revised manuscript.

      (41) Page 14: "To investigate whether Rps26 haploinsufficiency affects ribosome dynamics at stop codons, we knocked down Rps26 from HEK293 cells using shRNA (Figure S7A)". Haploinsufficiency properly refers to a heterozygous null/WT genotype, not shRNA knockdown.

      The Reviewer is correct in terms of haploinsufficiency. We have replaced the word of “haploinsufficiency” with “reduced Rps26 levels” in the revised manuscript.

      (42) Page 14: "The reciprocal change echoes the tissue-specific differences in initiation and termination (Figure 6A). " It's unclear why these peaks should be reciprocally related mechanistically, so examining changes in their ratio may not be incisive. Rps26 KD could reduce the efficiency of termination independently of pausing. And does Rps26 KD affect eRF1 occupancies in parallel with 80S occupancies?

      A prior study reported that Rps26 regulates translation initiation by recognizing Kozak sequence elements (Ferretti, et al. NSMB 2017). We therefore speculate that the role of Rps26 in termination might be correlated, although we don’t have direct evidence. We have further clarified this point in the discuss section of the revised manuscript.

      (43) Page 14: "The increased termination pausing, once again, primarily occurs at stop codons preceded by GA-rich sequences (Figure 7C)". No statistical analysis of replicates was done to see if the increase is significant, as it is quite small. They could have stratified mRNAs according to the number of base-pairs they can form with 18S rRNA rather than using this nebulous GA-richness, and see if the conclusion still holds.

      The metagene analysis shown in Figure 7C is standard for comparison of ribosome footprint distribution. We agree that the increase of termination peak at stop codons preceded by GA-rich sequences is not as striking as it should be, this is an underestimate because only a small fraction of ribosomes have sub stoichiometry of Rps26.

      (44) Page 14: "Remarkably, when mRNAs are stratified based on the sequence motif upstream of stop codons, we found that overexpression of Rps26 reduced the ribosome density (>50%) at stop codons preceded by the GA-sequence (Figure 7E)." They failed to normalize reads to the CDS occupancies to control for fewer ribosomes reaching the stop codons, especially considering that depletion of elongating 80S appeared to occur just upstream of stop codons on Rps26 OE. The same problem exists for the C-rich mRNAs. Also, their interpretation of the effects of Rps26 OE depends on there being Rps26-lacking 40S subunits in WT unstressed cells, which seems unlikely and has not been established directly. Finally, they didn't show increased Rps26 content in 40S subunits on Rps26 OE, which is also required.

      This question is the same as #7, which we have fully addressed in this letter (page 7).

      (45) Page 15: "To affirm the mechanistic connection between stop codon pausing and termination fidelity, we conducted HiBiT reporter assays that showed increased 3'UTR translation in cells with Rps26 overexpression (Figure 7F)." But both the C-rich and GA-rich reporters show increased expression on Rps26 OE. Why should that be if the C-rich sequences don't base pair with 18S rRNA in WT cells and are unaffected by Rps26 depletion? These data suggest that some other mechanism underlies the increased expression of the GA-rich reporters seen on Rps26 OE.

      The Reviewer’s concern is valid, and we agree that additional mechanisms might contribute to the increased reporter expression. The simplest explanation is that Rps26 overexpression promotes ribosome biogenesis, which globally increases mRNA translation. Supporting this notion, more polysome could be observed in cells with Rps26 overexpression (Figure S7E).

      (46) Page 15: "Without pausing at stop codons, terminating ribosomes are likely to undergo incomplete dissociation, resulting in continuous translation in 3'UTR." The language here is imprecise. Are they proposing reinitiation by unrecycled 80S ribosomes, or stop codon read-through with or without frameshifting, or both?

      This question is the same as #2, which we have fully addressed in this letter (page 3).

      (47) Page 15: "Importantly, lack of termination pausing leads to stop codon-associated random translation, giving rise to mixed C-terminal extension." Again, what does this mean? Read-through generally accompanied by frameshifting?

      Stop codon-associated random translation differs from ribosome readthrough, reinitiation, or frameshifting. We have extensively clarified this confusion in the revised manuscript.

      (48) Page 16: "For terminating ribosomes, the prolonged dwell time at stop codons offers an extended window for eRF1 loading, peptide cleavage, and ribosome recycling." This sentence is confusing because the eRF1-Seq data suggest that the pause occurs after eRF1 decodes the stop codon, with delayed peptide cleavage and recycling.

      We thank the Reviewer’s effort to improve our manuscript. We have rephrased the entire paragraph in the revised manuscript.

      Reviewer #3 (Recommendations for the authors):

      The manuscript is well-written, and the conclusions are overall well-supported by the data. I have only a few relatively minor questions and comments:

      (1) For termination sites overlapping with coding regions, the lack of 3-nt periodicity downstream of these sites could result from overlapping translation of multiple ORFs, rather than indicating that translation readthrough events can happen in multiple frames. Could the authors clarify this interpretation?

      We appreciate the Reviewer’s positive comments on our manuscript. The Reviewer is correct that overlapping ORFs would result in the lack of 3-nt periodicity. Although it is common for overlapping ORFs near the canonical start codons, ORFs overlapping the canonical stop codons are rare. Nevertheless, we have rephrased the statement in the revised manuscript.

      (2) The observation that multiple eRF1-seq peaks are located within CDS regions suggests that eRF1 may compete with A-site tRNAs during elongation. This is an interesting finding. Do the authors think this competition could lead to premature termination, or is it more likely to represent elongation pausing? Additionally, do the authors observe corresponding ribosome pausing peaks at these sites in conventional Ribo-seq data?

      The Reviewer’s comment on eRF1-seq peaks in CDS is insightful. We agree that pre-mature termination is possible because of competition. However, we do not observe corresponding ribosome pausing peaks in regular Ribo-seq, presumably due to low frequency of which events.

      (3) Regarding the regulation of ribosome pausing across tissue types, how robust are these results? For example, are the tissue-specific effects (such as stronger pausing in the testis) consistent among different mice or across age groups, given that many aspects of translational regulation are known to change with aging?

      We found that tissue-specific distribution of ribosome footprints is highly reproducible, especially liver and testis. Notably, the lack of termination peaks in liver is also reported by other independent studies (Gobert, et al. PNAS 2020), arguing that such effect is not a result of sequencing bias. We haven’t compared mice with different ages, but aging-associated translational regulation is an interesting topic awaits further investigation.

      Reviewer #4 (Recommendations for the authors):

      (1) Translation termination has been studied by ribose in several organisms, including mammalian cells and yeast. In those cases, what is analyzed is not the peak height at the stop codon, but rather the difference in the ribosome density before and after the stop. Thus, analyzing peak height is not validated. I understand that this is relevant only for the ribosome profiling experiments (and Ezra-seq), not the RF1 profiling. But the large majority of the data was acquired that way.

      With due respect, we disagree with the Reviewer’s point regarding how to study ribosome dynamics at stop codons. Comparing footprint density before and after stop codons does not infer dynamics of terminating ribosomes. By establishing eRF1-seq, we are for the first time able to analyze ribosome behaviors at stop codons, which represents a significant advancement of technological development.

      (2) Moreover, the data do not reproduce previous findings, and no attempt is made to connect them to previous data. Previous data have shown that stop codon efficacy varies. This is not reproduced (S1C). Similarly, an effect from the +1 residue is not reproduced. The data isn't stratified by different stop codons, and previous work has shown that different surrounding residues have different effects in the context of different stop codons. Thus, none of the sequencing data is validated or trusted and does not reproduce previous findings.

      We are certainly aware of previous findings regarding stop codon readthrough. We would like to emphasize that our findings do not contradict established principles of translation termination. Rather, enabled by the development of eRF1-seq, we provide new insights into termination dynamics that extend existing models.

      (3) The GA-rich sequence identified by Ezra-Seq and RF1 seq is not the same, and it differs from previous sequences (Wangen &Green).

      We don’t quite understand why the Reviewer is preoccupied with prior studies without accepting new results obtained from newly developed technology. The GA-rich sequences identified by Ezra-Seq and eRF1-seq are similar, albeit not identical. This is simply because eRF1-seq offers much higher resolution to reveal termination pausing than regular Ribo-seq.

      (4) The authors claim that the majority of Rf1 peaks are at stop codons, but that is not true. It is only about 30% of the peaks. Also, not all mRNAs have peaks at the stop codons. That is, at best, problematic. Finally, there are mRNAs that are known to "suffer" from NMD. What do these look like in the Ezra-Seq and RF1-Seq? How about mRNAs that have programmed frameshifts? The eRF1 data is invalid.

      The Reviewer is confused about the eRF1 peak density versus frequency, which has totally different meanings. Additionally, the Reviewer seems to be surprised that not all mRNAs have peaks at the stop codons. The differential ribosome dynamics at stop codons is an exciting feature previously unappreciated, rather than problematic. Regarding programmed frameshifting, we argue that such events are rare in mammalian cells.

      (5) Figure 4 has many flaws; it is hard to know where to start. First, instead of the M/P ratio, one should analyze M/M+P, to normalize out differences in the loading and effects from collisions, which are guaranteed to occur here, but not considered or analyzed. Second, the data are analyzed as if what matters are codons in the P and E site (and beyond, where there are definitely NOT recognized codons). While there is evidence for some interactions, one would think that an additional analysis based on sequence would be helpful. Also, the supplemental data indicate that very rarely are there reciprocal changes (as should be the case), as seen for stop codons. Thus, the assay is at best questionable and likely worse.

      The Reviewer appears to be unfamiliar with massively parallelled assay, which has been widely used to uncover sequence elements crucial in translational regulation. We urge the Reviewer to read our prior study using MPRA to investigate alternative translation initiation (Jia, et al. NSMB 2020). The similar approach has also been used to decipher 5’ UTR sequence elements in mRNA engineering (Sample, et al. Nat Biotech 2019).

      (6) Things do not look up for the HiBit reporter assay. The two sequences clearly have effects on translation without considering stop codon context (Figure 4C), which need to be taken into account. Also, the effect from the sequences varies in the context of the assay in 4C and 4D (2-fold vs. 5-fold), further questioning the assay. Moreover, the authors claim that re-initiation cannot account for Hibit levels, but that is clearly incorrect. The western in Figure 4E does not reproduce the data in 4D. While Hibit goes up (as in 4D, the putative GFP-fusion goes down. Finally, while the second reading frame should be more efficient, it is not explained and further argues for an artifact. Previous work (and work herein) suggests that read-through occurs equally in each reading frame.

      The Reviewer is confused about the HiBiT-based reporter assay shown in Figure 4C-4E. First, we have included important controls, i.e., same reporters without stop codons, to normalize sequence variation. Second, Figure 4C and 4D used totally different reporters and it is not appropriate to directly compare their values. Third, re-initiation events would not generate fusion proteins containing the N-terminal GFP. The Reviewer is encouraged to re-examine the results presented in Figure 4.

      (7) No controls for these assays are presented: e.g., stimulation by antibiotics, ABCE1 depletion, etc.

      We are not sure which assay the Reviewer is referring to. For reporter assays shown in Figure 4, we focused on effects of cis-sequence elements, rather than trans-acting factors. We thank the Reviewer for understanding.

      (8) Figure 5 has similar problems. I don't understand how Figure 5A is made, but when one overlays the cited structures on Rps26, the molecules are identical. I guess the authors chose to build non-existing sequences differently into the structure. There is no basis for that. In panel C, and the same in Figure 7, the number of analyzed mRNAs varies. This could influence the outcome, and the EXACT same set of mRNAs should be analyzed. But the main problem here is that the authors need to analyze readthrough and not peak height, as detailed above. Essential controls are missing that show what fraction of the 18S rRNA is mutated. Previous work has shown that 2 nt-truncated 18S rRNA is actively degraded. It is hard to believe how 15% of altered ribosomes can abolish 100% of the effect from the C-rich sequences. Important validation is missing: the authors should analyze rRNA sequences in their ribo-seq dataset to demonstrate that they have the mutated rRNAs, and that these enrich and de-enrich as predicted.

      The Reviewer’s comment on Figure 5A is baseless. As indicated in the Figure legend, Figure 5A was made from the existing cryoEM structure (PDB: 6ZMW). Regarding 18S rRNA mutants, we simply followed prior studies (Burman and Mauro. NAR 2012) and there is no evidence indicating degradation of such rRNA mutants. Given the low percentage of ribosomes incorporated with the rRNA mutants, the observed effect on termination pausing represent an underestimation, rather than an overstatement.

      (9) In Figures 5-7, the authors develop a model that the sequence selectivity arises from base pairing between 18S rRNA and the mRNA. If so, then they should really stratify the data by the number of WC pairs that can be formed. And only WC pairs, as GU pairs have a totally different geometry that will likely be discriminated against in this context. Also, the mutation is in a part of the helix that has no effect (Figure S3G). Thus, the data within the manuscript are inconsistent.

      As the Reviewer might be aware, GU pairs are commonly found in tRNA and rRNA structures. Since both WC and GU pairs contribute to mRNA:rRNA interaction, there is no point to stratify sequences based on different pairing format. Additionally, we would like to point out that the putative mRNA:rRNA interaction is not static, considering the continuous movement of mRNA along the ribosome channel.

      (10) Figure 6 does not agree with published data (Li et al., Nature 2022). Previous work did not show testis depletion of Rps26 in purified ribosomes. This is the critical difference, as the authors here did not purify ribosomes. Also, another Rps is an essential control, even if purified ribosomes are used. This dataset should not be shared. Depletion from polysomes is hard to believe, as overall, there is less signal in the polysomes.

      The Reviewer finally made a good point regarding Rps26 in testis. In our study, we did not separate different cell types such as spermatocytes and therefore we do not know which cell type dominantly influences termination pausing.

      Regarding varied Rps26 levels in different tissues, we noticed different polysome between testis and liver. Because the polysome volume is difficult to normalize, we used RACK1, a constitutive component of ribosome, to quantify the amount of polysome.

      (11) Figure 7 has similar problems to Figure 5. Different pools of mRNAs are analyzed; peak height is not validated. Overexpression of Rps26 is not shown, as only Myc is shown, not Rps26. Beyond that, increased occupancy in ribosomes needs to be shown for the effect to come from ribosomes. Given how sick the cells are, it is most likely that all effects are secondary and arise from whatever else is going on in the overexpression or depletion of Rps26. No controls are presented to show specific effects from Rps26.

      We are surprised that the Reviewer ignored the supplementary data that shows Rps26 levels. Regarding controls, it is not appropriate to use different ribosomal proteins because every ribosomal protein has its won functionality. We acknowledge that experiments by gene knockdown is not perfect, but the results are still informative especially when different mRNA pools from the same cells are compared.

      (11) The authors need to check Rli1/ABCE levels in their cells. Their data have features that are indicative of low ABCE1 levels. These include a very small effect from ABCE1 depletion. These could be responsible for some of the effects they observe.

      Once again, we are surprised that the Reviewer ignored the supplementary data that already shows ABCE1 levels in cells with or without ABCE1 knockdown (Figure S4A). Constantly addressing the Reviewer’s lack of careful reading of our manuscript is frustrateing. Nevertheless, we have thoroughly revised the entire manuscript by clarifying interpretations, moderating mechanistic claims, and expanding relevant discussion.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      Plasmodesmata are channels that allow cell-cell communication in plants; based on the functional similarities between facilitated transport within plasmodesmata and into the nucleus, the authors speculate that nuclear pore complex proteins might be involved in plasmodesmata function. If supported, this would transform our understanding of cell-to-cell communication in plants. The authors localize nuclear pore complex proteins to plasmodesmata using proteomics and heterologous overexpression; however, the data are incomplete since key controls for localization, functionality, and expression level of fluorescent protein fusions are absent.

      Thank you for the constructive reviews. We have tried to address the comments as outlined below. Specifically, we added new data to the manuscript with respect to the assessment of the protein levels of three independent stable Arabidopsis lines expressing NUP62-GFP from its own promoter using mass spectrometry quantification. These experiments were carried out to evaluate whether the observed PD localization of NUP62-GFP to peripheral puncta might be an artifact caused by inadvertent overexpression and resulting mistargeting. Quantitative analysis shows no indication for significant overexpression of NUP62-GFP.

      To assess whether the localization of NUPs is distinct from localization of an ER marker, we have now included a comparison of the NUP43-mVenus localization with that of the mCherry-HDEL luminal ER marker, revealing distinct localization patterns. The peripheral puncta thus do not appear to be due to simple ER accumulation.

      To evaluate whether the CPR5-mCitrine fusion is functional, we tested whether the fusion construct was able to complement the loss-of-function cpr5-1 mutant. In two independent complementation lines (cpr5-1/CPR5:CPR5-mCitrine), the roots of 14-d old seedlings were significantly longer compared to the cpr5-1 mutant, and four-week-old plants showed a more WT-like growth phenotype. Although we did not detect CPR5-mCitrine fluorescence, the construct appears to be able to restore the wild type phenotype, indicating that the lines express a functional CPR5 protein.

      We have restructured the figures and provided additional information in the figure legends.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Plasmodesmata are channels that allow cell-cell communication in plants; based on the functional similarities between facilitated transport within plasmodesmata and into the nucleus, the authors speculate that nuclear pore complex proteins might be involved in plasmodesmata function. In this manuscript, they localize nuclear pore complex proteins to plasmodesmata using proteomics and heterologous overexpression. They also document a possible plasmodesmata transport defect in a mutant affecting one nuclear pore complex protein.

      Strengths:

      The main strength of this manuscript is the interesting and novel hypothesis. This work could open exciting new directions in our understanding of plasmodesmata function and cell-cell communication in plants. They also localized many NUPs (12/35 Arabidopsis NUPs).

      Weaknesses:

      The main weakness of this manuscript is that the data are incomplete. While the authors appropriately and frequently acknowledge caveats to their data, two controls are essential to interpret the results that fluorescently-tagged NUPs localize to the plasmodesmata: (1) assessment of the expression level of these fluorescently-tagged NUPs to determine whether the plasmodesmata localization might be an overexpression artefact;

      As we outlined in the manuscript, we also considered the possibility that the peripheral localization could be a consequence of overexpression, in particular in the transient expression system. To be able to control the levels, NUP genes were expressed under the control of the b-estradiol-inducible XVE promoter which allows for b-estradiol dose dependent gene expression (Bashandy et al., 2015; Schlücking et al., 2013). We assessed the dependence of localization on expression levels by studying NUP localization under conditions of reduced estradiol concentrations for induction and shortened incubation time. We validated that the fluorescence was substantially reduced relative to the standard estradiol concentration experiments, however we still detected both nuclear and peripheral localization of the NUPs (Figure 4C-F).

      We also considered that in stable transformants the expression of one extra copy of a NUP62-GFP fusion under the control of the native promoter could cause a moderate overexpression and as a consequence lead to artifactual accumulation in the periphery (Figure 3C-E).

      To evaluate the level of NUP62-GFP fusion protein relative to untransformed controls, we quantified the levels of NUP62 in three independent transgenic fluorescent WT/NUP62p:NUP62-GFP Arabidopsis lines and in Arabidopsis WT using mass spectrometry (new Figure 3F). The new data indicate that there is no significant increase in NUP protein amounts in the lines expressing the fusion construct relative to WT.

      We now write in the revised manuscript (line 200-205):

      “NUP62 protein abundance in two-week-old cotyledons of the stable NUP62p:NUP62-GFP transformants was not statistically different to NUP62 protein levels in WT (Figure 3F). Notably, the punctate fluorescence at the cell periphery, encompassing both PD-associated and non-PD-associated localization, were not detectable or absent in roots and young leaves of four-day-old seedlings (Figure 3D). However, it cannot be excluded that the GFP fusion impacts NUP62 localization.” We provide a new Method section for the mass spec analysis of the cotyledons in lines 582-590.

      The use of antibodies in wild type tissue would be a potential way to avoid overexpression when trying to detect the localization of NUPs in planta. To investigate the localization of NUPs at physiological expression levels, we attempted to immunolocalize NUPs using antibodies. However, the anti-NUP antibodies available to us were not optimized for immunolocalization and we were unable to detect any fluorescence in the cells at the NPC nor the periphery.

      (2) assessment of the function of the fluorescently-tagged NUPs, either by molecular complementation of a knockout mutant phenotype or by biochemical methods to test whether the fluorescently-tagged NUP incorporates into nuclear pore complexes. Conducting these experiments for even one fluorescently-tagged NUP would substantially strengthen this manuscript.

      We agree with the reviewer that validation of the functionality of NUP fusion proteins would be valuable. Previously, C-terminally fused Arabidopsis NUPs, such as NUP93a-GFP, GP210-GFP, NUP58-GFP were reported to localize to the nuclear envelope when stably expressed in transgenic Arabidopsis lines (Tamura et al., 2010). As reported for transmembrane NUP GP210 and CPR5 fusion proteins (Gu et al., 2016; Tamura et al., 2010), C-terminally fused GP210 and CPR5 localized to the nuclear envelope but not to the nucleoplasm when expressed heterologously in N.benthamiana (see Figure 3-figure supplement 1). We found several soluble NUPs to also localize to the nucleoplasm (PpNUP98.1, PpNUP62, AtNUP62, AtHOS1) (Figure 1-figure supplement 1, Figure 3, Figure 3-figure supplement 1). Previous studies have reported that several FG NUPs (i.e. NUP98a/b or NUP62) and Y-complex NUPs (i.e. HOS1, NUP96, and NUP107) have been found to also localize in the nucleoplasm rather than specifically to the nuclear envelope when expressed as fusion proteins (Chen et al., 2023; Gallemí et al., 2016; Huang et al., 2024; Lazaro et al., 2012). Of note, for NUP98a, Gallemi and colleagues (2016) discussed the localization to the nucleoplasm as confirmation that, like vertebrate NUP98, Arabidopsis NUP98a is a dynamic NUP rather than just a key structural element of the NPC. HOS1 was reported to interact with ICE1, CO, FVE, and HDA6 in the nucleoplasm (Dong et al., 2006; Jung et al., 2012; Lazaro et al., 2012), indicating that HOS1 might dynamically shuttle between the nuclear pore and nucleoplasm, which could also explain the observed nucleoplasmic localization. In Drosophila, the FG-NUPs NUP98, NUP62, and NUP50 localized in the NPC, and also in the nucleoplasm and interacted with genes (Kalverda et al., 2010). The nucleoplasmic localization could thus have a functional relevance. Yet we cannot rule out, whether soluble NUPs mislocalize in overexpression conditions as we state multiple times in the manuscript.

      For this revision, we generated two new independent transgenic Arabidopsis lines stably expressing CPR5-mCitrine under control of its own promoter in the cpr5-1 mutant background (cpr5-1/CPR5p:CPR5-mCitrine). The roots were significantly longer in the two independent transgenic cpr5-1/CPR5p:CPR5-mCitrine Arabidopsis lines compared to the cpr5-1 mutant, and four-week-old plants showed a more WT-like growth phenotype (new Figure 7-figure supplement 1, G–I). However, we could not detect fluorescence in the 10-14 day old seedlings, which could be due to a variety of reasons, such as cleavage of the FP and degradation of the FP without accumulating elsewhere in the cells.

      In the new manuscript we write in lines 275-283:

      “To assess whether the CPR5-mCitrine fusion protein is functional in Arabidopsis, we tested whether CPR5p:CPR5-mCitrine (including all introns) expression in the cpr5-1 mutant background results in a rescue of the severe growth phenotype of the cpr5-1 loss-of-function mutant (Bowling et al., 1997). Indeed, roots were significantly longer in the two independent transgenic cpr5-1/CPR5p:CPR5-mCitrine Arabidopsis lines compared to the cpr5-1 mutant, and four-week-old plants showed a more WT-like growth phenotype (Figure 7-figure supplement 1, G–I). However, we could not detect fluorescence in 10-14 day old seedlings, which could be due to a variety of reasons, such as cleavage of the FP and degradation of the FP without accumulating elsewhere in the cells. The lack of fluorescence in the transgenic lines requires further investigation.“

      Reviewer #2 (Public review):

      Summary:

      The authors aim to address whether nuclear pore complex components localize and function at PD in plant cells to mediate cell-to-cell communication.

      Strengths:

      (1) Novelty and Significance:

      The core hypothesis, drawing parallels between PD and NPC transport, is highly original and addresses a critical gap in understanding plant intercellular communication. The idea that phase-separated domains formed by FG-NUPs could act as diffusion barriers at PD offers a plausible and sophisticated explanation for their complex transport properties, including size exclusion and facilitated translocation. This could fundamentally change how we view PD function.

      (2) Comprehensive Evidence:

      The study employs a rigorous and diverse set of experimental approaches, including a comprehensive bioinformatic analysis of both moss and Arabidopsis NUPs in available PD proteomic datasets, extensive imaging analysis of Nup localization in vivo, and functional transport assays using a loss-of-function nup mutant (cpr5). The transport assay is particularly important to provide functional evidence linking CPR5 to PD-mediated transport. The finding that callose levels were not significantly different in cpr5 mutants under these conditions is helpful and supports a distinct, callose-independent mechanism of transport regulation.

      (3) Objectivity:

      The authors are forthright in discussing the limitations and potential artifacts of their own data, clearly distinguishing between observations and definitive conclusions.

      Weaknesses:

      While the claims are generally justified as hypotheses or consistent observations, the authors themselves extensively detail the caveats, which are worth reiterating for clarity:

      (1) Potential Overexpression Artifacts in Localization:

      Although efforts were made to control expression levels, the authors acknowledge that transient overexpression could still lead to NUP accumulation at PD, either as a physiologically relevant accumulation under excess conditions or due to mis-targeting, or even as storage depots. The resolution of confocal microscopy also does not allow for a definitive conclusion on the nature of the location.

      We would like to add that in addition to the experiments using estradiol-controlled transient overexpression for localizing NUP fusions, we also provided localization data obtained from Arabidopsis transformants that stably express one extra copy of a NUP62-GFP fusion under the control of the native promoter. In cotyledons, NUP62-GFP localized to the nucleus and in the periphery, and in many cases to PD (Figure 3C-E). In the course of the revision we tested whether the extra copy of NUP62 could cause overexpression that might lead to artifactual accumulation in the periphery.

      To evaluate the level of NUP62-GFP fusion protein relative to untransformed controls, we quantified the levels of NUP62 in three independent transgenic fluorescent WT/NUP62p:NUP62-GFP Arabidopsis lines and in Arabidopsis WT using mass spectrometry (new Figure 3F). The new data indicate that there is no significant increase in NUP protein amounts in the lines expressing the fusion construct relative to WT.

      We now write in the revised manuscript (lines 200-205):

      “NUP62 protein abundance in two-week-old cotyledons of the stable NUP62p:NUP62-GFP transformants was not statistically different to NUP62 protein levels in WT (Figure 3F). Notably, the punctate fluorescence at the cell periphery, encompassing both PD-associated and non-PD-associated localization, were not detectable or absent in roots and young leaves of four-day-old seedlings (Figure 3D). However, it cannot be excluded that the GFP fusion impacts NUP62 localization.“ We provide a new Method section for the mass spec analysis of the cotyledons in lines 582-590.

      (2) Proteomics Purity:

      The authors note that the presence of NUPs in PD fractions/proteomics cannot definitively rule out contamination, as PD cannot currently be purified to absolute homogeneity and is often contaminated with other organelles, including the nucleus.

      We would like to add that despite their low abundance in plant cells, NUPs were found to be enriched in cell wall, and PD fractions relative to total cell extracts (revised Figure 2-supplement 2). To evaluate whether NUP enrichment might be a consequence of contamination by nuclear fractions, for the revision, we evaluated the enrichment of nucleolar proteins and histones. As shown in the revised Figure 2–figure supplement 2, other nuclear proteins did not show a significant enrichment, supporting the notion that NUPs were specifically enriched in PD fractions, consistent with the localization of NUP-FP fusions. We note however, that these data do not demonstrate unambiguously that NUPs are bona fide PD components.

      (3) CPR5 Mutant Interpretation:

      While cpr5 mutants exhibited reduced macromolecular transport, the authors state that they cannot exclude that the reduced transport is due to secondary effects in the cpr5 mutants, which show rather severe phenotypic defects. This is an important distinction, as CPR5 has known roles in defense responses and hormone signaling that could indirectly influence PD integrity, independent of callose deposition. The lack of effect on small molecule transport is a good control, but the broader pleiotropic effects of cpr5 mutants remain a consideration.

      We agree with the assessment of the reviewer. The mutant is compromised in many ways and thus the effects we observe could be indirect. This is stated also in the manuscript (lines 314-317).

      (4) Conceptual Distinction between NPC and PD:

      The authors correctly point out that while similarities exist, the physical assembly of NUPs at PD must differ from that at the NPC due to the presence of the desmotubule and smaller cytoplasmic sleeve width at PD. Moreover, nucleocytoplasmic transport depends on karyopherin proteins that interact with the NPC central channel to complete the transport. Yet the role of karyopherins in this case is not clear. Therefore, the proposed "PD pore complex" may bear some NPC features, but not be identical.

      Reviewer 2 summarized the key concerns that we highlighted and discussed in the manuscript, which addressed differences in PD and NPC architecture. In particular, we noted that one of the major differences in PD is the presence of the desmotubule (in lines 370-372). We also highlighted that we did not detect all NUPs at PD (in lines 375-376). While a negative result, this observation may also be consistent with differences regarding the assembly of NUPs in or near PD vs the NPC. We fully agree with the reviewer that the proposed “PD pore complex” may be not identical to the NPC, and we also discussed that the NUPs seen at PD could represent sites of accumulation in the ER near PD.

      Reviewer #3 (Public review):

      Summary:

      This manuscript presents a step towards testing the hypothesis that plasmodesmata have homology to nuclear pores. The similarities between the two structures have long been noted as both structures allow the transport of proteins and nucleic acids, and both structures are composed of curved membranes. The manuscript has identified nuclear pore proteins (NUPs) in plasmodesmal protein fractions and uses live imaging in a non-endogenous system and functional assays of a mutant to propose that this might be a bona fide association.

      The conclusions the authors seek to draw are that: NUPs are present in plasmodesmal protein fractions; NUPs localise at plasmodesmata; NUPs might form a pore-gating complex at plasmodesmata, regulating non-specific (2xGFP) and specific (SHR) transport through plasmodesmata

      The authors then use these conclusions to propose the possibility that phase separation mediates transport through plasmodesmata. If there is phase separation at plasmodesmata or a nuclear pore-like complex, it would revolutionise the community. However, this data is insufficient to act as a cornerstone for such a discovery.

      Strengths:

      The strength of the manuscript lies in the boldness and novelty of the idea.

      Weaknesses:

      The weaknesses lie in the lack of informative controls. The authors' own assessments of their data suggest they agree with this - in their abstract alone, they point out that the transport defects they observe might be off-target effects, and suggest there is a requirement in the future to determine whether the NUPs are bona fide PD components.

      Across the proteomic and live imaging experiments, the conclusions could be stronger if they compared the NUP localisation and accumulation with ER proteins - the question of whether NUPs behave like other ER proteins is not addressed. As NUPs reside in the nuclear envelope, continuous with the ER, and the ER traverses plasmodesmata, a comparison between the NUPs and ER proteins would be extremely informative.

      We agree with the comments of the reviewer. To assess whether NUPs show localization patterns that are similar to ER proteins, we transiently co-expressed NUP43-mVenus fusions with the mCherry-HDEL luminal ER marker in N.benthamiana. Comparison of the localization patterns reveals distinct patterns of NUP43-mVenus and mCherry-HDEL (see the new Figure5, new Figure 5-figure supplement 1). NUP43-mVenus appears to be associated with the ER, however restricted to subregions that partially overlay with aniline blue-labeled pit fields (new Figure 5, new Figure 5-figure supplement 1).

      In the new version of the manuscript, we write (lines 209-214):

      “We assessed whether NUP localization is distinct from ER localization in N. benthamiana leaves that heterologously co-expressed NUP43-mVenus and the ER luminal marker mCherry-HDEL. The localization patterns of NUP43-mVenus and of the mCherry-HDEL luminal ER marker were clearly distinct (Figure 5, Figure 5-figure supplement 1). NUP43-mVenus may be associated to the ER, however restricted to subregions of the ER, which partially overlay with aniline blue-labeled pit fields (Figure 5, Figure 5-figure supplement 1).”

      Regarding the proteomic identification of NUPs in plasmodesmal fractions, the authors place significant weight on their own metric for PD enrichment, the PD score. As I understand it, this a metric derived from addition of two factors: a two component enrichment score that is the difference between intensity of peptides of a given protein in the PD fraction and cell wall fraction, added to the difference between intensity of peptides of a given protein in the PD fraction and total cell fraction, and a feature score that is a factor that describes representation of protein domains contained in said given protein in the plasmodesmal fraction relative to the representation of that domain in proteins in the whole proteome. The features chosen for analysis are not indicated, and the feature factor, as I understand it, is a score common to all proteins with a given feature. While each of the factors carries a measure of meaning and information, I do not understand how adding them is mathematically or biologically meaningful.

      The feature score was defined based on PD proteome analysis previously described (Gombos et al., 2023). Features of known PD proteins were extracted and weighted against the entire Arabidopsis proteome. Structural features included Pfam domains PF00722 (GHL), PF06955 (XET_C), PF08372 (PRT_C), PF00335 (Tetraspanin), and PF00168 (C2 domain). Subcellular localization features included plasma membrane (PM), endoplasmic reticulum (ER), extracellular space (EX), and cell wall (CW). Functional features were assigned according to MapMan categories bin 10, 15, 26, and 30. To clarify the approach, we added a more detailed explanation to the feature score in the Methods of the revised manuscript.

      We agree with the reviewer that experimental values and feature factors represent two distinct, independent parameters. The PD score aims to identify proteins that are not only experimentally enriched in the plasmodesmal fraction but also share structural features characteristic of bona fide plasmodesmata-associated proteins, reducing the number of false positive candidates driven by either parameter alone in PD proteome lists. From a mathematical standpoint, we combined the two normalized factors in the PD score by summation, treating them as contributing equally to a protein’s PD association tendency.

      Conclusion:

      The conclusions of the study are not fully supported in the absence of ER controls. Of note, the imaging is ambiguous because the proteins do not show a discrete plasmodesmal association. This is a localisation reminiscent of cortical ER association and needs to be further investigated to determine whether it is a true and specific plasmodesmal association.

      We agree with the reviewer’s comments. In the revised version of the manuscript, we have now included a comparison of the NUP43-mVenus localization with that of the mCherry-HDEL luminal ER marker, which reveals distinct localization patterns (see new Figure5, new Figure 5-figure supplement 1). NUP43-mVenus may be associated with the ER; however, NUP43 is restricted to subregions of the ER, which partially overlay with aniline blue-labeled pit fields (new Figure 5, new Figure 5-figure supplement 1). Whether NUP localization is distinct from cortical ER requires further investigation.

      The conclusions drawn from Figure 1, Figure Supplement 4 are confusing. The text describing this data says that "NUPs were enriched in cell wall and PD fractions compared to total cell extract, while the abundance of other nuclear envelope proteins was unaffected by the PD purification and showed no enrichment in PD fractions". However, the data show that there is no difference in the normalised protein intensity for the NUPs across TC, CW, and PD fractions. The only sample that shows enrichment in PDs is the PDLP/MCTPs.

      To address this point, we rephrased the text (line 146-152). Among all NUPs identified in our PD proteome, 75% were more abundant in PD fractions (Figure 2-figure supplement 2), exceeding the proportions observed in TC (60%) and CW (~50%) fractions. In contrast, other nuclear proteins such as nuclear envelope proteins, nucleolar proteins, or histones showed PD intensities that fell within or overlapped the ranges observed in TC or CW. The native abundance of NUPs was lower compared to that of proteins from other compartments, which may explain why the enrichment significance was not statistically significant (p = 0.24 for PD vs. TC). By comparison, the corresponding p-values for other nuclear compartment proteins were higher, ranging from 0.5 to 0.9.

      Regarding the possibility that there is a pore-gating complex at plasmodesmata. If NUPs are specifically located at plasmodesmata, this is a strong hypothesis. The authors approach this functionally by assaying for protein and dye movement through plasmodesmata in the cpr5 mutants. These experiments suggest that cpr5 mutants have reduced transport through plasmodesmata for both proteins, but not for a smaller dye. They infer that the latter finding suggests that the cpr5 mutant has no alterations in plasmodesmal number, but this is completely unsupported - in their introduction, the authors identify how PD structure can modify transport capacity, so there are many technical and biological phenomena that could explain these data.

      We wrote in the manuscript: “The cpr5 mutants showed no detectable defect in small molecule transport indicative of WT-like PD density and preservation of the capability to mediate small molecule transport as shown by ‘Drop-ANd-See’ trans-leaf diffusion assays.”

      Indeed, we did not study PD density by e.g. quantification of a PD-marker fluorescence. Theoretically, PD density might be changed and permeabilities adjusted by unknown mechanisms to allow for WT-like small molecule transport. Strikingly, we observed transport differences for larger cargo. As we cannot exclude potential changes in PD density, we have rewritten and deleted the conclusion on PD density and now write: “The cpr5 mutants showed no detectable defect in small molecule transport indicative of preservation of the capability to mediate small molecule transport as shown by ‘Drop-ANd-See’ trans-leaf diffusion assays”. (Lines 310-312)

      I note for their DANS assays that the diffusion of dye from ad- to abaxial surface varies in the path followed (indicated by the asymmetry of the surfaces) and is not consistent within a leaf, let alone between leaves. This presents challenges in quantification and data interpretation that have not been addressed, and so the data cannot be confidently concluded to be an indicator of a different phenomenon rather than a less sensitive measure of the same.

      Indeed, in our hands, the spread of the small molecule dye did not proceed radially and was very often asymmetrical. Therefore, we quantified the fluorescent area by identifying pixels with fluorescence above a threshold, instead of determining a diameter of the fluorescent area. We describe the analysis in the figure legend and briefly mention it in the method section.

      “Fluorescent areas on the abaxial side were identified using auto threshold and Fiji YEN-algorithm with user modifications. The same threshold setting was used for the adaxial side. The extent of dye diffusion was quantified by the ratio between the areal spread of fluorescence on the abaxial side and the areal spread of fluorescence on the adaxial side.” (Figure 7)

      Furthermore, to avoid any positional artifacts in the comparison between different plants and genotypes, we only assessed the 4th leaf and 24 hours later the 5th leaf with the same labelling position on the leaf.

      Further, as the authors themselves acknowledge, altered protein movement might also arise from an off-target developmental phenotype. Many proteins have been shown to have no association with plasmodesmata but an indirect effect on their function. This hasn't been investigated and so cannot be ruled out.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This is a really interesting hypothesis, but the support is incomplete.

      (1) P. 5 "Although the single insertion Arabidopsis lines tested here should have FG-NUP62-GFP levels closer to native conditions than the heterologous overexpression of FG-NUP62-mVenus in N. benthamiana, it cannot be excluded that the levels in tested lines are still higher than the native levels, or that the fluorescence protein fused to the NUP affects localization." I appreciate the authors' cautious interpretation of their results, but they could exclude both of these possibilities. The first is relatively easy: test the expression level of the transgene compared to endogenous NUP expression; although transcript and protein levels are not tightly correlated, this can give some estimate of whether the transgene is overexpressed. The second would be to conduct complementation assays of a knockout mutant. I understand that this would be difficult if nup mutants are lethal, but it is pretty common practice to transform heterozygotes and isolate homozygotes expressing fluorescent protein to conduct complementation assays. Anyhow, there is a defect in the cpr5 mutants that the authors could assess in complementation assays. Alternatively, the authors could use biochemical approaches to determine whether FP-tagged NUPs are incorporated into nuclear pore complexes. These three experiments, even for only one NUP, would provide compelling evidence that the authors are localizing a functional NUP fusion protein at near-native expression levels. This is essential to support their speculation that NUPs play a biological role in PD.

      Thank you for these three important recommendations: the quantification of NUP FP expression, the complementation of a mutant phenotype with NUP FP expression, and the assessment whether NUP FPs are incorporated into the NPC.

      First, to evaluate the abundance of NUP62-GFP fusion protein relative to untransformed controls, we quantified the abundance levels of NUP62 in three independent transgenic fluorescent WT/NUP62p:NUP62-GFP Arabidopsis lines and in Arabidopsis WT using mass spectrometry (new Figure 3F). The new data indicate that there is no significant increase in NUP protein amounts in the lines expressing the fusion construct relative to wild type.

      We now write in the revised manuscript (line 200-205):

      “NUP62 protein abundance in two-week-old cotyledons of the stable NUP62p:NUP62-GFP transformants was not statistically different to NUP62 protein levels in WT (Figure 3F). Notably, the punctate fluorescence at the cell periphery, encompassing both PD-associated and non-PD-associated localization, were not detectable or absent in roots and young leaves of four-day-old seedlings (Figure 3D). However, it cannot be excluded that the GFP fusion impacts NUP62 localization.” We provide a new Method section for the mass spec analysis of the cotyledons in lines 582-590.

      Second, to tested whether a NUP fusion is functional we assessed whether CPR5-mCitrine can complement the cpr5-1 mutant phenotype in complementation lines. We generated two new independent transgenic Arabidopsis lines stably expressing CPR5-mCitrine under control of its own promoter in the cpr5-1 mutant background (cpr5-1/CPR5p:CPR5-mCitrine). The roots were significantly longer in the two independent transgenic cpr5-1/CPR5p:CPR5-mCitrine Arabidopsis lines compared to the cpr5-1 mutant, and four-week-old plants showed a more WT-like growth phenotype (new Figure 7-figure supplement 1, G–I). However, we could not detect fluorescence in the 10-14 day old seedlings, which could be due to a variety of reasons, such as cleavage of the FP and degradation of the FP without accumulating elsewhere in the cells.

      In the new manuscript we write in lines 275-283:

      “To assess whether the CPR5-mCitrine fusion protein is functional in Arabidopsis, we tested whether CPR5p:CPR5-mCitrine (including all introns) expression in the cpr5-1 mutant background results in a rescue of the severe growth phenotype of the cpr5-1 loss-of-function mutant (Bowling et al., 1997). Indeed, roots were significantly longer in the two independent transgenic cpr5-1/CPR5p:CPR5-mCitrine Arabidopsis lines compared to the cpr5-1 mutant, and four-week-old plants showed a more WT-like growth phenotype (Figure 7-figure supplement 1, G–I). However, we could not detect fluorescence in 10-14 day old seedlings, which could be due to a variety of reasons, such as cleavage of the FP and degradation of the FP without accumulating elsewhere in the cells. The lack of fluorescence in the transgenic lines requires further investigation.”

      Third, to assess whether NUP FP fusions are also detectable specifically in nuclei, we have provided example images for potential nuclear localization of NUP62-GFP in the stable Arabidopsis line (Figure 3C), and for AtGP210-mVenus, AtNUP98b-mVenus, AtCPR5-mCitrine, and At NUP43-mCitrine in transient expression experiments in N. benthamiana (Figure 3-figure supplement 1).

      (2) The rationale for experiments was sometimes unclear. For example, why study Physcomitrium NUPs, then switch to Arabidopsis? Why use heterologous overexpression lines for SIM, rather than the stable Arabidopsis line for NUP62-GFP?

      Our initial work focused on the PD proteome in Physcomitrium patens. We had identified NUPs in PD-enriched fractions of the moss (Gombos et al., 2023). To evaluate whether this was a specific feature of the moss, or a technical artifact of PD enrichment in moss extracts, we extended the study to Arabidopsis thaliana and subsequently focused on the higher plant. The text in the manuscript reflects this flow.

      The NUP62-GFP stable transgenic Arabidopsis line was generated after the SIM experiments with CPR5-mCitrine. We plan to follow the suggestion of the reviewer to perform SIM experiments with the stable Arabidopsis NUP62p:NUP62-GFP lines.

      (3) The organization of the figures was confusing. Why present transient Physco NUP localization, and also Arabidopsis proteomics in Figure 1? Why split the results on transient localization of Arabidopsis NUPs in benth across Figures 2 & 3?

      We reorganized the Figures and created a separate proteome main figure (now Figure 2 with 2 figure supplements). We classified Arabidopsis NUPs in FG-NUPs and structural NUPs. Thus, we present the data also in two separate Figures: Figure 3 and supplements, dedicated to FG-NUPs, and Figure 4, dedicated to structural NUPs. According to the NPC, FG-NUPs play a direct role in transport facilitation, setting them apart from the structural NUPs.

      (4) Why are several NUPs localized to the interior of the nucleus and not restricted to the nuclear membrane (e.g., Figure 1 Sup 1 top two rows, Figure 2)? How does this unusual nuclear localization alter the authors' interpretation of their results?

      We observed that the transmembrane NUPs tested localized to the nuclear envelope and not to the nucleoplasm (see Figure 3-figure supplement 1 for example AtGP210 and AtCPR5). We found several soluble NUPs to also localize to the nucleoplasm (PpNUP98.1, PpNUP62, AtNUP62, AtHOS1). Previous studies had reported that several FG NUPs (i.e. NUP98a/b or NUP62) and Y-complex NUPs (i.e. HOS1, NUP96, and NUP107) also localized in the nucleoplasm rather than specifically to the nuclear envelope when expressed as fusion proteins (Chen et al., 2023; Gallemí et al., 2016; Huang et al., 2024; Lazaro et al., 2012). Of note, for NUP98a, Gallemi and colleagues (2016) discussed the localization to the nucleoplasm as confirmation that, like vertebrate NUP98, Arabidopsis NUP98a is a dynamic NUP rather than just a key structural element of the NPC. HOS1 was reported to interact with ICE1, CO, FVE, and HDA6 in the nucleoplasm (Dong et al., 2006; Jung et al., 2012; Lazaro et al., 2012), indicating that HOS1 might dynamically shuttle between the nuclear pore and nucleoplasm, which could also explain the observed nucleoplasmic localization. In Drosophila, the FG-NUPs NUP98, NUP62, and NUP50 localized in the NPC, and also in the nucleoplasm and interacted with genes (Kalverda et al., 2010). The nucleoplasmic localization could thus have a functional relevance. Yet we cannot rule out, whether soluble NUPs mislocalize in overexpression conditions as we state multiple times in the manuscript.

      (5) Figure legends are insufficiently detailed. Figure legends should be sufficiently detailed to explain the figure without consulting the main text. For example,

      (a) Figure 1A, 3C don't describe the cell type or even the organism that is being imaged. Are Physco proteins expressed in Physco? Arabidopsis? Benth? Leaves?

      We added the missing information including cell types and organism.

      (b) In Figure 1 Supplement 3, many abbreviations are not defined (HC, MC, etc).

      We now define the abbreviations in the figure legend.

      (c) In Figure 2B, the legend says "At least 15 images from 3 biological replicates were analyzed for each NUP", but there are MANY more than 15 datapoints in Figure 2B. What do the points represent?

      We obtained at least three independent replicates for each data set we show here. We analyzed 15 ROIs derived from three biological replicates of AtNUP50b. In the other cases, a larger number of experiments was performed resulting in more ROIs being analyzed.

      (d) For all microscopy images, are they single images or reconstructions (e.g., maximum projections)?

      We now specify single confocal optical section or maximum projections.

      Reviewer #2 (Recommendations for the authors):

      (1) PD index shall be measured for data in Figures 3D and 3E.

      To address this question, we have performed PD index quantification for the data in Figures 4D and 4E and added the information to the main text (lines 178-184):

      “In leaves transiently expressing NUP43-mCitrine or CPR5-FP fusions, the fluorescence intensity correlated with the estradiol concentration used, with decreased fluorescence intensity for samples where 2µM estradiol was applied versus the intensity in samples exposed to 20µM estradiol (Figure 4 D,E). Notably, the fluorescence ratio between periphery and nucleus did not differ significantly after expression induction by 2 µM compared to 20 µM β-estradiol (Figure 4F) and PD localization was not eliminated (example for localization of NUP43-mCitrine in Figure 4C; PD index(NUP43, 2µM) = 1.42, PD index(CPR5, 2µM) = 1.40).”

      (2) The expression level of the native promoter-driven Nup62-GFP shall be measured and compared with the native level using RT-qPCR. Even if this turns out to be an overexpression line, it would still be useful to support the hypothesis.

      To evaluate the level of NUP62-GFP fusion protein relative to untransformed controls, we quantified the levels of NUP62 in three independent transgenic fluorescent WT/NUP62p:NUP62-GFP Arabidopsis lines and in Arabidopsis WT using mass spectrometry (new Figure 3F). The new data indicate that there is no significant increase in NUP protein amounts in the lines expressing the fusion construct relative to WT.

      We now write in the revised manuscript (line 200-205):

      “NUP62 protein abundance in two-week-old cotyledons of the stable NUP62p:NUP62-GFP transformants was not statistically different to NUP62 protein levels in WT (Figure 3F). Notably, the punctate fluorescence at the cell periphery, encompassing both PD-associated and non-PD-associated localization, were not detectable or absent in roots and young leaves of four-day-old seedlings (Figure 3D). However, it cannot be excluded that the GFP fusion impacts NUP62 localization.“ We provide a new Method section for the mass spec analysis of the cotyledons in lines 582-590.

      (3) Last sentence in the introduction: Nup136 has been considered as the plant homolog of Nup153.

      In the manuscript we wrote:

      “The majority of the FG-NUPs were conserved, with only three FG-NUPs lost in the green lineage (NUP153, POM121, NUP358).“

      As the FG-NUP136 is the plant homolog to NUP153, we now write (lines 90-92):

      “The majority of the FG-NUPs were conserved, with two FG-NUPs apparently lost in the green lineage (POM121, NUP358).“

      Reviewer #3 (Recommendations for the authors):

      (1) Generally, my interpretation of the images in this manuscript is that many of the localisations are not clean and discrete plasmodesmal associations and are rather more consistent with cortical ER association. As the ER is a component of plasmodesmata, the ER is continuous with the nuclear envelope, and the authors also predict and show ER localisation of one of their key NUPs, CPR5 in Figure 4B. This is not necessarily surprising. However, what becomes essential is that the authors need to determine whether NUPs behave any differently from other ER proteins. To that end, I think co-localisations with ER-located proteins would be helpful in interpreting these ambiguous localisations.

      To address this point, we performed additional colocalization experiments using an ER marker. In the new version of the manuscript, we now include a comparison of the NUP43-mVenus localization with that of the mCherry-HDEL luminal ER marker, which reveals distinct localization patterns (see new Figure5, new Figure 5-figure supplement 5-1). NUP43-mVenus may be associated with the ER; however, NUP43 is restricted to subregions of the ER, which partially overlay with aniline blue-labeled pit fields (new Figure 5, new Figure 5-figure supplement 1).

      (2) The super-resolution images of CPR5 show some clear structures peripheral to plasmodesmata. However, again, I would like to see what an ER protein looks like at this location, as the ER feeds into the plasmodesmata. Is this a specific structure or a general feature of the localisation of an ER protein?

      Since mCherry-HDEL (see above) did not show a similar localization or enrichment at PD, we did not perfrom SIM analyses with the marker.

      (3) The authors support their use of the PD score using validated PD proteins as the positive control and contaminants from mitochondria and other organelles as the negative control. No mention is made of where ER proteins are classified. The ER passes through plasmodesmata but might also represent a contaminating pool. As NUPs reside in the nuclear envelope, continuous with the ER, a comparison between the NUPs and ER proteins would be extremely informative.

      To evaluate a potential enrichment of ER proteins in the plasmodesmata fraction, we analyzed ER protein enrichment and added the new data as a graph in Figure 2-figure supplement 2. ER-resident proteins did not show significant enrichment in the cell wall fraction relative to total cell extract, while displaying a slight but consistent enrichment in the plasmodesmata fraction. Notably, NUPs enrichment was higher in both cell wall fraction and plasmodesmata fraction compared to transmembrane ER-resident proteins. While ER membrane co-purification cannot be entirely excluded, the enrichment of NUPs in the plasmodesmata fraction may not be due to desmotubule membrane carryover alone. The analysis was incorporated into the revised manuscript (lines 152-155).

      (4) Regarding the data analysis and use of the Kruskal-Wallis test, the Kruskal-Wallis test tests differences in the distribution of the data, not differences in the mean or median values. In many cases, it can be inferred that the median changes when the data distribution does, but this is not as confident an inference for means. There are other methods available to compare the means of such datasets.

      We used the Kruskal–Wallis test for statistical comparison of more than two nonparametric data sets. However, we did not state in the manuscript that we performed a Dunns´ test for the post hoc pairwise comparison after the Kruskal-Wallis test. In the revised manuscript, we added this information in the Methods, Results and Figure legends. For the bombardment experiment data, we now added mean bootstrapping, as used previously in this context (Johnston and Faulkner, 2021). Mean bootstrap analysis for the bombardment data set was performed with n=5000 resamples and we provide the p values and confidence intervals in the figure legend (Figure 7 B):

      “Mean fluorescent cell counts: n<sub>(WT)</sub> = 2.67, n<sub>(cpr5-T3)</sub> = 1.59, n<sub>(cpr5-1)</sub> = 0.68; median fluorescence cell counts: n<sub>(WT)</sub> = 2, n<sub>(cpr5-1)</sub> = 0, n<sub>(cpr5-T3)</sub> = 1. Based on Bonferroni-corrected Dunn´s test for pairwise comparison after Kruskal-Wallis test: a indicates significant difference to WT with p(<sub>cpr5-1</sub>) < 10<sup>-15</sup>; b indicates significant difference to WT with p<sub>(cpr5-T3)</sub> = 0.0004; c indicates p(cpr5-1 vs. cpr5-T3) = 0.0002. Mean bootstrap analysis according to (Johnston and Faulkner, 2021) with 95% confidence interval (CI) and bootstrap resampling of B = 5000: CI<sub>WT vs. cpr5-1</sub> [1 x 10<sup>-5</sup> , 0.001], p<sub>(cpr5-1)</sub> = 0002 ; CI<sub>WT vs. cpr5-T3</sub> [1 x 10<sup>-5</sup> , 0.001], p<sub>(cpr5-T3)</sub> = 0.0002; CI<sub>cpr5-1 vs. cpr5-T3</sub>, p<sub>(cpr5-1 vs. cpr5-T3)</sub> = 0.0002 [1 x 10<sup>-5</sup> , 0.001].“

      (5) The comments that estradiol induction prevents over-expression, or allows for controlled expression, are not experimentally supported or widely established outside this manuscript. I suggest they tone this claim down.

      As outlined above the reduction in estradiol concentration lead to reduced fluorescence intensity for the NUP-FP fusions as one would expect; here notably with a reduction at both nuclei and periphery (Figure 4C-F). The system has been used previously in the Simon lab, from whom we obtained the constructs. There is substantial literature regarding the use of the b-estradiol-inducible XVE promoter system, specifically for b-estradiol dose-dependent gene expression in N. benthamiana leaves (Bashandy et al., 2015; Bleckmann et al., 2010; Borghi, 2010; Schlücking et al., 2013). We assessed the dependence of localization on expression levels by studying NUP localization with a lower estradiol concentration for induction and shortened incubation time. Interestingly, despite the apparent lower expression, we still find NUPs at PD.

      References

      Bashandy H, Jalkanen S, Teeri TH. 2015. Within leaf variation is the largest source of variation in agroinfiltration of Nicotiana benthamiana. Plant Methods 11:47. DOI: https://doi.org/10.1186/s13007-015-0091-5

      Bleckmann A, Weidtkamp-Peters S, Seidel CAM, Simon R. 2010. Stem Cell Signaling in Arabidopsis Requires CRN to Localize CLV2 to the Plasma Membrane. Plant Physiology 152:166–176. DOI: https://doi.org/10.1104/pp.109.149930

      Borghi L. 2010. Inducible gene expression systems for plants. In: Hennig L, Köhler C (Eds). Plant Developmental Biology: Methods and Protocols. Humana Press. p. 65–75. DOI: https://doi.org/10.1007/978-1-60761-765-5_5

      Bowling SA, Clarke JD, Liu Y, Klessig DF, Dong X. 1997. The cpr5 mutant of Arabidopsis expresses both NPR1-dependent and NPR1-independent resistance. The Plant Cell 9:1573–84.

      Chen G, Xu D, Liu Q, Yue Z, Dai B, Pan S, Chen Y, Feng X, Hu H. 2023. Regulation of FLC nuclear import by coordinated action of the NUP62-subcomplex and importin β SAD2. Journal of Integrative Plant Biology 65:2086–2106. DOI: https://doi.org/10.1111/jipb.13540

      Dong C-H, Agarwal M, Zhang Y, Xie Q, Zhu J-K. 2006. The negative regulator of plant cold responses, HOS1, is a RING E3 ligase that mediates the ubiquitination and degradation of ICE1. Proceedings of the National Academy of Sciences 103:8281–8286. DOI: https://doi.org/10.1073/pnas.0602874103

      Gallemí M, Galstyan A, Paulišić S, Then C, Ferrández-Ayela A, Lorenzo-Orts L, Roig-Villanova I, Wang X, Micol JL, Ponce MR, Devlin PF, Martínez-García JF. 2016. DRACULA2 is a dynamic nucleoporin with a role in regulating the shade avoidance syndrome in Arabidopsis. Development 143:1623–1631. DOI: https://doi.org/10.1242/dev.130211

      Gombos S, Miras M, Howe V, Xi L, Pottier M, Kazemein Jasemi NS, Schladt M, Ejike JO, Neumann U, Hänsch S, Kuttig F, Zhang Z, Dickmanns M, Xu P, Stefan T, Baumeister W, Frommer WB, Simon R, Schulze WX. 2023. A high-confidence Physcomitrium patens plasmodesmata proteome by iterative scoring and validation reveals diversification of cell wall proteins during evolution. New Phytologist 238:637–653. DOI: https://doi.org/10.1111/nph.18730

      Gu Y, Zebell SG, Liang Z, Wang S, Kang B-H, Dong X. 2016. Nuclear pore permeabilization is a convergent signaling event in effector-triggered immunity. Cell 166:1526-1538.e11. DOI: https://doi.org/10.1016/j.cell.2016.07.042

      Huang P, Zhang X, Cheng Z, Wang X, Miao Y, Huang G, Fu Y-F, Feng X. 2024. The nuclear pore Y-complex functions as a platform for transcriptional regulation of FLOWERING LOCUS C in Arabidopsis. The Plant Cell 36:346–366. DOI: https://doi.org/10.1093/plcell/koad271

      Johnston MG, Faulkner C. 2021. A bootstrap approach is a superior statistical method for the comparison of non-normal data with differing variances. New Phytologist 230:23–26. DOI: https://doi.org/10.1111/nph.17159

      Jung J-H, Seo PJ, Park C-M. 2012. The E3 ubiquitin ligase HOS1 regulates Arabidopsis flowering by mediating CONSTANS degradation under cold stress. Journal of Biological Chemistry 287:43277–43287. DOI: https://doi.org/10.1074/jbc.M112.394338

      Kalverda B, Pickersgill H, Shloma VV, Fornerod M. 2010. Nucleoporins directly stimulate expression of developmental and cell-cycle genes inside the nucleoplasm. Cell 140:360–371. DOI: https://doi.org/10.1016/j.cell.2010.01.011

      Lazaro A, Valverde F, Piñeiro M, Jarillo JA. 2012. The Arabidopsis E3 ubiquitin ligase HOS1 negatively regulates CONSTANS abundance in the photoperiodic control of flowering. The Plant Cell 24:982–999. DOI: https://doi.org/10.1105/tpc.110.081885

      Schlücking K, Edel KH, Köster P, Drerup MM, Eckert C, Steinhorst L, Waadt R, Batistič O, Kudla J. 2013. A new β-estradiol-inducible vector set that facilitates easy construction and efficient expression of transgenes reveals CBL3-dependent cytoplasm to tonoplast translocation of CIPK5. Molecular Plant 6:1814–1829. DOI: https://doi.org/10.1093/mp/sst065

      Tamura K, Fukao Y, Iwamoto M, Haraguchi T, Hara-Nishimura I. 2010. Identification and characterization of nuclear pore complex components in Arabidopsis thaliana. The Plant Cell 22:4084–4097. DOI: https://doi.org/10.1105/tpc.110.079947

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      Some of the authors proposed in a PNAS paper in 2016 the occurrence of the Entner-Doudoroff (ED) pathway in cyanobacteria and plants, on the basis of several lines of biochemical and genetic evidence. However, more recent results indicated that one of the two specific enzymes of the ED pathway (EDD) is missing in Synechocystis PCC 6803. The authors carried out additional experiments, which demonstrated that EDD is missing, and one of the enzymes (ED aldolase) is a promiscuous enzyme which seems to be involved in proline metabolism and is not actually participating in the ED pathway as initially believed. The results described in this paper are strong evidence that this new interpretation is appropriate, and therefore, it corrects the previous proposal, providing an honest description of the reasons why the authors had reached the wrong conclusion about the existence of the ED pathway in cyanobacteria and plants.

      We thank Reviewer 1 for the summary and comments. We found that EDA is a promiscuous aldolase that, in addition to the cleavage of KDPG to GAP and pyruvate (a reaction of the ED pathway) catalyzes other reactions in vitro. Based on the in vitro results obtained, potential in vivo functions of EDA are proposed, including its involvement in proline metabolism. However, these assumptions require further experimental testing. We do not yet have definitive findings regarding the function of the promiscuous aldolase EDA in Synechocystis in vivo, but respective studies are currently underway.

      Strengths:

      Thorough reanalysis of the experimental results obtained in previous studies, which led to the publication of the PNAS paper in 2016.

      New experimental evidence to confirm that enzymes previously considered as participating in the ED actually are not catalyzing the ED biochemical reactions, but are involved in other metabolic pathways. Also, the authors completely discarded the occurrence of the GDH/GK shunt in Synechocystis PCC 6803. Generally speaking, the manuscript is very clearly written, with a precise description of the previous findings, the mistakes which took place in the 2016 paper, and the strategies they have used to address those issues, in order to reach a thoroughly revised vision of the glucose metabolic pathways in Synechocystis PCC 6803. In this regard, the drawings shown in Figures 1 and 7 are very helpful for the reader to follow the story and understand the possible metabolic transformations depending on the working hypothesis.

      Also, I commend the authors for openly describing previous mistakes. In this paper, they reassess past observations in light of more recent findings and to integrate the information in this manuscript. The scientific conclusions are solid and very interesting, and besides, they use the opportunity to offer valuable advice to researchers. This is especially focused on the importance of careful biochemical characterization of enzymes, which should always be carried out when studying proteins which have been identified as a specific enzyme on the basis of sequence homology. In a similar way, they found that an insertional mutant was the cause of the absence of specific metabolites, which had been attributed to particularities of a metabolic pathway in that mutant, when it was actually due to a nucleotide insertion; this could have been easily prevented by confirming the correct generation of the mutant by DNA sequencing.

      We agree that biochemical characterization of enzymes as well as DNA sequencing to check deletion mutants, are important and valuable tools. As outlined in the manuscript and additionally in more detail in a recently submitted article, which is available at bioRxiv (Theune et al. 2026, doi: https://doi.org/10.64898/2026.04.08.717167) and is currently under review at PLOS One, we suggest that genome sequencing of deletion mutants in combination with complemented strains as controls are required to minimize the risk of misinterpretation based on secondary mutations (1). During the early stages of our research on the ED pathway, and later as well when we were already trying to resolve the conflicting results that had accumulated concerning the ED pathway, genome sequencing for Synechocystis mutants was not affordable as a routine procedure (2-4). Therefore, we could not have easily prevented this misconception based on this technique at that time. However, we strongly encourage genome sequencing of deletion mutants in combination with complemented strains as routine procedures these days (1).

      Weaknesses:

      The authors propose that EDA might be involved in the PEP-pyruvate-OAA node, or in the proline metabolism, but this requires further experimental work for clarification; what their results indicate clearly is that this enzyme is not actually catalyzing the transformation of KDPG to GAP, which is the second specific enzyme of the ED pathway. But the real physiological function in this cyanobacterium is still unconfirmed.

      As stated above and in the manuscript, we agree that the in vivo role of EDA requires further experimental work which is in progress. However, our results demonstrate that EDA splits KDPG into GAP and pyruvate in vitro, but we assume that this reaction does not play a role in vivo due to the absence of its substrate.

      Another aspect which could be improved is that the recombinant expression of some genes was carried out in E. coli; even if this is a useful and valid research strategy, in studies like this (where there is a strong focus on the physiological function of enzymes in the original organism, Synechocystis PCC 6803), I think it would have been more appropriate to express the 6803 genes in another cyanobacterium easily amenable for genetic transformation and gene expression, which would produce the protein in a physiological environment more similar to another cyanobacterium (compared to E. coli, which is an heterotrophic bacterium). I am not sure this would change any of the obtained results, but it certainly would confer additional robustness to the enzymatic results.

      Synechocystis is easily amendable to genetic manipulation, and we agree that expression and purification of all enzymes from this host would have been ideal. However, the first characterization of Synechocystis EDA was performed with proteins that were purified from Synechocystis and showed activity on KDPG at comparable rates as proteins that were purified from E. coli in this study (2). Moreover, most biochemical characterizations of EDAs from archaea, bacteria and plants were performed after recombinant expression in E. coli and yielded highly active enzyme as in the case of Synechocystis is this study (5-7). Therefore, we currently have no reason to worry that the expression in E. coli might affect the enzymatic activity of EDA. The main reason for utilizing E. coli as an expression strain in this study was to gain higher yields of protein for in-depth analyses.

      Bibliography:

      I think the list of papers used in this manuscript is complete and up to date. However, I do miss recent papers which addressed one aspect that was proposed in the original 2016 PNAS paper: the authors wrote, "We therefore suggest that Prochlorococcus might oxidize glucose via the ED pathway under mixotrophic conditions, as shown for Synechocystis." Recent studies checked this hypothesis and have shown that the ED pathway seems to be also missing in Prochlorococcus and marine Synechococcus, and I think this manuscript is a good place to cite them, since these results are consistent with the findings of this paper.

      We will include a references from Moreno-Cabezuelo et a. 2023 (DOI: 10.1128/spectrum.03275-22) in which the proteomes of three marine Prochlorococcus and three marine Synechococcus strains were investigated upon exposure to glucose (8). Protein levels of EDA were either downregulated or not affected while proteins involved in OPP pathway and CBB cycle were upregulated. The authors of this study conclude that this indicates that the latter processes rather than the ED pathway are involved in photomixotrophy in these strains. However, flux analyses are still missing. 

      Reviewer #2 (Public review):

      Summary:

      The study presents novel results on the presence of the Entner-Doudoroff pathway in Synechocystis sp. PCC 6803. In contrast to an earlier study, compelling evidence is given that this strain lacks both an ED pathway and a glucose dehydrogenase/glucokinase bypass but contains a promiscuous aldolase, which also decarboxylates oxaloacetate and cleaves 2-keto-4-hydroxyglutarate (as it occurs in proline degradation). The study concludes with successfully reconciling data from different studies and with lessons learned from the previous misconception.

      Strengths:

      Solid biochemical data are presented to reconcile contradicting data of earlier studies and to serve as a basis for disclosing possible functions of a promiscuous aldolase. Earlier misconceptions and lessons to be learned are well discussed.

      Weaknesses:

      The materials and methods section is rather lengthy, suffering from a lack of conciseness and repetition, and nevertheless misses some specifications.

      We thank Reviewer 2 for the summary and comments and will improve the materials and methods part accordingly in a revised version.

      (1) M. Theune et al., Easy-to-use whole-genome sequencing workflows and standardized practices to uncover hidden genetic variation in Synechocystis PCC 6803 wild-type and knock-out strains. bioRxiv 10.64898/2026.04.08.717167, 2026.2004.2008.717167 (2026).

      (2) X. Chen et al., The Entner–Doudoroff pathway is an overlooked glycolytic route in cyanobacteria and plants. Proceedings of the National Academy of Sciences 113, 5441-5446 (2016).

      (3) D. Schulze et al., GC/MS-based 13C metabolic flux analysis resolves the parallel and cyclic photomixotrophic metabolism of Synechocystis sp. PCC 6803 and selected deletion mutants including the Entner-Doudoroff and phosphoketolase pathways. Microbial Cell Factories 21, 69 (2022).

      (4) A. Makowka et al., Glycolytic Shunts Replenish the Calvin–Benson–Bassham Cycle as Anaplerotic Reactions in Cyanobacteria. Molecular Plant 13, 471-482 (2020).

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      Summary:

      The authors describe the results of a single study designed to investigate the extent to which horizontal orientation energy plays a key role in supporting view-invariant face recognition. The authors collected behavioral data from adult observers who were asked to complete an old/new face matching task by learning broad-spectrum faces (not orientation filtered) during a familiarization phase and subsequently trying to label filtered faces as previously seen or novel at test. This data revealed a clear bias favoring the use of horizontal orientation energy across viewpoint changes in the target images. The authors then compared different ideal observer models (cross-correlations between target and probe stimuli) to examine how this profile might be reflected in the image-level appearance of their filtered images. This revealed that a model looking for the best matching face within a viewpoint differed substantially from human data, exhibiting a vertical orientation bias for extreme profiles. However, a model forced to match targets to probes at different viewing angles exhibited a consistent horizontal bias in much the same manner as human observers.

      Strengths:

      I think the question is an important one: The horizontal orientation bias is a great example of a low-level image property being linked to high-level recognition outcomes, and understanding the nature of that connection is important. I found the old/new task to be a straightforward task that was implemented ably and that has the benefit of being simple for participants to carry out and simple to analyze. I particularly appreciated that the authors chose to describe human data via a lower-dimensional model (their Gaussian fits to individual data) for further analysis. This was a nice way to express the nature of the tuning function, favoring horizontal orientation bias in a way that makes key parameters explicit. Broadly speaking, I also thought that the model comparison they include between the view-selective and view-tolerant models was a great next step. This analysis has the potential to reveal some good insights into how this bias emerges and ask fine-grained questions about the parameters in their model fits to the behavioral data.

      Weaknesses:

      I will start with what I think is the biggest difficulty I had with the paper. Much as I liked the model comparison analysis, I also don't quite know what to make of the view-tolerant model. As I understand the authors' description, the key feature of this model is that it does not get to compare the target and probe at the same yaw angle, but must instead pick a best match from candidates that are at different yaws. While it is interesting to see that this leads to a very different orientation profile, it also isn't obvious to me why such a comparison would be reflective of what the visual system is probably doing. I can see that the view-specific model is more or less assuming something like an exemplar representation of each face: You have the opportunity to compare a new image to a whole library of viewpoints, and presumably it isn't hard to start with some kind of first pass that identifies the best matching view first before trying to identify/match the individual in question. What I don't get about the view-tolerant model is that it seems almost like an anti-exemplar model: You specifically lack the best viewpoint in the library but have to make do with the other options. Again, this is sort of interesting and the very different behavior of the model is neat to discuss, but it doesn't seem easy to align with any theoretical perspective on face recognition. My thinking here is that it might be useful to consider an additional alternate model that doesn't specifically exclude the best-matching viewpoint, but perhaps condenses appearance across views into something like a prototype. I could even see an argument for something like the yaw-averages presented earlier in the manuscript as the basis for such a model, but this might be too much of a stretch. Overall, what I'd like to see is some kind of alternate model that incorporates the existence of the best-match viewpoint somehow, but without the explicit exemplar structure of the view-specific model.

      The design of the view-tolerant model aligned with the requirements of tolerant recognition and revealed the stimulus information enabling to abstract identity away from variations in face appearance. However, it did not involve the notion that such ability may depend on a prototype or summary representation of face identity built up through varied encounters (Burton, Jenkins, & Schweinberger, 2011; Burton et al., 2016; Jenkins et al., 2011; Menon, Kemp, & White, 2018; Mike Burton, 2013).

      We agree with the Reviewer that the average of the different views of a face is a good proxy of its central tendency (i.e., stable identity properties; Figure 1). We thus followed their suggestion and included an additional model observer that compared specific views to full-spectrum view-averaged identities. The examination of the orientation tuning profile of this so-called view-average model observer confirmed the crucial contribution of horizontal identity cues to view-invariant recognition as the horizontal range best predicted the average summary of full-spectrum face appearances across views. This additional model observer is now presented in the Discussion and Supplementary files 2 and 3.

      Besides this larger issue, I would also like to see some more details about the nature of the cross-correlation that is the basis for this model comparison. I mostly think I get what is happening, but I think the authors could expand more on the nature of their noise model to make more explicit what is happening before these cross-correlations are taken. I infer that there is a noise-addition step to get them off the ceiling, but I felt that I had to read between the lines a bit to determine this.

      In the Methods section, we now provide detailed information about the addition of noise to model observer cross-correlations: ‘In a pilot phase, we measured the overall identification performance of each model. Initially, the view-selective model performed at ceiling, yielding a correlation of 1 since there was an exact target-probe match across all trials. To avoid ceiling effects and to keep model performance close to human levels (Supplementary File 2), we thus decreased the signal-to-noise ratio (SNR) of the target and probe images to .125 by combining each with distinct noise patterns (face RMS contrast: .01; noise RMS contrast: .08). Each trial (i.e. target-probe pairing) was iterated ten times with different random noise patterns.’

      We also added a supplemental with the graphic illustration of the d’ distributions of each model and human observers: ‘Sensitivity d’ of the view-tolerant model was much lower than view-selective model and human sensitivity (Supplementary File 2), even without noise. The view-tolerant model therefore processed fully visible stimuli (SNR of 1). This decreased sensitivity in the view-tolerant compared to the view-selective model is expected, as none of the probes exactly matched the target at the pixel level due to viewpoint differences. In contrast to humans who rely on internally stored representations to match identity across views, the model observer lacks such internal representations and entirely relies on (less efficient) pixelwise comparisons.’

      Another thing that I think is worth considering and commenting on is the stimuli themselves and the extent to which this may limit the outcomes of their behavioral task. The use of the 3D laser-scanned faces has some obvious advantages, but also (I think) removes the possibility for pigmentation to contribute to recognition, removes the contribution of varying illumination and expression to appearance variability, and perhaps presents observers with more homogeneous faces than one typically has to worry about. I don't think these negate the current results, but I'd like the authors to expand on their discussion of these factors, particularly pigmentation. Naively, surface color and texture seem like they could offer diagnostic cues to identity that don't rely so critically on horizontal orientations, so removing these may mean that horizontal bias is particularly evident when face shape is the critical cue for recognition.

      Our stimuli were originally designed by Troje and Bulthoff (1996). These are 3D laser scans of white individuals aged between 20 and 40 years, posing with a neutral expression. Different views of the faces were shot under a fixed illumination. Ears and a small portion of the neck were visible while the hair region was removed. All face images had a normalized skin color and we further converted them to grayscales

      While we agree that this stimulus set offers a restricted range of within- and between-identity variations compared to what is experienced in natural settings, we believe that the present findings generalize to more ecological viewing conditions. Indeed, past evidence showed that the recognition of face pictures shot under largely variable pose, age, expression, illumination, hair style is tuned to the horizontal range of the face stimulus (Dakin & Watt, 2009; Dumont, Roux-Sibilon, & Goffaux, 2024). In other words, our finding that view-tolerant identity recognition is mainly driven by horizontal face information would likely replicate with the use of a more ecological stimulus set.

      Moreover, the skin color normalization and grayscale conversion, while limiting the range of face variability, did not eliminate the contribution of surface pigmentation in our study. It is thus unlikely that our findings exclusively reflect the orientation dependence of face shape processing. Pigmentation refers to all surface reflectance properties (Russell et al., 2006) and hue (color) is only one among others. The grayscaled 3D laser scanned faces used here contained natural variations in crucial surface cues such as skin albedo (i.e., how light or dark the surface appears) and texture (i.e., spatial variation in how light is reflected); they have actually been used to disentangle the role of shape and surface cues to identity recognition (e.g., Jiang et al., 2009; Russell et al., 2007; Russell et al., 2006; Troje & Bulthoff, 1996; Vuong et al., 2005). Moreover, a past study of ours demonstrated that the diagnosticity of the horizontal range of face information is not restricted to face shape cues; the specialized processing of face shape and surface both selectively rely on horizontal information (Dumont, Roux-Sibilon, & Goffaux, 2024).

      For these reasons, the present findings are unlikely to be fully determined by shape processing, and we expect them to generalize to more ecological stimulus sets. We discuss these aspects in the revised manuscript.

      Reviewer #2 (Public review):

      This study investigates the visual information that is used for the recognition of faces. This is an important question in vision research and is critical for social interactions more generally. The authors ask whether our ability to recognise faces, across different viewpoints, varies as a function of the orientation information available in the image. Consistent with previous findings from this group and others, they find that horizontally filtered faces were recognised better than vertically filtered faces. Next, they probe the mechanism underlying this pattern of data by designing two model observers. The first was optimised for faces at a specific viewpoint (view-selective). The second was generalised across viewpoints (view-tolerant). In contrast to the human data, the view-specific model shows that the information that is useful for identity judgements varies according to viewpoint. For example, frontal face identities are again optimally discriminated with horizontal orientation information, but profiles are optimally discriminated with more vertical orientation information. These findings show human face recognition is biased toward horizontal orientation information, even though this may be suboptimal for the recognition of profile views of the face.

      One issue in the design of this study was the lowering of the signal-to-noise ratio in the view-selective observer. This decision was taken to avoid ceiling effects. However, it is not clear how this affects the similarity with the human observers.

      In the Methods section, we now provide detailed information about the addition of noise to model observer cross-correlations: ‘In a pilot phase, we measured the overall identification performance of each model. Initially, the view-selective model performed at ceiling, yielding a correlation of 1 since there was an exact target-probe match across all trials. To avoid ceiling effects and to keep model performance close to human levels (Supplementary File 2), we thus decreased the signal-to-noise ratio (SNR) of the target and probe images to .125 by combining each with distinct noise patterns (face RMS contrast: .01; noise RMS contrast: .08). Each trial (i.e. target-probe pairing) was iterated ten times with different random noise patterns.’

      We also added a supplemental with the graphic illustration of the d’ distributions of each model and human observers.

      Another issue is the decision to normalise image energy across orientations and viewpoints. I can see the logic in wanting to control for these effects, but this does reflect natural variation in image properties. So, again, I wonder what the results would look like without this step.

      All stimuli were matched for luminance and contrast. It is crucial to normalize image energy across orientations as natural image energy is disproportionately distributed across orientations (e.g., Hansen et al., 2003). Images of faces cropped from their background as used here contain most of their energy in the horizontal range (Goffaux & Greenwood, 2016; Keil, 2008, 2009). If not normalized after orientation filtering, such uneven distribution of energy would boost recognition performance in the horizontal range across views. Normalization was performed across our experimental conditions merely to avoid energy from explaining the influence of viewpoint on the orientation tuning profile.

      We were not aware of any systematic natural variations of energy across face views. To address this, we measured face average energy (i.e., RMS contrast) in the original stimulus set, i.e., before the application of any image processing or manipulation. Background pixels were excluded from these image analyses. Across yaws, we found energy to range between .11 and .14 on a 0 to 1 grayscale. This is moderate compared to the range of energy variations we measured across identities (from .08 to .18). This suggests that variations in energy across viewpoints are moderate compared to variations related to identity. It is unclear whether these observations are specific to our stimulus set or whether they are generalizable to faces we encounter in everyday life. They, however, indicate that RMS contrast did not substantially vary across views in the present study and suggest that RMS normalization is unlikely to have affected the influence of viewpoint on recognition performance.

      In the revised methods section, we explicitly motivate energy normalization: ‘Images of faces cropped from their background as used here contain most of their energy in the horizontal range (Goffaux, 2019; Goffaux & Greenwood, 2016; Keil, 2009). Across yaws, we found face energy to range between .11 and .14 on a 0 to 1 grayscale, which is moderate compared to the range of face energy variations we measured across identities (from .08 to .18). To prevent energy from explaining our results, in all images, the luminance and RMS contrast of the face pixels were fixed to 0.55 and 0.15, respectively, and background pixels were uniformly set to 0.55. The percentage of clipped pixel values (below 0 or above 1) per image did not exceed 3%.’.

      Despite the bias toward horizontal orientations in human observers, there were some differences in the orientation preference at each viewpoint. For example, frontal faces were biased to horizontal (90 degrees), but other viewpoints had biases that were slightly off horizontal (e.g., right profile: 80 degrees, left profile: 100 degrees). This does seem to show that differences in statistical information at different viewpoints (more horizontal information for frontal and more vertical information for profile) do influence human perception. It would be good to reflect on this nuance in the data.

      Indeed, human performance data indicates that while identity recognition remains tuned to horizontal information, horizontal tuning peak shows some variation across viewpoints. We primarily focused on the first aspect because of its direct relevance to our research objective, but also discussed the second aspect: with yaw rotation, certain non-horizontal morphological features such as the jaw line or nose bridge, etc. may increasingly contribute to identity recognition, whereas at frontal or near frontal views, features are mostly horizontally-oriented (e.g., Keil, 2008, 2009). In the revised Discussion, we directly relate the modest fluctuations of peak location to yaw differences in face feature appearance.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Based on a discussion with the reviewers, we integrated the recommendations and reached a consensus on the eLife assessment. To move from a "solid" to a "compelling/convincing" strength-of-evidence rating, please address the reviewers' comments. Key points are to clarify and test the plausibility of the models (e.g., effects of different noise-addition steps, inclusion/exclusion of specific orientation channels in the view-dependent comparison, and alternative decision criteria), and to address or discuss the limitations of the stimulus set in capturing recognition under more naturalistic scenarios, for example, including texture cues.

      Reviewer #1 (Recommendations for the authors):

      I generally found the paper to be very well-written, so I have only a few minor comments here.

      (1) I didn't really follow why the estimation of the Gaussian functions described in the text was preferred over a simpler ML framework. Do these approaches differ that much? I see references to prior studies in which these were applied, so I can certainly go check these out, but I could see value in adding just a bit of text to briefly make the case that this is important.

      Employing a simpler linear framework, i.e. a linear model predicting d’ from the interaction between orientation and viewpoint, would result in an 8 (orientation) * 7 (viewpoint) design that is difficult to analyze. The interaction term would almost certainly reach significance but its interpretation would be limited. We would either have to rely on numerous local comparisons, which are not particularly informative for our research objectives (e.g., knowing whether d’ differs significantly between two adjacent orientations at a given viewpoint is of little relevance), or to use a polynomial contrast approach (testing the linear, quadratic, … up to the 7th order trends), which would also be difficult to interpret. For such complex, approximately Gaussian-shaped data, the highest-order polynomial trend would likely provide the best fit, but without offering meaningful insight.

      In contrast, a nonlinear approach appears more appropriate. The Gaussian model we used allows us to characterize the parameters of the tuning profile, namely, peak location, peak amplitude, standard deviation (or bandwidth) and base amplitude. These parameters are not merely statistical parameters. Rather, they are directly interpretable in cognitive/functional terms. The peak location corresponds to the orientation at which the Gaussian curve is centred, i.e. the preferred orientation band for identity recognition. The standard deviation represents the width of the curve, reflecting the strength or selectivity of the tuning. The base amplitude is the height of the Gaussian curve base, indicating the minimum level of sensitivity, typically found near vertical orientation. Finally, the peak amplitude refers to the height of the Gaussian curve relative to its baseline, that is, it captures the advantage of horizontal over vertical orientations.

      Moreover, the use of a nonlinear, Gaussian model is motivated by past work that showed that the Gaussian function fits the evolution of recognition performance as a function of orientation (Dakin & Watt, 2009; Goffaux & Greenwood, 2016). Orientation selectivity at primary stages of visual processing has also been modelled using Gaussian (or Difference of Gaussians; Ringach, Hawken, & Shapley, 2003).

      We revised the data analysis section to include a justification for our use of a Gaussian model: “Therefore, fitting the human sensitivity data could be fitted using a simple Gaussian model. seemed most appropriate as it allows characterizing the parameters of the tuning profile, namely, peak location, peak amplitude, standard deviation and base amplitude, which are directly interpretable in cognitive/functional terms. Moreover, the use of a nonlinear, Gaussian model is motivated by past work that showed that the Gaussian function fits the evolution of recognition performance as a function of orientation (Dakin & Watt, 2009; Goffaux & Greenwood, 2016). Simpler frameworks, i.e. a linear model predicting d’ from the interaction between orientation and viewpoint, would result in an 8 (orientation) * 7 (viewpoint) design that is difficult to analyze and interpret.”

      (2) When reporting the luminance and contrast of your stimuli, please make clear what these units and measures are. This was a case where I had to take a second to assure myself that I knew what the values meant.

      We clarified that the luminance and contrast values reported in the manuscript are on a grey scale ranging from 0 to 1.

      (3) In your Procedure section, I think describing the familiarization task right away would help the text flow more clearly. At present, you began talking about the old/new task, and I was immediately wondering how familiarization worked!

      The procedure section now starts with the description of the familiarization task.

      (4) p. 3 - "Culminates" doesn't seem like the right word here.

      We agree and rephrased this way: ‘The tolerance of face identity recognition is stronger for familiar than unfamiliar faces’.

      (5) p. 5 - I think "with the multiple" shouldn't have "the".

      Indeed, we removed the “the”.

      Reviewer #2 (Recommendations for the authors):

      I enjoyed reading the manuscript, but thought the Introduction was a bit long. I wasn't sure about the relevance of the section on temporal contiguity. I think this might have been more relevant if this had been a manipulation in the design. So, I wonder if this might be shortened or removed to focus on the key questions. On the other hand, I found the overview of the view-selective and view-tolerant to be a bit brief. There is plenty of detail here, but I found it difficult to break down what was done when I first read it. It might be good to provide an overview in the Discussion too.

      While past research on the contribution of temporal contiguity to face identity recognition brings interesting insights into the nature of the visual experience leading to view-tolerant performance, we agree with the Reviewer that this aspect is not directly at stake here. We reduced the review of this literature in the Introduction.

      We clarified the description of the model observers as suggested by the reviewer and made sure to provide an overview of the model observers in the Discussion as well.

      References.

      Burton, A. M., Jenkins, R., & Schweinberger, S. R. (2011). Mental representations of familiar faces. Br J Psychol, 102(4), 943-958. https://doi.org/10.1111/j.2044-8295.2011.02039.x

      Burton, A. M., Kramer, R. S., Ritchie, K. L., & Jenkins, R. (2016). Identity From Variation: Representations of Faces Derived From Multiple Instances. Cogn Sci, 40(1), 202-223. https://doi.org/10.1111/cogs.12231

      Collin, C. A., Rainville, S., Watier, N., & Boutet, I. (2014). Configural and featural discriminations use the same spatial frequencies: a model observer versus human observer analysis. Perception, 43(6), 509-526. https://doi.org/10.1068/p7531

      Dakin, S. C., & Watt, R. J. (2009). Biological "bar codes" in human faces. J Vis, 9(4), 2 1-10. https://doi.org/10.1167/9.4.2

      Dumont, H., Roux-Sibilon, A., & Goffaux, V. (2024). Horizontal face information is the main gateway to the shape and surface cues to familiar face identity. PLOS ONE, 19(10), e0311225. https://doi.org/10.1371/journal.pone.0311225

      Goffaux, V., & Greenwood, J. A. (2016). The orientation selectivity of face identification [Article de recherche] [peer-reviewed]. Scientific Reports, 6(34204), 34204. https://doi.org/10.1038/srep34204

      Gold, J., Bennett, P. J., & Sekuler, A. B. (1999). Identification of band-pass filtered letters and faces by human and ideal observers. Vision Research, 39(21), 3537-3560. http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Citation&list_uids=10746125

      Hansen, B. C., Essock, E. A., Zheng, Y., & DeFord, J. K. (2003). Perceptual anisotropies in visual processing and their relation to natural image statistics. Network, 14(3), 501-526. http://www.ncbi.nlm.nih.gov/pubmed/12938769

      Jenkins, R., White, D., Van Montfort, X., & Mike Burton, A. (2011). Variability in photos of the same face. Cognition, 121(3), 313-323. https://doi.org/10.1016/j.cognition.2011.08.001

      Jiang, F., Dricot, L., Blanz, V., Goebel, R., & Rossion, B. (2009). Neural correlates of shape and surface reflectance information in individual faces. Neuroscience, 163(4), 1078-1091. https://doi.org/10.1016/j.neuroscience.2009.07.062

      Keil, M. S. (2008). Does face image statistics predict a preferred spatial frequency for human face processing? Proc Biol Sci, 275(1647), 2095-2100. https://doi.org/10.1098/rspb.2008.0486

      Keil, M. S. (2009). "I look in your eyes, honey": internal face features induce spatial frequency preference for human face processing. PLoS Comput Biol, 5(3), e1000329. http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Citation&list_uids=19325870

      Menon, N., Kemp, R. I., & White, D. (2018). More than a sum of parts: robust face recognition by integrating variation. R Soc Open Sci, 5(5), 172381. https://doi.org/10.1098/rsos.172381

      Mike Burton, A. (2013). Why has research in face recognition progressed so slowly? The importance of variability. Quarterly journal of experimental psychology, 66(8), 1467-1485. https://doi.org/10.1080/17470218.2013.800125

      Näsänen, R. (1999). Spatial frequency bandwidth used in the recognition of facial images. Vision Research, 39(23), 3824-3833. http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Citation&list_uids=10748918

      Oruc, I., Shafai, F., Murthy, S., Lages, P., & Ton, T. (2019). The adult face-diet: A naturalistic observation study. Vision Res, 157, 222-229. https://doi.org/10.1016/j.visres.2018.01.001

      Ringach, D. L., Hawken, M. J., & Shapley, R. (2003). Dynamics of orientation tuning in macaque V1: the role of global and tuned suppression [Research Support, Non-U.S. Gov't

      Research Support, U.S. Gov't, P.H.S.]. Journal of neurophysiology, 90(1), 342-352. https://doi.org/10.1152/jn.01018.2002

      Russell, R., Biederman, I., Nederhouser, M., & Sinha, P. (2007). The utility of surface reflectance for the recognition of upright and inverted faces. Vision Res, 47(2), 157-165. https://doi.org/10.1016/j.visres.2006.11.002

      Russell, R., Sinha, P., Biederman, I., & Nederhouser, M. (2006). Is pigmentation important for face recognition? Evidence from contrast negation. Perception, 35(6), 749-759. https://doi.org/10.1068/p5490

      Troje, N. F., & Bulthoff, H. H. (1996). Face recognition under varying poses: the role of texture and shape. Vision Res, 36(12), 1761-1771. https://doi.org/10.1016/0042-6989(95)00230-8

      Vuong, Q. C., Peissig, J. J., Harrison, M. C., & Tarr, M. J. (2005). The role of surface pigmentation for recognition revealed by contrast reversal in faces and Greebles. Vision Res, 45(10), 1213-1223. https://doi.org/10.1016/j.visres.2004.11.015

    1. Author response:

      The following is the authors’ response to the original reviews.

      In the revised manuscript, we have implemented several substantive changes. Most notably, we have revised the statistical reporting throughout to use Wald z statistics and GLMM-based contrasts, replacing the previously reported F statistics and figure caption t-tests. We have also expanded the Discussion to more explicitly acknowledge interpretational caveats regarding the null tuning width result and to address the alternative explanation of general alertness or motivational changes. Throughout the manuscript, we have revised our language to ensure that our conclusions are appropriately calibrated to the data.

      Reviewer #1 (Public review):

      Summary:

      The authors attempt to use a combination of behavioural and EEG analyses in order to investigate whether expectation of task difficulty influences spatial focus narrowing in the context of a spatially cued task, alongside an expected attention-related amplitude effect. This distinguishes the experiment from previous tasks, which looked at this potential spatial narrowing in the context of more non-cued diffuse attention tasks. The authors present two major findings:

      (1) Behaviourally, they analysed the effects of cue validity and difficulty expectation on response accuracy, and found that participants displayed an effect of difficulty expectation in validly cued trials, showing relatively enhanced behaviour to Hard Expectation trials, but no effect of expectation in invalidly cued trials.

      (2) Inverted encoding modelling on broadband EEG showed greater pre-target attentional processing in the Hard Expectation blocks. They go on to show that this enhancement comes in the form of greater amplitude of the Channel Tuning Functions (CTFs) approximately 300 to 400ms post-cue, in the absence of any spatial tuning specificity enhancement (as would be evident in a difference in CTF fit width).

      Together, these results provide valuable findings for those investigating the separable effects of expectation and attention on target detection in visual search.

      Strengths:

      (1) This is a very solidly performed experiment and analysis, with different streams of evidence convincingly pointing in the same direction, i.e. a gain effect of Expectation in the absence of a spatial tuning effect.

      (2) EEG is competently analysed and interpreted, and the paper is well written and simple in its motivation.

      (3) The authors report appropriately on the results in the Discussion, without overreaching. 

      Weaknesses:

      I mainly have a few minor issues for the authors to clarify, which I will leave to Recommendations. However, a few analyses need further work:

      We thank Reviewer 1 for the overall positive evaluation of our work and for the constructive and detailed feedback. The reviewer highlighted several strengths of the study, including the convergent evidence across behavioral and neural measures, the competent EEG analysis, and the appropriateness of the Discussion. In response to the specific recommendations, we have: clarified the type of EEG analysis in the Abstract; revised the description of the Serences et al. (2004) finding in the Introduction; added a Figure 1 reference in the relevant paragraph; clarified the logic of the planned comparisons; corrected and updated Figure 2 and its caption; added clarifying information about the EEG analysis in the Results; corrected the ambiguous reference to stimulus onset; clarified the status of edge-marked participants in Figure 4a; and added caveats and clarifications regarding the decoding analysis. We also address the two analytical concerns raised under Weaknesses below.

      (1) The GLMM method used has very large degrees of freedom (pages 6 and 7) of 34542. I assume this is the number of trials minus the number of parameters? This would imply that random slopes were not modelled in the analyses. However, looking at the Methods, it is reported that they were modelled. The authors should clarify exactly what was done here and why, including the LMM model. 

      We thank the reviewer for raising this point. The previously reported denominator degrees of freedom (e.g., 34,542) reflected the number of trial-level observations used in the model and arose from reporting Type III Wald F-tests. We agree that this reporting format may have been misleading in the context of generalized linear mixed-effects models (GLMMs), where inference does not rely on classical denominator degrees of freedom in the same way as traditional ANOVA.

      To improve clarity, we have revised the manuscript to report fixed effects using Wald z statistics derived from the model summary, which is the standard approach for binomial GLMMs implemented in lme4. We no longer report F statistics or denominator degrees of freedom. Importantly, all models included by-participant random intercepts and random slopes for all within-subject factors (Expectation, Search condition, and Cue validity), as specified in the Methods. These random effects account for the non-independence of trial-level observations within participants and ensure that statistical uncertainty is estimated at the participant level rather than the trial level. We have clarified the random-effects structure explicitly in the revised Methods section.

      The revised reporting yields the same overall pattern of results, with the key planned comparison remaining significant.

      (2) Figure 4 shows an "example CTF fit". Why only one? You could put transparent lines in the background for each individual fit, followed by the grand average, or show each fit in the supplementary section?

      We thank the reviewer for this suggestion. We would like to clarify that Figure 4 does not show an example single-subject CTF fit; it shows the CTF fit to the group-averaged data, i.e., the grand average across participants. The purpose of the figure is to illustrate the group-level tuning function. This is now clarified in the updated Figure caption.

      To convey individual differences, Figure 4a already presents the parameter estimates for each participant (width, amplitude, and baseline) as separate points, providing a clear view of variability across participants. We considered including individual CTF fits in the background, but this would make the figure crowded without adding interpretive value, since the individual parameters are already visualized.

      We could, if the reviewers prefer, include the individual fits in the Supplementary Material; however, we believe that the current presentation conveys both the group average and participant-level variation clearly.

      Reviewer #1 (Recommendations for the authors):

      (3) Specify what type of EEG results are found in the Abstract. It is broadband, but one might expect, e.g. Alpha analyses. 

      We thank the reviewer for this suggestion. We have added "broadband" to the Abstract when describing the EEG analysis approach, clarifying that the inverted encoding model was applied to broadband EEG data rather than a specific frequency band (e.g., alpha).

      “We applied inverted encoding models to broadband EEG data to reconstruct spatial channel tuning functions, enabling precise characterization of both the locus and breadth of attentional deployment.”

      (4) In the Intro, please clarify the Serences finding that they found enhanced activity at expected distractor locations. The interpretation is that this reflects preparatory tagging of where distractors will appear, possibly to facilitate their suppression once they arrive, rather than enhancement in the service of processing those locations. It is confusing as it is currently worded.

      We thank the reviewer for flagging this. We have revised the description of the Serences et al. (2004) finding to clarify that the enhanced activity at expected distractor locations is interpreted as preparatory tagging in service of subsequent suppression, rather than signal enhancement facilitating processing at those locations. The revised sentence now makes this interpretive distinction explicit.

      “Complementing these findings, Serences et al. (2004) used fMRI to show that preparatory attention when expecting high distractor interference selectively enhanced activity in early visual cortex at retinotopic locations corresponding to the expected distractor positions, an effect interpreted as preparatory tagging of distractor locations to facilitate their subsequent suppression.”

      (5) Page 6: refer to Figure 1 in the relevant paragraph.

      We thank the reviewer for this suggestion. We have added a reference to Figure 1 in the relevant paragraph to help orient the reader.

      (6) Page 7: I find the interaction confusing. The authors say there is an interaction of Expectation and Cue Validity, such that there is a larger cueing benefit when dense displays were expected. However, this leads one to expect planned comparisons between Valid vs Invalid for Easy then Hard expectations. However, that's not what is done, actually comparing Easy vs Hard for Valid then Invalid trials.

      We thank the reviewer for highlighting this potential source of confusion. We have clarified in the manuscript that the planned comparisons examined the effect of Expectation separately within valid and invalid trials, rather than comparing cueing effects (valid vs. invalid) within each Expectation level. This analytic approach was chosen to directly test our hypothesis regarding expectation-related modulation of performance at attended versus unattended locations. We hope this clarification makes the logic of the comparisons more transparent.

      “To identify the locus of this interaction, we examined the effect of Expectation separately within valid and invalid trials, allowing us to test whether expectations exerted their effects at both cued and uncued locations, or selectively at either cued or uncued locations.”

      (7) Page 7: Issue with asterisk in Figure 2. Text says it is not significant. Also, can you make the transparent grey lines more visible? Also, the inner plot shows two sets of lines, apparently easy and hard display results. Needs to be denoted.

      We thank the reviewer for these observations. We have made the following changes: (1) The pairwise comparisons reported in the figure caption have been replaced with contrasts derived from the GLMM using estimated marginal means, consistent with the statistical approach used throughout the manuscript. (2) We have corrected the asterisk annotation in Figure 2, which was incorrectly placed on a non-significant comparison. (3) We have increased the visibility of the transparent grey lines in the figure. (4) We have revised the figure such that it is visually clear that the legend from the main plot applies to the inset plot as well.

      (8) Page 8: Really need some info on the EEG analysis.

      We thank the reviewer for this suggestion. We have added a sentence to the Results section briefly explaining that CTF slope reflects the overall strength of spatially selective neural activity at the attended location, with steeper slopes indicating stronger spatial selectivity, before directing readers to the Methods for full technical details. We hope this provides sufficient context for readers less familiar with the IEM approach without overloading the Results with methodological detail.

      (9) Page 8: 100ms after stimulus onset = target or cue? From Figure 4, it seems to be a cue, but this really needs to be clarified.

      We thank the reviewer for catching this ambiguity. We have replaced "stimulus onset" with "cue onset" throughout the results section to make clear that the time course is locked to cue presentation rather than target onset.

      (10) Page 10: Figure 4a, are edge-marked participants outliers? Were they included in analyses?

      We thank the reviewer for this observation. The edge-marked data points in Figure 4a reflect the default matplotlib boxplot visualization, which flags points beyond 1.5 × IQR, and do not represent a formal outlier exclusion criterion. We have added a brief clarification to this effect in the figure caption. All participants were retained in the primary analyses. To confirm that these participants did not unduly influence the results, we conducted a sensitivity analysis excluding them. Notably, the flagged participant showed a pattern in the opposite direction to the group, and excluding this individual yielded a stronger and more consistent effect, suggesting that our primary analysis with all participants included represents a conservative estimate.

      (11) Page 11: Can't infer the same mechanism from the lack of decoding ability; it could be a signal-to-noise issue. However, one interesting question. How is it that the Encoding analysis worked out, but the Decoding analysis did not?

      We thank the reviewer for raising both points. We agree that chance decoding could in principle reflect limited sensitivity rather than a true null effect, and we have added a caveat acknowledging this in the manuscript. We have also added a clarifying sentence explaining the complementary nature of the IEM and decoding analyses: the IEM captures the strength of spatial tuning within each condition, whereas decoding tests whether spatial patterns differ between conditions. Amplitude modulation of a shared spatial pattern would not necessarily produce discriminable multivariate patterns, which explains why the IEM detected amplitude differences while decoding remained at chance. We hope this resolves the apparent paradox.

      Reviewer #2 (Public review):

      Summary:

      The authors set out to determine whether people can adjust how narrowly or broadly they focus attention in advance based on expectations about how difficult an upcoming visual task will be. Specifically, they aimed to test whether expecting a more demanding search leads to a narrower focus of attention or instead strengthens attention at the relevant location without changing its spatial extent.

      Strengths:

      The study addresses a timely and interesting question about how expectations influence the preparation of attention before a task begins. The experimental design is well-suited to isolating anticipatory effects by manipulating expectations about task difficulty independently of moment-to-moment stimulus information. The manuscript is clearly written, and the methods are described in sufficient detail to support transparency and reproducibility.

      Weaknesses:

      Despite the strengths of the design and the merit of the work, I have a few concerns regarding the analysis and the interpretation of the results.

      We thank Reviewer 2 for the positive assessment of the study and for the thoughtful and constructive feedback. The reviewer highlighted several strengths, including the timeliness of the research question, the suitability of the experimental design, and the clarity of the manuscript. In response to the concerns raised, we have: revised the statistical reporting throughout to use Wald z statistics and replaced figure caption t-tests with GLMM-based contrasts; added a caveat in the Discussion acknowledging that the absence of tuning width differences does not definitively rule out changes in attentional scope; and added a paragraph in the Discussion addressing the alternative explanation of general alertness or motivational changes. We address each concern in detail below.

      (1) I was somewhat confused by aspects of the behavioural analysis. I may be mistaken, but fixed effects in generalised mixed-effects models are more commonly reported using Wald statistics with beta coefficients rather than F statistics, and the very large degrees of freedom reported here are difficult to interpret. In particular, they appear closer to trial counts than to the number of participants, which raises questions about how statistical uncertainty is being estimated. This concern is compounded by the fact that different statistical approaches appear to yield different conclusions: the generalised mixed-effects models and the pairwise t-tests reported in the figure caption do not fully align. Moreover, the latter are not described in the Methods, and the justification for using them in the figure is not provided. Taken together, this makes it difficult to assess the strength of the behavioural evidence. The reported effects of expectation on behaviour also appear small, and there is no clear cost at uncued locations. This limited behavioural footprint makes it difficult to determine how robust the proposed preparatory mechanism is. It also complicates the interpretation of the neural findings as reflecting a general strategy for optimising task preparation.

      We appreciate this observation and agree that reporting Wald statistics is more appropriate for GLMMs. In the revised manuscript, we now report fixed effects as regression coefficients (β), standard errors, z values, and associated p values, rather than Type III F statistics. This reporting more directly reflects the estimation procedure used in lme4, where inference for binomial GLMMs is based on Wald z tests.

      We have also removed the reporting of large denominator degrees of freedom, which reflected the number of trial-level observations but may have been confusing in this context. All models included by-participant random intercepts and random slopes for the within-subject factors, ensuring that statistical uncertainty is appropriately estimated while accounting for the hierarchical structure of the data.

      Regarding the pairwise comparisons shown in the figure caption, these previously reflected conventional pairwise t-tests and have now been replaced with contrasts derived from the GLMM using estimated marginal means, consistent with the statistical approach used throughout the manuscript. We have clarified in both the Methods and Results sections that these contrasts are fully model-based and examine the effect of Expectation separately within valid and invalid trials.

      Overall, the revised reporting format aligns the statistical presentation more closely with current standards for GLMM analyses and improves interpretability, while leaving the substantive conclusions unchanged.

      (2) A central premise of the study is that, if observers proactively narrow their attentional focus when expecting difficult search, this should be reflected in sharper spatial tuning profiles. This prediction is presented as a diagnostic test of whether expectations modulate attentional scope. However, the absence of such sharpening is later taken as evidence that expectations do not alter spatial extent and instead operate exclusively through gain modulation. This inference may be stronger than the data allow. The lack of an observed difference in tuning width does not necessarily rule out changes in attentional scope, particularly if such changes are subtle, temporally limited, or not well captured by the spatial resolution of the approach. As a result, while the findings are consistent with a gain-based account, they do not definitively exclude the possibility that expectations also influence spatial extent, and the logic linking the original prediction to the final conclusion would benefit from a more cautious interpretation.

      We thank the reviewer for this important point. We agree that the absence of a tuning width difference does not definitively rule out changes in attentional scope, and we have added a caveat in the Discussion acknowledging that subtle or temporally limited changes may not be fully captured by the spatial resolution of the current approach. We have revised the relevant section to adopt a more cautious interpretation while maintaining that the current findings are most consistent with a gain-based account.

      “We note, however, that the absence of a tuning width difference should be interpreted with caution. Subtle or temporally limited changes in attentional scope may not be fully captured by the spatial resolution of the current approach, and we cannot definitively exclude the possibility that expectations also influence spatial extent under some conditions.”

      (3) The difference between easy and hard searches in the CTF slope is taken as evidence for enhanced preparatory spatial attention under high expected difficulty. However, these differences could also reflect broader changes in alertness or motivational state between blocks. The behavioural results show a small overall increase in accuracy in expect-hard blocks, which may be consistent with a more general increase in task engagement rather than a spatially specific preparatory mechanism. Although the authors decompose slope differences into amplitude and width parameters, the interpretation still relies on ruling out alternative, more global explanations for enhanced signal strength or reduced variability. This leaves some ambiguity as to whether the observed modulation reflects a specific adjustment of preparatory attention or a more general change in task state.

      We thank the reviewer for raising this important alternative explanation. We agree that a general increase in alertness or motivational state could in principle produce broader changes in neural signal strength. We have added a paragraph in the Discussion addressing this concern directly. We highlight two aspects of the data that argue against a purely global account: first, the behavioral benefit of expectation was selective to the cued location with no corresponding effects elsewhere, which is inconsistent with a global alertness account; second, multivariate decoding of expectancy condition remained at chance throughout the cue-target interval, indicating that the two conditions did not produce globally distinct patterns of broadband EEG activity. If general arousal were driving the amplitude differences, we would expect such global pattern differences to be detectable by the classifier. Together, these considerations suggest that the observed modulation reflects spatially specific preparatory gain enhancement rather than a general change in task state. We acknowledge, however, that we cannot fully rule out a contribution of motivational or arousal-related factors, and have added appropriate caveats to the Discussion.

      “A related concern is whether the amplitude enhancement observed in expect-hard blocks reflects a spatially specific preparatory mechanism or instead a more general change in alertness or motivational state. Several aspects of the data argue against a purely global account. First, the behavioral benefit of expectation was selective to the cued location, with no corresponding costs or benefits at uncued locations, suggesting that expectancy effects were spatially constrained rather than globally distributed. Second, if expect-hard blocks induced a broadly different neural state through general arousal or motivational engagement, this should manifest as a globally distinct pattern of broadband EEG activity that a multivariate classifier could detect. However, decoding accuracy remained at chance throughout the cue-target interval, indicating that the two expectancy conditions did not produce categorically different spatial patterns of neural activity. Together, these findings suggest that the observed amplitude modulation reflects spatially specific preparatory gain enhancement rather than a global change in task engagement.”

    1. Author response:

      [Note: The final version has been published in Brain, Behavior, and Immunity: https://doi.org/10.1016/j.bbi.2026.106473]

      eLife Assessment

      Rhis useful study raises interesting questions but provides inadequate evidence of an association between atovaquone-proguanil use (as well as toxoplasmosis seropositivity) and reduced Alzheimer's dementia risk. The findings are intriguing but they are correlative and hypothesis-generating with the strong possibility of residual confounding.

      We thank the editors and reviewers for characterizing our work as useful and for the opportunity to publish a Reviewed Preprint with a corresponding response. However, the statements in the Assessment characterizing the evidence as ‘inadequate’ and asserting a ‘strong possibility of residual confounding’ are factually incorrect as applied to our data and incompatible with the empirical findings presented in the manuscript. We have notified the editors of this factual inaccuracy. As the Assessment will be published as originally written, we provide clarification here to ensure an accurate scientific record for readers of the Reviewed Preprint.

      Our study shows that the association between atovaquone–proguanil (A/P) exposure and reduced dementia risk, first identified in a rigorously matched national cohort in Israel, is robustly reproduced across three independently constructed age-stratified cohorts in the U.S. TriNetX network (with exposure at ages 50–59, 60–69, and 70–79). In each cohort, individuals exposed to A/P were compared with rigorously matched individuals who received another medication at the same age and were then followed over a decade for incident dementia. Cases and controls were matched on all major established dementia risk factors: age, sex, race/ethnicity, diabetes, hypertension, obesity, and smoking status.

      Across all three strata, each containing more than 10,000 exposed individuals with an equal number of matched controls, we observed substantial and consistent reductions in cumulative dementia incidence (HR 0.34–0.51), extremely low P-values (10<sup>–16</sup> to 10<sup>–40</sup>), and continuously widening divergence of Kaplan–Meier curves over the follow-up period. To more rigorously exclude the possibility of unmeasured baseline differences in health status, we additionally performed, for the purpose of this response, comparative analyses of key indicators of frailty and clinical utilization, including emergency and inpatient encounters, as well as the prevalence of mild cognitive impairment prior to medication exposure (values provided below in response to Reviewer #2, Weakness 1). These analyses provide clear evidence showing no pattern suggestive of exposed individuals being medically or cognitively healthier at baseline.

      Taken together, these findings constitute a rigorously matched and independently replicated association across two national health systems, using TriNetX, the most widely cited real-world evidence platform in published cohort studies. Replication across three age strata, each with >10,000 exposed individuals, followed for a decade, and matched on all major known risk factors for dementia, meets the accepted epidemiologic definition of strong and reproducible evidence.

      Although we disagree with elements of the editorial Assessment that appear inconsistent with the empirical findings, we will proceed with publication of the current manuscript as a Reviewed Preprint in order to ensure timely dissemination of findings with meaningful implications for public health and dementia prevention. In this initial public version, the point-by-point responses below provide concise explanations addressing the critiques underlying the Assessment. A revised manuscript, incorporating expanded baseline comparisons across each TriNetX age stratum, additional stringent exclusions, and an expanded discussion that will address the remarks presented in this review, will be submitted shortly.

      Reviewer #1 (Public review):

      Summary:

      This useful study provides incomplete evidence of an association between atovaquone-proguanil use (as well as toxoplasmosis seropositivity) and reduced Alzheimer's dementia risk. The study reinforces findings that VZ vaccine lowers AD risk and suggests that this vaccine may be an effect modifier of A-P's protective effect. Strengths of the study include two extremely large cohorts, including a massive validation cohort in the US. Statistical analyses are sound, and the effect sizes are significant and meaningful. The CI curves are certainly impressive.

      Weaknesses include the inability to control for potentially important confounding variables. In my view, the findings are intriguing but remain correlative / hypothesis generating rather than causative. Significant mechanistic work needs to be done to link interventions which limit the impact of Toxoplasmosis and VZV reactivation on AD.

      We thank the reviewer for describing our study as useful and for highlighting several of its strengths, including the very large cohorts, sound statistical analyses, meaningful effect sizes, and the impressive CI curves. We also appreciate the reviewer’s recognition that our findings reinforce prior evidence linking VZV vaccination to reduced AD risk.

      Regarding the statement that the evidence remains incomplete due to “inability to control for potentially important confounding variables,” we refer to our introductory explanation above. As noted there, our analyses meet the accepted criteria for reproducible epidemiological evidence, and the assumption of uncontrolled confounding is contradicted by rigorous matching and by additional baseline evaluations. We fully agree that mechanistic work is warranted, and our epidemiologic findings strongly motivate such efforts.

      We address the reviewer’s specific comments in detail below.

      (1) Most of the individuals in the study received A-P for malaria prophylaxis as it is not first line for Toxo treatment. Many (probably most) of these individuals were likely to be Toxo negative (~15% seropositive in the US), thereby eliminating a potential benefit of the drug in most people in the cohort. Finally, A-P is not a first line treatment for Toxo because of lower efficacy.

      We agree that individuals in our cohort received Atovaquone-Proguanil (A-P) for malaria prophylaxis rather than for treatment of toxoplasmosis. However, this does not contradict our interpretation. Because latent CNS colonization by T. gondii is not currently considered clinically actionable, asymptomatic carriers are not offered treatment, and therefore would only receive an anti-Toxoplasma regimen unintentionally, through a medication prescribed for another indication such as malaria prophylaxis. Importantly, atovaquone is an established therapy for toxoplasmosis, including CNS disease, with documented efficacy and CNS penetration in current treatment guidelines. It is therefore reasonable to assume that, during the multi-week course typically administered for malaria prophylaxis, A-P would exert significant anti-Toxoplasma activity in individuals with latent CNS infection, potentially reducing or eliminating parasite burden even though the medication was not prescribed for that purpose.

      The reviewer notes that only ~15% of individuals in the U.S. are Toxoplasma-seropositive, based on surveys performed primarily in young adults of reproductive age (serologic testing is most commonly obtained in women during prenatal care). However, seropositivity increases cumulatively over the lifespan, and few reliable estimates exist for the age groups in which Alzheimer’s disease and dementia occur. Even if we accept the lower estimate of ~15% latent colonization in older adults, this proportion is still smaller than the lifetime cumulative incidence of dementia in the general population.

      Therefore, if latent toxoplasmosis contributes causally to dementia risk, and A-P is capable of eliminating latent Toxoplasma in the subset of individuals who harbor it, then a multi-week course of treatment—such as the one routinely taken for malaria prophylaxis—would be expected to produce a substantial reduction in dementia incidence at the population level, of the same order of magnitude reported here. A protective effect concentrated in a minority of exposed individuals is fully compatible with, and can mechanistically explain, the large overall reduction in risk that we observe.

      Finally, the reviewer notes that A-P is not a first-line treatment for toxoplasmosis due to assumed lower efficacy. This point does not undermine our results. Even a second-line agent, when administered over several weeks—as is routinely done for malaria prophylaxis—is expected to exert substantial anti-Toxoplasma activity. The long duration of exposure in large populations receiving A-P for travel provides a unique natural experiment that does not exist for other anti-Toxoplasma medications, which, when prescribed for their non-Toxoplasma indications, are not taken more than a few days. Thus, the widespread use of A-P for malaria prophylaxis allows a unique opportunity to evaluate long-term outcomes following inadvertent anti-Toxoplasma treatment.

      Moreover, “first line” recommendations in clinical guidelines refer to treatment of acute toxoplasmosis in immunosuppressed individuals, where tachyzoites are actively replicating. These guidelines do not consider efficacy against latent CNS colonization, which is dominated by bradyzoites, a biologically distinct form, in immunocompetent individuals. Therefore, the guideline hierarchy is not informative regarding which medication is more effective at clearing latent brain infection, the stage we consider most relevant to dementia risk.

      (2) A-P exposure may be a marker of subtle demographic features not captured in the dataset such as wealth allowing for global travel and/or genetic predisposition to AD. This raises my suspicion of correlative rather than casual relationships between A-P exposure and AD reduction. The size of the cohort does not eliminate this issue, but rather narrows confidence intervals around potentially misleading odds ratios which have not been adjusted for the multitude of other variables driving incident AD.

      We agree that prior to matching, A-P exposure may be associated with demographic features such as health or to travel internationally. However, this does not apply after matching. In all age-stratified analyses, exposed and control individuals were rigorously matched on all major risk factors known to influence dementia risk, including age, sex, race/ethnicity, smoking status, hypertension, diabetes, and obesity. Owing to the extremely large pool of individuals in TriNetX (~120M), our matching was performed stringently, producing exposed and unexposed cohorts that are near-identical with respect to the established determinants of dementia risk.

      The reviewer correctly identifies that large cohorts alone do not eliminate confounding; however, confounding must still be biologically and epidemiologically plausible. Any hypothetical confounder capable of producing a 50–70% reduction in dementia incidence over a decade would need to: (1) produce a very large protective effect against dementia; (2) be strongly associated with A-P exposure; and (3) remain entirely uncorrelated with age, sex, race/ethnicity, smoking, diabetes, hypertension and obesity, which have been rigorously matched. No such factor has been proposed. The suggestion that an unspecified ‘subtle demographic feature’ could produce effects of this magnitude remains hypothetical, and no such factor has been described in the dementia risk literature.

      If a specific evidence-supported confounder is proposed that meets these criteria, we would be pleased to test it empirically in our cohorts. In the absence of such a proposal, the interpretation that the association is merely “correlative rather than causal” remains speculative and does not negate the strength of a replicated, rigorously matched, long-term association across large cohorts in two national health systems.

      (3) The relationship between herpes virus reactivation and Toxo reactivation seems speculative.

      We respectfully disagree with the characterization of the herpesvirus–Toxoplasma interaction as speculative. The mechanism we describe is biologically valid, based on established virology and parasitology literature showing that latent T. gondii infection can reactivate from its bradyzoite state under inflammatory or immune-modifying conditions, including viral triggers. A published clinical report has documented CNS co-reactivation of T. gondii and a herpesvirus, explicitly noting that HHV-6 reactivation can promote Toxoplasma reactivation in neural tissue (Chaupis et al., Int J Infect Dis, 2016).

      Moreover, this mechanism is the only currently evidence-supported explanation that simultaneously and parsimoniously accounts for all of the epidemiologic observations in our study:

      (1) Substantially higher cumulative incidence of dementia in individuals with positive Toxoplasma serology, indicating that latent infection is a risk factor for subsequent cognitive decline;

      (2) Strong protective association following A-P exposure, a medication with established activity against Toxoplasma gondii, including in the CNS;

      (3) Independent protection conferred by VZV vaccination, observed consistently for two vaccines with distinct formulations (one live attenuated, one recombinant protein), whose only shared property is suppression of VZV reactivation;

      (4) Greater protective effect of A-P among individuals who were not vaccinated against VZV, consistent with a model in which dementia risk requires both herpesvirus reactivation and persistent latent Toxoplasma infection—such that reducing either factor alone (via VZV vaccination or anti-Toxoplasma suppression) substantially lowers risk.

      Taken together, these observations are difficult to reconcile under any alternative hypothesis.  

      To date, we are unaware of any other biologically coherent mechanism that can explain all four findings simultaneously. We would welcome any alternative explanation capable of accounting for these converging epidemiologic signals, as such a proposal could meaningfully advance the scientific discussion. In the absence of a competing explanation, the interaction between latent toxoplasmosis and herpesvirus reactivation remains the most parsimonious hypothesis supported by current knowledge.

      Finally, while observational studies are inherently limited in their ability to provide causal inference, the mechanism we propose is biologically grounded and experimentally testable. Our results provide a strong rationale for mechanistic studies and clinical trials, and warrant publication precisely because they generate a verifiable hypothesis that can now be evaluated directly.

      (4) A direct effect on A-P on AD lesions independent on infection is not considered as a hypothesis. Given the limitations above and effects on metabolic pathways, it probably should be. The Toxo hypothesis would be more convincing if the authors could demonstrate an enhanced effect of the drug in Toxo positive individuals without no effect in Toxo negative individuals.

      A direct effect of A-P on AD established lesions is indeed possible, and this hypothesis would be of significant therapeutic interest. However, we did not consider it within the scope of our epidemiologic analyses because all cohorts explicitly excluded individuals with existing dementia. Under these conditions, proposing a disease-modifying effect on established Alzheimer’s lesions based on our data would itself be speculative. Evaluating such a mechanism would be better answered by mechanistic or interventional studies rather than inference from populations without baseline disease.

      We also agree that demonstrating a stronger protective effect among Toxoplasma-positive individuals would be informative. Unfortunately, this “natural experiment” cannot be performed using the available data: Toxoplasma serology is rarely ordered in older adults, and A-P exposure is itself uncommon, resulting in a cohort overlap far too small to yield valid statistical inference (n≈25 in TriNetX).

      Thus, while both proposed hypotheses are scientifically attractive and merit further study, neither can be resolved using currently available real-world clinical data. Our findings provide the rationale to investigate both hypotheses experimentally, and we hope our report will motivate such studies.

      Reviewer #2 (Public review):

      Summary:

      This manuscript examines the association between atovaquone/proguanil use, zoster vaccination, toxoplasmosis serostatus and Alzheimer's Disease, using 2 databases of claims data. The manuscript is well written and concise. The major concerns about the manuscript center around the indications of atovaquone/proguanil use, which would not typically be active against toxoplasmosis at doses given, and the lack of control for potential confounders in the analysis.

      Strengths:

      (1) Use of 2 databases of claims data.

      (2) Unbiased review of medications associated with AD, which identified zoster vaccination associated with decreased risk of AD, replicating findings from other studies.

      We thank the reviewer for the thoughtful assessment and for noting key strengths of our work, including (1) the use of two large national databases, and (2) the unbiased discovery approach that replicated the widely reported association between zoster vaccination and reduced Alzheimer’s disease (AD) risk. We agree that these features highlight the validity and reproducibility of the analytic framework.

      Below we respond to the reviewer’s perceived weaknesses.

      Weaknesses:

      (1) Given that atovaquone/proguanil is likely to be given to a healthy population who is able to travel, concern that there are unmeasured confounders driving the association.

      We agree that, prior to matching, A-P exposure may correlate with demographic or health-related differences (e.g., ability to travel). However, this potential bias was explicitly controlled for in the study design. Across all three age-stratified TriNetX cohorts, exposed and unexposed individuals were rigorously matched on all major established dementia risk factors: age, sex, race/ethnicity, smoking status, obesity, diabetes mellitus, and hypertension. Comparative analyses confirm that these risk factors are equivalently distributed at baseline.

      As noted in our response to Reviewer #1, for any hypothetical unmeasured confounder to explain the results, it would need to satisfy three conditions simultaneously:

      (1) Be capable of producing a 50–70% reduction in dementia incidence sustained over a decade and across three distinct age strata (ages 50–79);

      (2) Be strongly associated with likelihood of receiving A-P;

      (3) Remain entirely uncorrelated with age, sex, race/ethnicity, smoking, diabetes, hypertension, or obesity, all of which were rigorously matched and balanced at baseline.

      No such factor has been proposed in the literature or by the reviewer. Thus, the concern remains hypothetical and unsupported by any measurable demographic or biological mechanism.

      Importantly, empirical evidence contradicts the notion of a “healthy traveler” bias:

      Emergency and inpatient encounter rates prior to exposure were comparable between A-P users and controls. Across the three age-stratified cohorts, emergency visits were similar or slightly higher among A-P users (EMER: 19.6% vs 16.4%, 19.9% vs 14.2%, 22.0% vs 14.8%), and inpatient encounters were effectively equivalent (IMP: 14.8% vs 15.2%, 17.7% vs 17.6%, 22.1% vs 22.2%). These patterns directly contradict the suggestion that A-P users were a healthier or less medically burdened population at baseline.

      Prevalence of mild cognitive impairment was not lower among A-P users and was, in fact, slightly higher in the oldest cohort. Across the three age groups, baseline diagnoses of mild cognitive impairment (MCI) were comparable or slightly higher among exposed individuals (0.1% vs 0.1%, 0.3% vs 0.2%, 1.1% vs 0.6%). These data contradict the suggestion that A-P users had superior baseline cognition.

      The strongest protective association occurred in the youngest stratum (age 50–59; HR 0.34). At this age, when nearly all individuals are sufficiently healthy to travel internationally, A-P uptake is the least likely to confound health status. A frailty-based “healthy traveler” hypothesis would instead predict the opposite pattern, with older adults showing the greatest apparent benefit, since health limitations are more likely to restrict travel in later life. In contrast, the protective association weakens with increasing age, empirically contradicting any explanation based on differential travel capacity.

      In conclusion, the empirical evidence directly contradicts the existence of a ‘healthy traveler’ effect.

      (2) The dose of atovaquone in atovaquone/proguanil is unlikely to be adequate suppression of toxo (much less for treatment/elimination of toxo), raising questions about the mechanism.

      A few important points should address the reviewer’s concern:

      In our cohorts, A-P was prescribed for malaria prophylaxis, as correctly noted. In this setting, it is taken for the entire duration of travel, plus several days before and after, typically resulting in many weeks of continuous exposure. This creates an unintentional but scientifically valuable natural experiment, in which a CNS-penetrating anti-Toxoplasma agent is administered for long durations.

      Atovaquone is an established treatment for CNS toxoplasmosis, has strong CNS penetration, and is included in current clinical guidelines for acute toxoplasmosis in immunocompromised patients, although at higher doses. Because latent, asymptomatic CNS colonization is not treated in clinical practice, there are currently no data establishing the dose required to eliminate bradyzoite-stage Toxoplasma in immunocompetent individuals.

      Our observations concern atovaquone–proguanil (A-P), a fixed-dose combination of atovaquone with proguanil, a DHFR inhibitor targeting a key metabolic pathway shared by malaria parasites and T. gondii. The combination has well-established synergistic effects in malaria prophylaxis and the same mechanism would be expected to enhance anti-Toxoplasma activity. This fixed-dose regimen has never been formally evaluated for toxoplasmosis treatment at prolonged durations or against latent bradyzoite infection.

      Our hypothesis does not require or imply complete eradication of Toxoplasma. A clinically meaningful reduction in latent cyst burden among the subset of colonized individuals may be sufficient to alter long-term disease trajectories. Thus, a population-level decrease in dementia incidence does not require universal clearance of infection, but only partial suppression or reduction of parasite load in susceptible individuals, which is entirely compatible with the known pharmacology and duration of A-P exposure.

      (3) Unmeasured bias in the small number of people who had toxoplasma serology in the TriNetX cohort.

      The relatively small number of older adults with Toxoplasma serology stems from current clinical practice: serologic testing is mostly performed in women during reproductive years due to risks in pregnancy, whereas in older adults a positive result has no clinical consequence and therefore testing is rarely ordered.

      Importantly, the seropositive and seronegative groups were drawn from the same underlying population of individuals who underwent serology testing, and the only difference between groups is the test result itself. Because the decision to order a test is made prior to and independent of the result, there is no plausible rationale by which the serology outcome (positive or negative) would introduce a bias favoring either group beyond the result of the test itself.

      Furthermore, the two groups were here also rigorously matched on all major dementia risk factors, including age, sex, race/ethnicity, smoking, diabetes, hypertension, and BMI, and these characteristics are similarly distributed between groups. A small sample size does not imply bias; it simply reduces statistical power. Despite this limitation, the observed association (HR = 2.43, p = 0.001) remains strongly significant.

      Finally, this result is consistent with multiple published studies reporting higher rates of Toxoplasma seropositivity among individuals with Alzheimer’s disease, dementia, and even mild cognitive impairment, such that our finding reinforces a broader and independently observed epidemiologic pattern. Importantly, in our cohort the serology testing clearly preceded dementia diagnosis, which supports the plausibility of a causal rather than merely correlative relationship between latent toxoplasmosis and cognitive decline.

      To conclude our provisional response, we thank the editor and reviewers for raising points that will be further addressed and expanded upon in the discussion of the forthcoming revision. We welcome transparent scientific dialogue and acknowledge that, as with all observational research, residual confounding cannot be eliminated with absolute certainty. However, we disagree with the overall Assessment and emphasize that our findings—reproduced independently across two national health systems and three age-stratified cohorts, each rigorously matched on all major determinants of dementia risk, meet, and in many respects exceed, current standards for high-quality observational evidence.

      Assigning the results to “residual confounding” requires more than speculation: it requires identification of a confounding factor that is (1) anchored in established dementia risk literature, (2) empirically plausible, and (3) quantitatively capable of generating a sustained ~50 percent reduction in dementia incidence over a decade. No such factor has been identified to date. We note that the assertion of “residual confounding” has not been supported by a specific, quantitatively plausible mechanism. A hypothetical bias that is both extremely large in effect and uncorrelated with all major risk factors is not statistically or biologically credible.

      The explanation we propose, reduction in dementia risk through elimination of latent Toxoplasma gondii, is biologically grounded, directly supported by independent epidemiologic literature, and uniquely capable of accounting for all convergent observations in our data. No alternative hypothesis has been put forward that can plausibly explain these findings.

      A revised version of the manuscript will be submitted shortly, incorporating expanded baseline analyses, with the strictest possible exclusion criteria (including congenital, vascular, chromosomal, and neurodegenerative disorders such as Parkinson’s disease), and complete tabulated comparisons. These data will further reinforce that the observed protective associations are not attributable to any measurable confounding. We also plan to enhance the discussion in order to address the points raised by the reviewers.

      In light of the expanded analyses, any reservations expressed in the initial Assessment can now be re-evaluated on the basis of the empirical evidence. The findings reported in our study meet, and in several respects exceed, current epidemiologic standards for high-quality observational research, clearly warrant publication, and provide a robust scientific foundation for future mechanistic and interventional studies to determine whether elimination of latent toxoplasmosis can prevent or treat dementia.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Torro et al. presented CellDetective, an open-source software designed for a user-friendly execution of single-cell segmentation, tracking, and analysis of time-lapse microscopy data. The authors demonstrated the applications of the software by measuring NK cell spreading events acquired with reflection interference contrast microscopy (RICM), as well as detecting target cell death events and their interaction with neighboring NK cells in a multichannel widefield microscopy dataset.

      Strengths:

      The segmentation (StarDist, Cellpose) and tracking (bTrack) modules implemented were based on existing and published software packages. The authors added the event detection, classification, and analysis modules to enable an end-to-end time-lapse microscopy data processing and analysis pipeline, complete with a graphical user interface (GUI). This minimizes the coding experience required from the user. The documentation that accompanies CellDetective is also adequate.

      Weaknesses:

      Given that the software was designed to improve user experience, such an approach also limits its scope and functionality and is currently capable of handling very specific types of experiments. Additionally, this reviewer has also encountered many technical difficulties (see documented bugs/crashes below) that have prevented an extensive exploration of all the functionality of CellDetective.

      We thank the reviewer for recognizing the interest of the end-to-end pipeline design and the value of the graphical interface for non-coding users.

      Scope and technical difficulties

      We acknowledge the technical difficulties encountered during the review and sincerely apologize for the inconvenience. Since v1.3.9, we have invested substantial effort into stability, testing, and documentation. All reported bugs have been corrected and the software has been extensively tested (see points 4–7 below). Furthermore, in response to the concern about the software being limited to specific experiments, we note that Celldetective has since been successfully applied to other biological contexts beyond the immunological assays presented in the article, including microbiology (10.1128/mbio. 03342-25) and stem-cell-related studies (10.3390/jimaging11100371) (see also the positive remarks of Reviewer #2 regarding applicability). We also point the reviewer to the expanded documentation, which now includes modality-agnostic how-to guides.

      Additionally, model transfer has been improved: retraining now freezes most layers by default, accelerating convergence and stabilizing fine-tuning for new datasets.

      Specifics:

      (1) The software can only handle 2D 'widefield' time-lapse imaging datasets. It should be noted that many studies that examine cell-cell interactions in vitro also used confocal microscopy and acquired the time-lapse images in 3D z-stacks to enable the reconstruction of entire cell volumes from multiple optical sections along the z-axis.

      Given that almost all of the implemented segmentation (StarDist, Cellpose) and tracking (bTrack) packages already support the handling of 3D datasets, it is unclear why CellDetective was designed to only work with 2D datasets.

      As noted above, extending the support for 3D images would allow the scope and utility of this software to be further extended for imaging studies acquired in z-stacks. As an example, the dense clustering of effector cells in Figure 4 had prevented accurate segmentation due to the 2D nature of the experimental dataset. More importantly, support for a 3D dataset could also allow for the tracking of fluorescent protein-based sub-cellular as well as membrane protein localization during cell-cell interactions.

      Furthermore, it also widens the potential applicability for analyzing datasets from 3D organoid imaging and perhaps even intravital two-photon microscopy.

      Scope and technical difficulties

      We thank the reviewer for this suggestion and maintain our position that Celldetective is purposefully designed for high-throughput, high-temporal-resolution 2D imaging. We have now articulated this rationale more clearly in the revised manuscript (see Discussion lines 414-417).

      Specifically, we emphasize that Celldetective's two core strengths — harnessing the statistical power of cell populations together with multiplexing biological conditions, and dynamic analysis of fast cellular events — both benefit from maximizing temporal resolution and field-of-view throughput. In our experience, Z-stack acquisition would reduce the achievable time resolution and throughput (in terms of captured events and parallel conditions) below acceptable levels for the minute-scale dynamics relevant to immunological assays.

      That said, the modular architecture and the choice of 3D-compatible backends (StarDist, Cellpose, bTrack) leave the door open for community-driven 3D extensions in the future. We note in the revised manuscript that Celldetective is "specifically optimized for highthroughput, high-temporal-resolution imaging of quasi-2D systems" and that "by prioritising temporal sampling over Z-axis depth, Celldetective enables the capture of rapid biological dynamics that are often the focal point of interaction studies, where Z-stacking would otherwise limit throughput or resolution."

      (2) The software in its current form only allows the broad demarcation of the cells examined into two populations: targets and effectors. This limits the number of cell populations that can be examined for their interactions. It might be more useful to just allow multiple user-defined populations instead of restricting the populations to target and effector cells only.

      Extension to more than 2 custom populations

      This has been fully implemented. Starting with version 1.4, Celldetective supports an arbitrary number of user-defined cell populations with user-chosen names. The restriction to "targets" and "effectors" has been removed. When creating a new experiment, users can now define any number of populations with custom labels (e.g., nk, rbc, macrophages, tumor_cells). The experiment configuration file stores this information in a generic [Populations] section. The control panel, segmentation, measurement, tracking, event detection, and neighbourhood modules all operate on these user-defined populations.

      We illustrate this in the documentation with a figure showing a 3-population configuration (see the "How to create a new experiment" guide).

      Neighbourhood analysis (interactions) currently supports pairwise interactions between any two populations (including same-population neighbourhoods), which covers the most common biologically relevant scenarios. We note that three-way (multipartite) interactions involve a substantially higher level of complexity and are, to our knowledge, not currently addressed by biologists in this context.

      (3) Similarly, subsetting of each of the populations could be made more intuitive. Although it is possible to define subsets of cells using the "Custom classification" function under the "Measure" module with user-defined parameters, visualization of multiple groups remains unintuitive and it appears that only one custom classified group can be selected and visualized at any given time in the Signal Annotator under Measurement instead of allowing visualization of multiple (custom defined) groups of cells in different colors. It is also unclear how, if possible at all, to visualize a custom group of cells in the Signal Annotator under the Detect Events module.

      Subsetting and visualization of multiple groups

      The reviewer noted that defining cell subgroups through the Custom Classification was unintuitive, and that only one classified group could be visualized at a time.

      We first clarify that Celldetective distinguishes between two visualization tools: the static measurement annotator (under Measure), which displays groups and characteristic groups on a per-frame basis, and the Event Annotator (under Detect Events), which displays event classes along temporal signal traces. The reviewer's request to visualize multiple customdefined groups in different colors falls under the measurement annotator.

      We have addressed this concern with a multi-label characteristic group workflow: users perform successive threshold classifications to isolate individual phenotypes of interest (e.g., "spread", "dead", "high-intensity"), then merge these binary columns into a single characteristic group via the table view (Math → Merge states…). Each combination of states is automatically mapped to a distinct label and color. This merged column can then be explored in the measurement annotator, effectively displaying all subgroups simultaneously in different colors.

      For more complex classification logic, the classification tool supports logical AND/OR operators for composing conditions, enabling flexible definition of subgroups without scripting.

      The Event Annotator, by contrast, operates on a single event class at a time by design, as it is intended for reviewing and annotating individual event types along temporal signal traces; multi-group visualization is not applicable in this context.

      Software issues:

      (4) When initially tested on v1.3.9, the Segment module could not be initiated (with the error message AttributeError: 'WindowsPath' object has no attribute 'endswith' when attempting to run segmentation).

      Update: this has been fixed in v1.3.9.post4 dated February 7th, 2025.

      (5) Further testing was then performed by downgrading the software to v1.3.1. While testing the ADCC demo experiment (https://celldetective.readthedocs.io/en/latest/adcc-example.html), the workflow was stuck at attempts to initiate the Detect Events step:

      AssertionError: No signal matches with the requirements of the model ['dead_nuclei_channel_mean', 'area']. Please pass the signals manually with the argument selected_signals or add measurements. Abort.

      (Update: fixed in the latest v1.3.9.post4 version dated February 7th, 2025)

      (6) Random bugs causing the software to crash. Example: switching characteristic to 'status_color' in the Signal Annotator under Measurement caused the software to crash (v1.3.9.post4):

      TypeError: ufunc 'isnan' is not supported for the input types, and the inputs could not be safely coerced to any supported types according to the casting rule 'safe'

      (7) Overall, when exploring the functionality of the software, there have been multiple instances of software crashes when clicking/switching around to show different parameters, etc.

      This reviewer understands the difficulties and time involved in bug fixing and hopes that the experience could have been much smoother and that the software behaves much more stably in order to maximize its useability.

      General stability — bug fixes and crash instances

      We have made comprehensive improvements to software stability since the review period:

      100+ bug fixes across v1.4.0–v1.5.0, systematically addressing crashes, edge cases, and error handling throughout the GUI.

      Expanded automated test suite: the project now includes 43 test files (26 GUI-level tests + 17 unit test files) covering segmentation, tracking, measurements, event detection, filters, preprocessing, neighbourhoods, viewers, table operations, and more.

      These tests run automatically via CI/CD on every commit.

      Lazy imports for heavy dependencies (e.g., TensorFlow) to reduce startup time and potential import-order crashes.

      Improved error handling: informative error messages instead of silent crashes; graceful fallbacks when optional dependencies are missing.

      Usage and stability can be verified via GitHub traffic statistics and CI/CD action metrics.

      Reviewer #2 (Public review):

      Summary:

      Immune assays enable the analysis of immune responses in vitro. These assays generate time series image data across several experimental conditions. The imaging parameters such as the imaging modality and the number of channels can vary across experiments. A challenge in the field is the lack of (open source) tools to process and analyze these data. R. Torro, et. al. developed an open source end-to-end pipeline for the analysis of image data from these immune assays. The pipeline is designed with a GUI and is suited for experimental biologists with no coding experience. The authors have incorporated several existing methods and tools for individual tasks such as for segmentation and cell tracking, and incorporated them with custom methods where necessary such as for tracking cell state transitions.

      Strengths:

      (1) The tool is extremely well-documented and easy to install.

      (2) Applicable to a wide variety of imaging modalities and analysis.

      (3) There are several different options for each step, such as segmentation using traditional methods or deep learning methods, and all the analysis steps are integrated in one place with a GUI. The no-coding requirement makes this a very powerful tool for biologists and has the potential to enable a wide variety of analyses.

      We are grateful for the recognition of the tool's documentation quality, ease of installation, and versatility.

      Weakness:

      (1) It would be good to provide documentation on how to make the tool applicable for applications and analysis other than for immune profiling since most methods integrated here are applicable well beyond immune profiling. For example, a user might want to use the tool just for the segmentation of their IF microscopy-images.

      Documentation for non-immune applications

      We have undertaken a major documentation overhaul following the Diátaxis framework (Tutorials, How-to Guides, Explanations, Reference). The documentation now includes:

      24 How-to guides covering individual tasks (segmentation, tracking, measurements, background correction, texture analysis, spot detection, channel alignment, survival analysis, interactions, event annotation, etc.), written in a modality-agnostic manner so that users from any application domain can follow them.

      Concept pages explaining key abstractions (data organization, population-specific segmentation, single-cell events, survival, neighbourhoods) without assuming an immunology context.

      Expanded tutorials, including the RICM spreading assay and the ADCC co-culture assay, which serve as worked examples that can be adapted to other biological systems.

      The overview now presents Celldetective as "an open-source Python platform designed for biologists to study interacting cell populations in multimodal time-lapse microscopy", explicitly broadening the scope beyond immune profiling.

      Additionally, the user-defined population naming (see Reviewer #1, point 2) naturally makes the tool more accessible to non-immunology users, as they are no longer constrained by "target/effector" terminology. The following articles from the literature refer to Celldetective in microbiology (10.1128/mbio.03342-25), for stem cells (10.3390/ jimaging11100371), or for CAR-T cells (10.1101/2025.06.24.661290v1, 10.1101/2025.07.25.666844v1), beyond the applications of this manuscript.

      (2) They applied Celldetective to two immune assays. The authors present the results from these assays and use the results to validate their assay. However, they have not included data that demonstrates results obtained via this pipeline are comparable to results obtained with other pipelines and/or if these results are consistent with what is expected in the literature.

      Comparison with other pipelines / literature validation

      We emphasize that most of the presented data are original and do not have published equivalents, making direct pipeline-to-pipeline comparison impossible in many cases. We note that, to our knowledge, no existing open-source pipeline performs the complete endto-end analysis that Celldetective offers (from preprocessing through segmentation, tracking, event detection, neighbourhood analysis, to population-level survival curves), making a head-to-head software comparison impractical. Nevertheless, some recent publications have tested the software for various features (10.1128/mbio.03342-25, 10.1101/2025.07.25.666844v1), and results are in line with existing solutions when comparison is possible.

      We reserve systematic comparison with traditional (non-microscopy-based) immunological assays for future dedicated studies, as we consider it out-of-scope for this software-focused manuscript.

      Additional items for the revised manuscript

      Manuscript changes (including private recommendations made by reviewers)

      Modifications or additions in text appear in red:

      Abstract: lines 15-17, 20-22, 24-26

      Introduction: lines 71-72

      Results: lines 91, 103, 127-137, 170-171, 196-201, 239-242, 250-252, 255-257, 261-264,

      266-269, 292-295, 303, 319-321

      Figure captions fig.1, fig. 2, fig. 3, fig. 5

      Discussion: 372-377, 384-387, 406-407, 414-417, 418

      Materials and Methods: lines 462-464, 542-546, 673-677, 684-685, 733-734

      Figure S10

      References have been updated.

      Article statistics (as of 30 Apr 2026)

      2799 views

      162 downloads

      7 citations

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor points:

      (1) For the study involving LAMP1 measurement, a representative image of LAMP1 antibody staining should be included.

      We added a new supplementary figure Fig S10 with reference to it on line 304

      (2) In Figure 5B, can the authors comment on the ostensibly higher effector velocity under HER2+ target conditions? Is this caused by variation within assay, and whether they have been confirmed in independent wells/experiments with the same conditions?

      Thanks to the reviewer for this remark; We have added a comment on lines 320-322

      As we don’ t have systematic replicates for this effect, we tentatively attribute it to a variation in the target cell coverage.

      (3) It is not clear why in the Signal Annotator under Measurement, the movie playback is performed with a user-draggable slider but in the Signal Annotator under Detect Events, the movie plays with only a Play/Stop button with no options to modify the playback speed or to advance the movie frame-byframe.

      This has been addressed in Celldetective v1.5.0. The Signal Annotator for event detection now provides frame-by-frame navigation buttons, an autoplay mode for natural playback of dynamics, and on-the-fly animation speed control, matching the functionality available in the Measurement viewer.

      Reviewer #2 (Recommendations for the authors):

      (1) Main text

      One major comment throughout the manuscript is that the experimental setup sections (ie lines 143149; 221-226 and interspersed in the section 'From a single time point to a dynamic readout of effector-target interactions') are hard to follow. It is clear that the tool can achieve what is described, but it is hard to follow why the experiments were set up that way and it took digging through methods and figure captions to understand what the setup was in terms of antibodies (what are the specificity, why are they chosen, what are they proxies for). Each section has some of that information but not jointly. It would help to have a high-level description of the experimental aim and then how this has been achieved in the setup with details on the antibodies, including targets and what their role would be. This would help with for instance understanding what the purpose of the PI stain as a measure introduced in line 245 is, or how the antibodies in experiment one relate to the ones in experiment 2, etc. The hardest part to parse was the section on effector-target interactions, specifically how the simulation is set up and why. The clarity of the manuscript could really benefit from a reworking of these paragraphs.

      The mentioned paragraphs have been rewritten following the reviewer’s suggestions.

      Lines 127-137: The RICM assay intro has been substantially rewritten in red, providing high-level experimental aim, bsAb function, surface preparation, and RICM rationale — all in a single coherent block.

      Lines 196-201: The ADCC assay intro is rewritten in red with clear description of bsAb purpose, cell types, HER2 variation, PI monitoring, and fluorescent labels.

      In the same spirit, we also added a biological context to introduce the last section of results on lines 261-264.

      It is not clear what implications the statement in line 267 has for the user.

      A comment was added on lines 239-242

      In line 191, it is stated that the position-based approach showed a spike that was not observed in the mask-based approach. It is not clear what the spike means, is it an artifact or a real phenomenon discovered by the position-based approach; this is important as the t_spread definition would differ depending on which segmentation is used

      A comment was added on line 167. It does not impact the definition of t_spread since the peak is observed during the spreading phase.

      In the co-culture assay, StarDist approach is used to segment the MCF7 cell line while Cellpose is to segment the NK cells. Please provide a rationale for selecting these differing approaches for segmentation.

      A justification was added on lines 265-267

      The impact of cell density was looked at for 32 micrometers, however, it is not clear why this cut-off was chosen.

      A justification was added on lines 250-252

      (2) Methods

      Lines 743/744: what type of manual adjustments? If important for usable, should be described in detail.

      Details have been added on lines 674-678

      If specifying what software was used for plots, then also mention which ones are used for exceptions.

      Details are provided in a new dedicated paragraph, lines 735-739

      (3) Discussion

      Conclusion in line 404 - direct protective effect, or just sampling effect?any data for either, or too strong a conclusion otherwise.

      We have added a short discussion on this topic, lines 372-377.

      Preliminary analysis of ADCC rates stratified by local target density and number of effector neighbours suggests that both factors contribute (unpublished data), and Celldetective's neighbourhood analysis module provides the tools to perform such stratified survival studies.

      I don't understand the implications in line 412, maybe just the wording choice. Prior studies in T cells could not resolve, but would now be feasible with celldetective? Or for T cells this is still not possible due to other experimental constraints?

      Thanks for this remark; indeed it could not be resolved yet for T cells, to our knowledge, but would be facilitated by celldetective.

      A comment was added on lines 385-388

      (4) Figures

      (a) Font sizes in all figures are generally too small.

      Fonts in all figures have been enlarged.

      (b) Figure 2

      F, G, H: clarify caption.

      F: single cells grey traces, average colored line?

      G: what's the confidence/error interval?

      H: State the statistic and meaning of the qualitative assessment.

      DONE

      (c)Figure 3:

      F/G: choice of 3.5 as neighbouring cell is not motivated; mode would have been at 4 and choosing a non-integer for cell counts seems strange from a biological perspective.

      A comment has been added in the caption.

      E/G: what is the error/confidence interval?

      DONE

      (d) Figure 5:

      A: error bars?

      ADDED

      (5) Minor typos/word choices

      (a) Typo in line 59 - double the.

      OK

      (b) Typo in line 214 - upper case U in middle of sentence.

      OK

      (c) Typo/word choice in 516/17 - cells were split? Kept instead of keep.

      OK

      Typo 706; missing space between time and using.

      OK

      All corrected

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper investigates the physical basis of epithelial invagination in the morphogenesis of the ascidian siphon tube. The authors observe changes in actin and myosin distribution during siphon tube morphogenesis using fixed specimens and immunohistochemistry. They discover that there is a biphasic change in the actomyosin localization that correlates with changes in cell shapes. Initially, there is the well-known relocation of actomyosin from the lateral sides to the apical surface of cells that will invaginate, accompanied by a concomitant lengthening of the central cells within the invagination, but not a lot of invagination. Coincident with a second, more rapid, phase of invagination, the authors see a relocalization of actomyosin back to the lateral sides of the cells. This 2nd "bidirectional" relocation of actin appears to be important because optogenetic inhibition of myosin in the lateral domain after the initial invaginations phase resulted in a block of further invagination. Although not noted in the paper, that the second phase of siphon invagination is dependent on actomyosin is interesting and important because it has been shown that during Drosophila mesoderm invagination that a second "folding" phase of invagination is independent of actomyosin contraction (Guo et al. elife 2022), so there appear to be important differences between the Drosophila mesoderm system and the ascidian siphon tube systems.

      Using the experimental data, the authors create a vertex model of the invagination, and simulations reveal a coupled mechanism of apicobasal tension imbalance and lateral contraction that creates the invagination. The resultant model appears to recapitulate many aspects of the observed cell behaviors, although there are some caveats to consider (described below).

      We thank the reviewer for the insightful summary and for bringing the important study by Guo et al. (2022) to our attention. We have now added a dedicated comparison with Drosophila ventral furrow invagination in the Discussion, explicitly highlighting that the second rapid folding phase in Drosophila does not require lateral contractility, whereas in our system lateral contractility is obligatory for the accelerated invagination stage.

      Strengths:

      The studies and presented results are well done and provide important insights into the physical forces of epithelial invagination, which is important because invaginations are how a large fraction of organs in multicellular organisms are formed.

      Thank you for this positive assessment and for recognizing the significance of our work in elucidating the physical mechanisms underlying fundamental morphogenetic processes. We have striven to provide a comprehensive and rigorous analysis, and are grateful for this encouraging feedback.

      Weaknesses:

      (1) This reviewer has concerns about two aspects of the computational model. First, the model in Figure 5D shows a simulation of a flat epithelial sheet creating an invagination. However, the actual invagination is occurring in a small embryo that has significant curvature, such that nine or so cells occupy a 90-degree arc of the 360-degree circle that defines the embryo's cross-section (e.g., see Figure 1A). This curvature could have important effects on cell behavior.

      Thank you for bringing up the issue of tissue curvature. In the initial version of our model, we treated the tissue as flat based on the local geometry of the anterior epidermis. Although the embryo at 13 hpf indeed possesses significant curvature, its overall transverse cross-section is approximately elliptical, and the region undergoing invagination is situated in a relatively low-curvature zone, occupying only a 30° ∼ 40° arc of the entire tissue. More importantly, the embryo undergoes anisotropic elongation and expansion, becoming significantly flattened during the accelerated invagination stage, eventually adopting a very flat geometry by 18 hpf. We have now included Figure 5—figure supplement 1 to clarify these global morphological transitions.

      Nevertheless, the curvature does exist during the early stages, and we agree that clarifying its potential role is essential. Therefore, in the revised manuscript, we have updated our vertex model to incorporate a simplified circular geometry. Furthermore, unlike Drosophila ventral furrow formation (Guo et al., eLife, 2022), the invagination here eventually forms a hollow tubular structure, which led us to introduce a surface bending stiffness term into the mode. Although global tissue growth is not explicitly modeled, we explored the impact of curvature by varying the initial system size. Our results demonstrate that the invagination process, driven by apico-basal tension imbalance and lateral contraction, is highly localized and remains robust across different curvatures.

      (2) The second concern about the model is that Figure 5 D shows the vertex model developing significant "puckering" (bulging) surrounding the invagination. Such "puckering" is not seen in the in vivo invagination (Figure 1A, 2A). This issue is not discussed in the text, so it is unclear how big an issue this is for the developed model, but the model does not recapitulate all aspects of the siphon invagination system.

      Thank you for pointing out this. In our experiments, the similar "puckering" shape is observed during the early stages of morphogenesis (~17 hpf, as seen in Figure 1A) when the tissue size is relatively small. However, this feature rapidly disappears as the tissue grows and the overall geometry becomes flatter. This suggests that "puckering" is more pronounced in highly curved epithelia, a phenomenon that aligns with mechanical expectations. Previous vertex models of Drosophila ventral furrow formation do not exhibit this effect (Brodland et al., 2010; Polyakov et al., 2014), because they modeled cells within a rigid unmovable boundary. However, in our system of siphon morphogenesis, a tubular structure ultimately forms in the epithelium without strong boundary constraints. Thus, the mechanical boundary conditions are basically different.

      Also, the formation of a hollow tubular structure—supported by strong F-actin accumulation at the tissue surface—indicates a bending stiffness of surface tissue (Figure 1), which we have incorporated into the model. This bending term enforces smooth curvature transitions, which can manifest as a "puckering" shape surrounding the invagination. In our previous flat-geometry model, this significant bending stiffness led to a "puckering" effect surrounding the invagination. In our updated curved vertex model, this phenomenon also exists and is found to be related to tissue curvature. By simulating a larger system with low curvature (N = 324 cells in Figure 6D), we find that this puckering is significantly reduced. This confirms that the shape discrepancy is a size-dependent effect of the bending constraints within a fixed system size that did not account for tissue growth. In biological development, continuous growth and flattening of the embryo diminish this effect (Figure 5—figure supplement 1), aligning our model's predictions.

      Furthermore, we note that the cell-cell adhesion between the surface epithelium and the internal bulk cells (a factor not explicitly captured in our current model) likely further suppresses such evagination in vivo, as outward puckering would necessitate the coordinated deformation of the underlying tissues. We aim to investigate the interplay between global growth and local active forces in future work. We have added a detailed description and mechanical explanation of these simulated shapes in the revised manuscript.

      (3) In Figure 2A, Top View, and the schematic in Figure 2C, the developing invagination is surrounded by a ring of aligned cell edges characteristic of a "purse string" type actomyosin cable that would create pressure on the invaginating cells, which has been documented in multiple systems. Notably, the schematic in Figure 2C shows myosin II localizing to aligned "purse string" edges, suggesting the purse string is actively compressing the more central cells. If the purse string consistently appears during siphon invagination, a complete understanding of siphon invagination will require understanding the contributions of the purse string to the invagination process.

      Thank you for this excellent observation. We agree that the ring-like actomyosin structure is a prominent feature during the initial stages of invagination, and its potential role warrants discussion. We carefully re-examined our data. Our analysis confirms that this myosin ring is most pronounced during the early initial invagination stage. This inward compression from the periphery would work in concert with apical constriction to help shape the initial invagination. However, this ring-like myosin pattern significantly diminishes during the accelerated invagination stage, indicating that sustained compression from the purse string is not required for the entire process. We have added a discussion of this point in the revised manuscript. We also agree with that future experiments using laser ablation or optogenetic inhibition specifically targeting this actomyosin ring would be valuable to further dissect its precise contribution during the early invagination stage, and we have noted this as a future direction in the Discussion.

      (4) The introduction and discussion put the work in the context of work on physical forces in invagination, but there is not much discussion of how the modeling fits into the literature.

      We thank the reviewer for this suggestion. We have now incorporated additional references and discussion regarding existing theoretical models and the physical forces involved in tissue invagination. These previous studies provided the foundational framework for our updated curved vertex model. We have also added an explanation of how our model differs from these existing works and discussed potential future directions for further investigation.

      Reviewer #2 (Public review):

      Summary:

      The authors propose that bidirectional translocation of actomyosin drives tissue invagination in Ciona siphon tube formation. They suggest a two-stage model where actomyosin first accumulates apically to drive a slow initial invagination, followed by translocation to lateral domains to accelerate the invagination process through cell shortening. They have shown that actomyosin activity is important for invagination - modulation of myosin activity through expression of myosin mutants altered the timing and speed of invagination; furthermore, optogenetic inhibition of myosin during the transition of the slow and fast stages disrupted invagination. The authors further developed a vertex model to validate the relationship between contractile force distribution and epithelial invagination.

      Thank you for your thoughtful and accurate summary of our work and for your constructive critique.

      Strengths:

      (1) The authors employed various techniques to address the research question, including optogenetics, the use of MRLC mutants, and vertex modelling.

      (2) The authors provide quantitative analyses for a substantial portion of their imaging data, including cell and tissue geometry parameters as well as actin and myosin distributions. The sample sizes used in these analyses appear appropriate.

      (3) The authors combined experimental measurements with computer modeling to test the proposed mechanical models, which represents a strength of the study. It provides a framework to explore the mechanical principles underlying the observed morphogenesis.

      We are grateful for your positive assessment of the multidisciplinary approaches, quantitative analyses, and the integration of modeling with experiments.

      Weaknesses:

      (1) The concept of coordinated and sequential action of apical and lateral actomyosin in support of epithelial folding has been documented through a combination of experimental and modeling approaches in other contexts, such as ascidian endoderm invagination (PMID: 20691592) and gastrulation in Drosophila (PMIDs: 21127270, 22511944, 31273212). While the manuscript addresses an important question, related findings have been reported in these previous studies. This overlap reduces the degree of novelty, and it remains to be clarified how their work advances beyond these prior contributions.

      We thank the reviewer for raising this important point. In the revised Introduction and Discussion, we have explicitly distinguished our findings from prior studies. Specifically: (1) Unlike ascidian endoderm invagination, where actomyosin shifts from apical to basolateral (Sherrard et al., 2010), our system exhibits a bidirectional redistribution between apical and lateral domains, with the basal domain playing a passive role. (2) Unlike Drosophila ventral furrow invagination, where lateral contractility is not essential for the second folding phase (Guo et al., 2022), our optogenetic inhibition demonstrates that lateral contractility is obligatory for the accelerated invagination stage. These comparisons, now clearly stated in the Introduction and Discussion, establish bidirectional actomyosin redistribution as a distinct mechanical paradigm for sequential morphogenesis. We believe these revisions adequately clarify how our work advances beyond prior contributions.

      (2) One of the central statements made by the authors is that the translocation of actomyosin between the apical and lateral domains mediates invagination. The use of the term "translocation" infers that the same actomyosin structures physically move from one location to another location, which is not demonstrated by the data. Given the time scale of the process (several hours), it is also possible that the observed spatiotemporal patterns of actomyosin intensity result from sequential activation/assembly and inactivation/disassembly at specific locations on the cell cortex, rather than from the physical translocation of actomyosin structures over time.

      We thank the reviewer for this important point. We agree that our data do not demonstrate physical translocation of actomyosin structures, and that the observed patterns could arise from sequential assembly/disassembly over time. To avoid overinterpretation, we have replaced “translocation” with “redistribution” throughout the manuscript (including the title) and toned down the language in the Results and Discussion.

      (3) Some aspects of the data on actomyosin localization require further clarification. (1) The authors state that actomyosin translocation is bidirectional, first moving from the lateral domain to the apical domain; however, the reduction of the lateral actomyosin at this step was not rigorously tested. (2) During the slow invagination stage, it is unclear whether myosin consistently localizes to the apical cell-cell borders or instead relocalizes to the medioapical domain, as suggested by the schematic illustration presented in Figure 2C. (3) It is unclear how many cells along the axis orthogonal to the furrow accumulate apical and lateral myosin.

      Thank you for your insightful comments, which will help us significantly improve the clarity and rigor of our actomyosin localization analysis. To address the points raised, we undertake several key revisions: First, we have added new quantitative analyses of active myosin intensity from earlier time points (14-15 hpf) to rigorously support the initial lateral-to-apical redistribution phase (Figure 2B). Second, the schematic in Figure 2C has been corrected to show myosin at the apical cell‑cell borders. We have clarified that redistribution occurs in a domain of approximately 15‑20 cells (the invagination primordium), not only the center cell.

      (4) The overexpression of MRLC mutants appears to be rather patchy in some cases (e.g., in Figure 3A, 17.0 hpf, only cells located at the right side of the furrow appeared to express MRLC T18ES19E). It is unclear how such patchy expression would impact the phenotype.

      Thank you for your observation. We acknowledge that mosaic expression is common in Ciona electroporation. For all quantitative analyses, we only selected embryos in which the central cell, along with more than half of the surrounding cells in the primordium, showed clear expression of the plasmid. This selection criterion has been added to the Materials and Methods section.

      (5) In the optogenetic experiment, it appears that after one hour of light stimulation, the apical side of the tissue underwent relaxation (comparing 17 hpf and 16 hpf in Figure 4B). It is therefore unclear whether the observed defect in invagination is due to apical relaxation or lack of lateral contractility, or both. Therefore, the phenotype is not sufficient to support the authors' statement that "redistribution of myosin contractility from the apical to lateral regions is essential for the development of invagination".

      We have performed the additional immunostaining experiment of myosin II. The new data (Figure 4—figure supplement 2) showed that light stimulation specifically reduced lateral myosin intensity without significantly affecting apical myosin compared to the dark control. Therefore, the observed block of invagination is primarily due to loss of lateral contractility.

      (6) The vertex model is designed to explore how apical and lateral tensions contribute to distinct morphological outcomes. While the authors raise several interesting predictions, these are not further tested, making it unclear to what extent the model provides new insights that can be validated experimentally. In addition, modeling the epithelium as a flat sheet and not accounting for cell curvature is a simplification that may limit the model's accuracy. Finally, the model does not fully recapitulate the deeply invaginated furrow configuration as observed in a real embryo (comparing 18 hpf in Figure 5D and 18 hpf in Figure 1A) and does not fully capture certain mutant phenotypes (comparing 18 hpf in Figure 5F and 18 hpf in Figure 3B right panel).

      Thank you very much for these helpful and constructive comments. We have addressed your concerns through the following model updates and clarifications.

      First, we have reformulated our vertex model from a flat sheet to a curved geometry that incorporates initial tissue curvature. We found that the core mechanical mechanism, mediated by the coupling of apical and lateral active contraction, consistently recapitulates the experimental invagination process. By independently inhibiting apical or lateral contractions in the model, we further clarified their distinct mechanical contributions to tissue bending and cell shortening.

      Regarding the model predictions concerning the apical-to-lateral redistribution of actomyosin in the original version (previously shown in Figure 6E-H), we agree that these lacked direct experimental validation in the current study and may have strayed from the primary focus on the invagination mechanism itself. Therefore, we have removed these predictive components from the revised manuscript. Instead, we have refocused our analysis on the robustness of the localized active process across tissues of varying sizes and curvatures, particularly because the in vivo invagination is accompanied by global tissue growth and geometry changes.

      Finally, we acknowledge that the simulated final shapes do not perfectly match the experimental geometry in every detail. We attribute these discrepancies to the omission of global tissue growth and the simplification of cell-cell adhesions between the surface epithelium and internal bulk cells. While these factors are not the primary drivers of the invagination, they undoubtedly refine the local morphology. We have added discussions of these limitations in the revised manuscript and aim to incorporate precise experimental measurements of tissue growth and inter-layer interactions in future modeling efforts.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript by Qiao et al., the authors seek to uncover force and contractility dynamics that drive tissue morphogenesis, using the Ciona atrial siphon primordium as a model. Specifically, the authors perform a detailed examination of epithelial folding dynamics. Generally, the authors' claims were supported by their data, and the conceptual advances may have broader implications for other epithelial morphogenesis processes in other systems.

      Thank you for your positive summary and for recognizing the broader implications of our work.

      Strengths:

      The strengths of this manuscript include the variety of experimental and theoretical methods, including generally rigorous imaging and quantitative analyses of actomyosin dynamics during this epithelial folding process, and the derivation of a mathematical model based on their empirical data, which they perturb in order to gain novel insights into the process of epithelial morphogenesis.

      Thank you for highlighting the strengths of our multidisciplinary methodology.

      Weaknesses:

      There are concerns related to wording and interpretations of results, as well as some missing descriptions and details regarding experimental methods.

      We have revised the manuscript to address your concerns regarding the wording and the details of the methodology.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Based on the feedback from the reviewers, a focus on the following major points has the potential to improve the overall assessment of the significance of the findings and the strength of the evidence:

      (1) It would be helpful to clearly articulate how these findings advance the field beyond what has already been demonstrated or suggested in other systems.

      We thank the editor for this helpful suggestion. To better articulate how our findings advance the field, we have revised both the Introduction and Discussion to explicitly contrast our system with previously studied invagination models. Specifically, we highlight that our work demonstrates a bidirectional redistribution of actomyosin between apical and lateral domains, which differs from the apical-to-basolateral shift reported in ascidian endoderm invagination. Moreover, we emphasize that lateral contractility is obligatory for the accelerated invagination stage in our system, whereas in Drosophila ventral furrow invagination the second folding phase can proceed without it. These comparisons have been clearly presented in the revised manuscript. We think our findings represent a distinct mechanical paradigm for sequential epithelial morphogenesis.

      (2) It would be helpful to clarify the meaning of "translocation" and more explicitly describe the temporal and spatial patterns of active myosin localization during the two steps of invagination.

      We have replaced the term “translocation” with “redistribution” throughout the manuscript, including the title. We have also added new quantitative analyses of active myosin intensity from earlier time points (14–15 hpf) to rigorously support the initial lateral-to-apical redistribution phase (Figure 2B). High-resolution top-view images have been included to show the ring‑like localization of myosin at the apical cell‑cell junctions during the initial stage (Figure 2A). The schematic in Figure 2C has been corrected to accurately reflect the predominant localization of active myosin at the apical cell‑cell borders.

      (3) It would be helpful to explain how the optogenetic data support the conclusion that "redistribution of myosin contractility from the apical to lateral regions is essential for the development of invagination".

      We have performed additional experiments combining optogenetic inhibition with subsequent immunostaining of active myosin II (anti-pS19 MRLC). We quantitatively compared the distribution of actomyosin in light‑stimulated versus dark‑control embryos. The new data show that after light exposure, lateral myosin intensity is significantly reduced compared to the dark control, whereas apical myosin levels decrease similarly in both groups. This indicates that the optogenetic manipulation effectively attenuates lateral contractility during the accelerated invagination stage without affecting concurrent apical contractility changes. These results directly support the conclusion that lateral contractility acquisition is essential for invagination progression. (Figure 4—figure supplement 2)

      (4) It would be helpful to describe how the modeling work fits within the existing literature on modeling epithelial folding and to address discrepancies between the model and the actual biological observations, such as tissue curvature, limited invagination depth in the model, and the "puckering" surrounding the invagination. In addition, certain descriptions of the modeling results should be clarified, as suggested by Reviewer #3.

      We thank the referees for the detailed and constructive comments on our modeling work. In response to these suggestions, we have significantly updated the theoretical section of the manuscript. Specifically, we have reformulated the vertex model within a curved geometry that represents the entire tissue, and revised the subsequent analyses to better clarify the mechanical principles driving the observed morphogenesis. We have added relevant references and discussed the mechanistic connections and distinctions between our model and previous studies on epithelial invagination. We hope that our point-by-point responses of the modeling work and the corresponding revisions in the manuscript adequately address the reviewers’ concerns.

      (5) It would be helpful to elaborate on the methods for quantitative image analysis and statistical tests.

      We have thoroughly expanded the Materials and Methods section by adding a dedicated subsection “Quantification and statistical analysis”. This subsection provides step‑by‑step descriptions of how apical, lateral, and basal domains were defined (segmented line, width 1 μm), how normalization was performed (basal intensity set to 1), how center cell height, invagination depth, and lateral cell distance were measured (referencing Figure 1B), and what statistical tests were used (two‑tailed Student’s t‑test, with significance levels indicated). (see revised Materials and Methods, “Quantification and statistical analysis” subsection)

      Reviewer #1 (Recommendations for the authors):

      (1) This reviewer has concerns about two aspects of the model. First, the model in Figure 5D shows a simulation of a flat epithelial sheet creating an invagination. However, the actual invagination is occurring in a small embryo that has very significant curvature, such that nine or so cells occupy a 90-degree arc of the 360-degree circle that defines the embryo's section (e.g., see Figure 1A). This curvature could potentially have important effects on cell behavior. Ideally, the developed model would reflect the actual geometry of the observed behavior. A more nuanced analysis would provide important insight into whether the embryo's curvature makes a difference. Importantly, any result comparing the planar versus curved system would be interesting because if the model worked equally well in the high curvature or planar systems, the model is robust, or if invagination requires different strategies for high curvature and for planar systems, this is an important finding that reveals the importance of local geometries. I don't think the consideration of invagination from a planar vs curved epithelium has been previously modeled.

      We fully agree with the reviewer that comparing planar versus curved systems provides valuable insights into the invagination mechanism. As we addressed in our response to Reviewer #1 (Public Review) - Weakness (1), we have now updated our vertex model to incorporate curved geometries and introduced surface bending stiffness to better reflect the embryo's actual shape. Our systematic comparison reveals that the invagination process, driven by apico-basal tension imbalance and lateral contraction, is indeed highly localized and remains robust across different initial curvatures. We have added Figure 5—figure supplement 1 and corresponding discussions in the revised manuscript to highlight these findings on model robustness and the role of local geometry.

      (2) The second concern about the model is that Figure 5D shows the vertex model developing significant "puckering" (evagination) surrounding the invagination. Such "puckering" is not seen in the in vivo invagination (Figures 1A, 2A). This issue is not discussed in the text, so it is unclear how big an issue this is for the developed model. A discussion of this issue in the text would be appropriate. Maybe puckering goes away if a curved epithelium is modeled?

      Thank you for this comment. In our model, the "puckering" effect naturally arises due to the presence of surface bending stiffness and the absence of rigid boundary constraints, which resembles the tissue morphology observed at 17 hpf in our experiments. However, our updated simulations show that this effect significantly diminishes as the tissue curvature decreases. We have addressed this concern in detail in our response to Reviewer #1 (Public Review) - Weakness (2) and have included the relevant analysis and discussions in the revised manuscript.

      (3) Because of the puckering, it is unclear in the model what measurement is being used to define the invagination depth in Figure 5E. Is the depth from the maximal height of the surrounding epithelial cells? Or the location of the apical surface before invagination begins? It would be helpful to have that parameter better defined, and it would also be helpful to add a line to Figure 5D showing how the reference point for invagination depth.

      Thank you for your suggestion. We measured the vertical distance from the baseline connecting the maximal height of apical midpoints of the surrounding cells to the apical surface of the center cell, which is consistent with our experimental measurements. We have now added a schematic line and indicators to Figure 5D.

      (4) In Figure 2A Top View, as well as the schematic in Figure 2C, the developing invagination is surrounded by a ring of aligned cell edges characteristic of a "purse string" type actomyosin cable that would create pressure on the invaginating cells, which have been documented in multiple systems. Notably, the schematic in Figure 2C shows myosin II localizing to aligned "purse string" edges, suggesting the purse string is actively compressing the more central cells. If the purse string consistently appears during siphon invagination, a complete understanding of siphon invagination will require understanding the contributions of the purse string to the invagination process. For this paper, a discussion of the possible involvement of a purse string would be helpful for the readers, but follow-up work could include laser cutting or optogenetic blockage of the purse string contractility.

      Thank you for your suggestion. We agree that the ring-like actomyosin structure is a prominent feature during the initial stages of invagination, and its potential role warrants discussion. We carefully re-examined our data. Our analysis confirms that this myosin ring is most pronounced during the early initial invagination stage (Figure 2A). This inward compression from the periphery would work in concert with apical constriction to help shape the initial invagination. However, this ring-like myosin pattern significantly diminishes in the accelerated invagination stage. We propose that the purse string may play a collaborative role in the early phase. We agree that follow‑up work (e.g., laser cutting or optogenetic manipulation) would be valuable and have noted this as a future direction in the Discussion.

      (5) The introduction and discussion put the work in the context of work on physical forces in invagination, but there is not much discussion of how the modeling fits into the literature. Did the current work advance the state of modeling of such phenomena? What were the strengths and limitations of the modeling in this paper compared to what has been done previously?

      Thank you for this suggestion. While we have incorporated additional literature in the revised manuscript as mentioned in our response to Reviewer #1 (Public Review) - Weakness (4), we would like to further clarify the specific advances and limitations of our modeling framework. Our updated vertex model builds upon established foundational frameworks but advances the state of modeling by: (i) incorporating dynamic apico-lateral tension variations coupled with actomyosin signals, and (ii) achieving localized, activity-mediated morphogenesis without the need for external rigid boundary constraints—a feature that distinguishes it from many classical models. We also recognize the model's current limitations. Specifically, it does not explicitly account for compressive stress and global geometric changes induced by tissue growth. The mechanical interactions between surface epithelial cells and the underlying internal bulk cells are also simplified. These factors represent important directions for our future work. We have added a dedicated paragraph in the Modeling and Discussion sections to contrast our model with existing literature and to explicitly state these strengths and limitations.

      (6) Figure 4D. Minor point, but the labeling on the X-axis is out of register with the bar graphs.

      We have corrected the alignment of the X‑axis labels with the bar graphs in Figure 4D. The figure has been updated accordingly.

      (7) Figure 4B does not have a scale bar.

      We have added a scale bar to Figure 4B (10 μm).

      Reviewer #2 (Recommendations for the authors):

      (1) Live imaging is necessary to demonstrate bidirectional translocation by visualizing the movement of the actomyosin network between the apical and lateral domains. Alternatively, a term other than "translocation" should be used to describe the observation.

      We agree that live imaging of actomyosin movement would be ideal but is technically challenging in this system. Instead, we have replaced the term “translocation” with the more accurate and conservative term “redistribution” throughout the manuscript, including the title, to avoid implying physical movement of the same molecules. This addresses the reviewer’s concern.

      (2) The optogenetic tool could be used to its full potential by manipulating myosin spatially or temporally, for example, by inhibiting myosin at various stages or subcellular locations, which would provide an opportunity to thoroughly test the domain and stage-specific needs for actomyosin. That said, I recognize that such experiments may be challenging in the model system used in this study.

      We thank the reviewer for this suggestion. We have indeed attempted spatially restricted optogenetic activation in the Ciona atrial siphon system, but found it technically very challenging due to tissue geometry and light scattering. We appreciate the reviewer's understanding of these technical limitations.

      (3) Some additional characterization of the optogenetics tool, such as the distribution of active myosin and F-actin post-stimulation, could further strengthen the interpretation of the inhibitory effect on invagination.

      We thank the reviewer for this suggestion. After optogenetic inhibition, we fixed and stained embryos for active myosin II. The results (Figure 4—figure supplement 2) show that light exposure significantly reduces lateral myosin intensity compared to the dark control, while apical myosin decreases similarly in both groups. This confirms that the optogenetic manipulation selectively attenuates lateral contractility without affecting apical changes. We have added this data to the Results section.

      (4) It would be helpful to address how heterogeneity in MRLC mutant overexpression might impact the interpretation of the outcome.

      We acknowledge that mosaic expression is common in Ciona electroporation. For all quantitative analyses, we only selected embryos in which the center cell and more than half of the surrounding cells in the primordium showed clear expression of the plasmid. This selection criterion has been added to the Materials and Methods section.

      (5) For Figure 2, it would be helpful to include the en face view of the cells at different apical-basal depths to better demonstrate the changes in the subcellular localization of myosin at different stages.

      We have added top‑view images in Figure 2A at both the apical and a deeper (lateral) plane. These images clearly show the ring‑like localization of active myosin at the apical cell‑cell junctions during the initial stage. Together with the cross‑sectional views, they adequately demonstrate the subcellular localization changes.

      (6) The Methods section should include more detailed descriptions of image quantification procedures. For example, for Figure 2B, how were the apical and lateral signals defined, and how were background intensities determined? In addition, the methods used for statistical tests should be clearly stated.

      We agree that detailed quantification procedures are essential. We have therefore expanded the Materials and Methods with a new subsection “Quantification and statistical analysis”. This subsection includes precise definitions of apical, lateral, and basal domains (segmented line, width 1 μm), background subtraction (region outside the tissue), normalization (basal intensity set to 1), and descriptions of how cell height, invagination depth, and lateral distance were measured (referencing Figure 1B). Statistical tests (two‑tailed Student’s t‑test) and significance levels are clearly stated.

      (7) The discrepancies between the model and experimental data, as described above, should be acknowledged. Commentary on how the model's assumptions and setup might contribute to these differences would be helpful.

      We thank the reviewer for this suggestion. As detailed in our response to Reviewer #2 (Public Review) - Comment (6), we have included the discrepancies between the model and experimental results in the Modeling and Discussion sections. We have added comments explaining how our key modeling assumptions might contribute to these differences. Specifically, while we have updated the model to a curved geometry, the omission of continuous global tissue growth and expansion could affect the final invagination depth and shape. Meanwhile, the neglect of mechanical interactions between the surface epithelium and the internal bulk cells prevents the model from fully capturing the constraints that refine the local furrow configuration in vivo. By clarifying these limitations, we now provide a more balanced view of the model's scope and its role in identifying the primary mechanical drivers of invagination.

      Reviewer #3 (Recommendations for the authors):

      General comments:

      (1) Methods: More information is needed to describe how imaging and quantification were performed. A couple of examples:

      (a) In Figure 1, how were the apical and basal surface area of the center cell quantified?

      (b) In Figure 1, Supplement 1, how was fluorescence intensity measured? Was there a constant area or volume that was quantified between samples? This is important because a decreasing apical surface can cause the signal to appear "concentrated" and increased.

      We thank the reviewer for this important suggestion. We have added a dedicated subsection “Quantification and statistical analysis” in the Materials and Methods. This subsection includes precise definitions of apical, lateral, and basal domains (segmented line, width 1 μm), background subtraction (region outside the tissue), normalization (basal intensity set to 1), and descriptions of how cell height, invagination depth, and lateral distance were measured (referencing Figure 1B). Statistical tests (two‑tailed Student’s t‑test) and significance levels are also stated.

      (2) The manuscript could use some editing and proofreading for grammar.

      The manuscript has been carefully edited for grammar and clarity. We thank the reviewer for the suggestion.

      Specific points:

      (1) Figure 1A: Could the authors please annotate the location of the center cell throughout the time course? This would make it easier for the reader to understand what is being quantified.

      We have added arrows to indicate the center cell at each time point in Figure 1A. This makes it easier for readers to follow the quantification.

      (2) Figure 1 Supplement 1A, Line 143, "...before 15 hpf, F-actin concentration decreased at the lateral domains..."

      It is not clear that the graph shows a decrease in the lateral domains when taking the error bars into account. It is possible that the F-actin concentration is stable in the lateral domains before 15 hpf. Are there some statistical analyses that can be performed?

      We re-analyzed the F-actin data and agree that the change before 15 hpf is not statistically convincing given the error bars. However, we have added new quantitative analysis of active myosin (p-MLC) at 14–15 hpf (Figure 2B), which shows a clear and significant shift from lateral to apical enrichment during this early phase. This myosin dynamic strongly supports our hypothesis of bidirectional redistribution. The corresponding text has been updated in the Results section.

      (3) Figure 1 Supplement 1A, Line 147-148, "...after 16 hpf, during which apical F-actin levels showed a gradual decline." Based on the graph, it does not appear that apical F-actin levels show a gradual decline after 16 hpf; rather, they may be steady or slightly increase.

      We agree with the reviewer. Our original statement was inaccurate. What we intended to emphasize was that at 16 hpf, the F-actin level at the lateral domain exceeded that at the apical domain. The detailed changes of F-actin after 16 hpf were not a focus of our discussion. We have revised the text accordingly to avoid any misinterpretation. The correction has been made in the Results section.

      (4) Figure 2C Hypothesis and line 169-170, "Initially, actomyosin translocated from the lateral regions to the apical domains..."

      Related to the comment above, it is not clear that one can state that the actomyosin "translocated". The quantification does not necessarily demonstrate a loss of actin at the lateral domain in the initial stage, and even if there was a loss of lateral actomyosin, one would require experiments (perhaps photoconversion experiments) to demonstrate that machinery from the lateral region was transferred to the apical surface, rather than a process of new assembly at the apical surface.

      We fully agree with the reviewer. We have replaced the term “translocation” with “redistribution” throughout the manuscript, including the title, to avoid implying physical movement of the same actomyosin structures. The text in the Results and Discussion has been revised accordingly.

      (5) A similar comment is relevant to the subsequent statement in line 175, "actomyosin translocated from the apical domains to the lateral regions." Without direct experiments to demonstrate movement of the actomyosin machinery, it is possible that there is de novo assembly of actomyosin in the lateral region rather than translocation.

      This wording ("translocation") becomes important primarily because it is in the title and appears to be one of the authors' major conclusions.

      We fully agree with the reviewer that the wording is critical given our main conclusion. We have therefore systematically replaced “translocation” with “redistribution” across the manuscript (title, results, and discussion).

      (6) Figure 4, Lines 215-216, "These results confirm that the redistribution of myosin contractility from the apical to lateral regions is essential for the development of invagination."

      This experiment did not specifically test the redistribution of myosin; rather, the authors demonstrated that myosin contractility globally is necessary for invagination. In these experiments, is it known where the myosin is?

      We have performed additional immunostaining experiments (new Figure 4—figure supplement 2) to directly examine myosin distribution after optogenetic inhibition. The results show that light exposure specifically reduces lateral myosin intensity compared to the dark control, while apical myosin decreases similarly in both groups. This demonstrates that the optogenetic manipulation selectively attenuates lateral contractility. We have revised the conclusion to state that the acquisition of lateral contractility is essential for invagination progression. The new data and revised text are in the Results section.

      (7) Figure 4B, minor point: It would be helpful if the authors included a timestamp for the bottom row images (Dark 1 h).

      Thank you for pointing out this typo. Timestamps have been added to the bottom row images (Dark 1 h) in Figure 4B.

      (8) Figure 5E, F, minor point: It seems that the label on the red curve has a typo; it should be T18ES19E (rather than T18AS19E).

      Thank you for pointing out this typo. We have corrected it in the revised manuscript (now Figure 6A, B).

      (9) Figure 5F and corresponding text: Can the authors please clarify what is meant by "Coupled mode" as marked in the schematic? Is this meant to refer to simultaneous apical constriction and lateral contraction? Or sequential?

      We thank the reviewer for this question. By "coupled mode," we refer to the mechanical synergy between apical and lateral contractions in driving the final invagination. As observed in our experimental data and recapitulated in the model, these two processes occur sequentially rather than simultaneously. We have revised the corresponding text to explicitly clarify this sequential process.

      (10) Figure 6A, B, Lines 274-275: "...the invagination depth increased significantly under higher alphaa (Figure 6A), while the central height remained relatively independent of alphaa (Figure 6B)." This caused me some confusion until I realized that "Figure 6B" might be a typo and should be Figure 6C.

      We sincerely apologize for this confusion. In the revised manuscript, this specific section and the corresponding figures have been updated.

      (11) Line 287, typo: I believe that "Figure 5B" should be Figure 6B.

      We sincerely apologize for this confusion. In the revised manuscript, this specific section and the corresponding figures have been updated.

      (12) Figure 6A, B, comparing invagination depth with varying apical or lateral actomyosin intensity: The authors state that "invagination depth increased significantly under higher alphaa", but describe "mild invagination depth variation" with varied lateral actomyosin intensity. The graphs seem to suggest that there is increased invagination depth when either apical or lateral actomyosin intensity is increased, and that the increase is to a similar extent. Can the authors comment on what they think the differences are, if the apical effect is "significant" but the lateral effect is "mild"?

      We thank the reviewer for this meticulous observation. We agree and feel sorry that our original description was not sufficiently precise. In the revised manuscript, we have re-analyzed the distinct contributions of apical and lateral tensions using the updated curved vertex model, which provides a more accurate mechanical decoupling. We have accordingly replaced the previous wording with a more rigorous description of the simulations and streamlined the corresponding figures to ensure the conclusions are clearly supported.

      (13) Figure 6H, Lines 307-309, "...stronger regional translocation and redistribution contribute to the rapid reduction in height of invaginating cells..."

      It appears from the graph that this is really only apparent at high alpha (total actomyosin); at empirically determined levels (alpha = 1), the effect of varying ratio is less dramatic. Can the authors comment on how significant they consider this effect?

      We thank the reviewer for this insightful comment. We agree that the theoretical predictions regarding translocation strength in the original model lacked sufficient experimental validation. To maintain the scientific rigor of our study, we have removed the sections concerning the translocation ratio and the corresponding Figure 6H from the revised manuscript. Instead, we now refocus our analysis on the core mechanical drivers of invagination that are directly supported by our observations. We also have added discussions acknowledging other factors not fully captured in the current model (e.g., tissue growth), which we aim to investigate in future work.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #4 (Public review):

      Summary:

      The authors demonstrate a computational rational design approach for developing RNA aptamers with improved binding to the Receptor Binding Domain (RBD) of the SARS-CoV-2 spike protein. They demonstrate the ability of their approach to improve binding affinity using a previously identified RNA aptamer, RBD-PB6-Ta, which binds to the RBD. They also computationally estimate the binding energies of various RNA aptamers with the RBD and compare against RBD binding energies for a few neutralizing antibodies from the literature. Finally, experimental binding affinities are estimated by electrophoretic mobility shift assays (EMSA) for various RNA aptamers and a single commercially available neutralizing antibody to support the conclusions from computational studies on binding. The authors conclude that their computational framework, CAAMO, can provide reliable structure predictions and effectively support rational design of improved affinity for RNA aptamers towards target proteins. Additionally, they claim that their approach achieved design of high affinity RNA aptamer variants that bind to the RBD as well or better than a commercially available neutralizing antibody.

      Strengths:

      The thorough computational approaches employed in the study provide solid evidence of the value of their approach for computational design of high affinity RNA aptamers. The theoretical analysis using Free Energy Perturbation (FEP) to estimate relative binding energies supports the claimed improvement of affinity for RNA aptamers and provides valuable insight into the binding model for the tested RNA aptamers in comparison to previously studied neutralizing antibodies. The multimodal structure prediction in the early stages of the presented CAAMO framework, combined with the demonstrated outcome of improved affinity using the structural predictions as a starting point for rational design, provide moderate confidence in the structure predictions.

      We thank the reviewer for this accurate summary and for recognizing the strength of our integrated computational–experimental workflow in improving aptamer affinity.

      Weaknesses:

      The experimental characterization of RBD affinities for the antibody and RNA aptamers in this study present serious concerns regarding the methods used and the data presented in the manuscript, which call into question the major conclusions regarding affinity towards the RBD for their aptamers compared to antibodies. The claim that structural predictions from CAAMO are reasonable is rational, but this claim would be significantly strengthened by experimental validation of the structure (i.e. by chemical footprinting or solving the RBD-aptamer complex structure).

      The conclusions in this work are somewhat supported by the data, but there are significant issues with experimental methods that limit the strength of the study's conclusions.

      (1) The EMSA experiments have a number of flaws that limit their interpretability. The uncropped electrophoresis images, which should include molecular size markers and/or positive and negative controls for bound and unbound complex components to support interpretation of mobility shifts, are not presented. In fact, a spliced image can be seen for Figure 4E, which limits interpretation without the full uncropped image.

      Thank you for your valuable comments and careful review.

      In response to your suggestion, we have now provided all uncropped electrophoresis raw images corresponding to the results in the main figures and supplementary figures (Fig. 2F, 3D, 3E, 4E, S9A, S10 and S11 of the original manuscript) in the revised version. Regarding the spliced image in Fig. 4E, the uncropped raw gel image clearly shows that the two C23U samples were run on an adjacent lane of the same gel due to the total number of samples exceeding the well capacity of a single lane. All samples were electrophoresed and signal-detected under identical experimental conditions in one single experiment, ensuring the validity of direct signal intensity comparison across all samples. These complete uncropped raw images have been supplemented in the revised manuscript as Fig. S12.

      The following highlighted words have been added to the revised manuscript.

      “All uncropped raw gel images corresponding to these EMSA experiments are provided in Supplementary Fig. S12.”

      Additionally, the volumes of EMSA mixtures are not presented when a mass is stated (i.e. for the methods used to create Figure 3D), which leaves the reader without the critical parameter, molar concentration, and therefore leaves in question the claim that the tested antibody is high affinity under the tested conditions.

      Thank you for your valuable comment on this oversight.

      For the EMSA assay in Fig. 3D, the reaction mixture (10 μL total volume) contained 3 μg of RBD protein and 3 μg of antibody (40592-R001), either individually or in combination, with incubation at room temperature for 20 minutes. Based on the molecular weights (35 kDa for RBD and 150 kDa for the IgG antibody), the corresponding molar concentrations in the mixture were calculated as 8.57 μM for RBD and 2 μM for the antibody. To ensure consistency, clarity and provide the critical molar concentration parameter, we have revised the legend of Fig. 3D, replacing the mass values with the calculated molar concentrations as you suggested.

      The following highlighted words have been added to the revised manuscript.

      “(D) Binding ability of the commercial antibody (40592-R001) to RBD was assessed by native-PAGE. The reaction mixture (10 μL) contained 8.57 μM RBD protein and 2 μM antibody, incubated individually or combined, followed by Coomassie brilliant blue staining.”

      Additionally, protein should be visualized in all gels as a control to ensure that lack of shifts is not due to absence/aggregation/degradation of the RBD protein. In the case of Figure 3E, for example, it can be seen that there are degradation products included in the RBD-only lane, introducing a reasonable doubt that the lack of a shift in RNA tests (i.e. Figure 2F) is conclusively due to a lack of binding.

      We sincerely appreciate your careful evaluation of our work, which helps us further clarify the experimental details and data reliability.

      First, we would like to clarify the nature of the gel electrophoresis in Fig. 3E: the RBD protein was separated by native-PAGE rather than denaturing SDS-PAGE. The RBD protein used in all experiments was purchased from HUABIO (Cat. No. HA210064) with guaranteed quality, and its integrity and purity were independently verified in our laboratory via denaturing SDS-PAGE (see revised Fig. S11), which showed a single, intact band without any degradation products. The ladder-like bands observed in the RBD-only lane of the native-PAGE gel are not a result of protein degradation. Instead, they arise from two well-characterized properties of recombinant SARS-CoV-2 Spike RBD protein expressed in human cells: intrinsic conformational heterogeneity (the RBD domain exists in multiple dynamic conformations due to its structural flexibility) (Cai et al., Science, 2020; Wrapp et al., Science, 2020) and heterogeneity in N-glycosylation modification (variable glycosylation patterns at the conserved N-glycosylation sites of RBD) (Casalino et al., ACS Cent. Sci., 2020; Ives et al., eLife, 2024), both of which could cause distinct migration bands in native-PAGE under non-denaturing conditions.

      Second, to ensure the reliability of the RNA-binding results, the EMSA experiments for determining the binding affinity (K<sub>d</sub>) of RBD to Ta, Tc and Ta variants were performed with three independent biological replicates (the original manuscript includes all replicate data in Fig. 2F and S9). Consistent results were obtained across all replicates, which effectively rules out false-negative outcomes caused by accidental absence or loss of functional RBD protein in the reaction system. In addition, our gel images (Fig. 2F and S9 in original manuscript) and uncropped raw images of all EMSA gels (Fig. S12 in revised manuscript) show no significant signal accumulation in the sample wells, confirming the absence of RBD protein aggregation in the binding reactions—an issue that would otherwise interfere with RNA-protein interaction and band shift detection.

      New results for RBD analysis by denaturing SDS-PAGE, along with the associated discussion, have been added to the revised manuscript (Fig. S11).

      References

      Cai, Y. et al. Distinct conformational states of SARS-CoV-2 spike proteins. Science 369, 1586-1592 (2020).

      Casalino, L. et al. Beyond shielding: the roles of glycans in the SARS-CoV-2 spike protein. ACS Cent. Sci. 6, 1722-1734 (2020).

      Ives, C.M. et al. Role of N343 glycosylation on the SARS-CoV-2 S RBD structure and co-receptor binding across variants of concern. eLife 13, RP95708 (2024).

      Wrapp, D. et al. Cryo-EM structure of the 2019-nCoV spike in the prefusion conformation. Science 367, 1260-1263 (2020).

      The following highlighted words have been added to the revised manuscript.

      “The integrity and purity of the RBD protein were confirmed by denaturing SDS-PAGE (Fig. S11), showing a single intact band without degradation. The multiple bands observed in native PAGE (e.g., Fig. 3E) are due to conformational and glycosylation heterogeneity [63–66] rather than protein degradation. To rule out non-specific aptamer–protein interactions, BSA was additionally included as a non-target protein control in EMSA assays; the wild-type Ta, the negative control Tc, and the optimized Ta<sup>G34C</sup> all showed only weak, comparable background signals with BSA but distinct target-specific binding to RBD (Fig. S10). Uncropped EMSA gel images (Fig. S12) and consistent results from three biological replicates (Fig. 2F and S9) confirm the absence of protein aggregation and ensure data reliability.”

      Finally, there is no control for nonspecific binding, such as BSA or another non-target protein, which fails to eliminate the possibility of nonspecific interactions between their designed aptamers and proteins in general. A nonspecific binding control should be included in all EMSA experiments.

      Thank you for this constructive comment.

      Following your recommendation, we have supplemented the EMSA assays with BSA as a non-target protein control to rule out non-specific binding between our designed aptamers (Ta, Tc and Ta<sup>G34C</sup>) and exogenous proteins. The results revealed that all three aptamers (Ta, Tc and Ta<sup>G34C</sup>) exhibited only weak and comparable background signals with BSA (Fig. S10), which may originate from BSA itself or trace contaminating proteins in the protein sample (Fig. S11). The similar intensities of these background signals across Ta, Tc, and Ta<sup>G34C</sup> indicate a comparable, low level of non-specific binding among these aptamers (Fig. S10). In sharp contrast, RBD displayed markedly stronger binding toward Ta<sup>G34C</sup> than Ta, while no detectable binding was observed with the negative control Tc (Fig. S10). Collectively, these results verify that the aptamer–RBD interactions characterized in this study are target-specific and exclude non-specific aptamer–protein interactions.

      All the new experimental data of the non-specific binding controls have been integrated into the revised manuscript (Fig. S10) and the corresponding results and Methods have been updated accordingly. The following highlighted words have been added to the revised manuscript:

      “To further exclude non-specific aptamer–protein interactions, we performed parallel EMSA assays using bovine serum albumin (BSA) as a non-target protein control for Ta, Tc, and the optimized Ta<sup>G34C</sup> (see Fig. S10). Only weak, comparable background signals were observed for all three aptamers with BSA. Such minor non-specific binding may originate from BSA itself or trace contaminating proteins in the BSA samples (Fig. S10). In contrast, markedly stronger binding was detected between RBD and Ta or Ta<sup>G34C</sup>, whereas no detectable binding was observed with the negative control Tc (Figs. 4E, S10). Such distinct binding profiles of aptamers with RBD and BSA confirm that the aptamer–RBD interactions characterized in this study are target-specific.”

      “To rule out non-specific aptamer–protein interactions, BSA was additionally included as a non-target protein control in EMSA assays; the wild-type Ta, the negative control Tc, and the optimized Ta<sup>G34C</sup> all showed only weak, comparable background signals with BSA but distinct target-specific binding to RBD (Fig. S10).”

      (2) The evidence supporting claims of better binding to RBD by the aptamer compared to the commercial antibody is flawed at best. The commercial antibody product page indicates an affinity in low nanomolar range, whereas the fitted values they found for the aptamers in their study are orders of magnitude higher at tens of micromolar. Moreover, the methods section is lacking in the details required to appropriately interpret the competitive binding experiments. With a relatively short 20-minute equilibration time, the order of when the aptamer is added versus the antibody makes a difference in which is apparently bound. The issue with this becomes apparent with the lack of internal consistency in the presented results, namely in comparing Fig 3E (which shows no interference of Ta binding with 5uM antibody) and Fig 5D (which shows interference of Ta binding with 0.67-1.67uM antibody). The discrepancy between these figures calls into question the methods used, and it necessitates more details regarding experimental methods used in this manuscript.

      Thank you for your insightful comments, which have helped us refine the rigor of our study. We address each of your concerns in detail below:

      First, we agree with your observation that the commercial neutralizing antibody (Sino Biological, Cat# 40592-R001) is reported to bind Spike RBD with low nanomolar affinity on its product page. However, this discrepancy in affinity values (nanomolar vs. micromolar) stems from the use of distinct analytical methods. The product page affinity was determined via the Octet RED System, a technique analogous to Surface Plasmon Resonance (SPR) that offers high sensitivity for kinetic and affinity measurements. In contrast, our study employed EMSA, a method primarily optimized for semi-quantitative assessment of binding interactions. The inherent differences in sensitivity and principle between these two techniques—with Octet RED System enabling real-time monitoring of biomolecular interactions and EMSA relying on gel separation—account for the observed variation in affinity values.

      Second, regarding the competitive binding experiments, we appreciate your note on the critical role of reagent addition order and equilibration time. To eliminate potential biases from sequential addition, we clarify that Cy3-labeled RNAs, RBD proteins, and the neutralizing antibody were added simultaneously to the reaction system. We have revised the Methods section to provide a detailed protocol for the EMSA experiments, to ensure full reproducibility and appropriate interpretation of the results.

      Third, we acknowledge and apologize for a critical error in the figure legends of Fig. 3E: the concentrations reported (5 μM aptamer and antibody 40592-R001) refer to stock solutions, not the final concentrations in the EMSA reaction mixture. The correct final concentrations are 0.5 μM for aptamer Ta, and 0.5 μM for the antibody. This correction resolves the apparent inconsistency between Fig. 3E and Fig. 5D, as the final antibody concentration in Fig. 3E is now consistent with the concentration range used in Fig. 5D. We have updated the figure legends for Fig. 3E and revised the Methods section to explicitly distinguish between stock and final reaction concentrations, ensuring clarity and internal consistency of the results.

      We sincerely thank you for highlighting these issues, which have prompted important revisions to improve the clarity, accuracy, and rigor of our manuscript.

      The following highlighted words have been added to the revised manuscript.

      “For competitive binding experiments, Cy3-labelled RNAs, RBD proteins, and neutralizing antibody 40592-R001 were added simultaneously to the EMSA buffer and incubated at room temperature for 20 min.”

      “(E) The RBD binding abilities of the aptamer Ta and commercial antibody 40592-R001 were compared by EMSA competitive binding experiments. The aptamer-RBD complex bands were shown after running on an agarose gel following the incubation of 40 μM RBD protein, 0.5 μM aptamer Ta, and 0.5 μM antibody 40592-R001 (final concentrations in the reaction mixture).”

      “(D) EMSA images of competitive binding experiments to characterize the RBD binding abilities of RNA aptamers (WT Ta and Ta<sup>G34C</sup>) and the commercial monoclonal SARS-CoV-2 neutralizing antibody 40592-R001. The aptamer-RBD complex bands were showed by running an agarose gel after incubation of 40 μM of RBD protein and 0.5 μM indicated aptamer with varying concentrations of the antibody 40592-R001. Final antibody concentrations ranged from 0 to 1.67 μM in the reaction mixtures. Results showed that Ta<sup>G34C</sup>, but not WT Ta, exhibited a higher binding affinity to the RBD proteins than that of the antibody.”

      (3) The utility of the approach for increasing affinity of RNA aptamers for their targets is well supported through computational and experimental techniques demonstrating relative improvements in binding affinity for their G34C variant compared to the starting Ta aptamer. While the EMSA experiments do have significant flaws, the observations of relative relationships in equilibrium binding affinities among the tested aptamer variants can be interpreted with reasonable confidence, given that they were all performed in a consistent manner.

      We sincerely appreciate your valuable concerns and constructive feedback, which have greatly facilitated the improvement of our manuscript. Regarding the flaws of the EMSA experiments you pointed out, we have provided a detailed response to clarify the related issues and supplemented necessary experimental details to enhance the rigor and reproducibility of our work (see corresponding answers in the point-to-point response letter). It is worth noting that EMSA remains a classic and widely used technique for studying biomolecular interactions, and its reliability in qualitative and semi-quantitative analysis of binding events has been well recognized in the field. Furthermore, we fully agree with and are grateful for your view that, since all tested aptamer variants were analyzed using a consistent experimental protocol, the observations on the relative relationships of their equilibrium binding affinities can be interpreted with reasonable confidence. This recognition reinforces the validity of the relative affinity improvements we observed for the G34C variant compared to the parental Ta aptamer, which is a key finding of our study.

      (4) The claim that the structure of the RBD-Aptamer complex predicted by the CAAMO pipeline is reliable is tenuous. The success of their rational design approach based on the structure predicted by several ensemble approaches supports the interpretation of the predicted structure as reasonable, however, no experimental validation is undertaken to assess the accuracy of the structure. This is not a main focus of the manuscript, given the applied nature of the study to identify Ta variants with improved binding affinity, however the structural accuracy claim is not strongly supported without experimental validation (i.e. chemical footprinting methods).

      We thank the reviewer for this comment and agree that experimental validation would be required to establish the structural accuracy of the predicted RBD–aptamer complex. We note, however, that the primary aim of this study is not structural determination, but the development of a general computational framework for aptamer affinity maturation. In most practical applications, experimentally resolved structures of aptamer–protein complexes are unavailable. Accordingly, CAAMO is designed to operate under such conditions, using computationally generated binding models as working hypotheses to guide rational optimization rather than as definitive structural descriptions. In this context, the predicted structure is evaluated by its utility for affinity improvement, rather than by direct structural validation. We have revised the manuscript to clarify this scope.

      The following highlighted words have been added to the revised manuscript.

      “We note that CAAMO is not intended to establish experimentally validated complex structures, but rather to provide preliminary binding models that enable rational affinity maturation of aptamers in scenarios where structural information is limited or unavailable.”

      “Overall, these results indicate that the proposed binding conformation of the aptamer Ta to the RBD serves as a plausible working binding model for structure-guided aptamer optimization, and demonstrate the great potential of our CAAMO framework in aptamer design and optimization.”

      “which supports the robustness of our approach in generating informative binding models for comparative analysis and affinity optimization of an RNA aptamer with a target protein.”

      “We believe that the predicted binding conformation represents a plausible member of the predicted ensemble that is functionally informative for guiding structure-based aptamer optimization, although it may not correspond to the exact native structure.”

      (5) Throughout the manuscript, the phrasing of "all tested antibodies" was used, despite there being only one tested antibody in experimental methods and three distinct antibodies in computational methods. While this concern is focused on specific language, the major conclusion that their designed aptamers are as good or better than neutralizing antibodies in general is weakened by only testing only three antibodies through computational binding measurements and a fourth single antibody for experimental testing. The contact residue mapping furthermore lacks clarity in the number of structures that were used, with a vague description of structures from the PDB including no accession numbers provided nor how many distinct antibodies were included for contact residue mapping.

      We thank the reviewer for this important comment regarding language precision, experimental scope, and clarity of the antibody dataset used in this study. We agree that the phrase “all tested antibodies” was imprecise and could lead to overgeneralization. We have carefully revised the manuscript to use more accurate and explicit wording throughout, clearly distinguishing between experimentally tested antibodies, computationally analyzed antibodies, and antibody structures used for large-scale contact analysis.

      Specifically, the experimental comparison in this study was performed using one commercially available SARS-CoV-2 neutralizing antibody, whereas free energy–based computational analyses were conducted on three representative neutralizing antibodies with available structural data. We have revised the text to explicitly state these distinctions and have avoided general statements referring to neutralizing antibodies as a class.

      Importantly, the residue-level contact frequency analysis was not based solely on these individual antibodies. Instead, this analysis leveraged a comprehensive set of experimentally resolved SARS-CoV-2 RBD–antibody complex structures curated from the Coronavirus Antibody Database (CoV-AbDab), a publicly available and actively maintained resource developed by the Oxford Protein Informatics Group. CoV-AbDab aggregates all published coronavirus-binding antibodies with associated PDB structures and provides a systematic and unbiased structural foundation for antibody–RBD interaction analysis. All available high-resolution RBD–antibody complex structures indexed in CoV-AbDab at the time of analysis were included to compute contact residue frequencies across the structural ensemble. We have now explicitly stated this data source, clarified the number and nature of structures used, and added the appropriate citation (Raybould et al., Bioinformatics, 2021, doi: 10.1093/bioinformatics/btaa739).

      Finally, we have revised the conclusions to avoid claims that extend beyond the scope of the data. The comparison between aptamers and antibodies is now framed in terms of representative antibodies and consensus interaction patterns derived from a large structural ensemble, rather than as a general statement about all neutralizing antibodies. These revisions improve the clarity, rigor, and reproducibility of the manuscript, while preserving the core conclusion that the CAAMO framework enables effective structure-guided affinity maturation of RNA aptamers.

      The following highlighted words have been added to the revised manuscript.

      “Notably, the aptamer Ta<sup>G34C</sup> exhibited the highest binding affinity to the RBD, outperforming the tested neutralizing antibodies in competitive binding assays.”

      “Since we determined the most probable binding model of the aptamer Ta to the RBD, comparing the binding properties of the aptamer Ta with those of representative neutralizing antibodies to the RBD is both feasible and meaningful.”

      “To further explore this, we analyzed the contact ratios of residues on the RBD bound to ACE2 (derived from MD simulations), to the aptamer Ta (derived from MD simulations), or to the neutralizing antibodies (derived from all available experimentally resolved SARS-CoV-2 RBD–antibody complex structures curated in the Coronavirus Antibody Database, CoV-AbDab [35]). CoV-AbDab is a publicly available, curated database that aggregates all published coronavirus-binding antibodies with associated structural information, providing a comprehensive and unbiased structural ensemble for contact frequency analysis.”

      “Notably, the Ta-RBD complex formation remained unchanged after adding the antibody (Fig. 3E), suggesting that the aptamer Ta exhibits binding capability comparable to the tested monoclonal neutralizing antibody.”

      “neutralizing antibodies (derived from all available SARS-CoV-2 RBD–antibody complex structures curated in CoV-AbDab).”

      “Our computational and experimental studies showed that the aptamer Ta has comparable binding abilities to the RBD compared to representative neutralizing antibodies analyzed in this study.”

      Overall, the manuscript by Yang et al presents a valuable tool for rational design of improved RNA aptamer binding affinity toward target proteins, which the authors call CAAMO. Notably, the method is not intended for de novo design, but rather as a tool for improving aptamers that have been selected for binding affinity by other methods such as SELEX. While there are significant issues in the conclusions made from experiments in this manuscript, the relative relationships of observed affinities within this study provide solid evidence that the CAAMO framework provides a valuable tool for researchers seeking to use rational design approaches for RNA aptamer affinity maturation.

      Recommendations for the authors:

      Reviewer #4 (Recommendations for the authors):

      The computational aspects seem to be the strength of this manuscript, however there remain some issues with experimental approaches. The previous reviewers concern with non-specific binding remains an issue that should be dealt with through additional experimentation. The indication of Tc showing no binding is a good control for nonspecific RNA binding by RBD, but does not address nonspecific protein binding by Ta or its derivatives. For example, if a variant of Ta bound strongly to hydrophobic or highly charged patches in binding sites, they could also bind strongly to hydrophobic or highly charged patches in other proteins. As such, a non-specific binding test should be included for all tested variants to show target-specific binding.

      Thank you for your constructive suggestion. To address the concern of non-specific binding, we have supplemented a dedicated control experiment using bovine serum albumin (BSA) as the non-specific protein target. The results demonstrated that Ta and its derivatives exhibited specific binding to the RBD protein. Detailed experimental procedures and corresponding results for this control assay are provided in our response to your first comment in this point-by-point response letter.

      There is a serious concern to me that all data (i.e. the triplicate EMSAs claimed in your study) are not shown, with only one EMSA replicate shown for each variant in the supplemental materials. Additionally, the manuscript does not include unedited gel images, with apparent splicing of images in Figure 4E. All raw data should be available for review, which includes unedited images of the entirety of each gel electrophoresis experiment. Moreover, internal controls (positive of Ta+/-RBD, negative of Tc+/-RBD, and aptamer+/-non-RBD-protein) should be included and shown in every EMSA experiment.

      Thank you for raising these critical concerns regarding the rigor and completeness of our EMSA experimental data. We highly appreciate your attention to detail, which helps us improve the quality and transparency of our manuscript.

      First, regarding the number of EMSA replicates, we have indeed performed triplicate EMSA experiments for each variant, and all three replicates are provided in the supplementary materials (Fig. S9 of the original manuscript). We have added explicit labels for each replicate in the revised Fig. S9 to avoid confusion, ensuring the reproducibility of our results is clearly demonstrated.

      Second, concerning unedited gel images, we fully agree with the importance of providing uncropped, raw gel images for peer review. In the revised manuscript, all unedited, full-length raw images of each gel electrophoresis experiment have been included in Supplementary Fig. S12, with clear annotations to correspond to the cropped images in the main text.

      Third, with respect to internal controls, we acknowledge the necessity of comprehensive internal controls for EMSA experiments to validate specific binding. For the EMSA assays of RBD with Ta and its variants (Fig. 4E), we have already included the full set of internal controls, namely the Ta-RBD positive control, Tc-RBD negative control, and non-RBD protein control. Notably, the K<sub>d</sub> values of RBD binding to Ta, Tc, and Ta variants are consistent with the signal intensity exhibited in the EMSA images, which further corroborates the reliability of our binding results. In addition, we have supplemented non-specific binding control data in the revised Supplementary Fig. S10, which fully validates the binding specificity between Ta/its derivatives and RBD and effectively rules out non-specific binding.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We have addressed all the reviewers’ comments through new experiments, additional analyses, or, in some cases, additional text. Below is a summary of the major changes in the manuscript.

      (1) We have added a considerable amount of new characterization of the biochemical enrichment of the ribosome clusters, including EM of the ribosome clusters, UV absorbance profiles, immunoblots of additional targets, and additional replicates (new Figure 1). In summary, we provide better evidence that (i) the biochemical enrichment is working and (ii) that the loss of FMRP has no effect on this biological enrichment of ribosomal clusters.

      (2) We have now reanalyzed all of the data in Figs. 5-8 using only the data after removing PCR duplicates from the RPFs. Other than the comparison between the nuclease treatments (Fig. 3), only this data is now used. Moreover, we have reanalyzed this data using suggestions from the reviewers, including providing PCA analysis (Fig S5-1), GSEA analysis (Fig 5), and normalizing for group size when comparing significance to total mRNAs, (Fig 6-7). We now also include a new analysis (Fig S7-1) to better explain how the loss of FMRP affects mainly FMRP targets defined by CLIP, but not all mRNAs resistant to run-off.

      (3) We are now more conservative in our nomenclature; we use "pellet" instead of "RNA granule (RG)" and "fraction 5/6" instead of "ribosome clusters (RC)". We have added a section to the discussion about the relationship between the RNA granules measured using imaging of hippocampal neurites and the biochemical purification of ribosome clusters in the pellet, as requested by the reviewers.

      (4) We have made many other minor changes to the text and analysis, which can be found in the specific response to the reviewers.

      (5) One major additional requested change that was not implemented was to repeat our experiments at different time points. We have added a paragraph to the discussion outlining (i) why this was not done and (ii) the caveats of our conclusions without this data being present.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors have investigated the role of FMRP in the formation and function of RNA granules in mouse brain/cultured hippocampal neurons. Most of their results indicate that FMRP does not have a role in the formation or function of RNA granules with specific mRNAs, but may have some role in distal RNA granules in neurons and their response to synaptic stimulation. This is an important work (though the results are mostly negative) in understanding the composition and function of neuronal RNA granules. The last part of the work in cultured neurons is disjointed from the rest of the manuscript, and the results are neither convincing nor provide any mechanistic insight.

      Strengths:

      (1) The study is quite thorough, the methods and analysis used are robust, and the conclusion and interpretation are diligent.

      (2) The comparative study of Rat and Mouse RNA granules is very helpful for future studies.

      (3) The conclusion that the absence of FMRP does not affect the RNA granule composition and many of its properties in the system the authors have chosen to study is well supported by the results.

      (4) The difference in the response to DHPG stimulation concerning RNA granules described here is very interesting and could provide a basis for further studies, though it has some serious technical issues.

      Thank you for these positive comments on the paper.

      Weaknesses:

      (1) The system used for the study (P5 mouse brain or DIV 8-10 cultured neuron) is surprising, as the majority of defects in the absence of FMRP are reported in later stages (P30+ brain and DIV 14+ neurons). It is important to test if the conclusions drawn here hold good at different developmental stages.

      Unfortunately, myelin strongly interferes with the ability to use this protocol to purify ribosome clusters in older brains (See Khandjian et al., 2004). It is possible to redo the ribopuromycylation results at later times in culture, but since we cannot compare this to a comparable time in the brain, we have chosen not to do this experiment. We acknowledge this limitation in the discussion, noting that our results are only a snapshot of development and that different results may be observed at different times.

      (2) The term 'distal granules' is very vague. Since there is no structural or biochemical characterization of these granules, it is difficult to understand how they are different from the proximal granules and why FMRP has an effect only on these granules.

      We agree with the reviewer and have removed all references to distal granules. We clarified that we did not measure RPM puncta close to the neuron because the much stronger RPM signal made defining puncta more difficult, and thus, we cannot determine if there are differences between proximal and distal puncta.

      (3) Since the manuscript does not find any effect of FMRP on neuronal RNA granules, it does not provide any new molecular insight with respect to the function of FMRP

      We would respectfully disagree that the study does not provide molecular insight into the function of FMRP, as disproving that FMRP is important for stalling and determining the position of stalling would remove one of the major hypotheses about the function of FMRP, and showing that a major hypothesis in the literature is unlikely to be correct, is at least to me, providing insight. Moreover, we do show an effect of the loss of FMRP on the RPM puncta that represent neuronal RNA granules containing stalled ribosomes. This also provides insight.

      Reviewer #2 (Public review):

      In the present manuscript, Li et al. use biochemical fractionation of "RNA granules" from P5 wildtype and FMR1 knock-out mouse brains to analyze their protein/RNA content, determine a single particle cryo-EM structure of contained ribosomes, and perform ribo-seq analysis of ribosome-protected RNA fragments (RPFs). The authors conclude from these that neither the composition of the ribosome granules, nor the state of their contained ribosomes, nor the mRNA positions with high ribosome occupancy change significantly. Besides minor changes in mRNA occupancy, the one change the authors identified is a decrease in puromycylated punctae in distal neurites of cultured primary neurons of the same mice, and their enhanced resistance to different pharmacological treatments. These results directly build on their earlier work (Anadolu et al., 2023) using analogous preparations of rat brains; the authors now perform a very similar study using WT and FMR1-KO mouse brains. This is an important topic, aiming to identify the molecular underpinnings of the FMRP protein, which is the basis of a major neurological disease. Unfortunately, several limitations of this study prevent it from being more convincing in its present form.

      In order to improve this study, our main suggestions are as follows:

      (1) The authors equate their biochemically purified "RG" fraction with their imaging-based detection of puromycin-positive punctae. They claim essentially no differences in RGs, but detect differences in the latter (mostly their abundance and sensitivity to DHPG/HHT/Aniso). In the discussion the authors acknowledge the inconsistency between these two modalities: "An inconsistency in our findings is the loss of distal RPM puncta coupled with an increase in the immunoreactivity for S6 in the RG." and "Thus, it may be that the RG is not simply made up of ribosomes from the large liquid-liquid phase RNA granules."

      How can the authors be sure that they are analysing the same entities in both modalities? A more parsimonious explanation of their results would be that, while there might be some overlap, two different entities are analyzed. Much of the main message rests on this equivalence, and I believe the authors should show its validity.

      Thank you for your comments. We have been more conservative in the revised paper, referring to the pellet fraction as the pellet fraction rather than the RNA granule fraction to acknowledge the possibility that these two modalities differ. However, we would respectfully disagree that our main message requires RPM-labeled RNA granules in neurites and the ribosome clusters isolated by sedimentation to be “equivalent”. We do believe they are related and added a section in the discussion on this important point.

      (2) The authors show that increased nuclease digestion (and magnesium concentration) led to a reduction of their RPF sizes down to levels also seen by other researchers. Analyzing these now properly digested RPFs, the authors state that the CDS coverage and periodicity drastically improved, and that spurious enrichments of secretory mRNAs, which made up one of the major fractions in their previous work, are now reduced. In my opinion, this would be more appropriately communicated as a correction to their previous work, not as a main Figure in another manuscript.

      We have removed all discussion of the secretory mRNAs, as our attempts to obtain independent evidence for this finding by examining ribophorin enrichment in the pellet across different Mg<sup>2+</sup> concentrations did not support this interpretation (data not shown in the paper). I understand that the change in nuclease is somewhat out of place narratively, but it is clearly relevant to this work. We would disagree with our previous work requiring a ‘correction’. We believe that the nuclease resistance of the mRNA at the entrance site is important. We reproduce our results from rats with similar nuclease treatment in mice as seen in our previous publication; thus, this work is not wrong. We have a paper in preparation that suggests the secondary structure of the mRNA at this location may be important for stalling and thus feel strongly that this result should remain in the manuscript.

      (3) The fold changes reported in Figure 7 (ranging between log2(-0.2) and log2(+0.25)) are all extremely small and in my opinion should not be used to derive claims such as "The loss of FMRP significantly affected the abundance and occupancy of FMRP-Clipped mRNAs in WT and FMR1-KO RG (Fig 7A, 7B), but not their enrichment between RG and RCs".

      We agree that the changes are small and indeed did not appear in the DEG analysis. However, because we are analyzing a large set of mRNAs in this analysis, the results are highly significant and remain significant when using the new statistical tests suggested by the reviewer below. We now emphasize that these are small changes and remind readers that none of the individual mRNA changes were significant in the DEG analysis.

      (4) Figure 8 / S8-1 - The authors show that ~2/3 of their reads stem from PCR duplicates, but that even after removing those, the majority of peaks remain unaltered. At the same time, Figure S8-1 shows the total number of peaks to be 615 compared with 1392 before duplicate removal. Can the authors comment on this discrepancy? In addition, the dataset with properly removed artefacts should be used for their main display item instead of the current Figure 8.

      We now use only the data after removing PCR duplicates for all the analyses except in Figure 3. The number of peaks observed is determined mainly by the threshold used, as stated in the methods “To be identified as a peak, the zenith of an abundance site for the reads must be 4x higher of the average of the total transcript.” Due the lower number of reads after the PCR duplicates fewer peaks reached this threshold.

      (5) Figure 9 / S9-1, the density of punctae in both WT and FMR1-KO actually increases after treatment of HHT or Anisomycin (Figure S9-1 B-C). Even if a large fraction would now be "resistant to run-off", there should not be an increase. While this effect is deemed not significant, a much smaller effect in Figure 9C is deemed significant. Can the authors explain this? Given how vastly different the sample sizes are (ranging from 23 neurites in Figures S91 to 5,171 neurites in Figure 9), the authors should (randomly) sample to the same size and repeat their statistical analysis again, to improve their credibility.

      The box and whisker plots emphasize the median and not the average. We now also show the averages in Figure S9-1, which indicate a slight decrease for both HHT and anisomycin.

      We apologize for the typo in the figure legend in Figure 9, 171, not 5171. We now use random sampling in Figures 6 and 7, where the sample sizes differ substantially.

      Reviewer #3 (Public review):

      Summary:

      Li et al describe a set of experiments to probe the role of FMRP in ribosome stalling and RNA granule composition. The authors are able to recapitulate findings from a previous study performed in rats (this one is in mice).

      Strengths:

      (1) The work addresses an important and challenging issue, investigating mechanisms that regulate stalled ribosomes that are part of stress granules, and focusing on the role of FMRP. This is a complicated problem, given the heterogeneity of the granules and the challenges related to their purification. This work is a solid attempt at addressing this issue, which is widely understudied.

      (2) The interpretation of the results could be interesting if supported by solid data. The idea that FMRP could control the formation and release of stress granules, rather than the elongation by stalled ribosomes, is of high importance to the field, offering a fresh perspective into translational regulation by FMRP.

      (3) The authors focused on recapitulating previous findings, published elsewhere (Anadolu et al., 2023) by the same group, but using rat tissue, rather than mouse tissue. Overall, they succeeded in doing so, demonstrating, among other findings, that stalled ribosomes are enriched in consensus mRNA motifs that are linked to FMRP. These interesting findings reinforce the role of FMRP in the formation and stabilization of RNA granules. It would be nice to see extensive characterization of the mouse granules as performed in Figure 1 of Anadolu et al., 2023.

      (4) Some of the techniques incorporated aid in creating novel hypotheses, such as the ribopuromycilation assay and the cryo-EM of granule ribosomes.

      Thank you for these positive comments. We have now added a more extensive characterization in Figure 1.

      Weaknesses:

      (1) The RNA granule characterization needs to be more rigorous. Coomassie is not proper for this type of characterization, simply because protein weight says little about its nature. The enrichment of key proteins is not robust and seems not to reach significance in multiple instances, including S6 and UPF1. Furthermore, S6 is the only proxy used for ribosome quantification. Could the authors include at least 3 other ribosomal proteins (2 from the small, 2 from the large subunit)?

      We have increased N to improve the robustness of the enrichment analysis and added several additional RBPs. Along with Coomassie we now include analysis of UV absorbance and include EMs from these fractions showing the presence of 80S ribosomal clusters in the fractions we are using.

      (2) Page 12-13 - The Gene Ontology analysis is performed incorrectly. First, one should not rank genes by their RPKM levels. It is well known that housekeeping genes, such as those related to actin dynamics, molecular transport, and translation, are highly enriched in sequencing datasets. It is usually more informative when significantly different genes are ranked by p-adjust or log2 Fold Change, then compared against a background to verify enrichment of specific processes. However, the authors found no DEGs. I would suggest the removal of this analysis and the incorporation of a gene set enrichment analysis (ranked by p-adjust). I further suggest that the authors incorporate a dimensionality reduction analysis to demonstrate that the lack of significance stems from biology and not experimental artifacts, such as poor reproducibility across biological replicates.

      Thank you for the suggestion. We now use GSEA analysis to examine differences in gene sets between WT and FMR1- mice and find some significant changes (new Fig. 5). The old analysis is still included for comparison to our earlier paper as a supplemental figure. We have now included a PCA analysis (FigS5-1) to show reproducibility across biological replicates.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) RNA sequencing comparison between WT and FMR1 KO mice should be carried out at a later developmental stage, which may provide a better difference between these two groups

      There are a number of studies that have already done this analysis and in specific brain regions 10.1016/j.neuron.2017.07.013; 10.7554/eLife.46919; 10.3389/fnmol.2017.00340; https://doi.org/10.1016/j.neuron.2023.06.009. The main goal of our RNA-seq was to standardize for the RPF studies, not to identify differences in RNA-seq between WT and FMRP. In the response to public review point 1 we explain why we do not look at later developmental timepoints.

      (2) The same is true in characterizing the effect of FMRP on the RNA granules.

      See response to public review point 1, which addresses this point.

      (3) No evidence is provided for the effectiveness of DHPG stimulation in DIV8-10 neurons; this is needed for justification using neurons at this stage.

      We have previously shown that DHPG stimulation in these neurons at this developmental time from cultures made from rat brain is sufficient to decrease the number of RPM puncta and to induce an increase in the synthesis of proteins in an initiation resistant manner (Graber et al, 2013; Graber et al, 2017). This is now more clearly stated in the manuscript. Moreover, here we replicate the result of DHPG in WT mice at reducing the number of RPM puncta.

      (4) In Figure 9 B, it is not clear whether the neurites indicated are axons or dendrites. Since neurons are still in the early stages of dendritogenesis/synaptogenesis, it is important to make that distinction.

      We have previously characterized RNA granules in axons and dendrites in hippocampal cultures from rats at this time (Miller et al, 2009, MCN 40:485-495)) and they are similar. While it is likely that the vast majority of the neurites at this time are dendrites, since we did not use markers, we conservatively just use the term neurites.

      (5) In Figure 1 (and elsewhere), fraction 5/6 is used as a polysome or RNA cluster. The authors have not provided a UV absorption profile and only have s6 as evidence to say this polysome. In the Coomassie gel, this fraction is any different than fractions 7/7 or 9/10; what is the justification for using this fraction?

      The main justification for these fractions is to be consistent with our previous paper (Anadolu et al, 2023) and the Khandian study comparing polysomes to pellet using the same fractionation protocol (El-Fatimy et al, 2016). We now provide a UV absorption profile (Fig. 1C) and EM pictures (Fig. 1D) to show the ribosome clusters in this fraction. We do not believe our results would be fundamentally different from those obtained if we had used other heavy fractions.

      Minor comments

      (1) The font size very small in the figures, please increase it.

      We have worked hard to increase the font size in all the figures.

      (2) In the result section for Figure 3B - it is written 'majority of these mRNA are non-coding mRNA' - this doesn't make sense.

      Corrected

      Reviewer #2 (Recommendations for the authors):

      (1) There are lots of mistakes (e.g. word omissions or duplications, grammatical errors) throughout the text, too many to list here.

      We have carefully edited the text to try to minimize these mistakes.

      (2) In many positions related to their improved nuclease digestion protocol, samples are labelled "M ...", which apparently stands for "high magnesium and high nuclease treatment group". I would suggest switching to something more intuitive, such as "... (improved digestion)".

      We have removed most of the comparisons between these samples. What remains (Figure 3), we just use Low Nuclease when we refer to the sample with low Magnesium and low nuclease.

      (3) Figure 1,3 - It would be tremendously illuminating to see a polysome trace (UV260 absorbance) in addition to Coomassie-stained SDS-PAGE to underscore the interpretation of the different fractions by the authors. As it stands, there is no way of telling whether there are any polysomes present at all. This can also be done by hand using a UV absorption reader if no built-in device is available to the authors.

      We have now done this (Fig. 1C) and also provided EM of this fraction to show the presence of ribosomes in this fraction.

      (4) I don't understand why the authors switched from calling fraction 5/6 the "polysome fraction" in their previous work to calling it "ribosome cluster fraction" in this work. The argument given "[...] due to its structural similarity to ribosomes in RNA Granules (Anadolu et al., 2023), we conservatively call this the ribosome cluster fraction (RC)." does not instill confidence that these two fractions are indeed distinct.

      We agree with the reviewer and regret this decision. We now call the pellet, the pellet and Fraction 5/6, fraction 5/6.

      (5) Figure 1C - There are clear scanning or compression artefacts in the blot images (most prominently in the eEF2 lanes) that should be corrected.

      We have replaced all images in Figure 1 and have increased the N of this experiment considerably.

      (6) Figure 1C - The authors claim that WT mouse RG is enriched in FMRP compared to RC or starter fraction, but there is also a lot more protein loaded in the RG (especially when compared to RC). It is also hard to believe from the Coomassie staining that despite the much stronger presence of low MW bands (which is where ribosomal proteins migrate) in fraction 5/6, the s6 western blot signal is actually comparable between RC and RG. Can the authors please provide more detail on the loading of these fractions and supply quantification of FMRP in all three fractions, normalized by total protein? This might also be the source of their discrepancy, stating that contrary to their expectation, ribosomes (as measured by s6 signal / s6 signal in starter fraction) are actually increased in FMR1-KO brains.

      We have repeated all of these experiments and changed our method of quantification (See methods). We no longer use the starting material in our quantification. Indeed, with the additional data and change in method, we no longer see an increase in S6 in the FMR1- pellet fraction.

      (7) Figure 1 - I believe "D-F)" should only read "D-E)" based on the axis titles, and instead "FG)" should be added before the next sentence. Instead of "Staufen" it should be specified in the Figure that "Stau2" was quantified. "Staufen (59kd)" should read "Stau2 (59 kDa)" and "anti-Staufen (52kb)" should read "anti-Stau2 (52 kDa)" and the same for all other similar instances. It is further hard to believe that e.g., "Staufen2 (59kd)" (see above) is not significantly enriched with N=5, a very low spread, and over 1.5x enrichment. The authors should double-check that the appropriate statistical test was employed.

      Figure 1 has been completely redone, and the two Staufen bands are enriched in this new analysis.

      (8) Figure S4-2 - Most of the detail in the corresponding figure legend should be moved to the Materials and Methods section.

      Details relevant to the methods in this figure legend have been now moved to the Material and Methods section.

      (9) Figure 4A - The displayed/segmented tRNA densities appear unusually distorted. I would recommend displaying segmented densities of the original homogeneous reconstructions, not of separated and later fused partial maps.

      Figure 4 was modified according to the suggestions of this reviewer.’

      (10) Figure 9 C-D, S9-1 B-E - Are not all conditions also including puromycin as in B above? If so, it should be added to both the figure and the figure legend.

      The reviewer is correct and the figure and legend has been changed to reflect this.

      Reviewer #3 (Recommendations for the authors):

      (1) "Loss of FMRP causes Fragile X syndrome. In humans, the loss of FMRP occurs due to the expansion of a CGG repeat in the 5' untranslated region (UTR) of the gene, leading to excessive methylation and transcriptional inhibition."

      Comment: Genes don't have 5'UTR, but exons encoding 5'UTR. I suggest rephrasing this statement.

      This sentence has been rephrased.

      (2) "Several of these functions have been implicated in Fragile X syndrome, including FMRP's regulation of miRNA repression, splicing, translation initiation, and translational elongation".

      Comment: Is this a typo? miRNA instead of mRNA?

      No, this is correct. FMRP has been implicated in the regulation of microRNAs (miRNAs) in a number of studies.

      (3) "elongation rates are also increased in mouse models of FMRP".

      Comment: Mouse models of Fragile X?

      This has been corrected.

      (4) "Parts of this work were included in the Master's thesis of the first author (Li, 2024)."

      This has been removed.

      (5) Comment: Graphs in Figure 1 need proper y-axis labeling. What is the normalization method? What are the values presented in the y-axis?

      Figure 1 has been completely changed and the Y-axes are now clear in this new version.

      (6) "Thus, by looking at the percentage of puromycylation present in the presence of anisomycin, we can estimate the number of ribosomes in this state. "

      Comment: Are the authors really estimating the number of ribosomes in a resistant state? One could argue that they are collecting populational information regarding resistance to anisomycin.

      We have rephrased this sentence to be more conservative about what we are measuring.

      (7) Comment: Page 11 - Why did the authors assume magnesium would affect the conformation state of the ribosomes? What is the rationale behind increasing the [Mg2+]?

      Most preparations using ribosomes use 10 mM MgCl<sub>2</sub>. However, most neuroscientists use physiological buffers that contain 2.5 mM MgCl<sub>2</sub>. In bacteria, this makes a large difference, but evidence from eukaryotes is not clear. Since this is a collaboration between these two schools of thought, we decided to switch to 10 mM MgCl<sub>2</sub>, since in the EM, there were some free 60S ribosomes (Anadolu et al, 2024).

      (8) Page 11- "In other words, high Mg2+ decreased the abundance of mRNAs normally cotranslationally inserted into the ER which are unlikely to be components of transporting RNA granules containing stalled ribosomes and solidified our focus on the M protocol in the analyses below."

      We have removed this from the paper, as additional experiments aimed to solidify this interpretation failed to detect an effect on secretory mRNAs.

      (9) Comment: The whole "abundance", "enrichment", and "occupancy" nomenclature is hard to follow.

      We have rewritten this section.

      (10) Page 13 - "There were only 2 protein coding genes that were significantly different between the abundance of FMR1-KO and WT in protein coding genes - FMR1 and Wdfy1 (Extended Data Table 5-2). There were no significantly different genes between WT and FMR1-KO occupancy and enrichment. Thus, no difference rose to significance, given the large number of mRNAs used in this analysis."

      Comment: It seems like this is repeating the same information three times.

      This has been changed.

      (11) Page 13 - "Similar to previous experiments with rats, the most abundant mRNAs resistant to run off were significantly abundant, occupied and enriched in both WT and FMRP RPFs (Fig 6)"

      The Shah et al dataset we use was based on the most abundant mRNAs resistant to run-off. While we agree it is not surprising that they are also abundant in the pellet we observe, this would not necessarily be true unless the pellet is actually enriched in stalled mRNAs.

      (12) Page 14 - "These mRNAs had been identified by cross-linking FMRP with mRNA, fragmenting the mRNA, immunoprecipitating the mRNA still associated with FMRP and sequencing this mRNA."

      We shortened this description.

      (13) Page 14 - "Interestingly, while still significant, there appeared to be a decrease in the relative abundance of these mRNAs in the FMR1-KO RG (Fig 6B)"

      Comment: It is hard to observe this decrease in the boxplots. Second, the statistical tests for the bioinformatics analyses are not the most appropriate, given the large discrepancy in the number of mRNAs present in the experimental group ("All mRNAs") and the filtered groups.

      We have redone the statistics using multiple random sampling of all the mRNAs such that the total number of mRNAs in the group was the same. This lowered the significance for some groups, but they are mostly still highly significant. This analysis has also been affected by switching to using the data from the PCR-subtracted RPFs. The changes we now observe are more evident in the whisker box plots due to this improvement in the data.

      (14) Page 16 - "To rule out that peaks were due to amplification artifacts in the preparation of RPFs we repeated these analyses after removing PCR duplicates (Fig. S8-1; Extended Data Table S8-3) and found over 95% of the peaks identified without removing PCR duplicates were defined as a peak in at least one of the biological replicates after removing duplicates. More importantly, we found similar results with enrichment of FXS motif and enrichment of negatively charged amino acids in the FMR1-KO only, WT only and both peaks after removing PCR duplicates (Fig. S8-1; Extended Data Table S8-3)."

      Comment: It is unclear why the authors needed to include the analysis without PCR duplicate removal. This is an essential step to guarantee the robustness of ribo-seq findings. I recommend removing the whole analysis from Figure 8 from the manuscript and including only the post-duplicate removal analysis.

      As mentioned above, we completely agree with this statement and now show only this data and moreover have redone all the figures with only this data (except for Fig. 3).

      (15) Figure 9 - I am unsure that the data is convincing enough to demonstrate reinitiation of mRNA granules induced by DHPG. I suggest a colocalization experiment with another protein well known to be localized to RNA granules, such as G3BP1. In addition, repeat the experiment with an additional group where elongation is blocked after the addition of DHPG, which presumably would prevent the reduction in the WT puncta density.

      These are interesting additional experiments, but outside the scope of what we can manage. We have previously shown colocalization of Staufen, FMRP and UPF1 to these puncta (Graber et al, 2013; Graber et al, 2017) and shown that these puromycylated puncta also colocalize with nascent peptides detected using the Sun-Tag technique. While we think doing the experiment in the presence of an elongation inhibitor would be interesting, we disagree that it would prevent the reduction in WT puncta density, since we believe what is happening is the loss of the liquid-liquid phase separation of the ribosome clusters due to dephosphorylation of RBPs like FMRP and UPF1 (Graber et al, 2017), and this would reduce the puncta density whether or not the ribosomes were activated for translation.

      Nevertheless, we have tried to temper the conclusions made from this result, emphasizing what we know (RPM puncta are decreased) as opposed to actual reactivation of stalled polysomes which we are not measuring.

      Discussion - Page 18 - "Nevertheless, if FMRP binding was the critical determinant for presence in neuronal RNA granules, we would have expected to observe more differences." This is not true. If the data is poorly collected, you will not see differences.

      This statement was removed.

      (16) "A proportion of the stalled ribosomes that are not stored in large RNA granules may still be pelleted in the sucrose gradients. This fraction may be greater in the absence of FMRP."

      Comment: The authors are right about this and touch on my original point that the characterization of the biochemical fractionation is not convincing enough. I'd suggest probing against more proteins that are contained in RNA granules.

      We have added several proteins to the biochemical characterization shown in Figure 1. We have added a discussion about the relationship between neuronal RNA granules and the sedimented pellet fraction in the discussion section.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We incorporated Reviewer #2’s suggestion to change the name of mll-1 because of overlap with a human gene. We used the updated gene names in our responses below to minimize confusion. Below are the updated gene names for the toxin-antidote system we described.

      tmrl-1 - Toxin-induced Maternal Rod Lethality (formerly mll-1). After we establish that B0250.8 is also a toxin, we refer to this gene as the “N2 tmrl-1 allele”.

      amrl-1 - Antidote of Maternal Rod Lethality (formerly smll-1)

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The article by Zdraljevic et al. reports the discovery of a third toxin-antidote (TA) element in C. elegans, composed of the genes mll-1 (toxin) and smll-1 (antidote). Unlike previously characterized TA systems in C. elegans, this element induces larval arrest rather than embryonic lethality. The study identifies three distinct haplotypes at the TA locus, including a hyper-divergent version in the standard laboratory strain N2, which retains a functional toxin but lacks a functional antidote. The authors propose that small RNA-mediated silencing mechanisms, dependent on MUT-16 and PRG-1, suppress the toxicity of the divergent toxin allele. This work provides insights into the evolutionary dynamics of TA elements and their regulation through RNA interference (RNAi).

      Overall, there are many things to like about this paper and only a few small quibbles, which will not require more than a little rewriting or relatively minor analyses.

      Strengths:

      (1) The discovery of a maternally deposited TA element with delayed toxicity due to delayed mRNA translation of the maternally deposited toxin mRNA is a significant addition to the literature on selfish genetic elements in metazoans.

      (2) Identifying three haplotypes at the TA locus provides a snapshot of potential evolutionary trajectories for these elements, which are often inferred but rarely demonstrated in naturally occurring strains. The genomic analysis of 550 wild isolates contextualizes the findings within natural populations, revealing geographic clustering and evolutionary pressures acting on the TA locus.

      (3) The study employs various techniques, including CRISPR/Cas9 knockouts, FISH, long-read RNA sequencing, and population genomics. The use of inducible systems to confirm toxicity and antidote functionality is particularly robust. This multifaceted approach strengthens the validity of the findings.

      (4) The authors provide compelling evidence that small RNA pathways suppress toxin activity in strains lacking a functional antidote. This highlights an alternative mechanism for neutralizing selfish genetic elements.

      Weaknesses:

      (1) The introduction focuses strongly (for good reason) on bacterial TA systems and then jumps to TA systems in C. elegans. It's unclear why TA systems in other eukaryotes are not discussed.

      We briefly introduced bacterial TA systems because of their ubiquitousness and focused on C. elegans TA systems. We chose certain aspects of previously described Caenorhabditis TA elements that were relevant to the narrative we presented. Furthermore, we have extensively reviewed TA systems previously and have added a citation to that review in the revised manuscript (Burga et al. 2020).

      (2) Similarly, there is a missed opportunity to discuss an analogy between the suppressor mechanism discovered here and the hairpin RNA suppressors of meiotic drive identified by Eric Lai and colleagues. Discussing these will provide a fuller context of the present study's findings and will not affect their novelty.

      Thank you for pointing this out. We added a mention of the Stellate and Dox systems in our discussion.

      (3) While the evidence for RNAi-mediated suppression is strong, the claim that positive selection drove diversification at piRNA binding sites requires further discussion and clarification. The elevated dN and dS are unusual (how unusual relative to other genes in vicinity? What is hyper-divergent statistically speaking?), but there is no a priori reason that there would be selection on piRNA binding sites within the mll-1 transcript to facilitate its recognition by endogenous RNAi machinery; what is the selective pressure for mll-1 to do so? Most TA systems would like to avoid being suppressed by the host. One cannot make the argument that this was motivated by the loss of the antidote because the loss of the antidote would be instantly suicidal, so the cadence of events described requiring hypermutation of the mll-1 transcript does not work.

      We largely agree with the reviewer’s point, which we believe is based on the following sentence in the discussion: “We propose that positive selection for piRNA binding sites in the tmrl-1 transcript drove the diversification of this gene toward the N2 version.” We have removed this argument from the discussion in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      In the manuscript by Walter-McNeill, Kruglyak, and team, the authors provide solid evidence of another toxin-antidote (TA) system in C. elegans. Generally, TA systems involve selfish and linked genetic elements, one encoding a toxin that kills progeny inheriting it, unless an antidote (the second element) is also present. Currently, only two TA systems have been characterized in this species, pointing to the importance of identifying new instances of such systems to understand their transmission dynamics, prevalence, and functions in shaping worm populations.

      Strengths:

      This novel TA system (mll-1/smll-1) was identified on LGV in wild C. elegans isolates from the Hawaiian islands, by crossing divergent strains and observing allele frequency distortions by high-throughput genome sequencing after 10 generations. These allele frequency distortions were subsequently confirmed in another set of crosses with a separate divergent strain, and crosses of heterozygous males or hermaphrodites resulted in a pattern of L1 lethality in progeny (with a rod arrest phenotype) that suggested the maternal transmission of this TA system from the XZ1516 genetic background. By elegantly combining the use of near-isogenic lines, CRISPR editing to generate knock-outs, and a transgene rescue of the antidote gene, the authors identified the genes encoding the toxin and the antidote, which they refer to as mll-1 and smll-1. Moreover, the specific mll-1 isoform responsible for the production of the toxin was identified and mll-1 transcripts were observed by FISH in early and late embryos, as well as in larvae. Inducible expression of the toxin in various strains resulted in larval arrest and rod phenotypes. The authors then characterized the genetic variation of 550 wild isolates at the toxin/antidote region on LGV and distinguished three clades: (1) one with the conserved TA system, (2) one having lost the toxin and retaining a mostly functional antidote, and (3) one having lost the antidote and retaining a divergent yet coding toxin (this includes the reference strain Bristol N2, in which the homologous toxin gene has acquired mutations and is known as B0250.8). Further, the authors show that this region is under positive selection. These data are compelling and provide very strong evidence of a new TA system in this species.

      Weaknesses:

      The question remained as to how one clade, including N2, could retain the toxin gene but not possess a functional antidote. In the second part of the manuscript, the authors hypothesized that small RNA targeting (RNAi) of the toxin transcript could provide the necessary repression to allow worms to survive without the antidote. Through a meta-analysis of multiple small RNA datasets from the literature, the authors found evidence to support this idea, in which the toxin transcript is targeted by 22G siRNAs whose biogenesis is dependent on the Mutator foci protein, MUT-16. They note that from previous studies, mut-16 null mutants displayed a varied penetrance of larval arrest. In their own hands, mut-16 mutants displayed 15% varied larval arrest and 2% rod phenotypes. In an attempt to link B0250.8 to mut-16/siRNAs, they made a double mutant and examined body length as a proxy for developmental stage. Here, they observed a partial rescue of the mut-16 size defect by B0250.8 mutation. Finally, the authors also highlight data from further meta-analysis, which predicts the recognition of B0250.8 by several piRNAs. Also based on existing data from the literature, the authors link loss of Piwi (PRG-1), which binds piRNAs, to a depletion of 22G-RNAs targeting B0250.8 and an upregulation of B0250.8 expression in gonads, suggesting that piRNAs are the primary small RNAs that target B0250.8 for downregulation. The data in this portion of the manuscript are intriguing, but somewhat preliminary and incomplete, as they are based on little primary experimentation and a collection of different datasets (which have been acquired by slightly different methods in most cases). This portion of the study would require subsequent experimentation to firmly establish this mechanistic link. For example, to be able to claim that "the N2 toxin allele has acquired mutations that enable piRNA binding to initiate MUT-16-dependent 22G small RNA amplification that targets the transcript for degradation" the identified piRNA sites should be mutated and protein and transcript levels analysed in wild-type and in the strain with mutated piRNA sites. At a minimum, the protein levels in wild-type and mut-16, prg-1, and/or wago-1 mutants should be measured by western blot and/or by live imaging (introducing a GFP or some other tag to the endogenous protein via CRISPR editing) to show that the toxin is not accumulated as a protein in wt, but increases in levels in these mutants. mRNA levels in Figure S5A suggest there is still some expression of the B0250.8 transcript in a wild-type situation.

      We thank the reviewer for their thoughtful assessment of our manuscript, and we appreciate that they recognized that the data linking the small RNA machinery to B0250.8 suppression is intriguing. While the reviewer claims our analysis is preliminary and incomplete, we believe we present an appropriate multi-faceted approach for establishing the small RNA-mediated suppression mechanism we describe. 

      First, the reviewer states that we rely on “little primary experimentation”. Our primary experiments show that loss of the N2 tmrl-1 allele partially rescues ∆mut-16 developmental delay and arrest phenotypes. Therefore, we provide direct evidence that the N2 tmrl-1 functionally contributes to the ∆mut-16 phenotype. Furthermore, we overexpressed the N2 tmrl-1 allele to show that this gene is a toxin.

      It is true that we use previously published datasets to establish a small RNA-mediated mechanism that likely explains our observations. The reviewer suggests that our claims are weakened by relying on a “collection of different datasets (which have been acquired by slightly different methods in most cases)”. We believe instead that evidence collected from multiple labs using an array of different techniques strengthens our conclusions. We show that N2 tmrl-1-targeting small RNAs have been identified across multiple datasets (references 26, 32, 33, 34). Taken together, these datasets support a mechanistic framework for the suppression of the N2 tmrl-1 that involves PRG-1-dependent piRNA binding, MUT-16-dependent 22G siRNA, and the secondary Ago WAGO-1 binding. 

      The reviewer suggests several experiments, but we do not view them as essential to support our claims. 

      (a) piRNA site mutatagenesis: we present multiple lines of evidence that the N2 tmrl-1 transcript is post-transcriptionally targeted by small RNAs in a piRNA-mediated manner, not that specific piRNA sites are necessary and sufficient for this silencing. The suggested experiment would be valuable for future work, but is beyond the scope of our study.

      (b) Characterization of TMRL-1 protein levels: We agree that this experiment would provide definitive evidence of complete small RNA-mediated suppression of the N2 tmrl-1 transcript. As we explain above, however, we do show that removing the N2 tmrl-1 allele partially rescues the ∆mut-16 growth defect, demonstrating that when this gene’s regulation is disrupted, it induces toxicity. Importantly, we observed no tmrl-1-induced toxicity when we overexpressed a version of this gene with a stop codon, indicating that it acts as a protein.

      Finally, the reviewer questions our claim that: "the N2 toxin allele has acquired mutations that enable piRNA binding to initiate MUT-16-dependent 22G small RNA amplification that targets the transcript for degradation."

      We agree that this statement is too definitive given our current data. We have revised it to: "Multiple lines of evidence suggest that the N2 tmrl-1 allele is recognized by piRNAs, leading to MUT-16-dependent 22G siRNA production and post-transcriptional silencing of the transcript."

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The paper suggests that antidote pseudogenization occurred because RNAi replaced its function, but does not explore whether this process is ongoing or complete across all N2-like strains.

      We explored this possibility, but we realize that we did not explicitly state so in the manuscript. The B0250.4 (amrl-1) gene is pseudogenized in all strains within the N2 clade. We have modified the following sentence in the results section to explicitly state this observation:

      “While the previously described C. elegans TA elements are characterized by their absence in susceptible strains (2, 3), all members of the N2-like susceptible clade harbor a divergent allele of tmrl-1 with an intact coding sequence, as well as a pseudogenized version of amrl-1.”

      (2) Some figures (e.g., allele frequency distortions) could benefit from additional annotations to guide interpretation. In general, the figures make the reader work harder than they need to.

      We attempted to add clarity to figure captions for clarity.

      Although mll-1 and smll-1 were identified as toxin and antidote genes, their molecular mechanisms remain unclear and are very interesting.

      We agree that identifying the molecular mechanism associated with the toxin and antidote would be of interest, but is beyond the scope of the current paper.

      Reviewer #2 (Recommendations for the authors):

      (1) Because the rod phenotype was important in identifying the TA system, it seems important to include representative images of this phenotype throughout the paper.

      We added a supplemental figure showing the resulting self progeny from a QX1211/XZ1516 heterozygote: Fig S1B

      (2) In Figure 2A, we were confused as to why there were so few reads of mll-1. We may be misunderstanding something, so could the authors explain this to us? We would have expected more reads of mll-1, given the diagram showing that the breakpoints of the NIL were beyond (closer to the right end of) the mll-1 locus, and the phenotype correlates with the presence of the toxin (frequency of .20 L1 arrest).

      The lack of sequencing depth arises because the sequence divergence between QX1211 and XZ1516 is too high to accurately map short sequencing reads derived from QX1211 to the XZ1516 genome. We added the following sentence to the figure caption to add clarity:

      “The XZ1516 and QX1211 genome are so diverged that short reads derived from QX1211 don’t align to the XZ1516 genome in the 200 bp windows with no corresponding read depth, as indicated by a lack of a gray bar.”

      (3) The use of TOF in Figure 4 as a proxy of animal length instead of directly indicating or measuring animal length hinders the comparison of these results with other studies (i.e., most often in the literature, we see images of worms and measurements of their sizes or use of some other morphological marker to demonstrate the proportion of worms in a particular developmental stage). Nonetheless, we think the approach is clever and certainly enables analysis of a large sample population. However, a wild-type control is missing from these experiments to give a sense of the typical distribution one would expect. Without this, one interpretation of the B0250.8 knock out data shown in B is that loss of B0250.8 results in ~10% arrested larval, which seems higher than would be expected for a wild type N2 strain, and should be explained-but again, if the wild type control showed the same pattern, that would be useful to know. The title for Figure 4 should be revised, as this figure suggests, but does not provide definitive evidence that B0250.8 is suppressed by sRNAs/sRNA pathways. See the next point for providing more definitive data to support this model.

      There is a long list of publications that rely on the large particle sorter to infer how growth rate is affected in various mutants and environmental conditions (See Andersen et al. 2015, ref 28 in the manuscript, and the papers that reference this work). As the reviewer pointed out, the use of time of flight, which is simply the amount of time an object obstructs a laser at a constant flow rate, enables accurate measurement of tens of thousands of individual animals for comparison. 

      The reviewer is correct to point out that without a wild type N2 control, it is impossible to tell what a typical distribution looks like. However, the experiment includes all strains necessary to make the comparisons that enable us to draw the conclusion that the N2 tmrl-1 allele contributes to larval arrest in the absence of MUT-16.

      We agree with the reviewers point that this figure does not provide evidence that B0250.8 is suppressed by small RNAs and we have therefore changed the figure title.

      The new figure title: The N2 tmrl-1 allele contributes to larval arrest in the absence of MUT-16

      (4) To be able to claim that "the N2 toxin allele has acquired mutations that enable piRNA binding to initiate MUT-16-dependent 22G small RNA amplification that targets the transcript for degradation" the identified piRNA sites should be mutated and protein and transcript levels analysed in wild-type and in the strain with mutated piRNA sites. At a minimum, the protein levels in wild-type and mut-16, prg-1, and/or wago-1 mutants should be measured by western blot and/or by live imaging (introducing a GFP or some other tag to the endogenous protein via CRISPR editing) to show that the toxin is not accumulated as a protein in wt, but increases in levels in these mutants. mRNA levels in Figure S5A suggest there is still some expression of the B0250.8 transcript in a wild-type situation.

      The reviewer makes several good suggestions for experiments to determine whether the conclusions we make from publicly available high-throughput sequencing datasets apply in our context. However, we disagree that the quoted statement “the N2 toxin allele has acquired mutations that enable piRNA binding to initiate MUT-16-dependent 22G small RNA amplification that targets the transcript for degradation” is not supported by the evidence we present from Reed et al. 2020. The data presented by Reed et al. clearly show that the N2 tmrl-1 transcript is heavily targeted by 22G siRNAs, and that the accumulation of these siRNAs depends on the presence of MUT-16 and PRG-1. The dependence on PRG-1 implicates piRNAs involvement in the mounting of a 22G response.

      (5) Importantly, it is not the mll-1/B0250.8 transcript itself that was not shown to interact with WAGO-1 in the Seroussi et al. eLife paper (Lines 257-259). This study investigated sRNAs associated with every AGO, and computationally inferred the targets of each AGO using those enriched sRNA sequences. Therefore, it is the siRNAs antisense to mll-1/B0250.8 that were detected in association with WAGO-1, making it likely that WAGO-1 is the secondary AGO that targets this transcript. The argument the authors make holds true, but the authors should revise how they describe the evidence supporting that argument to accurately reflect the existing data.

      Thank you for catching this mistake. We have updated the text to accurately reflect the results from the Seroussi et al 2023 publication:

      “Recent work has shown that the N2 tmrl-1 transcript-derived small RNAs co-immunoprecipitated with WAGO-1, providing additional evidence that this transcript is regulated by the endogenous RNAi machinery”

      (6) It seems likely that the authors explored the possibility that another antidote may be present in the third clade. Could they discuss what they did to rule out this explanation in lieu of piRNA/siRNA regulation?

      We did not look for another antidote in the third clade because this clade is defined by the presence of an antidote and the absence of a toxin. Figure 3C shows the result of a cross between a third clade strain (NIC195) and XZ1516. The conclusion we draw from this experiment is that the antidote present in NIC195 provides near complete resistance to the XZ1516 toxin.

      (7) Line 156, legend of Figure S3, and line 273: There was no marker used to indicate that these are the primordial germ cells. Best practices would indicate using a fluorescent marker (e.g., PIE-1 GFP or PGL-1 GFP or PRG-1 GFP, etc.) to definitively identify these as PGCs.

      We agree with the reviewer’s point. As we do not have the perfect experiment, we do not definitively state that tmrl-1 transcripts localize in the primordial germ cells. 

      Minor comments:

      (1) A minor suggestion: incorporating some of the results now shown in the supplementary figures - Figures S1, S3, and S4 - into the main figures may make the manuscript easier to read.

      We constructed the manuscript in a way we thought was straightforward. The figures listed by the reviewer are supplemental to the main conclusions of the manuscript, so we decided to leave them as supplemental figures.

      (2) Line 87, Figure S1A: include numbers in the y-axis.

      The numbers are included on the y-axis and we explain the x-axis tick marks in the figure caption.

      (3) Figures 1B, 2B, 3C, 4B, S1B, S4: statistical analyses missing.

      We have added a summary of the statistical analysis to the captions of Figures 1B, 2B, 3C, and S1B. We added more detail from the analysis of 4A, which is the figure we draw conclusions from. Figure S4 is observational data, and the only conclusion drawn from that figure is that the N2 tmrl-1 gene encodes a toxin. It is toxic in 100% of individuals we looked at and therefore doesn’t warrant statistics. 

      (4) Line 100, "The rod progeny were all homozygous for QX1211 alleles at the locus on the right arm of chromosome V that displayed the allele frequency distortion in the mapping populations". Is this supported by data? While there is strong evidence to suggest it, the way it is currently written makes it seem that the rod progeny have been genotyped (by sequencing or PCR?). Is this the case? If not, the authors should revise the statement accordingly.

      Yes, this is indeed the case and we have updated the text to reflect that we performed PCR of a QX1211-specific indel to verify the genotypes on the right arm of chromosome V.

      (5) Figure 2A: lower panel missing x axis label.

      The top panel is a cartoon representation of a NILs, and the x axis is labeled for the top panel, highlighting the mapped element. 

      (6) Line 140 to 148: The authors should provide data to support these statements.

      Realizing i skipped this one – these are the lines they are referring to -> Long-read RNA sequencing revealed two distinct mll-1 isoforms, a short isoform with three predicted exons and a long isoform with eight predicted exons (Fig. S2A). We constructed plasmids with inducible versions of each mll-1 isoform. When we injected susceptible strains with the short mll-1 isoform array, every F1 individual carrying the array died, with 64% of larvae exhibiting the rod phenotype, indicating that uninduced expression levels of the short mll-1 isoform are sufficient to induce lethality. By contrast, we were able to isolate susceptible strains that maintained the long mll-1 isoform array or a short mll-1 isoform array with a premature stop codon in mll-1. We observed no rod progeny upon induction of these arrays, indicating that the short isoform encodes the functional toxin, and that the toxin acts as a protein.

      (7) Line 193: It would be interesting to see if there is structural conservation between mll-1 and B0250.8 using alpha-fold. Have the authors done this?

      We did attempt to look for structural conservation but we found the confidence in the structural predictions to be very low, which didn’t warrant a comparison.

      (8) Line 206-207: Could the authors explain why the frequency of the rod phenotype is so low when presumably over-expressing B0250.8? Does this indicate that B0250.8 is not as functional a toxin as mll-1, or is it sufficiently repressed by sRNAs and not actually overexpressed? Further, what are "abnormal" phenotypes? This should be clarified for the reader.

      It is likely that the overexpression and misexpression of toxic proteins is causing the abnormal phenotypes. The rod phenotype probably manifests when the gene is expressed at the appropriate developmental stage and tissue to cause the phenotype, whereas abnormal phenotypes manifest when the expression is not in the correct stage or location. A summary of the observed phenotypes is provided in Supplementary Table 7.

      (9) Line 216 and thereafter: indicate that B0250.8 is now referred to as mll-1.

      We incorporated this suggestion.

      (10) Line 228-231: missing to state that this is shown in Figures 4A-B.

      This and the following comment suggests that we did not provide enough clarity in this section. We modified the line to the following:

      Consistent with this report, in an agar plate-based preliminary assay we observed that ~15% of ∆mut-16 progeny arrest at various larval stages, and 2% of progeny are rod, which is suggestive of derepression of tmrl-1 in N2.

      This lets readers know that this initial characterization of the mut-16 knockout strain is different from the data presented in figure 4.

      (11) Line 230: the Figure shows ~25% of arrest for the deletion mutant of mut-16, but the text says ~15%.

      The 15% the reviewer points out was obtained from a preliminary agar plate-based experiment where we attempted to characterize the mut-16 deletion strains. We turned to a more high-throughput approach to screen through more animals for each genotype, which we report in figure 4.

      (12) Line 233: TOF, and not animal length, was compared. The authors should indicate that TOF is used as a proxy for animal length.

      We made the suggested change. The new sentences read:

      To do so, we compared time of flight (TOF) measurements—a proxy for animal length, developmental stage, and growth rate (28)—between a strain with a single knockout of mut-16 and one with a double knockout of mut-16 and the N2 tmrl-1 (a strain with a single knockout of the N2 tmrl-1 served as a negative control). We observed a reduction in TOF and an increase in the fraction of worms in larval stages in the mut-16 knockout strain, and these effects were partially rescued in the double knockout strain (Fig. 4).

      (13) Line 237-239: This claim may be overstated without additional data. Consider adding a "likely" to the statement.

      The line in question: 

      These results indicate that the reduced growth rate observed in the mut-16 knockout strain is partially mediated by derepression of the N2 mll-1 allele.

      We modified it to reflect the reviewer’s concern: 

      These results indicate that the reduced growth rate observed in the mut-16 knockout strain is partially mediated by the presence of the N2 tmrl-1 allele, likely because tmrl-1 is derepressed in mut-16 knockout strains.

      (14) Line 257: Figure S5C should be moved to line 259.

      We made the suggested move. 

      (15) Is the name mll-1 firmly established? We ask because MLL1 is a human mutation commonly associated with leukemia, and it may lead to some confusion in the field. This is a minor point, but we wanted to bring it forth.

      This name was not firmly established. We modified the names to not overlap with known gene names:

      tmrl-1 - Toxin-induced Maternal Rod Lethality

      amrl-1 - Antidote of Maternal Rod Lethality

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the manuscript "Conformational Variability of HIV-1 Env Trimer and Viral Vulnerability", the authors study the fully glycosylated HIV-1 Env protein using an all-atom forcefield. It combines long all-atom simulations of Env in a realistic asymmetric bilayer with careful data analysis. This work clarifies how the CT domain modulates the overall conformation of the Env ectodomain and characterizes different MPER-TMD conformations. The authors also carefully analyze the accessibility of different antibodies to the Env protein.

      Strengths:

      This paper is state-of-the-art given the scale of the system and the sophistication of the methods. The biological question is important, the methodology is rigorous, and the results will interest a broad elife audience. The authors also establish strong connections to previous literature and acknowledge the limitations of the CT-truncated protein construct, which enhances the manuscript's relevance to the community.

      Reviewer #2 (Public review):

      In this work, the authors elucidate how a viral surface protein behaves in a membrane environment and how its large-scale motions influence the exposure of antibody-binding sites. Using long-timescale, all-atom molecular dynamics simulations of a fully glycosylated, full-length protein embedded in a virus-like membrane, the study systematically examines the coupling between ectodomain motion, transmembrane orientation, membrane interactions, and epitope accessibility. Multiple model variants differing in cleavage state, initial transmembrane configuration, and presence of the cytoplasmic tail are compared to identify general features of protein-membrane dynamics relevant to antibody recognition.

      A major strength of this study is the scope and ambition of the simulations. The authors perform multiple microsecond-scale simulations of a highly complex, biologically realistic system that includes the full ectodomain, transmembrane region, cytoplasmic tail, glycans, and a heterogeneous membrane. The finding that the ectodomain explores a wide range of tilt angles while the transmembrane region remains more constrained, with limited correlation between the two, offers useful conceptual insight into how global motions may be accommodated without large rearrangements at the membrane anchor. The explicit consideration of membrane and glycan steric effects on antibody accessibility further strengthens the study.

      The main limitations relate to sampling and model dependence inherent to simulations of this size and complexity. The analysis of antibody accessibility is based on geometric and steric criteria, which do not capture potential conformational adaptations of antibodies or membrane remodeling during binding; the authors have appropriately noted this as a limitation.

      In the revised manuscript, the authors have addressed all previously raised concerns. Time series plots of the tilt angles have been added, figure captions and visual encodings have been clarified, quantitative descriptions of angular distributions have been strengthened, and the distance metric for MPER exposure is now accompanied by temporal data. The overall presentation is substantially improved, and the conclusions are well supported by the data as presented.

      Reviewer #3 (Public review):

      Summary:

      This study uses large-scale all-atom molecular dynamics simulations to examine the conformational plasticity of the HIV-1 envelope glycoprotein glycoprotein (Env) in a membrane context, with particular emphasis on how the transmembrane domain (TMD), cytoplasmic tail (CT), protomer cleavage, and membrane environment influence ectodomain orientation and antibody epitope exposure. By comparing Env constructs with and without the CT, explicitly modeling glycosylation, and embedding Env in an asymmetric lipid bilayer, the authors aim to provide an integrated view of how membrane-proximal regions and lipid interactions shape Env antigenicity, including epitopes targeted by MPER-directed antibodies.

      Strengths:

      The authors have made a genuine effort to address the concerns raised in the first round of review, and the revised manuscript is substantively improved. The addition of dynamical cross-correlation maps, expanded citation of prior computational work, clarification of the membrane composition rationale, data deposition to Zenodo, and the new discussion contextualizing the independence of ectodomain and TMD motions are all welcome. Several scientifically interesting aspects of the work merit highlighting before the remaining concerns are addressed.

      A key strength of this work remains the scope, scale, and realism of the simulation systems. The authors construct a very large, nearly complete-Env-scale model that includes a glycosylated Env trimer embedded in an asymmetric bilayer, enabling analysis of membrane-protein interactions that are difficult to capture experimentally. The inclusion of specific glycans at reported sites, and the focus on constructs with and without the CT or cleavage, are well motivated by existing biological and structural data.

      The observation that R696 orientation and its interacting partners give rise to asymmetric protomer conformations and distinct TMD tilts is a notable finding. The statement that interactions between R696 and lipid headgroups or CT residues can be strong enough to introduce a kink into the TMD is well-supported by representative snapshots and consistent with prior isolated-TMD simulations. The use of two initialization depths ("high" and "low") to probe R696 leaflet preference is methodologically interesting and the authors' interpretation - that there is a slight bias toward cytoplasmic leaflet interactions, but that these contacts could be highly dynamic over the course of viral entry - is appropriately cautious. It would be valuable to explicitly frame this as a hypothesis with testable predictions that future experimental or enhanced-sampling work could address. Similarly, the equilibration-driven kinking of the TMD core, consistent with prior isolated-TMD studies, represents a useful validation that extends those earlier observations to the intact trimeric context.

      The simulations reveal substantial tilting motions of the ectodomain relative to the membrane, with angles spanning roughly 0-30° (and up to ~40° in some analyses), while the ectodomain itself remains relatively rigid. This framing, that much of Env's conformational variability arises from rigid-body tilting rather than large internal rearrangements, is an important conceptual contribution. The authors also provide interesting observations regarding asymmetric bilayer deformations, including localized thinning and altered lipid headgroup interactions near the TMD and CT, which suggest a reciprocal coupling between Env and the surrounding membrane.

      The analysis of antibody-relevant epitopes across the prefusion state, including the V1/V2 and V3 loops, the CD4 binding site, and the MPER, is another strength. The study makes effective use of existing experimental knowledge in this context, for example by focusing on specific glycans known to occlude antibody binding, to motivate and interpret the simulations.

      Finally, the revised discussion provides more context that situates the study's findings and discrepancies within the broader literature, strengthening the manuscript's clarity and interpretability.

      Weaknesses:

      The revised work is much improved, but still includes substantive issues with writing including organization, such as paragraph run-ons, and citation issues. Improving these would help readers make the most of this important study.

      The revised Introduction now includes a paragraph summarizing prior MD work, which is an Improvement. However, the paragraph remains structured around the limitations and setup of previous studies (e.g., "early studies were constrained by limited computational resources", short trajectory lengths, isolated constructs) rather than their findings. Readers benefit most from understanding what those studies showed - and where the present work confirms, extends, or diverges from those results. The current framing inadvertently positions prior work as deficient scaffolding rather than as independent data points converging on shared conclusions. The Introduction could be revised to briefly summarize the key biological conclusions from prior MD studies alongside their technical context, which could then be revisited in their appropriate place alongside key results.

      The authors have verified that PDB entries are cited at first mention, and this is noted. However, a recurring issue remains: key literature-supported conclusions appear in the Results and Discussion sections without accompanying citations at each point of use. Passages that summarize experimental or computational findings - particularly those used to validate or contextualize the authors' own results - require citation at every point of claim, not only at first introduction of a reference. This is not a minor stylistic preference. Downstream readers, systematic reviewers, and automated tools that map literature to claims (e.g., scite) rely on co-occurrence of claims and citations within the same passage. A citation appearing several paragraphs earlier does not carry attribution forward. As a practical example: the statement that "MPER-targeting antibodies bind effectively only after the gp120-gp41 trimer undergoes major conformational rearrangements toward a fusion-intermediate or post-fusion state (Frey et al., 2008; Alam et al., 2009; Chen et al., 2014; Lee et al., 2016)", which is appropriate. That same standard of inline attribution should be applied throughout - including in Results and Discussion subsections where prior experimental findings are mentioned without citation.

      Additionally, cited literature should be framed to highlight convergence with the authors' conclusions, not primarily to limitations of previous studies. Where prior studies independently support a finding, this should be stated explicitly. Independent replication across methods and systems is one of the strongest arguments for ground truth; treating it as such would improve the manuscript's scientific standing.

      Finally, the dynamical cross-correlation maps assess ectodomain-TMD coupling, and the authors appropriately acknowledge that microsecond simulations capture only the closed ground state. However, the revised manuscript does not address the question raised in the first review regarding CT-TMD and CT-ectodomain correlations. The Results section states that "very weak correlations between the ectodomain and the TMD" were found, but it is not clear whether the CT was included in this analysis or whether analogous correlation maps for CT-TMD and CT-ectodomain pairs were computed for the full-length systems. Additional analyses of the authors' deposited MD trajectories-such as probing for exposure of cryptic epitopes and potential allosteric coupling-could serve as valuable extensions of this work.

      We thank the Reviewer for the further comments and suggestions. We have revised the manuscript accordingly.

      The observation that R696 orientation and its interacting partners give rise to asymmetric protomer conformations and distinct TMD tilts is a notable finding. The statement that interactions between R696 and lipid headgroups or CT residues can be strong enough to introduce a kink into the TMD is well-supported by representative snapshots and consistent with prior isolated-TMD simulations. The use of two initialization depths ("high" and "low") to probe R696 leaflet preference is methodologically interesting and the authors' interpretation - that there is a slight bias toward cytoplasmic leaflet interactions, but that these contacts could be highly dynamic over the course of viral entry - is appropriately cautious. It would be valuable to explicitly frame this as a hypothesis with testable predictions that future experimental or enhanced-sampling work could address. Similarly, the equilibration-driven kinking of the TMD core, consistent with prior isolated-TMD studies, represents a useful validation that extends those earlier observations to the intact trimeric context.

      At the end of the subsection “The energetically unfavorable R696 in the hydrophobic core results in asymmetric, kinked TMD conformations and disrupts membrane integrity” we have added

      “Taken together, these observations suggest that interactions of R696 with lipid headgroups and CT residues may modulate TMD tilt and kink formation during viral entry. However, whether the orientation of R696 dynamically switches between the two leaflets over longer timescales and whether a preference exists for either leaflet remain to be examined in future experimental and/or enhanced sampling simulation studies.”

      The revised Introduction now includes a paragraph summarizing prior MD work, which is an improvement. However, the paragraph remains structured around the limitations and setup of previous studies (e.g., "early studies were constrained by limited computational resources", short trajectory lengths, isolated constructs) rather than their findings. Readers benefit most from understanding what those studies showed - and where the present work confirms, extends, or diverges from those results. The current framing inadvertently positions prior work as deficient scaffolding rather than as independent data points converging on shared conclusions. The Introduction could be revised to briefly summarize the key biological conclusions from prior MD studies alongside their technical context, which could then be revisited in their appropriate place alongside key results.

      We have modified the original fifth paragraph in the Introduction section and subdivided it into two separate paragraphs to emphasize the key biological conclusions in prior simulation studies.

      “Molecular dynamics (MD) simulations have been employed to investigate the stability and conformational properties of monomeric and trimeric TMD. An early study of the trimeric TMD established a foundational understanding of the domain's stability, though it was limited by the computational resources available at the time (Kim et al., 2009). Subsequent work utilizing metadynamics found that the monomeric TMD is characterized by significant conformational plasticity and multiple metastable states, with the individual helix tilting in the bilayer and the midspan arginine (R696) interacting with lipid headgroups in either leaflet (Gangupomu et al., 2010; Baker et al., 2014). Baker et al. also simulated the monomeric TMD on Anton supercomputers, extended sampling to the multi-microsecond time scale, and demonstrated that TMD tilting and the interaction of R696 with lipids lead to local membrane thinning and water defects (Baker et al., 2014). Hollingsworth et al. modeled and simulated trimeric TMD in an asymmetric membrane and observed that TMD tilting and membrane thinning also occurred for the trimeric helical bundle, where water and ions permeated to stabilize the three positively charged R696 residues (Hollingsworth et al., 2018).

      Piai et al. determined the NMR structure of a construct comprising the MPER, TMD, and CT, which currently serves as the only PDB structure to include the majority of the CT residues. They complemented this structural work with MD simulations to assess the structural stability of the trimeric MPER–TMD–CT complex (Piai et al., 2021). Recently, Majumder et al. simulated the same MPER–TMD–CT complex and applied a machine learning-based approach to classify the diverse conformational ensemble of the MPER-TMD-CT (Majumder et al., 2025). Maillie et al. combined conventional MD, steered MD, and coarse-grained simulations to demonstrate that interactions between MPER-targeting antibodies and membrane lipids are critical for effective epitope recognition (Maillie et al., 2025). In addition, MD simulations have been extensively applied to characterize the well-studied ectodomain.”

      The authors have verified that PDB entries are cited at first mention, and this is noted. However, a recurring issue remains: key literature-supported conclusions appear in the Results and Discussion sections without accompanying citations at each point of use. Passages that summarize experimental or computational findings - particularly those used to validate or contextualize the authors' own results - require citation at every point of claim, not only at first introduction of a reference. This is not a minor stylistic preference. Downstream readers, systematic reviewers, and automated tools that map literature to claims (e.g., scite) rely on co-occurrence of claims and citations within the same passage. A citation appearing several paragraphs earlier does not carry attribution forward. As a practical example: the statement that "MPER-targeting antibodies bind effectively only after the gp120-gp41 trimer undergoes major conformational rearrangements toward a fusion-intermediate or post-fusion state (Frey et al., 2008; Alam et al., 2009; Chen et al., 2014; Lee et al., 2016)", which is appropriate. That same standard of inline attribution should be applied throughout - including in Results and Discussion subsections where prior experimental findings are mentioned without citation.

      Additionally, cited literature should be framed to highlight convergence with the authors' conclusions, not primarily to limitations of previous studies. Where prior studies independently support a finding, this should be stated explicitly. Independent replication across methods and systems is one of the strongest arguments for ground truth; treating it as such would improve the manuscript's scientific standing.

      In addition to summarizing the biological conclusions from prior simulation studies in our response to the previous comment, we have also added the following citations.

      “Human immunodeficiency virus type 1 (HIV-1) is the most prevalent strain of HIV responsible for the development of acquired immunodeficiency syndrome (AIDS) (Sharp et al., 2011). The HIV-1 envelop (Env) consists of a host cell-derived lipid membrane and viral glycoproteins that play a crucial role in mediating viral entry into host cells. The Env glycoprotein is initially synthesized in the endoplasmic reticulum (ER) as a precursor gp160 and cleaved by furin into two subunits, gp120 and gp41. The non-covalently associated gp120–gp41 complex is transported to the cell surface in the form of a trimer, where it is subsequently incorporated into the envelope of nascent virions during viral assembly (Wyatt et al., 1998). The exposure of Env protein is essential for binding to the primary receptor CD4 and the co-receptors CCR5 or CXCR4, triggering membrane fusion and viral entry (Dalgleish et al., 1984; Feng et al., 1996; Huang et al., 1996). However, this exposure also renders the virus susceptible to immune attack. In response to host immune pressure, Env is densely coated with N-linked glycans added during post-translational modification in the ER and Golgi apparatus, which effectively shield vulnerable epitopes from immune recognition (Wei et al., 2003).”

      “While MPER plasticity has been linked to its role in virus-host membrane fusion because it enables the ectodomain and TMD to adopt distinct orientations during large-scale structural rearrangements (Salzwedel et al., 1999), our results show that this flexibility is already inherently present in the prefusion state.”

      “However, transition among these three states occur on millisecond-to-second timescales (Munro et al., 2014).”

      Finally, the dynamical cross-correlation maps assess ectodomain-TMD coupling, and the authors appropriately acknowledge that microsecond simulations capture only the closed ground state. However, the revised manuscript does not address the question raised in the first review regarding CT-TMD and CT-ectodomain correlations. The Results section states that "very weak correlations between the ectodomain and the TMD" were found, but it is not clear whether the CT was included in this analysis or whether analogous correlation maps for CT-TMD and CT-ectodomain pairs were computed for the full-length systems. Additional analyses of the authors' deposited MD trajectories-such as probing for exposure of cryptic epitopes and potential allosteric coupling-could serve as valuable extensions of this work.

      We have updated the manuscript to address the correlations involving the CT. Figure 2—figure supplements 12 and 13 display the dynamical cross-correlation maps (DCCM) for the full-length systems (including the CT), which indicate low correlations between the ectodomain and the CT. We have modified the figure captions to explicitly state that the CT is included in these analyses. We have also clarified in the text that we do not further interpret the coupling of the CT with the other domains. As the Reviewer noted, the high structural heterogeneity of the CT makes defining consistent parameters (such as a tilt angle) impractical. Given this variability, along with the inherent uncertainty in the experimental structure of the CT, we believe it is important to avoid over interpreting these observations.

      “Although Figure 2—figure supplements 12 and 13 also show low correlations between the ectodomain and the CT, we do not further interpret the coupling of the CT with the other domains due to its structural heterogeneity and the inherent uncertainty in its experimental structure.”

      We have modified captions of Figure 2—figure supplements 10–13

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      The authors have made meaningful progress in addressing first-round concerns. The remaining issues center on how prior literature is framed and integrated - not just cited - throughout the manuscript, consistent attribution at each point of claim, clarification of the CT correlation analysis, and major writing improvements. Addressing these points would substantially strengthen the manuscript's contribution to the field.

      Abstract

      "knowledge of the cytoplasmic tail (CT) is virtually absent" is overstated. While structural data for the CT are limited and largely uncertain, the CT has been extensively studied functionally and some NMR structural data exist (Piai et al., 2021; Murphy et al., 2017). Suggest revising to reflect that high-resolution structural information for the CT in the context of the intact trimer remains limited

      We have revised the abstract according to the Reviewer’s suggestion.

      “While structural information is available for the membrane-proximal external region (MPER) and transmembrane domain (TMD), these regions remain comparatively understudied. Furthermore, high-resolution structural information for the cytoplasmic tail (CT), particularly within the context of the intact trimer, is limited and largely uncertain.”

      Introduction

      The first paragraph is unreferenced. Foundational claims about HIV-1 biology, Env processing, and glycan shielding should carry at least landmark citations for readers new to the field.

      We have added references to the first paragraph.

      “Human immunodeficiency virus type 1 (HIV-1) is the most prevalent strain of HIV responsible for the development of acquired immunodeficiency syndrome (AIDS) (Sharp et al., 2011). The HIV-1 envelop (Env) consists of a host cell-derived lipid membrane and viral glycoproteins that play a crucial role in mediating viral entry into host cells. The Env glycoprotein is initially synthesized in the endoplasmic reticulum (ER) as a precursor gp160 and cleaved by furin into two subunits, gp120 and gp41. The non-covalently associated gp120–gp41 complex is transported to the cell surface in the form of a trimer, where it is subsequently incorporated into the envelope of nascent virions during viral assembly (Wyatt et al., 1998). The exposure of Env protein is essential for binding to the primary receptor CD4 and the co-receptors CCR5 or CXCR4, triggering membrane fusion and viral entry (Dalgleish et al., 1984; Feng et al., 1996; Huang et al., 1996). However, this exposure also renders the virus susceptible to immune attack. In response to host immune pressure, Env is densely coated with N-linked glycans added during post-translational modification in the ER and Golgi apparatus, which effectively shield vulnerable epitopes from immune recognition (Wei et al., 2003).”

      A paragraph break after "... and cytoplasmic tail (CT), are relatively understudied" would improve readability by separating the general context from the MPER/TMD-specific discussion that follows.

      A paragraph break before "Similarly, there are different conclusions about" would separate the TMD oligomeric state discussion from the MPER conformation discussion and improve navigation.

      A paragraph break after "Despite these advances, it remains challenging to investigate the gp120-gp41 trimer as an intact entity considering its structural complexity" would clearly delineate the literature context from the description of the present work.

      We have introduced paragraph breaks as suggested to improve the flow and readability of the introduction.

      The biological rationale for simulating both cleaved and uncleaved systems should be stated explicitly in the Introduction. Readers unfamiliar with the furin cleavage biology and NFL trimer constructs will benefit from a sentence explaining why this comparison is informative.

      In the middle of the last paragraph in the Introduction section we have added

      “While host furin cleavage of the gp160 precursor into gp120 and gp41 is a prerequisite for viral infectivity (McCune et al., 1988), native virions also incorporate a fraction of uncleaved gp160 (Zhang et al., 2021). Furthermore, many current immunogen designs, such as NFL and UFO constructs, utilize a covalent linker to stabilize the metastable prefusion conformation (Sharma et al., 2015; Kong et al., 2016). Therefore, we simulated both cleaved and uncleaved trimers to explore how the absence of proteolytic cleavage impacts the conformational landscape.”

      The modifier “subsequently” in “Majumder et al. subsequently simulated...” implies temporal sequence and invites the inference that Majumder et al.'s work is less sophisticated or prior. Given that both works are recent and peer-reviewed, a neutral modifier such as “recently” or “independently” is more appropriate.

      We agree that a more neutral modifier is appropriate and have replaced “subsequently” with “recently” to avoid any unintended inference.

      In the beginning of the sixth paragraph in the Introduction section we have modified

      “Piai et al. determined the NMR structure of a construct comprising the MPER, TMD, and CT; to date, this is the only PDB structure including the majority of CT residues. They complemented this structural work with MD simulations to access the structural stability of the trimeric MPER–TMD–CT complex (Piai et al., 2021). Recently, Majumder et al. simulated the same MPER–TMD–CT complex and applied a machine learning-based approach to classify its conformational ensemble (Majumder et al., 2025).”

      The sentence "Moreover, we selected several bNAbs targeting the epitopes across different regions of the Env protein and demonstrate that the simulation trajectories can be used to assess the epitope accessibility" implies that simulations of antibody binding were performed. This should be rephrased, for example: "Moreover, we selected various epitopes across Env that are targeted by bNAbs and demonstrate that the MD simulation trajectories can be used to assess epitope accessibility."

      At the end of the Introduction section we have modified

      “Moreover, by analyzing epitopes targeted by various bNAbs, we demonstrate that the simulation trajectories can be leveraged to assess the epitope accessibility.”

      The revised Methods section now cites van Meer et al. (2008) and Sampaio et al. (2011) as primary experimental sources for plasma membrane composition, which is appropriate. However, the Introduction still contains the statement: "we built a model of full-length gp120-gp41 trimer embedded in a lipid bilayer mimicking the lipid composition of the mammalian plasma membrane (Pogozheva et al., 2022)". This cites only the authors' own prior simulation study. A primary experimental reference (van Meer et al., 2008 and/or Sampaio et al., 2011) should be added here as well, so that readers encountering the claim in the Introduction have direct access to the supporting evidence.

      At the beginning of the last paragraph in the Introduction section we have modified

      “In this work, we built a model of full-length gp120–gp41 trimer embedded in a lipid bilayer mimicking the lipid composition of the mammalian plasma membrane (van Meer et al., 2008; Sampaio et al., 2011; Ingolfsson et al., 2014; Pogozheva et al., 2022) (Figure 1).”

      Additionally, a brief note in the Introduction on the cell type specificity of the plasma membrane model used (or its absence) would be informative, as membrane composition varies substantially across mammalian cell types and the choice has potential consequences for the conclusions.

      We have added a brief note clarifying that differences between the model membrane and the native viral envelope may influence the study's conclusions, particularly regarding protein-lipid interactions.

      “We chose this composition as a representative baseline, though we acknowledge that the native viral envelope may exhibit a distinct lipid profile that could influence protein-lipid interactions.”

      Results

      Connecting the observed accessibility frequencies to known neutralization potency, breadth, or escape propensity for each antibody class (PGT128, PG9, VRC01, 35O22, 10E8, 4E10) would provide a mechanistic framework and substantially increase the impact of this section. Even a brief discussion of how glycan shielding dynamics relate to reported neutralization sensitivity data would add value.

      We have expanded the Results section to include a comparison between our computational accessibility frequencies and established experimental metrics (potency and breadth).

      At the end of each paragraph in the subsection “Ectodomain epitopes are conditionally accessible, whereas MPER epitopes are virtually inaccessible in the closed prefusion state” we have added

      “The high accessibility frequency observed for the PGT128 epitope aligns with its exceptional potency. As demonstrated by Walker et al., PGT128 is capable of neutralizing approximately 72% of global isolates with a median IC<sub>50</sub> of ~0.02 µg/mL. This potency is approximately 10-fold greater than that of PG9 and VRC01, though its breadth is lower than the 93% reported for VRC01 (Walker et al., 2011). This comparatively lower breadth may be attributed to strict sequence dependency. Because PGT128 recognition depends on the N332-centered glycan epitope, loss, truncation, or shifting of the N332 glycan to N334 prevents productive engagement regardless of local steric accessibility.”

      “This is consistent with the lower neutralization potency and moderate breadth of PG9, which exhibits a median IC<sub>50</sub> of ~0.22 µg/mL and a breadth of ~79% (Walker et al., 2009).”

      “This intermediate accessibility is consistent with the biological requirement of the CD4 binding site to remain periodically available for receptor engagement while maintaining a certain degree of glycan shielding to evade neutralization. The potency of VRC01 is even lower than that of PG9, with a reported median IC50 of ~0.32 µg/mL, but it possesses an exceptionally high breadth of ~93% (Wu et al., 2010; Walker et al., 2011).”

      “Altogether, these results demonstrate that epitope accessibility for this antibody is highly sensitivity to the membrane environment, glycan orientation and ectodomain tilting. This complex dependency provides a structural context for the experimental profile of 35O22, which exhibits high potency with a median IC<sub>50</sub> of ~0.03 µg/mL, but a relatively limited breadth of ~62% (Huang et al., 2014).”

      “Though differing in potency — with 10E8 exhibiting a median IC<sub>50</sub> of ~0.35 µg/mL compared to ~1.93 µg/mL for 4E10 — both antibodies demonstrate extremely high breadth of ~98% (Huang et al., 2012). This extensive breadth is primarily attributed to the high sequence conservation of the MPER across global isolates. The negligible epitope accessibility observed in the prefusion trimer supports the conclusion that these antibodies require the transition of the Env trimer into intermediate states to fully engage their epitopes (Frey et al., 2008).”

      The first paragraph of the Results section dives directly into trajectory notation without a brief summary of the simulation systems. A short opening paragraph (2-3 sentences) summarizing the number of systems, the variables tested (cleavage, CT presence, TMD position), and the total number of trajectories would orient the reader before the naming convention is introduced.

      We have moved the original first sentence in the Material and methods — Simulation details subsection to the beginning of the Results section. This sentence summarizes all the configurations we have considered and the number of independent trajectories for each configuration.

      “The combination of cleavage state (cleaved vs. uncleaved), sequence length (full-length vs. CT-truncated), and initial TMD position in the membrane (high vs. low) resulted in eight distinct configurations, and we performed three independent 1-μs all-atom MD simulations for each configuration.”

      The statement "very weak correlations between the ectodomain and the TMD" leaves open the question of CT-TMD and CT-ectodomain correlations. If a tilt angle cannot be defined for the CT due to its structural heterogeneity, this should be stated.

      We have updated the manuscript to address the correlations involving the CT. Figure 2—figure supplements 12 and 13 display the dynamical cross-correlation maps (DCCM) for the full-length systems (including the CT), which indicate low correlations between the ectodomain and the CT. We have modified the figure captions to explicitly state that the CT is included in these analyses. We have also clarified in the text that we do not further interpret the coupling of the CT with the other domains. As the Reviewer noted, the high structural heterogeneity of the CT makes defining consistent parameters (such as a tilt angle) impractical. Given this variability, along with the inherent uncertainty in the experimental structure of the CT, we believe it is important to avoid over interpreting these observations.

      “Although Figure 2—figure supplements 12 and 13 also show low correlations between the ectodomain and the CT, we do not further interpret the coupling of the CT with the other domains, considering its structural heterogeneity and the inherent uncertainty in its experimental structure.”

      We have modified captions of Figure 2—figure supplements 10–13

      Throughout the Results, several long paragraphs could be broken up. In particular, the TMD section and the MPER exposure section each contain dense multi-example run-on paragraphs that would benefit from subdivision.

      We agree with the Reviewer and have introduced multiple paragraph breaks in the Results section to improve the flow and readability. In instances where longer paragraphs remain, they have been intentionally preserved to maintain the logical integrity of closely linked results, ensuring the reader can follow a single cohesive argument without interruption.

      Discussion

      The statement "transition among three states occur on millisecond-to-second timescales" is an important claim that contextualizes the limitations of the microsecond simulations, but it is currently uncited. This should be attributed to the relevant experimental smFRET work (Munro et al., 2014 is cited in the preceding sentence, but not explicitly for this claim) and/or any additional literature that established these timescales for Env conformational switching.

      We have now explicitly attributed the claim regarding the millisecond-to-second timescales of Env conformational transitions to the relevant smFRET literature (Munro et al., 2014).

      In the middle of the second paragraph in the Discussion section we have added

      “However, transition among these three states occur on millisecond-to-second timescales (Munro et al., 2014).”

      The Discussion contains several extended paragraphs that could be subdivided to improve readability and help the reader navigate between distinct topics (e.g., MPER flexibility, CT effects, coupling, lipid composition, antibody accessibility).

      We have subdivided the Discussion section as suggested by the Reviewer to improved readability.

    1. Author response:

      The following is the authors’ response to the original reviews.

      In this revised manuscript, we added new analyses of the DNA-binding tail domain of Kid. AlphaFold 3 predictions suggested that dimeric Kid interacts more stably with double-stranded DNA than monomeric Kid. To experimentally test this prediction, we introduced a point mutation into a critical residue predicted to contribute to DNA binding. Consistent with the AlphaFold 3 model, this mutation abolished the interaction between Kid and DNA.

      We also extended our DNA transport assays by testing DNA substrates of different lengths. In addition to 100-bp double-stranded DNA, full-length Kid transported 1,000-bp and 2,000-bp DNA molecules along microtubules in vitro. These findings show that Kid can transport longer duplex DNA substrates than those initially tested, although these substrates do not fully recapitulate the organization of condensed chromatin.

      Furthermore, we performed dual-color imaging using independently purified Kid-mScarlet3 and Kid-mStayGold proteins. We consistently observed co-migration of the two fluorescently labeled Kid molecules along microtubules, supporting the conclusion that Kid forms dimers on microtubules.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Mitotic kinesins carry out crucial roles in intracellular motility and mitotic spindle organization. Although many mitotic kinesins have been extensively studied, a few conserved mitotic motors remain poorly explored, including chromosome-associated kinesins. Here, Furusaki et al reconstitute recombinant chromosome-associated kinesin or chromokinesin (Kid) and reveal processive plus-end motility along microtubules. The authors purify multiple versions of Kid, revealing dimeric organization and their processive microtubule plus-ended motility which depends on their conserved motor domains, neck linkers, and coiled-coil regions. The study reveals for the first time that KID can recruit and transport duplex DNA along microtubules using its conserved C-terminal DNA binding domain. The work provides crucial revised thinking about the mechanisms of Chromokinesins mitosis as physical processive motors that mobilize chromosomes towards the microtubule plus ends in early metaphase.

      Strengths:

      The authors reconstitute multiple chromosome-associated kinesin (KID) orthologs from Xenopus and humans with microtubules and determine their oligomerization. The study shows how coiled-coil and neck linker regions of KID are essential for its function as its deletion leads to non-processive motility. CHimeras placing the KID coiled-coil and neck linker on the KIF1A motor domain led to the production of a processive recombinant motor supporting the compatibility of their motility mechanisms. The KID c-terminal tail binds and transports only double-stranded DNA and its deletion or single-stranded DNA leads to defects in this activity.

      Thank you very much.

      Weaknesses:

      A minor weakness in the studies is that they do not resolve the mechanisms of KID in binding large duplex DNA molecules or condensed chromatin. The authors suggest a model in which KID forms multimers along large chromosomes that lead to their transport, but this model was not directly tested.

      We agree with the reviewer that our study does not directly resolve how Kid binds large duplex DNA molecules or condensed chromatin. In the revised manuscript, we have therefore softened our model and now present the idea that multiple Kid dimers act along chromosomes as a possible mechanism rather than a demonstrated conclusion. To strengthen the mechanistic basis of DNA binding, we added AlphaFold 3-based analysis of the Kid DNA-binding tail domain and experimentally tested a predicted DNA-binding residue. Mutation of this residue abolished Kid–DNA binding, supporting the proposed role of the tail domain in DNA engagement. We also added dual-color imaging experiments showing co-migration of independently purified Kid-mScarlet3 and Kid-mStayGold on microtubules, supporting dimer formation on microtubules. We now explicitly state that future studies using chromatinized DNA or chromosome-like substrates will be required to determine how Kid interacts with condensed chromatin in a cellular context.

      Reviewer #2 (Public review):

      Summary:

      Previous work in the field highlighted the role of the kinesin-10 motor protein Kid (KIF22) in the polar ejection force during prometaphase. However, the biochemical and biophysical properties of Kid that enabled it to serve in this role were unclear. The authors demonstrate that human and xenopus Kid proteins are processive kinesins that function as homodimeric molecules. The data are solid and support the findings although the text could use some editing to improve clarity.

      Strengths:

      A highlight of the work is the reconstitution of DNA transport in vitro.

      A second highlight is the demonstration that the monomer vs dimer state is dependent on protein concentration.

      Thank you very much.

      Weaknesses:

      The authors make several assumptions of the monomer vs dimer state of various Kid constructs without verifying the protein state using e.g. size exclusion chromatography and/or nanophotometry.

      We newly added mass photometry analysis in Figure 3 and Figure 5.

      They also make statements about monomer-to-dimer transitions on the microtubule without showing or quantifying the data.

      We performed dual color imaging to show the assembly of Kid monomers on microtubules.

      The discussion needs to better put the work into context regarding the ability of non-processive motors to work in teams (formerly thought to be the case for Kid) and how their findings on Kid change this prevailing view in the case of polar ejection force.

      We have revised the Discussion to better place our findings in the context of collective motor function and polar ejection force generation. Previous biochemical studies led to the prevailing model that Kid is a monomeric and non-processive chromokinesin. Under this model, sustained chromosome movement would require many Kid monomers distributed along chromosome arms to act collectively. Our findings revise this view. We show that full-length Kid forms homodimers, moves processively along microtubules, and directly transports double-stranded DNA. Thus, the elementary force-generating unit of Kid is unlikely to be a non-processive monomer. Instead, a single Kid dimer may act as a processive DNA-bound motor. In the context of mitotic chromosomes, multiple processive Kid dimers bound along chromosome arms could cooperate to generate chromosome-scale polar ejection forces. We have clarified in the Discussion that our model does not exclude ensemble behavior. Rather, it changes the nature of the proposed ensemble from many non-processive monomers to multiple processive dimers.

      The authors also do not mention previous work on kinesins with non-conventional neck linker/neck coil regions that have been shown to move processively. Their work on Kid needs to be put into this context.

      We thank the reviewer for this important suggestion. We have revised the Discussion to place Kid in the broader context of processive kinesins with non-conventional neck linker or neck coil regions. We now discuss previous work showing that neck-linker length strongly influences kinesin processivity, and that changes in neck-linker length alter the run length and motility properties of kinesin-1, kinesin-2, and other N-terminal kinesins (Shastry and Hancock, 2010; Shastry and Hancock, 2011).

      We also discuss studies showing that longer or non-conventional neck linker regions can provide additional functions beyond supporting processive stepping. For example, kinesin-2 can bypass Tau and other microtubule-bound obstacles by protofilament switching, and the neck linker of the mitotic kinesin KIF18A contributes to obstacle navigation within the mitotic spindle (Hoeprich et al., 2014; Malaby et al., 2019).

      In this context, we now emphasize that Kid has an exceptionally long and flexible neck linker, approximately four times longer than that of kinesin-1. Despite this non-canonical architecture, the Kid neck linker and coiled-coil region support processive motility, as shown by the processive movement of the KIF1A–Kid chimera. We therefore propose that Kid represents a non-conventional processive chromokinesin whose extended neck linker may help it move along crowded spindle microtubules while remaining attached to DNA or chromatin. We have also stated that this possibility remains to be tested directly.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Furusaki et al reconstitute effectively the chromosome-associated kinesin. The studies are well performed and effectively controlled with few minor suggestions

      The studies generally lack a few minor items that would improve the current work:

      (1) Alpha fold or coiled-coil predictions of the c-terminal region characterizing its organization or the nature of its interaction site with DNA. These should aid the presentation of the work and help refine the boundaries for coiled coils and the DNA binding domain.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we added AlphaFold 3-based structural predictions and coiled-coil predictions for the C-terminal region of Kid (Figure 7). These analyses helped define the predicted DNA-binding tail domain more clearly. The AlphaFold 3 model also suggested a potential DNA-interaction surface within the C-terminal DNA-binding region. We have incorporated these predictions into the revised figure and modified the text to clarify the domain organization of Kid.

      (2) The DNA transport motor activity is quite interesting and extending those studies to cover larger segments of DNA which may bind multiple kid motors would be very interesting.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we extended our DNA transport assays using longer double-stranded DNA fragments. In addition to the 100-bp DNA substrate, we tested 1,000-bp and 2,000-bp DNA fragments. Full-length Kid was able to transport both 1,000-bp and 2,000-bp double-stranded DNA along microtubules in vitro. These new data are now included in Figure 6F–I. Interestingly, the motile parameters of 1,000-bp and 2,000-bp DNA were comparable to those observed with 100-bp DNA. This result suggests that, under our reconstituted assay conditions, increasing DNA length does not substantially enhance the apparent transport velocity or run length. One possible explanation is that the interaction between Kid and naked DNA is relatively weak, and thus only one or a small number of Kid molecules productively engage each DNA molecule during transport. Alternatively, additional Kid molecules bound to longer DNA may not strongly affect the measured motility parameters under these assay conditions.

      We have added this point to the revised manuscript and now discuss that, in cells, additional factors such as chromatin proteins or chromosome-associated proteins may enhance the avidity or organization of Kid on chromosomes. Future studies using chromatinized DNA or chromosome-like substrates will be needed to determine how multiple Kid molecules engage large chromatin substrates during chromosome congression.

      (3) The final model regarding KID transporting chromosomes is probably oversimplified since there are few experiments with large stretches of DNA or chromatin that were not conducted. I suggest longer segments of DNA be studied or the model be redrawn to scale.

      We thank the reviewer for this important comment. We agree that the original model was oversimplified because naked DNA fragments do not fully recapitulate the size, structure, or mechanical properties of condensed chromatin or mitotic chromosomes. To address this concern experimentally, we extended our DNA transport assays to longer double-stranded DNA fragments. In addition to 100-bp DNA, we tested 1,000-bp and 2,000-bp DNA fragments and found that full-length hKid can transport both substrates along microtubules in vitro. These new data are now included in Figure 6F–I.

      However, we agree that these DNA substrates are still much simpler than condensed chromatin. We have therefore revised the final model to avoid implying that the transport of naked DNA fully explains chromosome-scale movement. The revised model now emphasizes that Kid dimers can directly couple DNA to microtubule-based motility, and that multiple Kid dimers may cooperate on chromosome arms to generate polar ejection forces. We state "This model is not drawn to scale and does not fully represent the structural complexity of condensed chromatin." in the revised legends.

      We also state explicitly in the Discussion that future experiments using chromatinized DNA or reconstituted chromosome-like substrates will be required to determine how Kid engages condensed chromatin and generates chromosome-scale forces.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) The authors state that XKid(1-437), which lacks the coiled-coil domain, did not show any processive runs yet Figure 3D does show short events that look like directed movement. They do not appear to be diffusive events as they are uni-directional. The authors need to quantify these results (motility, mean square displacement) as they are essential to their arguments about monomer vs dimer state and processive motility.

      We thank the reviewer for pointing this out. We agree that, in the original kymographs acquired at lower temporal resolution, some short XKid(1–437) events could appear as directional movements. To address this concern, we repeated the single-molecule motility assays with improved temporal resolution. In the revised manuscript, the kymographs for XKid(1–437) were generated from data acquired at 100 ms per pixel, instead of 3 s per pixel in the previous version. This higher temporal resolution more clearly shows that XKid(1–437) undergoes short, diffusion-like fluctuations rather than sustained unidirectional processive movement.

      We also quantified these trajectories by mean-square displacement analysis. XKid(1–495), which retains the coiled-coil domain, showed superlinear MSD scaling with an α value of approximately 1.6, consistent with persistent, directionally biased movement. In contrast, XKid(1–437), which lacks the coiled-coil domain, showed an α value of approximately 0.8, consistent with hindered or diffusion-like motion rather than sustained processive motility.

      We have added these higher-temporal-resolution data and MSD quantification to the revised Figure 3 and revised the text accordingly. We now state that XKid(1–437) lacks sustained processive runs, rather than implying that it shows no movement at all.

      The authors speculate that the lack of XKid(1-437) processive runs is due to it being unable to form a homodimer. To confirm that the coiled-coil domain is responsible for dimerization, they fuse the coiled-coil to a fluorescent protein. However, the authors should actually show that XKif(1-437) is a monomer by size exclusion chromatography and/or nanophotometry.

      We thank the reviewer for this important suggestion. We agree that directly determining the oligomeric state of XKid(1–437) is essential for interpreting the loss of processive motility. We therefore performed mass photometry to measure the molecular mass of purified XKid(1–437).

      The mass photometry analysis showed that XKid(1–437) was predominantly monomeric, with no detectable dimer population under the conditions tested. In contrast, XKid(1–495), which retains the coiled-coil domain, showed a minor dimer population, similar to full-length XKid. These results support the conclusion that deletion of the coiled-coil domain disrupts Kid dimerization.

      Together with the motility assays and MSD analysis, these data indicate that the coiled-coil domain is required for homodimer formation and sustained processive motility of Kid. We have added these mass photometry data to the revised Figure 3 and revised the text accordingly.

      (2) Likewise, the chimeric protein KIF1AMD-XKidSt shows processive motility (Figure 4), and thus authors conclude that it must be a dimer. This should be verified using size exclusion chromatography and/or nanophotometry.

      We agree that the oligomeric state of KIF1AMD–XKidSt should be directly examined. We therefore performed mass photometry analysis of purified KIF1AMD–XKidSt.

      Mass photometry showed that KIF1AMD–XKidSt behaved similarly to full-length XKid and XKid(1–495). Under the nanomolar concentrations used for mass photometry, KIF1AMD–XKidSt was predominantly monomeric but retained a detectable dimer population. This behavior is consistent with our analysis of full-length Kid and XKid(1–495), which form weak, concentration-dependent dimers. These results indicate that the XKid stalk region in the chimera can support dimer formation, although the dimer is weak under dilute solution conditions.

      (3) Lines 236-239, the authors state "in TIRF-based motility assays, although Kid predominantly dissociates into monomers in solution, its direct interaction with microtubules leads to an increased local concentration of Kid on the microtubule surface. As a result, this would facilitate the formation of Kid dimers on the microtubules, leading to processive motility." This statement implies that monomeric motors diffuse on the microtubule surface until they can associate and begin processive motion. Do the authors see such events (diffuse motion and/or association of single monomers on microtubules and a resulting change to processive motion? The kymograph in Figure 1C shows only static and motile events for XKid but hKid does appear to undergo diffusive motion. What is the percent of static vs diffusive vs processive events and how does this change with increased concentrations of XKid and HKid?

      We thank the reviewer for this important point. We agree that our original statement was too strong, because we did not directly observe monomeric Kid molecules diffusing on microtubules and then associating to initiate processive movement. We have revised the text to clarify that microtubule-dependent dimerization is a model.

      To test this model, we performed dual-color imaging using independently purified hKid–mScarlet3 and hKid–mStayGold. These proteins were mixed at 1 pM each, a concentration at which Kid is expected to be predominantly monomeric in solution. We observed co-migration of the two fluorescently labeled Kid proteins along microtubules, supporting the idea that Kid molecules can associate on microtubules and move together.

      However, because of the limited temporal resolution of our two-color TIRF system, we could not directly capture the transition from two monomers to a processive dimer on the microtubule surface. We therefore do not quantify the fraction of static, diffusive, and processive events as a function of concentration in this revised manuscript. Instead, we have softened the relevant statement and explicitly note this limitation in the Discussion.

      (4) Lines 171-172 - optimal length of neck linker for coordination of the two motor domains has only been shown for kinesin-1 and kinesin-2. In contrast, there are a number of kinesins that do not have typical neck linker domains yet can achieve processivity. The authors need to discuss this work and put their results with Kid into this context.

      As described above, we have revised the Discussion to place Kid in the broader context of processive kinesins with non-conventional neck linker or neck coil regions. We now discuss previous work showing that neck-linker length strongly influences kinesin processivity, and that changes in neck-linker length alter the run length and motility properties of kinesins (Shastry and Hancock, 2010; Shastry and Hancock, 2011).

      We also discuss studies showing that longer or non-conventional neck linker regions, such as those of kinesin-2 and KIF18, can provide additional functions beyond supporting processive stepping (Hoeprich et al., 2014; Malaby et al., 2019). In this context, we now emphasize that Kid has an exceptionally long and flexible neck linker, approximately four times longer than that of kinesin-1. We described a possibility that the extended neck linker of Kid may help it move along crowded spindle microtubules while remaining attached to DNA or chromatin while this possibility remains to be tested directly.

      Minor points:

      (5) Lines 68-69 should note that non-processive motors have been shown to move cargo if they are present in multiple copies of the cargo. This should also be discussed in the Discussion.

      We described it in the revised manuscript:

      “Under this model, sustained chromosome movement would require many Kid monomers distributed along chromosome arms to act collectively.”

      “This model preserves the likely importance of motor ensembles on large chromatin, but changes the nature of the ensemble from many non-processive monomers to multiple processive dimers.”

      (6) For Figure 4, does the KIF1AMD-XKidSt chimeric protein contain both the stalk (coiled-coil?) and tail (DNA binding?) regions of XKid or just the stalk as shown in the schematic?

      We included coiled-coil domain only.

      (7) For Figure 5, please provide a schematic for XKid(delta tail).

      We now added Alphafold 3 data.

      Senior Editor:

      Along the lines of reviewer #2's request to put the results in the context of existing knowledge, please consider whether you want to cite Pike et al. 2018 (https://doi.org/10.1126/scisignal.aaq1060; some evidence for dimerization in Fig. 4) and Walker et al. 2019 (https://pubs.acs.org/doi/10.1021/acs.biochem.9b00011).

      We have cited these papers in the revised manuscript. These are consistent with our finding that Kid can form dimer at higher concentration while dissociate to monomers in lower concentrations.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      This is a well-written and fully documented methods paper.

      The authors have established a clear rationale for their new packages, especially for real-time use, and demonstrate significant speed improvements that will likely appeal to many users of tools like DLC, SLEAP, and LightningPose. The inclusion of a graphical user interface will help make the package more accessible to neuroscientists with limited computational expertise. While it may be challenging to get users to switch from their established workflows for video analysis, the speed gains offered by this package make it worth considering. The hardware aspects of the project are well-documented, and the GitHub repository for this part of the setup is also thorough. Overall, this paper provides a clear summary of the tools, their uses, setup, and benefits.

      We thank this reviewer for the positive comments and have provided responses to the specific and constructive questions listed below.

      I have a few minor questions about the collective set of tools.

      First, the GitHub repository for SqueakPoseStudio appears to be missing a testing routine and associated badge, and the package has not been formally released. This means users would need to download the repository to install it, correct? I suggest the authors consider publishing a formal release of the package, making it installable via pip, and including a basic testing routine to clearly display the package's status on the repository page. Adding a DOI from Zenodo would also be helpful. A testing routine is especially useful when updates are made, as many users avoid repositories with failing tests.

      We thank the reviewer for this helpful suggestion. We agree that visible testing improves user confidence and reproducibility.

      SqueakPose Studio is currently distributed through a repository-based uv workflow rather than through PyPI alone. This is intentional. The application depends on platform-specific deep-learning libraries, and cloning the repository followed by uv sync provides a reproducible environment across Linux, macOS, and Windows while allowing the application to select CUDA, Apple MPS, or CPU execution at runtime. The written installation instructions now clearly describe this workflow.

      In response to the reviewer’s suggestion, we have added a unit-test suite covering the core helper modules used for label handling, dataset export, prediction, inference, and training logic. We have also added an automated GitHub Actions workflow that runs the tests on pushes and pull requests, together with a repository badge that displays the current test status.

      Second, the installation instructions simply state "Create a virtualenv and install:". This may not be sufficient for many researchers, as most neuroscientists are not experienced Python programmers and require clear guidance on the environment specific to this package. The installation instructions should be expanded to provide more detailed guidance and encourage more users. It would also be helpful to verify that the setups work across Windows, Mac, and Linux.

      We agree that installation guidance should be accessible to researchers who may not routinely manage Python environments. In addition to the existing video walkthrough, we have expanded the written GitHub documentation to provide a clearer, step-by-step installation checklist.

      The revised README now distinguishes required components from optional tools, explains the repository-based uv workflow, and provides the minimal commands needed to create the managed environment and launch the application.

      We have also clarified that an integrated development environment is optional. Although Visual Studio Code is used in the tutorial as a convenient interface for demonstrating the workflow, users may launch SqueakPose Studio directly from a terminal and are not required to use Visual Studio Code, Visual Studio, or any other editor.

      We have tested the application on Apple Silicon macOS systems, Windows systems, Linux systems, and NVIDIA GPU-enabled machines. SqueakPose Studio selects CUDA, Apple MPS, or CPU execution at runtime according to availability. Because accelerator support is partly determined by upstream packages such as PyTorch and Ultralytics, we have added links to the relevant compatibility documentation so that users can confirm whether their current hardware and driver configuration are supported.

      Third, the package defaults to UMAP for non-linear dimensionality reduction, which has some known issues. Can the package be modified to allow for alternative mapping methods, such as PaCMAP, PyDiffMap, or the more comprehensive topometry package?

      We agree with the reviewer that UMAP has limitations and that no single nonlinear dimensionality-reduction method is optimal for all pose datasets or behavioral questions.

      In SqueakPose Studio, the UMAP/HDBSCAN workflow is included as an accessible exploratory example for dimensionality reduction and clustering of pose-derived features. Our goal was not to designate UMAP as a preferred or definitive analysis method, but to provide an interpretable starting point that allows users to identify candidate clusters and inspect representative videos to evaluate what the embedding is capturing.

      We agree that supporting additional approaches, such as PaCMAP, PyDiffMap, or related tools, could be useful, and we will consider adding these as modular options in future versions. At the same time, SqueakPose Studio is not intended to replace specialized downstream behavioral-analysis packages or to adjudicate which embedding method is best for a particular dataset. Pose outputs can be exported for downstream analysis in other environments, including CEBRA, Keypoint-MoSeq, and packages implementing alternative clustering or dimensionality-reduction approaches.

      We have clarified in the documentation that the included UMAP/HDBSCAN workflow is intended as an exploratory demonstration rather than as a required or privileged analysis pipeline.

      Finally, what specific GPUs have been tested with the package, and are there any limitations based on the age of the video card or the available libraries for the deep learning component of the package?

      As noted above, GPU compatibility is determined by the deep-learning and hardware-acceleration libraries on which SqueakPose Studio depends, including PyTorch, Ultralytics, CUDA, Apple MPS, and ROCm. Our development ethos is to track current stable versions of these packages rather than maintain separate legacy dependency stacks. This improves performance, simplifies support, and allows users to benefit from ongoing improvements in upstream libraries, but it also means that older GPU architectures may lose support as they are deprecated by those upstream tools.

      For NVIDIA systems, the current package is indexed against CUDA 13.2. CUDA 13.x has deprecated support for some older GPU architectures, so users with older NVIDIA cards may need to use CPU inference or upgrade hardware. However, CUDA 13 is supported on GeForce RTX 20-series, 30-series, 40-series, 50-series, and professional equivalents. We made this clearer in the documentation and provided links to upstream CUDA, PyTorch, and Ultralytics compatibility resources so users can determine whether their hardware is supported.

      For Apple Silicon, the package can use PyTorch MPS acceleration, which supports M-series chips. For AMD GPUs, we do not currently maintain AMD-specific test hardware, but PyTorch supports ROCm on Linux for supported AMD GPUs. ROCm support is more limited on Windows, so AMD users should consult the current PyTorch ROCm compatibility documentation.

      Overall, our support commitment is to maintain compatibility with current upstream deep-learning frameworks rather than to guarantee support for all older or vendor-specific GPU configurations.

      Reviewer #2 (Public review):

      Summary:

      This work presents three tools: SqueakPose Studio, which is used for pose estimation; SqueakView, which is used for real-time video and sensor data capture and analysis; and MouseHouse, which is a behavioral and sensor suite for mouse experiments. Together, these tools provide a comprehensive behavioral platform for acquiring and analyzing video, sensor, and behavioral data. The work is open source and provided as a resource for the field.

      Strengths:

      (1) Squeakpose Studio was relatively easy to install and use. We were impressed that we were able to install it and test our own videos with minimal struggles. The authors provide installation tutorial videos that were very helpful.

      (2) The GUI environment for SqueakPose Studio was very usable, and the authors should be commended on the time and effort that went into improving the useability of their system. The keypoint and skeleton configuration was flexible, allowing us to define custom body part sets without modifying code directly. The pose estimation accuracy on our own videos was good right out of the box, without requiring fine-tuning or retraining. For a tool being evaluated for the first time, this was all very impressive!

      We thank this reviewer for the positive comments and have provided responses to the specific potential weaknesses noted below.

      Weaknesses:

      (1) While we were able to install and test Squeakpose Studio, it was not entirely seamless. The primary installation resource is a tutorial video, and we would recommend supplementing this with a written installation checklist that explicitly lists all required software dependencies (e.g. Python, UV, Visual Studio). The tutorial video was also at times unclear in distinguishing required from optional components. For example, Visual Studio is described as not necessary, yet the tutorial demonstrates the workflow entirely within that environment, so it may be challenging for a user to follow along without that. We recommend that the authors adopt a stricter, step-by-step installation guide that is prescriptive about required software and leaves little room for confusion.

      We thank the reviewer for this helpful feedback and agree that the installation workflow should distinguish more clearly between required and optional components. Our goal with SqueakPose Studio is to place as much functionality as possible in the GUI so that users are not required to rely on command-line tools for additional features or advanced use. For that reason, the command-line surface is intentionally minimal: after the repository is cloned and the UV-managed environment is created, almost all functionality is accessed through the graphical interface.

      We also appreciate the opportunity to clarify the point about Visual Studio. The tutorial video demonstrates the workflow using Visual Studio Code, not Visual Studio. Visual Studio Code is optional and is used in the video only as a convenient editor and interface for demonstrating the workflow. The GUI can also be launched directly from a terminal, and users may use any preferred editor or IDE, including VS Code, Zed, Cursor, Jupyter-based workflows, or no IDE at all.

      We have updated the written README and YouTube walkthrough to make this distinction clearer. Specifically, provided a stricter installation checklist that separates required components, such as Python and UV, from optional tools, such as VS Code or other editors. We also demonstrated launching SqueakPose Studio directly from a terminal so users can follow the workflow without relying on a specific IDE.

      (2) The paper also describes SqueakView and MouseHouse. Unfortunately, we were unable to evaluate these components as both require the MouseHouse hardware platform. Even without directly using MouseHouse, we noticed some incompleteness here, as we could not locate a bill of materials, component pricing, or assembly guide in the paper or associated GitHub repositories. Given that affordability and accessibility are central claims, a consolidated parts list, approximate costs, and a build guide or video would be necessary for most labs to realistically decide whether they plan to replicate the hardware and evaluate this functionality that the paper describes. In this regard, we felt that MouseHouse and potentially SqueakView were not sufficiently documented for publication.

      We agree with the reviewer that MouseHouse and SqueakView are more difficult to evaluate than SqueakPose Studio because they involve dedicated hardware, including an edge-compute platform. This is an unavoidable tradeoff for a system designed not only for offline pose estimation, but also for real-time acquisition and deployment. We recognize, however, that if the manuscript emphasizes affordability and accessibility, then users need a clear way to estimate cost, order components, assemble the system, and reproduce the hardware configuration.

      We have therefore added a consolidated bill of materials to the GitHub repository, including component names, approximate pricing, and suggested sources where appropriate. We now provide a complete guide for connecting the hardware and flashing the required firmware/software to the devices. This documentation makes clearer what is required for MouseHouse-specific functionality versus what can be used independently through SqueakPose Studio.

      We also note that edge-compute devices such as the Jetson Orin Nano are increasingly common in robotics and real-time computer-vision applications, but we appreciate that many behavioral neuroscience laboratories may not yet have this hardware in place. For some users, this paper may be their first exposure to this compute platform. For that reason, we agree that the repository should provide more complete onboarding materials for labs that wish to adopt the hardware ecosystem, and we now provide that.

      (3) The benchmarking comparison to DeepLabCut (DLC) introduced multiple challenges that left us unclear if the head-to-head comparison was appropriate as described. First, the dataset used for benchmarking was small and homogeneous, from the methods they used "10 min open-field tasks of single mice with bilateral photometry cables." As such, the claims about comparisons between SqueakPose Studio and DLC may be too broad, given this single test case. Specifically, this dataset does not test robustness across lighting conditions, coat colors, species, occlusions, different-shaped arenas, etc. Second, the comparison to DLC in Figure 1 does not include any quantitative statistical comparisons, which are needed to evaluate the claims that were made. For instance, the error in Figure 1e looks worse for their system than DLC, although statistical comparisons were not made. Third, there are many settings and optimizations that can be made for both systems. Without more detail, this makes it hard to know if the head-to-head comparison is really fair. Fourth - the metrics are given as very specific numbers from single runs, i.e., an inference time of 71.59 minutes in Figure 1d. This metric would be more meaningful if it reported the mean of multiple runs, with error estimation. Finally, while the code is available, the trained datasets are made available only on "reasonable request". Given the importance of these datasets to evaluating the method and allowing others to benchmark it against other systems, these should be made available on GitHub. Overall, I would recommend toning down the comparison to DLC and focusing on the strengths of Squeakpose Studio on its own merits.

      We appreciate the reviewer’s thoughtful comments about the benchmarking comparison. We agree that no single dataset can establish universal performance across all lighting conditions, coat colors, species, occlusion regimes, arena geometries, or camera configurations. Our intention was not to claim that SqueakPose Studio is superior to DeepLabCut under every possible condition, nor to present a comprehensive benchmark across the full space of pose-estimation use cases. Rather, the benchmark was included as an applied demonstration of performance in a representative behavioral neuroscience workflow involving mouse open-field videos with photometry cables.

      We also agree that users can substantially affect performance in any pose-estimation framework through model selection, training settings, hardware configuration, inference parameters, and optimization choices. For this reason, we view the comparison as a practical workflow benchmark rather than a definitive ranking of all possible DLC and SqueakPose Studio configurations. The primary contribution of SqueakPose Studio is not simply that it is faster in one head-to-head comparison, but that it provides an integrated GUI-based workflow for pose estimation, review, export, and real-time/edge-AI deployment.

      That said, the speed improvements are not incidental. They reflect deliberate architectural and deployment choices, including the use of modern object-detection/pose-estimation architectures and optimized inference workflows. In practice, these choices can substantially reduce inference time relative to workflows that were not designed around the same deployment constraints. We will be careful in our public response and documentation not to overstate this as a universal claim across every dataset or every possible DLC configuration.

      Regarding statistical comparisons and repeated runs, we agree that reporting means and variance across repeated benchmark runs can be useful. However, because this manuscript is primarily an applications and methods resource rather than a large-scale benchmarking study, we do not intend to benchmark every relevant dataset class or hardware configuration. We instead encourage users to evaluate SqueakPose Studio on their own videos and hardware, which is ultimately the most informative test for adoption in a given laboratory.

      Regarding the trained datasets and models, we agree with the reviewer that broad access improves reproducibility and benchmarking. The limitation is practical rather than philosophical: the full benchmark datasets are large and are not well suited for direct hosting in a GitHub repository. We currently make these data available upon reasonable request and have included a Zenodo repo explore more appropriate public hosting options for large files, such as an institutional repository, Zenodo, OSF, or another archival data platform. We will also clarify the availability of trained models and example data so users can more easily reproduce or extend the benchmarking workflow.

      Overall, we agree that SqueakPose Studio is strongest when evaluated on its own merits: accessibility, speed, GUI-based usability, flexible keypoint configuration, real-time deployment, and integration with acquisition and edge-compute workflows. We now frame the DLC comparison as a representative applied benchmark rather than as an exhaustive claim of general superiority.

      (4) The paper at times makes general statements that are beyond what is shown. For instance, discussions of use in human applications are aspirational and should be treated much more conservatively in the discussion, or possibly even removed. As it stands, the discussion implies that this system can already do "zero-shot tracking of human posture and movement", enabling "a bridge between preclinical and clinical behavioral analysis". In principle, this may be true, but even for a Discussion section, this goes far beyond the capabilities that the paper actually shows.

      We appreciate this comment and agree that the manuscript should distinguish more clearly between capabilities demonstrated in the present study and broader potential applications of the software architecture.

      SqueakPose Studio and SqueakView are not intrinsically mouse-specific. Users can define custom classes, keypoints, and skeletons, train compatible pose-estimation models for other organisms or experimental preparations, and deploy those models using the same acquisition and inference workflow.

      To make this technical capability concrete, the SqueakView repository now includes deployment-ready FP16 model packages for both the validated MouseHouse-specific pose model and a stock human-pose model. The included human-pose model demonstrates that the deployment architecture can support zero-shot human posture tracking without requiring changes to the underlying SqueakView pipeline.

      We agree, however, that this technical compatibility should not be interpreted as validation for clinical behavioral analysis. The experimental demonstrations in the present manuscript focus primarily on mouse behavioral datasets. Any clinical application would require separate benchmarking, validation, and domain-specific evaluation beyond the scope of the present manuscript.

      (5) While the comprehensive nature of the system and its 3 parts is impressive, I felt that it also detracted from the main focus of the paper, which was Squeakpose Studio. I might recommend dropping the other two parts, as they also require a much higher bar for a user to evaluate, and only present the Squeakpose Studio in this paper, presenting this as a general resource for pose estimation. This would also allow them more space to more comprehensively benchmark SqueakPose Studio.

      We appreciate this perspective and agree that SqueakPose Studio is the most immediately accessible component of the platform for many users. However, we respectfully disagree that MouseHouse and SqueakView should be removed from the paper. The motivation for developing SqueakPose Studio was not simply to create another offline pose-estimation and analysis tool, but to enable real-time behavioral detection and deployment on edge hardware. SqueakView and MouseHouse provide the acquisition and deployment context that motivated the software architecture and demonstrate how the platform can be used in closed-loop or real-time behavioral workflows.

      In developing the system, we recognized that SqueakPose Studio also functions as a user-friendly general pose-estimation interface, with features that may be useful even for laboratories that do not adopt the full MouseHouse/SqueakView ecosystem. For that reason, we presented it as both a standalone tool and as part of a broader acquisition and deployment platform.

      We agree that this makes the manuscript broader than a paper focused exclusively on pose-estimation benchmarking. However, we view that breadth as important: the paper is intended to serve as a central, peer-reviewed entry point for laboratories interested in deploying real-time pose estimation in behavioral experiments. The manuscript points users to the relevant repositories, documents the design rationale, and provides a source of peer-reviewed validation for the integrated workflow. We have clarified in our response and documentation that users can adopt SqueakPose Studio independently, while MouseHouse and SqueakView support the broader real-time hardware ecosystem.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, the role of the insulin receptor and the insulin growth factor receptor was investigated in podocytes. Mice, where both receptors were deleted, developed glomerular dysfunction and developed proteinuria and glomerulrosclerosis over several months. Because of concerns about incomplete KO, the authors generated and studied podocyte cell lines where both receptors were deleted. Loss of both receptors was highly deleterious with greater than 50% cell death. To elucidate the mechanism of cell death, the authors performed global proteomics and found that spliceosome proteins were downregulated. They confirmed this directly by using long-read sequencing. These results suggest a novel role for insulin and IGF1R signaling in RNA splicing in podocytes.

      This is primarily a descriptive study and no technical concerns are raised. The mechanism of how insulin and IGF1 signaling regulates splicing is not directly addressed but implicates potentially the phosphorylation downstream of these receptors. In the revised manuscript, it is shown that the mouse KO is incomplete potentially explaining the slow onset of renal insufficiency. Direct measurement of GFR and serial serum creatinines might also enhance our understanding of progression of disease, proteinuria is a strong sign of renal injury. An attempt to rescue the phenotype by overexpression of SF3B4 would also be useful but may be masked by defects in other spliceosome genes. As insulin and IGF are regulators of metabolism, some assessment of metabolic parameters would be an optional add-on.

      Significance:

      With the GLP1 agonists providing renal protection, there is great interest in understanding the role of insulin and other incretins in kidney cell biology. It is already known that Insulin and IGFR signaling play important roles in other cells of the kidney. So, there is great interest in understanding these pathways in podocytes. The major advance is that these two pathways appear to have a role in RNA metabolism.

      Latest comments:

      The new reviewer raised two major points, whether the KO effect on splicing is specific to IGF1 and whether the interpretation could be developmental rather than due to splicing. The reviewer raises some important issues but the evidence to suggest that this is specific is data in the literature that IR/IGF signaling is already known to regulate splicing and that splicing defects were not detected in other models that they have analyzed. I agree with the reviewer (and authors) that the incomplete floxing of the genes is a major complication. The point that there could be a developmental defect with mice being born with fewer podocytes and perhaps the authors should caveat this point. The fact that they mice are born with normal function, that renal function can be maintained with up to 80% loss of podocytes suggest that they are likely born with a good number of podocytes and the dysfunction that occurs at 6 months is due to a process, induced by the loss of IR/IGF signaling that is detrimental to the podocyte.

      Thank you for these insightful comments. We fully acknowledge that the mouse model will not have had full insulin receptor and IGF1R knockdown and that this is likely the reason it took time to develop and not give a prominent early phenotype. We agree with this reviewer and new reviewer 4 that if the model had facilitated near complete IR and IGF1R knockdown then likely a significant neonatal / embryonic phenotype would have been obvious. We considered using an inducible mouse model to allow normal development before cre-excision but our experience is that the CreER and RtTA-tet-on-cre system is less good at excising genes and hence did not pursue this (we show evidence of reduced excision with an inducible system in supplementary Figure 1D using a reporter mouse system [this was included in a previous response to the reviewers only]). This was rationale for making the immortalised podocyte floxed IR and IGF1R cell line to ensure near complete knockdown. This, not surprisingly, was highly detrimental. We then looked mechanistically for pathways (using agnostic proteomics and phospho-proteomics) and found spliceosomal involvement. From our studies we think this was also involved in our mouse model as SF3B4 was found to be significantly down regulated in the podocytes of double receptor knockdown transgenic mice (Figure 3F).

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, submitted to Review Commons (journal agnostic), Coward and colleagues report on the role of insulin/IGF axis in podocyte gene transcription. They knocked out both the insulin and IGFR1 mice. Dual KO mice manifested a severe phenotype, with albuminuria, glomerulosclerosis, renal failure and death at 4-24 weeks.

      Long read RNA sequencing was used to assess splicing events. Podocyte transcripts manifesting intron retention were identified. Dual knock-out podocytes manifested more transcripts with intron retention (18%) compared wild-type controls (18%), with an overlap between experiments of ~30%.

      Transcript productivity was also assessed using FLAIR-mark-intron-retention software. Intron retention w seen in 18% of ciDKO podocyte transcripts compared to 14% of wild-type podocyte transcripts (P=0.004), with an overlap between experiments of ~30% (indicating the variability of results with this method). Interestingly, ciDKO podocytes showed downregulation of proteins involved in spliceosome function and RNA processing, as suggested by LC/MS and confirmed by Western blot.

      Pladienolide (a spliceosome inhibitor) was cytotoxic to HeLa cells and to mouse podocytes but no toxicity was seen in murine glomerular endothelial cells.

      The manuscript is generally clear and well-written. Mouse work was approved in advance. The four figures are generally well-designed, bars/superimposed dot-plots.

      Methods are generally well described.

      Comments on previous version:

      Coward and colleagues have done an excellent job of responding to all the reviewer comments.

      Thank you.

      Reviewer #4 (Public review):

      Summary and background:

      This report entitled "The insulin/IGF axis is critically important (for) controlling gene transcription in the podocyte" from Hurcombe et al is based on a mouse double knockdown of the IR and IGF1R and a parallel cultured mouse podocyte model. Insulin/IGF signaling system in mammals evolved as three gene reduplicated peptides (insulin, IGF-1, and IGF-2) and their two receptors IR and IGF1R that cross-react to variable extents with the peptides, are ubiquitously expressed, and signal through parallel pathways. The major downstream effect of insulin is to regulate glucose uptake and metabolism, while that of the IGF pathways is to regulate growth and cell cycling in part through mTORC1. The GH-IGF-1-IGF1R pathway regulates post-natal growth. IGF-2 signaling is thought to play a major role in regulating intrauterine growth and development, although IGF-2 is also present at high levels in post-natal life. Thus, one would anticipate that reducing IR/IGF1R signaling in any cell would slow growth and cell cycling by reducing growth factor and metabolic mTORC1-mediated and other processes including the splicing of RNA for protein synthesis.

      Thank you for this clear overview. Of note the podocyte is a terminally differentiated cell so the growth / cell cycling elements may be different from more proliferative cell types in relation IR/IGF1R mediated signalling.

      Comments on revised version:

      The second sentence of the Summary reads "This study sought to elucidate the compound role of the insulin/IGF1 axis in podocytes using transgenic mice and cell culture models deficient in both receptors." The study design and rationale for the proteosome analysis described is predicated on the finding that podocyte-specific knockdown of the IR/IGF-1R in mice is associated with development of proteinuria and reduced eGFR by 20months of life. Since the IR/IGF-1R are critically required for normal development and growth of all cells and organs, the obvious explanation for the observation would be that the model system results in defective podocyte development and deployment (caused by reduced IR/IGF-1) that, in turn, causes subsequent development of proteinuria and glomerulosclerosis (that may be much less dependent on a normal level of IR/IGF-1R expression). Thus, the experimental design does not allow a distinction between podocyte development and steady state function which are different biologic processes. The data provided does not examine podocyte status immediately after birth to confirm that podocyte number and size and structure is normal in mice that subsequently develop proteinuria and glomerulosclerosis. The response to the reviewer suggests that since this would require additional mice it has not been undertaken in order to reduce animal usage. This is not a valid argument, particularly when the investigators have not even used state-of-the-art methods to measure podocyte number, size and density in adult mice, key parameters that would be required to interpret their data. Counting podocyte nuclear number in glomerular cross-sections is simply an inadequate method, even if it is used and reported in other journals, and particularly where the examples given to justify its use can hardly be viewed as representing first rate science.

      Thank you for these comments. As discussed above we agree that the mouse model was not optimal as despite using a good cre driver we did not consistently knock down both receptors. It was the reason that we made the IR/IGF1R knockdown cell line. Importantly we found with both receptors >80% knocked down that this was highly detrimental and evidence that spliceosomal dysfunction was prominent. Thank you for the comment about methodology of assessing podocyte number which we and other investigators use.

      If the absence of studies that would answer the above questions, the investigators should add a sentence to the Discussion dealing with study limitations as follows. "The study design does not allow us to determine whether the primary effect of reduced IR/IGF-1R expression on the phenotype is during in utero and post-natal podocyte development and deployment, during periods of rapid growth when IGF-1 levels are highest, in steady state adult podocytes, or under all of the above conditions".

      Thank you. We have added a section describing that we did not investigate the embryonic neonatal early phenotype for more subtle changes in our model. We have also added a sentence saying we would have liked to have used an inducible model but the cre driven excision is less than constitutional driver and we think would have shown either a very mild or no phenotype due to minimal excision.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      The authors address whether theta/beta ratio /TBR) can be used as a clinical biomarker for ADHD.

      Strengths:

      The data were acquired independently from 2 separate datasets, and there are sufficient subjects for adequate statistical power. The authors applied up-to-date EEG data preprocessing, state-of-the-art feature extraction, and statistical analyses, using a multiverse approach. By testing and comparing all meaningful approaches, defined a priori in the previous meta-analysis, the author convincingly demonstrates that TBR cannot be used as a clinical biomarker, and previous positive results can be explained by interactions between different factors (alpha peak frequency, aperiodic component, age).

      Weaknesses:

      There are no apparent issues with data, separate datasets, large sample sizes, and state-of-the-art data analysis.

      We thank Reviewer #1 for their positive evaluation of our manuscript and for the constructive recommendations. The reviewer did not raise additional comments requiring a point-by-point response beyond the recommendations addressed below.

      Reviewer #2 (Public review):

      Summary:

      This manuscript examines whether the theta-beta ratio as derived from EEG data relates to ADHD diagnoses. To do so, it performs a multiverse analysis across a large number of analytical choices, applied to a large EEG dataset, and corroborated in an additional validation set. The results overall show that the TBR is not a reliable indicator of ADHD diagnosis. In discussing the patterns of results across analytical choices, the authors also demonstrate some key points about what appears to be driving the ratio measures, noting that significant results appear to be driven by choices regarding aperiodic-correction and the use of individualized alpha frequencies, suggesting TBR measures can be affected by these features rather than reflecting theta and/or beta activity.

      Strengths:

      This manuscript addresses a clearly posed and important question in the literature, addressing a longstanding discussion on the relationship between TBR and ADHD, and uses a large dataset and an expansive analysis approach to provide a definitive answer. The strengths of the approach allow for a clear answer, providing a notable contribution to the field.

      Weaknesses:

      I find no notable weaknesses in the current manuscript nor any major issues that I think challenge the key findings of this manuscript.

      We thank Reviewer #2 for their positive evaluation of our manuscript and for the constructive recommendations. The reviewer did not raise additional comments requiring a point-by-point response beyond the recommendations addressed below.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Strzelczyk, Vetsch, and Langer tackle an incredibly important question in clinical neuroscience: the use of the theta/beta ratio as a biomarker of attention deficit hyperactivity disorder (ADHD). The theta/beta ratio is argued to be so reliable as an ADHD biomarker that, in the United States, the Food and Drug Administration has approved its use as a biomarker for ADHD diagnosis. However, there is mounting evidence that the theta/beta ratio is likely not really measuring the relative power between two oscillations - the theta rhythm and the beta rhythm - but rather reflects differences in a singular, non-oscillatory aperiodic process. In this very convincing study, Strzelczyk and colleagues take a "multiverse" analysis approach to show that aperiodic activity differences between healthy controls and people with ADHD are driving the apparent theta/beta ratio differences. While in a vacuum, where a measure is a measure and if it's related to a diagnosis it's still useful no matter what, this distinction might not seem important, from a neuroscientific perspective this is a critical distinction, because the ratio between two oscillations has fundamentally very different underlying physiological mechanisms than aperiodic differences, and this framing has a major impact on guiding research on the diagnosis and treatment of ADHD.

      Strengths:

      While smaller studies and analyses have already hinted at similar results as shown here, the current study's multiverse analysis approach is comprehensive, convincing, and very well done. The large sample size of 1,499 participants is very impressive, as is the use of an independent validation sample of 381 participants.

      Overall, the technical and statistical aspects are very well done: the multiverse approach, the validation set, the resampling methods, and even the shiny apps. The authors should be applauded for being so thorough and making their data and analyses publicly accessible.

      Weaknesses:

      To be clear, I see no breaking weaknesses in the theoretical foundations, methods, statistical analyses, or interpretations. All of my recommendations below are for the sake of clarity, which I believe is especially important because this is such an important paper that many people should read.

      Comments:

      (1) Some figures are mislabeled. For example, Supplementary Figure 1 says (C) are scalp topographies, but those are (A), while (C) shows power spectra, but it's unclear what (C) is. I assume it's only the aperiodic part of the spectrum (oscillations removed)? But it would be better to plot on a log-log scale if so. In fact, I recommend showing all spectra on a log-log scale.

      The reviewer is correct that the figure legend was mislabeled. Panel (A) shows the scalp topographies, panel (B) shows the 1/f-uncorrected power spectra, and panel (C) shows the reconstructed aperiodic signal with oscillations removed. We have corrected the figure legend accordingly. In addition, the power spectra and the reconstructed aperiodic signal are now plotted on log-log scales to improve readability and interpretability.

      (2) Supplementary Figure 6 is also mislabeled, saying (A) shows age (it does not) and so on.

      We thank the reviewer for noticing this error. We have revised the figure legend so that the panel descriptions now match the displayed plots.

      (3) In Supplementary Figure 7, is (B) the aperiodic-removed spectrum? The authors are very inconsistent with what they're showing in these spectral plots, and not actually explaining what they're showing: raw spectra, semi-logged or not, aperiodic-removed or oscillations-removed, etc.

      Panel (B) in Supplementary Figure 7 shows the aperiodic-adjusted spectrum. We have now corrected the figure labeling and revised the figure legend to explicitly state what is shown in each panel.

      (4) For the HBN data, it is said that, "electrode impedances were kept below 40 kΩ, lower than EGI's standard recommendation of 50 (Net Station Acquisition Technical Manual)." For the validation data: "... electrode impedances were maintained below 5 kΩ." These are big impedance threshold differences. Of course, these recommendations differ by recording system, the use of active electrodes, and so on. But such differences can certainly influence signal-to-noise. The fact that the results are so consistent between them is a strength that perhaps should be explicitly called out.

      We appreciate the reviewer’s suggestion. We now explicitly state in the discussion section that the consistency of the results across datasets with different EEG systems and impedance thresholds strengthens the generalizability of the findings. The revised text reads as follows:

      “Our multiverse results thus converge with this broader literature, providing further evidence that TBR lacks the reliability and discriminative validity required for clinical utility. Beyond methodological convergence across analytical frameworks, the consistency of results across two datasets differing substantially in EEG recording systems and impedance thresholds further strengthens the generalizability of these null findings, suggesting they are unlikely to reflect idiosyncrasies of a specific acquisition protocol.”

      (5) The authors cite a lot of foundational / related work here, such as Finley et al, but they should also cite several other highly relevant ones:

      Saad et al., "Is the Theta/Beta EEG Marker for ADHD Inherently Flawed?", J Atten Disord, 2015

      Donoghue, Dominguez, Voytek, "Electrophysiological frequency band ratio measures conflate periodic and aperiodic neural activity", eNeuro, 2020

      Karalunas et al., "Electroencephalogram aperiodic power spectral slope can be reliably measured and predicts ADHD risk in early development", Develop Psychobiol, 2022

      Donoghue, "A systematic review of aperiodic neural activity in clinical investigations", Eur J

      Neurosci 2025

      We thank the reviewer for pointing us to these additional relevant references. We have added the suggested references to the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) "Multiverse analysis was conducted in RStudio (R version 4.4.1) using the multiverse package (version 0.6.1; Sarma et al., 2021). T" Ok, cool, but it would be useful to explain what it does compared to running the standard stat analysis N times.

      We thank the reviewer for this helpful recommendation. We have now expanded the Methods section to clarify this point. The revised text reads as follows:

      “Multiverse analysis was conducted in RStudio (R version 4.4.1) using the multiverse package (version 0.6.1; Sarma et al., 2021). The multiverse framework differs from simply repeating the same statistical analysis multiple times, because it first requires the researcher to define a structured analysis space consisting of multiple defensible analytic decisions. These decisions are then expanded into all valid combinations, with each combination representing one complete analysis specification, or “universe”, providing a transparent and reproducible record of which analytic decisions were considered and how they were combined. In addition, the package reduces the need to manually write, modify, and track separate analysis scripts for each specification, which helps avoid inconsistencies or coding errors across universes. The results can then be extracted and summarized across the full set of universes to evaluate whether the conclusions are robust across reasonable analytic alternatives or depend on specific combinations of choices.”

      (2) I may have missed it, but how many subjects per group do you end up with after all the cleaning (not what is in Table 1, but like in each dataset you describe how many got removed at each step, so we are left wondering the final numbers).

      We thank the reviewer for pointing this out. The final group sizes after all cleaning and exclusion steps were not described in the original manuscript. We have therefore revised Table 1 so that it now reports only the remaining participants included in the final analyses after all exclusions were applied. The revised table shows the final sample sizes separately for the HC (N = 228), ADHD-Combined (N = 429), and ADHD-Inattentive (n = 465) groups, together with the corresponding demographic and clinical characteristics. We have also revised the accompanying text in the result section 3. 1. 1. The same changes were applied to the validation sample, which is reported in the Supplement.

      (3) Missing reference in my opinion. In the discussion, the sentence "as both oscillatory and aperiodic contributions vary systematically across the lifespan" could do with a reference or two about that

      We have now added references showing that developmental changes in EEG spectra involve both periodic/oscillatory and aperiodic components. The revised text reads as follows:

      “These dynamics may account for the recurring Age ’ IAF interactions observed in our multiverse analyses, as both oscillatory and aperiodic contributions vary systematically across the lifespan (Merkin et al. 2023; Tröndle et al. 2022; Tröndle et al. 2021; McSweeney et al. 2023; Hill et al. 2022; Stanyard et al. 2024).”

      (4) Now the big one: this is a cool visualization, and beta estimates from linear modeling do tell us the strength, BUT I would like to see raw effect sizes. It could be in a table or text, to go with the discussion. What was the theta, alpha, beta power raw or adjusted in each group, what about the aperiodic component - even maybe some violin plots to show canonical vs individual - my point is I am convinced from the frequency analysis since an entire subspace become significant and your interpretation that this is spurious is satisfactory but showing that this subspace as tiny effect sizes driven by interactions would be even more convincing in my opinion.

      To complement the regression coefficients from the multiverse models, we now additionally report descriptive standardized effect sizes across representative analytical subspaces. Specifically, we grouped analytical paths according to frequency band definition (IAF-relative vs canonical) and spectral representation (aperiodic signal, 1/f-uncorrected power, and aperiodic-adjusted power). Within each subspace, we computed Cohen’s d values for theta power, beta power, and TBR between ADHD and healthy control groups across all corresponding analytical paths.

      To visualize the distribution of effects across analytical paths, we added violin plots with overlaid individual paths and mean effect sizes with 95% confidence intervals. Importantly, even in subspaces where interaction effects frequently emerged in the multiverse analysis, the corresponding descriptive group differences remained small, supporting our interpretation that the observed significant effects are driven by subtle interactions and analytical choices rather than large underlying group differences.

      The added text in the Results 3. 1. 4. reads as follows:

      “To complement the regression coefficients from the multiverse models, we additionally examined descriptive standardized effect sizes across representative analytical subspaces. Analytical paths were grouped according to frequency band definition (IAF-relative vs. canonical) and spectral representation (aperiodic signal, 1/f-uncorrected power, and aperiodic-adjusted power). Within each subspace, Cohen's d was computed for theta power, beta power, and TBR for both the HC vs. ADHD-Inattentive and HC vs. ADHD-Combined comparisons. To visualize the distribution of effect sizes across the analytical space, violin plots were constructed with each data point representing the Cohen's d value of a single analytical specification (Figure 8). Across all subspaces and outcome measures, Cohen's d values were small for both comparisons, including subspaces in which interaction effects frequently reached statistical significance in the multiverse analysis. This pattern indicates that even where the multiverse revealed reliable significant effects, the underlying group differences in theta power, beta power, and TBR remained small in magnitude. These findings support the interpretation that the significant interactions observed across analytical specifications are driven by subtle moderation effects and analytical choices rather than large, robust group differences in neural activity.”

      Reviewer #2 (Recommendations for the authors):

      (1) As a minor clarification, the manuscript could specify if the calculation of aperiodic-adjusted power values was done as subtraction with linear or log power values.

      The aperiodic-adjusted power values were computed by subtracting the aperiodic fit from the observed power spectrum in log10 power space. Specifically, both the observed power spectrum and the estimated aperiodic component were log10-transformed, and the aperiodic-adjusted signal was obtained as the difference between these two quantities. The result was then transformed back to linear scale. We have clarified this in the revised manuscript. The revised text reads as follows:

      “The aperiodic component was reconstructed based on its fitted parameters and subtracted from the total power spectrum in log10 power space, resulting in an aperiodic-adjusted, 1/f-corrected power spectrum. The resulting values were then transformed back to linear scale and therefore represent power relative to the estimated aperiodic background.”

      (2) The last section of the abstract is a bit repetitive in stating the main finding of what drives the TBR, and this could be edited/condensed.

      We agree that the final part of the abstract repeated the main interpretation regarding the role of aperiodic activity and IAF. We have therefore condensed this section to avoid redundancy while preserving the central conclusion. The revised text reads as follows:

      Across the multiverse, we found that group differences in TBR were highly contingent on analytical choices, with no evidence for robust main effects of diagnosis, indicating no reliable differences between healthy controls, ADHD-inattentive, and ADHD-combined subtypes. Instead, significant effects emerged primarily as interactions with age and individual alpha frequency (IAF), particularly when TBR was derived from aperiodic-uncorrected power or from the aperiodic signal itself. These interaction patterns replicated across both independent samples and were observed using both categorical and dimensional definitions of ADHD. Together, these findings indicate that previously reported TBR effects are largely driven by variability in aperiodic activity and IAF rather than genuine differences in oscillatory theta-beta dynamics. Our results challenge the interpretation of TBR as a reliable standalone biomarker for ADHD and underscore the importance of multiverse approaches for evaluating candidate neurobiological markers in heterogeneous clinical populations.

      (3) As a minor literature note, the finding that ratio measures often largely reflect aperiodic activity rather than oscillatory theta and/or beta per se activity is consistent with a previous (non-clinical) investigation of band ratio measures in a previous report that should perhaps be cited as relevant prior work:

      Donoghue, T., Dominguez, J., & Voytek, B. (2020). Electrophysiological Frequency Band Ratio Measures Conflate Periodic and Aperiodic Neural Activity. eNeuro, 7(6),ENEURO.0192-20.2020. https://doi.org/10.1523/ENEURO.0192-20.2020

      We appreciate the reviewer’s suggestion. We have added this reference to the Discussion section, where we interpret the observed TBR effects as reflecting variability in the aperiodic background rather than genuine differences in oscillatory theta-beta dynamics. The revised text reads as follows:

      “These results suggest that apparent TBR differences may reflect properties of the aperiodic background signal interacting with individual variability in IAF rather than true oscillatory theta or beta activity. This interpretation is consistent with previous work showing that electrophysiological frequency-band ratio measures can conflate periodic and aperiodic neural activity, such that apparent changes in theta/beta or other band ratios may partly reflect changes in the aperiodic spectral component rather than narrowband oscillatory activity (Donoghue et al., 2020).”

      (4) In Figure 3, it may be useful to highlight the theta and beta ranges in panel B.

      We considered highlighting the theta and beta ranges in Figure 3B, but decided against it. In the multiverse analysis, theta and beta were defined using both canonical frequency bands and IAF-relative bands. The IAF-relative bands differ across participants, therefore marking only the canonical ranges could give the impression that these were the only frequency definitions used in the analyses. We therefore kept the spectra unmarked.

      (5) In Figure 5 (and other figures following this motif), it may be useful to color the significant results as green or red based on direction, to match Figure 4.

      We have updated Figure 5 and the corresponding figures so that significant positive effects are shown in green and significant negative effects are shown in red, matching the color scheme used in Figure 4.

      Reviewer #3 (Recommendations for the authors):

      (1) P10, L30: "Individualized bands were centered on the IAF, defined as theta = IAF-6 Hz to IAF-4 Hz"; why is theta defined using such a narrow, 2 Hz band here, when canonical theta is usually defined as a 4 Hz wide, 4-8 Hz band?

      The individualized theta band was chosen to follow the IAF-based frequency-band framework proposed by the seminal work of Wolfgang Klimesch (1999, 2012), rather than to reproduce the width of the canonical 4-8 Hz theta band. In this framework, frequency bands are defined relative to each participant’s individual alpha frequency. Theta is defined as the range from IAF-6 Hz to IAF-4 Hz, while lower alpha occupies the range closer to the individual alpha peak. The narrower individualized theta band is therefore intended to reduce overlap with lower-alpha activity and to account for inter-individual and developmental differences in alpha peak frequency. The 2020 guidelines from the International Federation of Clinical Neurophysiology (IFCN) reaffirmed Klimesch’s division of the alpha and theta bands (Babiloni, 2020). We have explained the frequency bands selection in more detail in the manuscript. The revised text reads as follows in Methods 2. 5. 7. Extraction of power for statistical analyses:

      The selection of these frequency bands is grounded in the seminal work of Wolfgang Klimesch (1999), who demonstrated that the alpha band can be divided into distinct lower and upper sub-bands. The lower alpha band extends up to 4 Hz below the IAF, covering a broader range of approximately 3.5 to 4 Hz, while the upper alpha band, which lies above the IAF, is narrower, spanning about 1 to 1.5 Hz. Klimesch also characterized the theta band as a frequency range that is approximately 2 Hz below the lower alpha band (Klimesch, 1999; Klimesch, 2012). The 2020 guidelines from the International Federation of Clinical Neurophysiology (IFCN) reaffirmed Klimesch’s division of the alpha and theta bands (Babiloni, 2020).

      (2) Figure 3 and Supplementary Figure 1, 7, 8: "Electrodes highlighted on the topographies..." means just the text labels, right? It might be better to show all electrodes as black dots and highlight the others with white dots or something.

      We have revised the figures to display all electrodes as black dots. In addition, we have clarified in the figure legends that the highlighted electrode labels correspond to the regions of interest used in the analyses.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript titled," Sleep-Wake Transitions Are Impaired in the AppNL-G-F Mouse Model of Early Onset Alzheimer's Disease", is about a study of sleep/wake phenomena in a knockin mouse strain carrying "three mutations in the human App gene associated with elevated risk for early onset AD". Traditional, in-depth characterization of sleep/wake states, EEG parameters, and response to sleep loss are employed to provide evidence, "supporting the use of this strain as a model to investigate interventions that mitigate AD burden during early disease stages". The sleep/wake findings of earlier studies (especially Maezono et al., 2020, as noted by the authors) were extended by several important, genotype-related observations, including age-related hyperactivity onset that is typically associated with increased arousal, a normal response to loss of sleep and to multiple sleep latency testing, and a stronger AD-like phenotype in females. The authors conclude that the AppNL-G-F mice demonstrate many of the human AD prodromal symptoms and suggest that this strain may serve as a model for prodromal AD in humans, confirming the earlier results and conclusions of Maezono et al. Finally, based on state bout frequency and duration analyses, it is suggested that the AppNL-G-F mice may develop disruptions in mechanism(s) involved in state transition.

      Strengths:

      The study appears to have been, technically, rigorously conducted with high quality, in-depth traditional assessment of both state and EEG characteristics, with the concordant addition of activity and temperature. The major strengths of this study derive from observations that the AppNL-G-F mice: (1) are more hyperactive in association with decreased transitions between states; (2) maintain a normal response to sleep deprivation and have normal MSLT results; and (3) display a sex specific, "stronger" insomnia-like effect of the knockin in females.

      Weaknesses:

      The weaknesses stem from the study's impact being limited due to its being largely confirmatory of the Maezono et al. study, with advances of importance to a potentially more focused field. Further, the authors conclude that AppNL-G-F mice have disrupted mechanism(s) responsible for state transition; however, these were not directly examined. The rationale for this conclusion is stated by the authors as based on the observations that bouts of both W and NREM tend to be longer in duration and decreased in frequency in AppNL-G-F mice. Although altered mechanism(s) of state transition (it is not clear what mechanisms are referenced here) cannot be ruled out, other explanations might be considered. For example, increased arousal in association with hyperactivity would be expected to result in increased duration of W bouts during the active phase. This would also predictably result in greater sleep pressure that is typically associated with more consolidated NREM bouts, consistent with the observations of bout duration and frequency.

      Reviewer 1 succinctly summarizes the advances of this study beyond the ground-breaking Maezono et al (2020) study of this “humanized” mouse model exhibiting amyloid deposition. Whereas Maezono et al. conducted sleep/wake studies on male App<sup>NL-G-F</sup> mice at 6 and 12 months of age, we had the unusual opportunity to study both sexes of homozygous App<sup>NL-G-F</sup> mice and WT littermates at 14-18 months of age and to conduct a longitudinal assessment of many of the same individuals at 18-22 months. In addition to baseline sleep/wake and EEG spectral analyses, we (1) measured subcutaneous body temperature and activity to obtain a broader picture of the physiology and behavior of this strain at advanced ages; (2) assessed baseline sleepiness in this strain using the murine version of the clinically-relevant Multiple Sleep Latency Test (MSLT); (3) evaluated the response of App<sup>NL-G-F</sup> mice and WT littermates to a 6-h perturbation of the sleep homeostat; (4) compared the sleep/wake characteristics of male vs. female App<sup>NL-G-F</sup> mice at 18-22 months; and (5) to assess the stability of the phenotypes, analyzed these data over a continuous 14-d recording rather than the conventional 24h recordings typical of most sleep/wake studies including Maezono et al. We found that a long wake/short sleep phenotype was characteristic of homozygous App<sub>NL-G-F</sub> mice at these advanced ages which is also evident in the Maezono et al. (2020) study at 12 months of age (but not at 6 months), although the authors do not comment on this phenotype and instead focus on the reduced REM sleep which is particularly evident in female App<sup>NL-G-F</sup> mice in our study. Remarkably, despite being awake ~20% longer per day, we find that App<sup>NL-G-F</sup> mice are no sleepier than WT mice as determined by the MSLT and that their sleep homeostat is intact when challenged by 6-h sleep deprivation. At both advanced ages, the long wake/short sleep phenotype is due primarily to longer Wake bouts and shorter bouts of both NREM and REM sleep during the dark phase. Moreover, hyperactivity develops in older App<sup>NL-G-F</sup> mice, particularly females, which contributes to this phenotype. We agree with Reviewer 1 that “hyperactivity would be expected to result in increased duration of W bouts during the active phase” and that this could result in more consolidated NREM bouts. Accordingly, we have added the following sentence to the Discussion subsection Impacts of pathology on sleep/wake and activity: “Thus, the hyperactivity evident in Figures 4D, 4D’, and 5D’ could drive the longer wake bouts evident in Figure 7A and result in the longer NREM and REM sleep bouts found in male App<sup>NL-G-F</sup> mice (Figure 12A’ and 12A”).”

      The suggestion of greater sleep pressure is not borne out by our MSLT studies as we did not observe the shorter sleep latencies nor increased sleep during the nap opportunities on the MSLT that we have observed in other mouse strains. Moreover, due to their short sleep phenotype, App<sup>NL-G-F</sup> mice should be entering the sleep deprivation study with a greater sleep debt than WT mice, yet we did not observe a stronger homeostatic response (i.e., enhanced EEG Slow Wave Activity) in this strain during recovery from sleep deprivation. Thus, we have suggested that App<sup>NL-G-F</sup> mice are unable to transition from Wake to sleep as readily as their WT littermates. Our observations summarized above set the stage for subsequent mechanistic studies in aged App<sup>NL-G-F</sup> mice, although realistically, mice of this age and genotype are a rare commodity.

      Reviewer #2 (Public review):

      Summary:

      The authors have used a knock-in mouse model to explore late-in-life amyloid effects on sleep. This is an excellent model as the mutated genes are regulated by the endogenous promoter system. The sleep study techniques and statistical analyses are also first-rate.

      The group finds an age-dependent increase in motor activity in advanced age in the NLGF homozygous knock-in mice (NLGF), with a parallel age-dependent increase in body temperature, both effects predominate in the dark period. Interestingly, the sleep patterns do not quite follow the sleep changes. Wake time is increased in NLGF mice, and there is no progression in increased wake over time. NREMS and REM sleep are both reduced, and there is no progression. Sleep-wake effects, however, show a robust light:dark effect with larger effects in the dark period. These findings support distinct effects of this mutation on activity and temperature and on sleep. This is the first description of the temporal pattern of these effects. NLGF mice show wake stability (longer bout durations in the dark period (their active period) and fewer brief arousals from sleep. Sleep homeostasis across the lights-on period is normal. Wake power spectral density is unaffected in NLGF mice at either age. Only REM power spectra are affected, with NLGF mice showing less theta and more delta. There are interesting sex differences, with females showing no gene difference in wake bout number, while males show a gene effect. Similarly, gene effects on NREM bout number seem larger in males than in females. Although there was no difference in homeostatic response, there was normalization of sleep-wake activity after sleep deprivation.

      Strengths:

      Approach (model extent of sleep phenotyping), analysis.

      Weaknesses:

      The weaknesses are summarized below and are viewed as "addressable".

      (1) The term insomnia. Insomnia is defined as a subjective dissatisfaction with sleep, which cannot be ascertained in a mouse model. The findings across baseline sleep in NLGF mice support increased wake consolidation in the active period. The predominant sleep period (lights on) is largely unaffected, and the active period (lights off) shows increased activity and increased wake with longer bouts. There is a fantastic clue where NLGF effects are consistent with increased hypocretinergic (orexinergic) neuron activity in the dark period, and/or increased drive to hypocretin neurons from PVH.

      Although the DSM-5 definition of Insomnia Disorder indeed emphasizes a subjective “complaint of dissatisfaction with sleep quantity or quality”, I think the Reviewer takes an unnecessarily narrow view of the term “insomnia”. Aside from cases of “psychological” insomnia in which there is a mismatch between subjective and objective measures of sleep, most sleep researchers would likely agree that insomnia is objectively characterized by a greater than normal wake time during the sleep period (i.e., low sleep efficiency) due to difficulty in either initiating or maintaining sleep. This view has led to efforts to identify not only the biological causes of insomnia but also animal models in which this disorder can be studied. A PubMed search on the terms “mouse” and “insomnia” retrieves 844 publications, including an authoritative 2023 review in J Sleep Research entitled "Animal Models of Human Insomnia" co-authored by a clinician-scientist who has done human sleep research throughout his career and is an authority on CBT-I, in particular. Similarly, a PubMed search on the terms “fly” and “insomnia” retrieves 18 publications. So, although our intent in the submitted version of the manuscript was to use “insomnia” as an operational term to succinctly mean “less sleep than usual”, in the revised manuscript, we have eliminated use of the term “partial insomnia” and replaced it with the term “insomnia-like phenotype”. In the Discussion section “Impacts of pathology on sleep/wake and activity”, we have revised the opening sentence to read “Insomnia in humans is typically characterized by subjective reports of reduced sleep quality and can be accompanied by objective measures of sleep fragmentation and reduced sleep amounts.”

      (2) Sleep-wake transitions are impaired: This should not be termed an impairment. It could actually be beneficial to have greater state stability, especially wake stability in the dark or active period. There is reduced sleep in the model that can be normalized by short-term sleep loss. It is fascinating that recovery sleep normalized sleep in the NLGF in the immediate lights-on and light-off period. This is a key finding.

      Due to the Reviewer’s objection regarding “impairment”, we have changed the title of the manuscript to “Long Wake/Short Sleep Bouts and Hyperactivity with Advanced Age in a Mouse Model of Early Onset Alzheimer’s Disease”. In Comments (1) and (2), Reviewer 2 suggests a provocative hypothesis to test. In the section “Impacts of pathology on sleep/wake and activity“, we previously stated “A hyperactive hypocretin/orexin or monoaminergic arousal system or a dysfunctional GABAergic sleep onset system could underlie the longer bouts of Wake in App<sup>NL-G-F</sup>mice.” We have now added this additional sentence: “Indeed, Hcrt neurons in aged mice have been shown to exhibit more frequent neuronal activity driving wake bouts and optogenetic stimulation of Hcrt neurons in aged mice results in prolonged wakefulness (Li et al., 2022).“

      Reviewer #3 (Public review):

      Summary:

      In this study, Tisdale et al. studied the sleep/wake patterns in the biological mouse model of Alzheimer's disease. The results in this study, together with the established literature on the relationship of sleep and Alzheimer's disease progression, guided the authors to propose this mouse model for the mechanistic understanding of sleep states that translates to Alzheimer's disease patients. However, the manuscript currently suffers from a disconnect between the physiological data and the mechanistic interpretations. Specifically, the claim of "impaired transitions" is logically at odds with the observed increase in wake-state stability or possible hyperactivity. Additionally, the description of the methods, the quantification, and the figure presentation could be substantially improved. I detail some of my concerns below.

      Strengths:

      The selection of the knock-in model is a notable strength as it avoids the artifacts associated with APP overexpression and more closely mimics human pathology. The study utilizes continuous 14-day EEG recordings, providing a unique dataset for assessing chronic changes in arousal states. The assessment of sex as a biological variable identifies a more severe "insomniac-like" phenotype in females, which aligns with the higher prevalence and severity of Alzheimer's disease in women.

      Weaknesses:

      The study seems to lack a clear hypothesis-driven approach and relies mostly on explorative investigations. Moreover, lack of quantitative analytical methods as well as shaky logical conclusions, possibly not supported by data in its current form, leaves room for major improvement.

      Since this paper studied sleep states, the "Methods" section is quite unclear on what specific criteria were used to classify sleep states. There is no quantitative description of classifying sleep based on clear, reproducible procedures. There are many reasonably well-characterized sleep scoring systems used in rat electrophysiological literature, which could be useful here. The authors are generally expected to describe movement speed and/or EMG and/or EEG (theta/delta/gamma) criteria used to classify these epochs. The subjective (manual) nature of this procedure provides no verifiable validation of the accuracy and interpretability of the results.

      This was an oversight: the “Classification of Arousal States” section has been modified accordingly.

      One of the bigger claims is that "state transition mechanism(s)" are impaired. However, Figure 7 shows that model mice exhibit significantly more long wake bouts (>260s) and fewer short wake bouts (<60s). Logically, an "impaired switch" (the flip-flop model, Saper et al., 2010) results in state fragmentation. The data here show the opposite: the wake state has become too stable. This suggests the primary defect is not in the transition mechanism itself, but possibly in a pathological increase in arousal drive (hyper-arousal), likely linked to the dark-phase hyperactivity shown in Figures 4 and 5. Also, a point to note is that this finding is not new.

      Reviewers 1 and 2 also make comments conisistent with the alternative interpretation that “the wake state has become too stable.” However, I think we are using different words to say the same thing: that the transition from wake to sleep is impaired whether it is due to hyperarousal or to a defect in the flip/flop switch that results in greater Wake stability. I hope the reviewer would agree that a switch can be impaired in two directions: either it could “flicker” as seems to be the case in narcolepsy or it could get stuck in one position, which is what we suggest here based on the data in Fig. 12A, A’ and A” which show longer bouts of all states (Wake, NREM and REM) in older males. Nonetheless, the hyperarousal hypothesis suggested by the Reviewer is certainly a reasonable alternative. Consequently, we have added the following sentence to the Discussion subsection Impacts of pathology on sleep/wake and activity: “Thus, the hyperactivity evident in Figures 4D, 4D’, and 5D’ could drive the longer wake bouts evident in Figure 7A and result in the longer NREM and REM sleep bouts found in male App<sup>NL-G-F</sup> mice.”

      Figure 3 heatmaps lack color bars and units. Spectral power must be quantitatively defined and methods well-explained in the Methods section. Without these, the reader cannot discern if the "reduced power" in females is a global suppression of signal or a frequency-specific shift. Additionally, the representative example used to claim shorter sleep bouts lacks the statistical weight required for a major physiological conclusion. How does a cooler color (not clear what range and what the interpretation is) mean shorter sleep bout in female mice? The authors should clearly mark the frequency ranges that support their claims. In this figure, there is a question mark following the theta/delta range. The authors should avoid speculation and state their claims based on facts. They should also add the theta and delta ranges in the plot, such that readers can draw their own conclusions.

      The Y-axis in the previous version of this figure was labelled 0-25 Hz. This figure was intended to be a descriptive illustration of how unusual the female App<sup>NL-G-F</sup> mice are relative to WT of either sex rather than a quantitative analysis of spectral power. As suggested by Reviewer 2, we have combined this figure with the previous Fig. 14 as the revised Fig. 3 and we have modified the Y-axis labels to more explicitly indicate EEG frequencies. The question mark was legacy text from an earlier version of the manuscript; sorry for the confusion!

      Figure 8 and the MSLT results show that model mice are "no sleepier than WT mice" and have a functional homeostatic rebound. This presents a logical flaw in the "insomnia" narrative. True insomnia in AD patients typically involves a failure of the homeostatic process or a debilitating accumulation of sleep debt. If these mice do not show increased sleepiness (shorter latency) despite ~19% less sleep, the authors might be describing a "reduced need" for sleep or a "hyper-aroused" state, possibly not a clinical insomnia phenotype.

      Both Reviewer 2 and 3 suggest that we are using “insomnia” incorrectly, which we have used as shorthand to denote less sleep per 24h period. Reviewer 2 states that “Insomnia is defined as a subjective dissatisfaction with sleep” per DSM-5 and Reviewer 3 suggests that the mechanism underlying insomnia in AD patients is “a failure of the homeostatic process or a debilitating accumulation of sleep debt” which is not in DSM-5. Our clinical colleagues tell us that this is not established fact; some argue that the homeostat is intact and that the input(s) to the homeostat are defective. We agree that less sleep in these mice could be due to a reduced need for sleep or to hyperarousal. Consequently, we have changed the title of the manuscript to eliminate “Sleep-Wake Transitions are Impaired…” to the more objective “Long Wake/Short Sleep Bouts and Hyperactivity with Advanced Age in a Mouse Model of Early Onset Alzheimer’s Disease”.

      In Figure 9, LFP power shown and compared in percentages is problematic, as LFP power distribution is known to be skewed (follows power law). This is particularly problematic here because all the frequencies above ~20 Hz seem to be totally flattened or nonexistent, which makes this comparison of power severely limited and biased towards the relative frequency in the highly skewed portion of the LFP power spectrum, i.e., very low frequency ranges like delta, theta, and possibly beta. This ignores low, mid, and high gamma as well as ripple band frequencies. NREM sleep is known to have relatively greater ripple band (100-250 Hz) power bursts in hippocampal regions, and REM sleep is known to have synchronous theta-gamma relationships.

      We completely agree with the reviewer. There are at least 3 ways that spectral power data can be presented: (1) absolute power; (2) relative power (normalized to a baseline); and (3) power density. In this study, we intentionally presented results in terms of spectral power density so that our results could be compared to those in Figure 3A and 3B of Maezono et al. (2020). This was important because Maezono et al. recorded from mice of 6 and 12 months of age whereas we recorded from older mice, which allowed us to determine which parameters are likely changing with age (and, presumably, greater Ab deposition).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) A key finding for the AppNL-G-F mouse model is the emergence of hyperactivity that may be responsible for the altered sleep architecture. Further investigation to help determine the mechanism(s) responsible might include cFos expression to help localize or provide evidence for the distributed neuronal activity increase in this model. Additionally, identification of overly active areas might provide targets for their manipulation to test the authors' hypothesis of the mechanism of the altered sleep architecture. Does chronic hyperactivity caused by other mechanisms (DREADDs, LOF of a K channel) mimic the AppNL-G-F mouse model sleep phenotype? These sorts of findings would impact the study's significance.

      We agree with the Reviewer that identifying the mechanism underlying the long wake/short sleep phenotype of aged App<sup>NL-G-F</sup>mice would increase the study’s significance. However, we want to underscore that the opportunity to study both sexes of homozygous App<sup>NL-G-F</sup> mice and WT littermates at 14-18 months of age and to conduct a longitudinal assessment of many of the same individuals at 18-22 months was very unusual. Our observations of the phenotype described in this manuscript set the stage for subsequent mechanistic studies in aged App<sup>NL-G-F</sup> mice, although realistically, mice of this age and genotype are a rare commodity.

      (2) A more technical area of improvement involves the presentation of the results and the associated critical statistical analyses. Relevant tables and statistics are not always reported (in the results) or properly referenced. In the mixed models, the repeated measures are "time of day", I presume.

      Tables 1-6 present statistical results; these 6 Tables are referred to in the Results section a total of 14 times. The text states “The larger sample size in Experiment 2 (N=31 mice) allowed a mixed-effects model ANOVA to be conducted with Genotype, Sex, and Time as factors”. Although “Time of Day” was specified several places in the Results, thank you for pointing out omission of “of Day” from the “Data Analysis and Statistics” section; we have added this information accordingly.

      (3) The model is presented as age-dependent, but there was little statistical support for this. The subjects spanned a considerable age range, and a direct quantifiable correlation between age and the various measured dependent variables could be helpful in this regard.

      The long wake/short sleep phenotype characteristic of homozygous App<sup>NL-G-F</sup> mice that we describe here is also evident in the Maezono et al. (2020) study at 12 months of age but not at 6 months in either the Maezono et al. (2020) or Calafete et al. (2023) studies, although the authors do not comment on this phenotype and instead focus on the reduced REM sleep. Thus, between these studies, there seems to be an age-dependent progression of the phenotype. We have thus added this sentence to the Discussion subsection Sleep/wake and activity phenotypes of 14-18 month vs. 18-22 month old App<sup>NL-G-F</sup> mice: “This long wake/short sleep insomnia-like phenotype is also evident at 12 months of age (Maezono et al., 2020) but not at 6 months (Calafate et al., 2023; Maezono et al., 2020), suggesting a progression in this symptomatology.”

      (4) Would a more advanced age point be helpful? Would sleep fragmentation be likely to appear with more advanced age?

      The text states “Recordings collected throughout the entire 14-day period when Cohort 2 App KI and App WT mice were 21.0-24.3 months of age”. Mice on a C57BL6/J background are considered old at 18-24 months. Fig. 6B’ shows a strong trend (p=0.0558) toward shorter NREM bouts in App KI mice at 18-22 months during the dark phase at the same time that long wake bouts are evident (Fig. 6A’), strongly indicative of sleep/wake fragmentation but not quite significant with the sample size measured.

      (5) How does the onset of sleep-architecture-related symptoms relate to the cognitive impairment onset in AppNL-G-F mice?

      We have added this sentence to the Conclusions: “In a fear conditioning paradigm, impaired learning ability has been correlated with REM sleep duration in 13 month old but not 7 month old App<sup>NL-G-F</sup> mice (Maezono et al., 2020).

      (6) It is importantly concluded that the AppNL-G-F mouse phenotype is "stronger" in females. What is meant here by "stronger" and can this be quantified?

      We have eliminated use of “stronger” and replaced with “more evident” or “more apparent”.

      (7) Would ovariectomized females still show partial insomnia?

      This is an interesting question, particularly because the hyperactivity evident in Figure 7C is most evident in females. The average age of cessation of estrus cyclicity in C57BL6/J mice occurs between 13-16 months of age (Nelson et al., 1982, Biol Reproduction). The female KI mice in Cohort 2 ranged from 21.0 to 23.3 months of age at the time of recording and thus can be expected to be functionally ovariectomized.

      (8) The statement, "...female AppNL-G-F mice exhibited the most wakefulness and the least amount of sleep each day", sounds like a tautology.

      It was an intentional statement to underscore the long wake/short sleep phenotype.

      Reviewer #2 (Recommendations for the authors):

      (1) Introduction:

      The authors might mention in paragraph 3 that because these studies each used a mutant protein on a powerful, and not the endogenous, promoter, the effects on sleep may be skewed by overexpression in specific brain areas. In addition, they might mention that sleep homeostasis and sleep changes relative to brain temp and activity have not been examined longitudinally.

      We have added the following sentences to the Limitations subsection of the Discussion: “Moreover, because studies of this strain used a mutant protein on a powerful exogenous promoter, the effects on sleep described by us and previous investigators may be skewed by overexpression in specific brain areas” and “Neither the present nor previous studies have assessed the effects of age-related changes in brain temperature on sleep/wake, sleep homeostasis or activity.”

      (2) Results:

      Figure 2: Images in 1B and 1B' look like IHC labeling in well over 1 and 2% of the brain for Iba-1. Are these images correct?

      The use of “%” on the Y-axis was inappropriate and has been corrected. Due to variation in Iba1 immunostaining across WT mice, Iba1 measurements were normalized to WT such that the mean Iba1 area coverage for WT mice within each region of interest was set to 1. The negligible 82E1 signal in WT mice obviated the need for normalization.

      Figure 3: I would move to incorporate into Figure 14 with spectra, as this is descriptive but nicely illustrates Figure 14.

      Done -- thank you for this excellent suggestion!

      Figure 10: The figure supports no significant estrus effects in either WT or NLGF. Could run the analysis, but important finding.

      Agreed but, as indicated in the response to Reviewer 1, the average age of cessation of cyclicity in C57BL6/J mice has been reported to occur between 13-16 months of age (Nelson et al., 1982, Biol Reproduction). The female mice in the older cohort that we recorded were 18-22 months of age.

      (3) Discussion:

      Page 11, last paragraph: It is hard to say whether activity caused more wake or response to wake is different in these mice (anxiety and hyperactivity are both seen in Alzheimer's disease).

      Hypocretin MCH is touched on but could be elaborated upon, given light/dark differences.

      We agree that the directionality is difficult to ascertain. As mentioned above, we have added a discussion on hyperactivity but, having not made any assessment of anxiety in the present study, we have refrained from further speculation.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 9: Y-axis labels are missing on several plots.

      Due to the density of info on this figure, Y-axis labels were intentionally omitted for those panels for which the Y-axis label of the panel to the left applied. Since the reviewer found this to be confusing, we have added Y-axis labels to all panels at the risk of making the figure even more dense!

      (2) Figure 14: x tick labels are perplexing - why would they be labelled in such arbitrary decimal points?

      As stated in the text, “EEG spectra for each state were analyzed in 0.061 Hz bins”. Consequently, X-axis labels are modulo 0.061 Hz.

      (3) Figure S1 is not aligned; some plots cannot even be read.

      Figure S1 has been reformatted to portrait mode from the previous landscape version (although no alignment issues were evident when viewed in landscape mode).

      (4) For some reason, Tables 1-3 are horizontal, which I couldn't read.

      Our apologies, some of the info in Table 1 was omitted during export. We have retained landscape mode for Table 1 and re-formatted Tables 2 and 3 in portrait mode for ease of accessibility.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Freas and Wystrach present a computational model of steering in insects. In this model, the central complex provides an error signal indicating the animal should turn left or right; this error signal biases the function of an oscillator composed of two mutually inhibiting self-exciting units. The output of these units generates a "steering signal" that is used both to set the direction and speed of the ant. Additionally, a separate module induces pauses, and an inverse relation between forward speed and turning speed is externally imposed. Statistics of the trajectories generated by the model are compared to the measured behaviors of ants.

      Strengths:

      While the model is very simple compared to state-of-the-art models, that simplicity makes it a potentially useful guide to researchers studying insect navigation. Some predictions that emerge from the model appear to be experimentally testable, although a more complete description of the model and its parameters, as well as an analysis of how this model's predictions differ from previous models' predictions, would be required to design these experiments.

      Weaknesses:

      I found it difficult to identify evidence in the paper supporting central elements of the abstract. Hopefully, these difficulties can be resolved with a clearer presentation and the addition of supporting detail, especially in the methods.

      (1) The model is not clearly described

      In the Materials and Methods, there is no description of the model, just "The computational model is presented in Figure 1." (This is probably a typo and may refer to Figure 2A-C), and a link to Matlab source code. It is inappropriate to ask readers or reviewers to examine source code in lieu of providing a method, but I attempted to do so anyway. 

      We have now added a full description of the model in the methods.

      To my eye, the source code does not match the model presented in 2A-C. For instance, in 2C, "Steering signal" inhibits "Freeze", but I couldn't find this in the source. "Freeze" is shown to inhibit "steering signal," but as "steering signal" is a signed quantity, it's not clear what this means. Literally, since "ang_speed_raw = L-R," it would seem to indicate the "freeze" would bias towards right turns. In the code, "freeze" appears to be implemented through the boolean variable "speed_inhibition_time." The logic controlled by this variable doesn't appear to inhibit the "steering signal" but instead (depending on control parameters) either reduces the movement speed and amplifies the turning rate, or it turns the angular speed output into a temporal integral of the control signal.

      We understand the confusion. Our neural implementation does not go downstream of the neural steering signal (Left and Right Descending neurons), and the way it is transformed into a movement (ang_speed_raw = L-R) is not modelled neurally (the formula is explicitly shown on the right hand side of Figure 2). Indeed, we did not attempt to put forward any assumption about neural implementation for our freezing signal (see our response to comment 2 below). To avoid confusion, we have now removed the reciprocal inhibition portion as it was previously drawn in Figure 2C, and replaced it by a non neural sign (a cross, indicating that the signal is blocked) acting between steering signal and movement.

      There are a number of parameters in the source code that aren't described at all in the paper, including the internal oscillator parameters.

      We now provide all the parameters in the methods, together with figures showing the dynamics of oscillations across parameter range, and a rationale for their choice (see Supplemental Figure 2).

      Together, these limitations make it difficult to understand what is being simulated, what parts of the model are tied to biology, and where the model improves on or departs from previous work.

      It is absolutely essential that authors fully describe the computational model, that they explain the meaning of all parameters of the model, and that they explain how the particular values of these parameters were chosen.

      This is now done in the methods section under the “Model Overview” subsection.

      (2) The biological inspiration is unclear

      A central claim of the paper is that the model is "biologically grounded." But some elements, for instance, using a signed quantity to represent left-right steering drive, are not biologically possible; at best, these are shorthand for biologically possible implementations, e.g., opposing groups of left-right driving neurons.

      The mechanism that produces fixations and saccades - the "freeze" module - is not tied to any particular anatomy of the insect brain. Initiation of a freeze occurs at a specific time coded into the model by the authors; it is not generated by an internal model signal. Release of a freeze is by drawing a random variable; there is no neural mechanism proposed to generate this signal.

      We now clarified what is neural from is not from the introduction onwards, for instance:

      “Because we did not want to form pre-assumptions for how such a ‘freeze signal’ could be implemented in the insect nervous system; in our model this was achieved using a simple external signal that halts forward motion at random intervals.”

      In some versions of the model, instead of directly controlling the signal, during fixations, the angular drive signal is integrated into a variable "cumul_drive." No neural substrate is proposed for this integrator. In the code, if cumul_drive passes a threshold, the angular heading of the ant changes (saccades), but only if this threshold is passed before the Poisson process ends the fixation. No neural substrate is proposed for any of this logic.

      This has now also be clarified in the introduction:

      “During scanning, real ants display rotational saccades of variable duration and angular magnitude (Figure 1A–C). To replicate this, we introduced a threshold-based mechanism: after each fixation (i.e., zero angular and forward speed), the underlying angular steering signal accumulates until surpassing a threshold, triggering a saccade. The resulting angular magnitude of the saccade corresponds to the sum of the angular drive accumulated during the fixation. Here also we stuck to a non-neural, straight-forward algorithmic level, as we did not want to make assumptions about how such a cumulate-and-release mechanism could be neurally implemented in the insect brain (see discussion for potential implementations).”

      The model steps forward in time by a fixed increment - the actual duration (in seconds) of this time step is not specified. From Figure 4F, G, it appears a simulation time step is meant to be about 10ms. This would imply an oscillator frequency of about 2 Hz (Fig 2B), that the heading oscillates at a similar frequency (2G), and that a forward crawling ant stops moving every 500 ms (2I). Are these plausible? Can they be compared to an experiment? Model parameters, including the ones that control the frequency of the oscillator, are non-dimensionalized. It is not possible to evaluate whether these parameters are biologically plausible or match experimental results.

      We now added a figure showing the oscillatory dynamics of the oscillator across parameter ranges (supplemental figure 2). The step increment (i.e., and thus the sampling rate along an oscillatory cycle) necessarily varies according to the inhibition strength and self decay parameter chosen (e.g., small parameter values will lead to small step increment, and thus a high sampling rate along the oscillatory cycle). We chose oscillatory parameters to ensure that the sampling rate will be high enough to resolve multiple saccades within one oscillatory cycle and that sampling rate is small enough for computation time to remain practical.

      Beyond these constraints, the oscillator parameters can be chosen arbitrarily, and a conversion of time step to actual time (ms) would be equally arbitrary and give the illusion that the model captures the data quantitatively. Because we did not model spiking neural dynamics (or brain region low field potential frequencies), we can not constrain our model through a temporal link between brain clock and behavioural speed. We thus prefer to stick to the true and non-dimensional label ‘time steps’ in our figures.

      (3) Claims that behaviors emerge from the model may be overstated

      The abstract claims that steering correction and fixations/saccades emerge naturally from the same model. But it appears to me that fixations/saccades are externally imposed by the specification of specific times for a "freeze." Faster angular rotation during saccades than during course correction is imposed and does not emerge naturally from neural simulations.

      The abstract now clarifies that what emerges spontaneously is not scannings per se (indeed, the inhibition of movement is externally imposed) but their dynamics. Note that our model captures many aspects of scanning dynamics that are not trivial and which results from the dynamical interactions and contingencies between modules (figure 3 to 7), hence justifying the word ‘emerge’ insofar as these behavioural dynamics cannot be reduced to one module or parameter. Regarding the faster angular rotation during scanning, we agree that its cause is rather straightforward to understand: it results from the added bodily constraints of forward speed to rotational movements. Nonetheless it is not ‘imposed’ during saccades in the sense that 1.) it is biologically/physically evident rather than cherry picked and 2.) it is continuously present in our model, even during forward navigation. We believe the new version of the manuscript now conveys this message in a transparent manner.

      (4) Citations to previous literature are difficult to follow, and modeling results are presented as though they are experimental data

      I would ask the authors to be much clearer in their description and citation of previous work. It should be clear whether the cited work was experimental or computational. To the extent possible, the actual measurement should be described succinctly. Instead of grouping references together to support a sentence with multiple claims, references should be cited for each claim. Studies of computational models should not be presented as proving a biological result.

      Indeed, This we now clearly separated citations referring to experimental evidence vs. modelling. See examples citations below

      For example:

      (a) Lines 141-146:

      "Previous studies have established many key components of insect navigation, including .... the intrinsic oscillatory dynamics in the lateral accessory lobes (LALs) that support continuous zigzagging locomotion (Clément et al., 2023; Kanzaki, 2005; Namiki and Kanzaki, 2016;

      Steinbeck et al., 2020)."

      The first reference is to one author's previous modeling work - it hypothesizes that oscillations in the LAL support zigzagging but includes no data that would "establish" the fact. Kanzaki et al. 2005 describes numerical modeling and simulation with a physical robot. Namiki and Kanzaki, 2016 is a review article that links the LAL to zigzagging behavior. It describes the LAL as a winner-take-all bistable network but does not describe or hypothesize that the LAL has intrinsic oscillatory dynamics. Steinbeck et al. 2020 is a more comprehensive review; it reinforces that the LAL is a winner-take-all bistable network that drives left-right steering, including during zig-zagging behavior. But in my reading, I could not find a statement that the LAL has intrinsic oscillatory dynamics (the closest is Steinbeck et al. saying the activity pattern switches regularly, as does the behavior; this doesn't imply that the LAL is intrinsically oscillatory.)

      It now reads:

      “Previous studies have established many key components of insect navigation, notably, how goal headings are set in the central complex (CX) (Fisher, 2022; Green and Maimon, 2018). Modelling efforts have shown that the CX circuitry can naturally accommodate innate and learnt guidance such as path integration, learn vectors, visual route following or homing as observed in ants and bees. In parallel, oscillatory dynamics in the lateral accessory lobes (LALs) - produced by reciprocal inhibition across both hemispheres and conveyed by so-called descending flip-flopping neurons - were shown to drive the spontaneous zigzags displayed by moths upon losing their pheromone plume (Kanzaki and Mishima, 1996; Mishima and Kanzaki, 1998, 1999; Wada and Kanzaki, 2005; Kanzaki et al., 2005; Iwano et al., 2010). Here also, subsequent modelling efforts have shown how these circuits can equally support the continuous lateral oscillations displayed by a wide range of insect species, including ants.”

      (b) Lines 701-703:

      "In plume-tracking moths, CX output has been shown to modulate LAL flip-flop neurons driving zigzagging (Adden et al., 2022)."

      This reads as though an experimental measurement was made, but in fact, this is modeling work.

      Yes, this could be clearer, it now reads: 

      “In moths, descending neurons in the LALs exhibit characteristic 'flip-flop' activity patterns that correlate with zigzagging maneuvers (Olberg, 1983; Kanzaki and Ikeda, 1994). Computational models suggest that having these LAL neurons modulated by the CX output can explain aspects of the moths’ plume-tracking behaviour (Adden et al., 2022).”

      (c) Lines 703-706:

      "In ants, strong goal signals in the CX - whether elicited by the path integrator or visual familiarity (Wehner et al., 2016; Wystrach et al., 2020b, 2015) do not only sharpen directional accuracy but also increase oscillation frequency (Clément et al., 2023)."

      Here again, modeling results are presented as though they were experimental data.

      Here, we are referring to the experimental part of these works, although this comment demonstrates that our statement should be more clear in stating what are biological results. It now reads: 

      “In ants, behavioural studies show that strong directional drives elicited by the path integrator or visual familiarity do not only gain behavioural weights and sharpen directional accuracy (Wehner et al., 2016; Wystrach et al. 2015, Legge et al. 2014) but also increase the ants’ oscillation frequency (Clément et al., 2023). Assuming that path integrator and visual familiarity modulate goal signals in the CX, as modelled here and elsewhere (Wystrach et al., 2020b, Stone et al., 2017) and that the intrinsic oscillator is in the LAL (Clément et al., 2023, Steinbeck et al., 2020), it suggests that CX output modulates the intrinsic oscillatory activity of the LAL”

      Reviewer #2 (Public review):

      Summary:

      The paper by Freas and Wystrach is an interesting computational study, exploring the detailed mechanisms of how simple neural circuits could explain complex behavioral patterns observed in navigating ants. The authors compare detailed, high-speed video recordings of Australian desert ants (Melophorus bagoti) with predictions made by their new computational model and find convincing similarities between the model and the behavioral data, at a level of detail not previously studied. Particularly interesting are emerging properties of the model, yielding behavioral motifs it was not designed to reproduce, but which occur in natural ant behavior.

      Strengths:

      A strength of the study is that the model is based on previous models, without making major novel explicit assumptions. It combines existing models of the insect central complex with a model of the lateral accessory lobe and adds a stochastic inhibition of forward velocity to the interaction of central complex and lateral accessory lobes. The central complex provides corrective steering signals when the goal direction and the current heading of an insect are not aligned, while the lateral accessory lobes provide an intrinsic oscillator underlying the behavioral oscillations shown by walking ants at all times. These background oscillations are modulated by the steering signals from the central complex. Depending on which phase of the intrinsic oscillations coincides with the corrective signals, and how fast the ant is moving forward during this time, a complex set of behaviors emerges. Most prominently, scanning behaviors, which are regularly carried out by the ants, are recapitulated in great detail by the model. Additionally, other behaviors, such as full loops, emerge naturally from the model. While computational models are not to be seen as definite evidence for any biological reality, they can provide strong support for particular neural implementations. The current study is an excellent example in that it provides evidence for a serial arrangement of central complex circuits upstream of the lateral accessory lobe circuits, modulated by speed-regulating input. While the latter is hypothetical, it yields a clear hypothesis that can be validated by connectomics studies and functional work in the future.

      The study shows that even complex behavioral motifs do not require dedicated neural modules, but can rather emerge from the interplay of already known circuits - highlighting the efficiency of insect brains and possibly providing the path towards embodied hardware solutions of such circuits in autonomous agents.

      Weaknesses:

      There are several weaknesses in the paper as it is.

      Firstly, the model is not described in the methods, but only found when following the link to the authors' GitHub repository. This is clearly not sufficient and prevents readers from evaluating the model's assumptions directly. Most importantly, how natural do the emerging properties indeed emerge from the model? What parameters need to be tuned to generate a match between data and model?

      We have now added a full description of the model in the Methods section.

      These include:

      Mathematical equations for model components

      Complete parameter table along with justifications

      Description of what is fitted vs. what emerges 

      Key assumptions and limitations

      Regarding the emergence of scanning properties: The model has two types of parameters:

      Parameters tuned to match general navigation behavior (independent of scanning):

      Motor gains (g_ang, g_fwd, k): adjusted to produce realistic continuous walking paths and species differences between desert ants and Myrmecia

      CX gain (g_CX = 0.5): set to produce appropriate corrective steering strength during continuous navigation

      Oscillator parameters (α, β, s): are taken from Clément et al. (2023)

      Parameters tuned to match scanning behavior:

      CPG angular threshold (θ_CPG = 2.0): adjusted to generate realistic saccade timing Scan termination probability (p_stop = 0.5/timestep): matched to the Poisson-like distribution of scan durations in M. bagoti

      Properties that emerge without specific tuning:

      Fixation-saccade alternation structure (emerges from angular drive accumulation mechanism)

      Directional reversals (arise from oscillator dynamics competing with CX steering)

      Corrective saccade amplitude increasing with angular deviation (Figure 3)

      Rare full-loop scans (emerge from CX signal shifting oscillator phase)

      The behavioral continuum from straight paths → oscillations → voltes → scans (Figure 8)

      We have clarified this distinction in the Methods section and emphasized that our goal was qualitative demonstration of emergence rather than quantitative parameter optimization.

      Second, it is often not entirely clear what is biological data and what is a computational model. This relates to figures, text, and references. As a reader, this makes it difficult to clearly judge what is new in the current paper, how it adds to previous models, and what the predictions and assumptions are for biology.

      Indeed, we have now clarified the manuscript, clearly separating when we refer to behavioural data, neurobiological data and modelling. In the figures, each panel now clearly indicates if it is model data or biological data so that any reader can immediately tell the data type.

      Third, while neural data from bees and flies are taken to motivate and design the computational model, the discussion and interpretation revolve almost exclusively around ants. For the most part, this is justified, as the behavioral data used to benchmark the model are taken from ants. Nevertheless, more broadly discussing the newly defined circuit in the context of flying insects would give a better idea of the broad relevance of the neural circuits predicted by the model.

      To address this suggestion we have now added two paragraphs in the discussion called: “Scanning in flying hymenopterans”.

      Also happy to add more to this section if requested.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      As mentioned in the public review, I suggest fixing the two concerns I have regarding methods and discussion.

      (1) Include a full description of the model in the methods, so that the model remains reproducible even if the GitHub repo is deleted in the future.

      True, the code’s internal explanations could indeed be removed from GitHub later. The model component overview are now included in text.

      (2) Include the relevance of the model for flying insects in the discussion more prominently. This seems to be an implicit assumption in the model, as neural data from bees and, more prominently, from Drosophila are used to motivate the model to explain ant data.

      Add an “Expression in flying hymenopterans” section at ~line 834.

      Minor points:

      (1) Line 207: I suggest adding the recent review by Collett, Graham, and Heinze (2025, Current Biology), as it proposes interactions between LAL and CX as well.

      Added

      (2) Figure 4: I'm interested in the conversion from steps in the model to real units (ms) in the ants. In Figures 4F and G, it seems that 5 model steps represent circa 100ms. Does this allow us to define the neuronal time constants of the model neurons? If so, are the resulting values biologically plausible? This seems important when describing real-world dynamics being created by a model circuit.

      No the model is time agnostic.

      (3) Figure 7: Font sizes of axis labels are much too small. Also applies to other figures. Please ensure that when printed, labels can be read.

      Enlarged axis labels in all figures. 

      (4) Line 645: proprieties -> properties?

      Fixed. Thanks!

      (5) Figure 7: The figure heading states: "Slow forward speed (Myrmecia) example". This sounds as if real data from ants are shown here, while these are modeling data. It is clear after reading the text and caption in detail, but I was taken off course briefly here. Please make sure that there is no possibility of being misled here.

      We have altered the subtitle to “Slow forward speed (Myrmecia Model) example”. 

      Additionally, we have added a Model tag under each of the model image labels so classification can be done at a glance.

      (6) General discussion: What about search dynamics, i.e., increasing loops when not finding the nest entrance after homing? Are those emerging from this circuit as well? Or would that need to be a separate module? There have been discussions about search emerging from the PI circuit, but as far as I know, this is not settled, and it would be good to know if the current circuit adds something useful to this aspect.

      Because we kept a fixed goal heading, our model does not bring insight about overall trajectories such as search pattern. We now mention in the discussion:

      “In our simulations, the CX goal representation remained fixed in both direction and strength throughout each trial. This simplification allowed us to isolate and compare the effects of different CX strengths on scanning behaviour (Figure 6). However, goal headings in the CX are likely to be updated continuously, including during scans, by novel input from visual recognition in the MB (ref). This would in turn bias saccades direction and duration. Exploring such dynamics lies beyond the scope of the present study but would represent an interesting direction for future work. Notably, our proposed CX-LAL-Body relationship could be implemented downstream of an existing path integration or visual-based model (or both) to form predictions about the occurrence and dynamic of scans along the path, as well as their impact on the emerging trajectories.”

      (7) Line 690: The modulation of PFL3 by PFL2 was presented as a hypothesis in Westeinde et al., consistent with the data, but as far as I know, this is not an established fact.

      You are correct. We have now softened the text, which now reads: “In Drosophila, it has been proposed that PFL2 neurons, which respond maximally when the fly faces away from the goal, modulate steering gain by converging with PFL3 neurons (which drive left or right turns) onto downstream descending neurons (Westeinde et al., 2024).”

      (8) Please ensure that Drosophila is consistently spelled with a capital D and in italics.

      Fixed throughout the text.

      (9) Line 702: Reference Adden et al 2022: This reference is a modeling paper; it sounds as if you are referring to an experimental moth paper, though. Rephrase to clarify.

      You are correct, this could be unpacked much better regarding what is modelled and what has been experimentally shown. Changed to:

      Descending neurons in the LALs exhibit characteristic 'flip-flop' activity patterns that correlate with the zigzagging maneuvers of plume-tracking moths (Olberg, 1983; Kanzaki and Ikeda, 1994). Recent computational models suggest that CX output directly modulates these LAL circuits to coordinate orientation (Adden et al., 2022). 

      (10) Line 761: I would assume that during scans, information is acquired that would decrease uncertainty and thus, as a result change the amplitude of the CX steering signal. Maybe I missed this, but is this closed-loop interaction integrated in the model?

      In our simulation the CX goal representation remains stable in direction and strength throughout the trial. This enabled us to compare neatly the effect of different CX strengths on scanning. However, we fully agree with you that goal headings in the CX might well be continuously updated, both during scans and between scans! The goal heading novel strength or direction may thus bias the scan further left, right, in front or in the back, and also up or down regulate scan duration in both directions. 

      Modelling this would require adding a layer of complexity to determine how the goal heading is updated, which is beyond the scope of the current work, but would form a remarkable project for the future. We now mention this in a dedicated paragraph in the discussion section “Model limitations and future directions”

      (11) Line 814: Please add 'fly' in front of larva. Other insect larvae have a fully developed CX.

      Corrected. Added fly to this sentence 

      (12) Line 815: Maybe add the recent review, Heinze 2025.

      Added this one (Heinze 2024) which seems to fit the best and the 2025 Curr Biol Review doesn't quite fit this line (cited elsewhere though): 

      Heinze, S. (2024). Variations on an ancient theme—the central complex across insects. Current Opinion in Behavioral Sciences, 57, 101390.

      (13) Methods: Subheading formatting should start with capital letters.

      Ah yes, the second level of subheadings got formatted weirdly. Fixed now.

    1. Author response:

      We thank the reviewers for their enthusiasm for the work as well as for their thoughtful and constructive comments, which will lead to many improvements in the manuscript. We will address their concerns/suggestions in the following ways:

      Reviewer 1

      (1) We will revise text to help the reader more intuitively understand how dendritic asymmetry can translate into alterations in receptive field location, as well as provide a better description of the cited portions of the Methods section.

      Reviewer 2

      (1) The simulations in the current version of the manuscript modeled a transient response via a single synaptic conductance in part because one can better visualize the interplay between synaptic inputs and voltage-gated ion channels across both time and dendritic space. However, we agree that it is also important to show how our results are impacted during ongoing trains of synaptic activity exhibiting short-term depression as documented in the literature. We will add an additional figure showing simulations employing realistic statistical patterns of presynaptic excitatory and inhibitory inputs with appropriate short-term plasticity characteristics. These simulations are already complete and show that the increased complexity minimally alters the location of modeled ITD curves of the cell population over a wide range of frequencies (250 Hz – 2 kHz).

      (2) The reviewer’s suggestion of sequentially pruning the different orders of dendritic branches is an excellent one. However, removal of dendrites also alters overall whole cell resistance and capacitance as well as the cable properties of the remaining dendrites. It is thus impossible to disentangle the branch-specific effects of synapse location from changing intrinsic electrical properties. However, the reviewer has inspired us to address their suggestion in a slightly different way: we will add (via a new figure) simulations that take place in the same dendritic arbor, but with inputs restricted to progressively lower orders of dendritic branches. Thus, the relative contributions of synapses onto higher order dendritic branches can be visualized without fundamentally changing the electrotonic structure of the simulated neurons across the different conditions. These simulations will be performed under the “in vivo-like” conditions described in the previous point. We think they will effectively address the essence of the reviewer’s suggestion.

      (3) We will add more specific information about animal ages in relevant figures, including Supplementary Figure 1. We will also indicate that all physiological recordings were performed near physiological temperature (35°C), which was unintentionally omitted.

      Reviewer 3

      (1) We will add more detail about the anatomical assumptions regarding spatial input patterns vs. higher order dendrites. We do not think that VGluT staining with dendritic labeling will be a productive experiment, since the thin sections that provide high quality labeling conditions also preclude following single dendrites for long distances. The distal portions, which are of particular interest, are most difficult to follow because of their smaller diameter and more extensive branching out of the plane of thin sections. Further, the work of Callan and colleagues (2021) has addressed axonal input patterns as well as dendritic coverage, documenting that single axon inputs follow dendrites for variable distances, and typically provide multiple synaptic contacts. This work also highlights the many challenges and large effort involved in documenting synaptic innervation patterns in single cells at the light microscopic level. Thus, we do not think we can improve upon existing anatomical descriptions without excessively expanding the scope of an already long study, which will have 9 figures after revision.

      (2) We have analyzed many other measures of dendritic complexity but for reasons of clarity and focus included the two measures that appeared most intuitive and impactful (length and surface area). We agree that access to other measures would be useful even if some are less intuitive, and thus we will provide a more comprehensive analysis of dendritic structure in a supplementary figure.

      References:

      Callan, A. R., Heß, M., Felmy, F., & Leibold, C. (2021). Arrangement of Excitatory Synaptic Inputs on Dendrites of the Medial Superior Olive. The Journal of neuroscience, 41(2), 269–283. https://doi.org/10.1523/JNEUROSCI.1055-20.2020

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The central pair apparatus of motile cilia consists of two singlet microtubules, termed C1 and C2, each of which is associated with a set of projections, referred to as the C1 and C2 projections. Each projection comprises multiple distinct structural domains, designated a, b, c, and so on. Biochemical studies combined with genetic analyses in Chlamydomonas identified three proteins as the major components of the C2a projection, and subsequent cryo-EM studies confirmed these findings.

      In this paper, the authors aim to study the homologues of these three proteinsCCDC108/CFAP65, CFAP70, and MYCBPAP/CFAP147-using knockout mouse models. Biochemical and cell biological analyses demonstrate that, as in Chlamydomonas, these proteins are components of the C2 projection and form a complex that depends on the presence of each other. In addition, the authors use affinity purification to identify two previously uncharacterized proteins and show that they are central pair apparatus proteins that associate with the aforementioned complex. Knockout mice lacking any of the three core proteins exhibit phenotypes consistent with primary ciliary dyskinesia (PCD).

      Overall, the manuscript is clearly written, and the data are convincing and support the authors' conclusions. However, given the previous findings in Chlamydomonas, this work provides limited conceptual advances to the field. Nonetheless, it represents a useful and well-documented resource for understanding the conserved organization of the central pair apparatus in motile cilia. It will be of interest to cell and developmental biologists, biochemists, and clinicians studying and treating human ciliopathies.

      We sincerely appreciate the positive feedback on our work.

      Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the protein composition and functional role of the C2a projection of the central apparatus (CA) in vertebrate motile cilia. Using three knockout mouse models (Ccdc108, Mycbpap, and Cfap70), the authors demonstrate that these genes - homologs of Chlamydomonas FAP65, FAP147, and FAP70 - are required for normal motile cilia function in ependymal and tracheal multiciliated cells. Specifically, the authors show that:

      (1) Knockout mice for each gene exhibit primary ciliary dyskinesia phenotypes (hydrocephalus and sinusitis), accompanied by abnormal ciliary motion and reduced ciliary beat frequency.

      (2) CCDC108, MYCBPAP, and CFAP70 physically interact and localize to the axonemal central lumen, consistent with the C2a projection.

      (3) Loss of any one of these proteins destabilizes the others and disrupts CA integrity in a tissue-specific manner.

      (4) ARMC3 and MYCBP are C2a-associated proteins.

      Strengths:

      (1) Clarity: the results are presented in a coherent sequence that facilitates understanding of both the rationale and conclusions.

      (2) Genetic rigor: three independent knockout mouse lines that exhibit consistent motile cilia phenotypes provide in vivo support for the proposed role of these proteins.

      (3) Integration of structural and functional analyses: combination of ultrastructural (TEM) and immunofluorescence data with CBF measurements provides convincing correlation between structural defects and impaired ciliary function.

      (4) Mutual dependency model: reciprocal destabilization of CCDC108, MYCBPAP, and CFAP70 supports their interdependence in the C2a assembly.

      (5) Expansion of the vertebrate C2a proteome: the identification of ARMC3 and MYCBP as C2a-associated proteins provides a foundation for future mechanistic studies.

      We appreciate the valuable comments and pertinent suggestions, which provide important guidance for revising and improving this manuscript.

      Weaknesses:

      (1) Mechanistic depth: the data show a convincing correlation between C2a and ciliary function, but the cell type-specificity of CCDC108, MYCBPAP, and CFAP70 knockout effects is underdeveloped. This is an interesting observation that raises mechanistic/structural questions not addressed in the study, such as what is the role of C2a in CP nucleation, maintenance, or mechanical stabilization? Is C2a composition different in different cell types?

      We appreciate this comment. Based on current knowledge, loss of proteins essential for CP nucleation, such as WDR47 and KIF27, typically causes severe CP loss defects [1,2]. However, only mild CP-loss defects were observed in Ccdc108, Mycbpap, or Cfap70 knockout (KO) mouse ependymal cells (mEPCs) serum-starved for 10 days (Figure 2E, F), indicating that C2a proteins are more likely to play a role in CP maintenance or mechanical stabilization. In the revision, we tested this hypothesis by examining the effects of C2a loss on CA ultrastructure in Ccdc108 KO mEPCs serum-starved for 5 days. The percentage of axonemes with defective CA decreased further (Figure 2—figure supplement 1C, D). These results further confirm that C2a proteins play a role in CP maintenance or mechanical stabilization but not in CP nucleation. We have included these results and expanded the related discussion in the revised manuscript.

      To assess whether C2a composition differs across cell types, we performed co-immunoprecipitation using lysates from mouse trachea and mEPCs. We found that, in both tracheal and mEPC lysates, CFAP70, ARMC3, and MYCBP were co-immunoprecipitated with MYCBPAP (Figure 6—figure supplement 6A, B), indicating that at least the C2a core components are conserved in vertebrate motile ciliated cells. We have included these results in the revised manuscript.

      (2) Cell model choice: co-immunoprecipitation was performed using mouse testis lysates. While this is a reasonable source of CA proteins from flagellated cells, the functional analyses in this study focus on ependymal and tracheal multiciliated cells. It would therefore be helpful for the authors to clarify the extent to which these interactions are expected to be conserved across ciliated cell types, and to discuss potential tissue-specific differences in CA assembly.

      We thank the reviewer for the insightful suggestion. Following the reviewer’s suggestion, we performed co-immunoprecipitation using lysates from mouse trachea and mEPCs. We found that, in both tracheal and mEPC lysates, CFAP70, ARMC3, and MYCBP were coimmunoprecipitated with MYCBPAP (Figure 6—figure supplement 1A, B), indicating that at least the interactions among the C2a core components are conserved in vertebrate motile ciliated cells. We have included this result in the revised manuscript.

      (3) Statistical analysis: the manuscript states "Statistical significance was defined as P < 0.5", which is likely a typo, but should be P < 0.05. In general, the statistical methods require more clarification. In several figures (e.g., 2B, 2D, 5J, 5K), multiple knockout genotypes are compared with WT, yet unpaired t-tests are reported. When more than two groups are analyzed, multiple pairwise t-tests inflate Type I error unless appropriately corrected; a oneway ANOVA with post hoc comparisons (e.g., Dunnett's test for WT-referenced comparisons) would be more appropriate. Furthermore, the analysis of ciliary movement modes (Figure 2D) involves categorical data, for which a t-test is not statistically appropriate. These comparisons could instead be evaluated using chi-square or Fisher's exact tests. Addressing these issues is important to ensure accurate statistical inference.

      We thank the reviewer for identifying the error and for their suggestions on the statistical analysis. We performed a one-way ANOVA with Dunnett’s test in Prism to re-evaluate the differences between WT and each KO sample. In the revised manuscript, we have updated the statistical results and revised the Methods section.

      (4) Methods section: does not sufficiently describe how image-based quantifications were performed. For example, the criteria used to define cilia number, basal body number, and rotational beating are not specified, nor is how CBF measurements were analyzed. The authors should also provide details regarding analysis software and imaging parameters used (and whether they were kept constant across genotypes).

      We apologize for omitting a detailed description of image-based quantifications. For counting cilia or basal bodies, cells were immunostained with acetylated α-tubulin and CEP164 antibodies to label cilia and basal bodies, respectively, and imaged using 3D-SIM. Using these super-resolution images, cilia and basal bodies were counted in each multiciliated cell. With highspeed live-cell imaging, ciliary movements were recorded and analyzed using ImageJ. mEPCs in which the majority of motile cilia displayed rotational motility were considered ‘cells with rotational cilia’. The CBF of each cilium was calculated from the total time of 10 beating cycles. In the revised manuscript, we have included these details in the related figure legends and the methods section.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 121: "appeared decreased" should be revised to "appeared to decrease."

      We thank the reviewer for the suggestion. In the revised manuscript, we have changed the text accordingly.

      (2) Figure 2 legend: The statement "Arrowheads indicate the C2 projections" is misleading. The arrowheads indicate the positions of the C2 projections, as the C2 projections are absent at the locations marked by the arrowheads.

      We appreciate this comment and have revised the text in accordance with the reviewer’ s suggestion.

      (3) Statistical analysis: For the statistical analyses shown in Figure 2 and the other figures, a t-test was used. In general, a t-test is appropriate for comparisons between two groups. When more than two groups are compared with a single factor, a one-way ANOVA should be used, followed by appropriate post-hoc tests.

      We thank the reviewer for pointing out this issue. In the revised manuscript, we have re-evaluated all statistical analyses in Figures 2 and 5 using Dunnett’s test to compare multiple treatment groups (Ccdc108 KO, Mycbpap KO, and Cfap70 KO) with a single control group (WT). We have also revised the corresponding figure legend and methods section.

      (4) Docking methodology: In Figures 1A and 5L, the molecular model of the C2a projection (PDB: 7SOM) is superimposed onto the cryo-EM density map. I was unable to find a detailed description of the method used for this docking and would appreciate clarification.

      We apologize for omitting a detailed description of the docking methodology. The visualizations in Figures 1A and 5L were generated using the following procedure:

      (1) Generation of the Complete CA Density Map: Following the hierarchical local refinement strategy and map integration methods described in previous high-resolution studies of the Chlamydomonas central apparatus (CA) [3,4], we utilized the published density maps of the C2 microtubule and its associated projections (EMD-24191) and the C1 microtubule and its projections (EMD-24207). These maps were aligned and stitched together in UCSF ChimeraX to reconstruct a complete C1-C2 repeating unit of the central apparatus.

      (2) Superimposition and Fitting (Figure 1A): To generate the molecular model shown in Figure 1A, the atomic model of the Chlamydomonas C2a projection (PDB: 7SOM) was docked into the corresponding region of the integrated C2 density map [4]. The docking was performed as a rigidbody fit using the "Fit in Map" tool in UCSF ChimeraX, which optimizes the correlation between the molecular model and the cryo-EM density.

      (3) Simulation of C2a Loss (Figure 5L): For Figure 5L, we simulated the results of C2a loss observed in our mutation experiments. Using the model established for Figure 1A as a template, we selectively removed the C2a-specific density and the corresponding superimposed atomic model to schematically illustrate the structural consequences of the mutations involved in this study.

      We have updated the Methods section of the revised manuscript to include these details regarding structural visualization and docking analysis.

      Reviewer #2 (Recommendations for the authors):

      (1) Lines 106-107: "frameshift mutation was created by introducing a 458-bp deletion of exons 6-8 in the mouse 107 Cfap70 (ENSMUST00000056073.14) (Figure 1B)". The figure indicates deletion of exons 3-8; please indicate which is correct.

      We apologize for the oversight and confirm that the deletion region encompasses exons 3-8 (as shown in Figure 1B). In the revised manuscript, we have updated the text accordingly.

      (2) Lines 119-121: "Genotyping at postnatal day 0 (P0) revealed that Ccdc108 KO pups, Mycbpap KO pups, and Cfap70 KO pups were all born at the expected Mendelian ratios; however, the ratio of Mycbpap KO mice at P7 appeared decreased (Figure 1E)". The authors can test whether the genotype distribution changes between P0 and P7 to directly support their claim of postnatal lethality.

      We appreciate the reviewer’s comments. The P0 genotyping results were obtained from P0 neonatal mice sacrificed for mEPC culture. Therefore, the P0 and P7 genotyping distributions were from different batches of mice. Re-doing the genotyping distribution analysis would require a large number of mice and considerable time. We hope the reviewer understands the difficulty and allows us to forgo this experiment.

      References

      (1) Liu, H., Zheng, J., Zhu, L., Xie, L., Chen, Y., Zhang, Y., Zhang, W., Yin, Y., Peng, C., Zhou, J., et al. (2021). Wdr47, Camsaps, and Katanin cooperate to generate ciliary central microtubules. Nat Commun 12, 5796. 10.1038/s41467-021-26058-5.

      (2) Park, H., Choi, M., Zhang, Y., Cheung, H.O., Makino, S., Yoshikawa, Y., Qi, H., Liu, Z., Lan, G., Fu, G., et al. (2025). The kinesin-4 protein KIF27 forms a cytoskeletal scaffold at the transi\on zone to promote mo\le cilia structural integrity. Proc Natl Acad Sci U S A 122, e2515392122. 10.1073/pnas.2515392122.

      (3) Han, L., Rao, Q., Yang, R., Wang, Y., Chai, P., Xiong, Y., and Zhang, K. (2022). Cryo-EM structure of an ac\ve central apparatus. Nat Struct Mol Biol 29, 472-482. 10.1038/s41594022-00769-9.

      (4) Gui, M., Wang, X., Dutcher, S.K., Brown, A., and Zhang, R. (2022). Ciliary central apparatus structure reveals mechanisms of microtubule paderning. Nat Struct Mol Biol 29, 483-492. 10.1038/s41594-022-00770-2.

    1. Author response:

      Response to the eLife Assessment

      We thank the Editors and the Reviewers for their helpful suggestions, which will help us strengthen and test the key conclusions of this study of condensate dynamics at atomic resolution. In response to the Editors, we will make clearer in the Results and Discussion how the present work advances beyond our initial study of MUT-16 condensates, the scaffold of Mutator foci (Gaurav K et al., Biophys. J. 2025; 124:3987–4004). That study used a multiscale approach — residue-level (CALVADOS2) and near-atomic (Martini3) coarse-grained simulations together with in vitro experiments — to establish that the foci-forming region (FFR) phase separates whereas the adjacent MUT-8-binding region (M8BR) does not, and used atomistic simulations of that non-phase-separating region to dissect client–scaffold recognition. In this way the multi-scale simulations helped to provide a molecular basis for previous in vivo observations by Uebel et al. (PLOS Genet. 2018; 14(7):e1007542). That study did not, however, resolve with atomic resolution the interactions within the phase-separated FFR condensate itself. The present study addresses precisely this gap: from 10 µs of atomistic molecular dynamics of the FFR condensate, we characterise the sub-µs contact dynamics and the protein–ion and protein–water interactions that govern the condensed phase at atomistic resolution — observables inaccessible to the coarse-grained models used previously, but key to understanding the properties of Mutator foci and ultimately how they underpin biological function in small RNA biology.

      Reviewer 1:

      (1) I have several questions regarding the system preparation that require clarification. The authors state that "65 copies of the coarse-grained MUT-16 FFR were embedded in a slab-shaped simulation," but it is not clear how this initial configuration was generated. Were the molecules randomly distributed in the simulation box, or were they initially arranged in a preformed condensate? Alternatively, were they randomly inserted and allowed to self-assemble into a condensate during NpT simulations? In Figure 1, the atomistic snapshot appears to show a well-defined condensate at the center of the simulation box. It would be important to clarify how this configuration was obtained: Was it generated from coarse-grained simulations starting from random initial conditions? Or was a preassembled condensate used as input? Related to this, how do the authors ensure that the simulations are equilibrated? While 20 μs appears to be a reasonably long simulation time for coarse-grained simulations, it would be useful to demonstrate equilibration explicitly. For example, the authors could plot the center-of-mass positions (in the long axis of the simulation box) of individual proteins over time to show that all molecules reach a steady state and remain within the condensate without systematic drift.

      We thank the reviewer for these important clarifying questions regarding system preparation and equilibration.

      The initial structure for the atomistic simulation was generated by randomly inserting 65 copies of the coarse-grained MUT-16 FFR into a slab-shaped simulation box using the gmx insert-molecules tool. The molecules were therefore not pre-arranged in a condensate; instead, they were allowed to spontaneously self-assemble from this random configuration during NpT simulations using the Martini3-IDP force field over 20 μs. The well-defined condensate visible in Figure 1 is thus the product of this unbiased self-assembly process.

      To make this workflow transparent to the reader, we will revise Figure 1 to include a two-panel illustration of the Martini3 simulation: a snapshot at t = 0 ns showing the randomly distributed chains, and a snapshot at t = 20 μs showing the assembled condensate, connected by an arrow indicating the subsequent backmapping step to the atomistic representation. We believe this will clearly communicate the sequential nature of the pipeline (random insertion → coarse-grained self-assembly → atomistic backmapping).

      We appreciate the concrete suggestion for demonstrating equilibration. We will add a supplementary figure showing the center-of-mass positions of individual protein chains along the long axis of the simulation box as a function of simulation time. This will allow readers to verify that molecules converge into the condensate phase and reach a steady state without systematic drift, providing explicit evidence that 20 μs coarse-grained simulation time is sufficient for equilibration under these conditions.

      (2) The authors experimentally observe UCST behavior for these condensates. Do the coarse-grained or atomistic simulations reproduce this behavior?

      While atomistic simulations may be too computationally demanding to systematically explore temperature dependence, coarse-grained simulations could be used to test whether condensates are stable at lower temperatures and dissolve at higher temperatures. Such an analysis would provide valuable support for the experimental observations.

      We thank the reviewer for this valuable suggestion. In previous coarse-grained simulations we have used a coarse-grained force field that does not capture UCST vs LCST behavior (Gaurav K et al. Biophys. J. 2025; 124:3987–4004). It will be very interesting to revisit these coarse-grained simulations with a coarse-grained simulation force field that can capture UCST and LCST behavior such as the Mpipi-T (Chakravarti & Joseph, Protein Sci 2025;34(10):e70284) and HPS-T models (Dignon GL et al. ACS Cent. Sci. 2019; 5(5):821–830). We plan to perform additional coarse-grained simulations at multiple temperatures using the HPS-T force field. The HPS-T model has been shown to capture UCST versus LCST behavior (Changiarath A et al. bioRxiv 2024) in accordance with previous in vitro experiments. These simulations will allow us to test whether the MUT-16 FFR condensates remain stable at lower temperatures and dissolve at higher temperatures, providing direct computational support for the experimentally observed UCST behavior. We will include this analysis in the revised manuscript.

      (3) Regarding the analysis of ions, several points could be clarified and extended:

      a) It would be helpful to report the total number of ions and quantify how many are located inside vs. outside the condensate. While qualitative trends can be inferred from density profiles, quantitative analysis would strengthen the conclusions.

      b) It would also be interesting to analyze the number of contact ion pairs (e.g., Na⁺-Cl⁻ pairs), as described in J. Chem. Phys. 156, 044505 (2022). It is known that some ion models tend to overestimate ion pairing and underestimate solubility (e.g., J. Chem. Phys. 153, 010903 (2020)).

      c) In this context, the use of scaled-charge models has been shown to improve the description of ionic solutions and biomolecular systems (e.g., J. Phys. Chem. Lett. 2019, 10, 23, 7531-7536). I would suggest that, at least for one trajectory, the authors perform a test simulation using scaled charges (e.g., scaling by ~0.8) to evaluate whether ion distributions and protein-ion interactions are significantly affected.

      We thank the reviewer for these insightful suggestions regarding the ion analysis. We agree that a more quantitative treatment of ion behavior would strengthen the manuscript. To address all three points collectively, we will expand the existing Figure S7 with additional panels. These will include quantitative counts of Na<sup>+</sup> and Cl<sup>-</sup> ions partitioning inside versus outside the condensate complementing the existing density profiles, the Na<sup>+</sup>–Cl<sup>-</sup> radial distribution functions to estimate contact ion pair populations following J. Chem. Phys. 156, 044505 (2022).

      Following the Reviewer suggestion we will run a simulation with scaled charges (~0.8 scaling factor, J. Phys. Chem. Lett. 2019, 10(23):7531–7536) to evaluate the sensitivity of our results to the choice of ion model. We will compare ion distributions obtained with standard versus scaled charges . We will discuss the contact ion pair results in the context of known force field limitations regarding ion pairing (J. Chem. Phys. 153, 010903 (2020)) and assess whether the scaled-charge treatment leads to any qualitatively different conclusions.

      (4) Finally, while the selected water model is known to be accurate, it would be useful to assess its performance for concentrated salt solutions. For example, the authors could estimate the density of a 6 m salt solution and compare it with experimental data or validated models (e.g., J. Chem. Phys. 151, 134504 (2019)). This would help clarify to what extent the conclusions depend on the chosen force field.

      We thank the reviewer for this important suggestion. We agree that while the chosen water model is well established for biomolecular simulations, its performance under concentrated salt conditions is a legitimate concern that is worth explicitly validating in the context of this work. We will perform a short bulk simulation of a 6 m NaCl solution and compute the solution density, comparing it to experimental data (J. Chem. Phys. 151, 134504 (2019)). This straightforward validation will allow us to quantify how well our water and ion force field combination reproduces the thermodynamic properties of concentrated salt solutions, and to transparently discuss any deviations and their potential implications for the ion partitioning and protein–ion interaction results presented in the manuscript. The results will be added to the supplementary information alongside the expanded ion analysis in Figure S7.

      (5) In the Introduction, it would be helpful to elaborate further on the possible driving forces of LLPS in this region. Are there prior hypotheses or evidence pointing to specific interactions (e.g., cation-π, π-π, electrostatic interactions)? While this work addresses these questions, a brief discussion of previous experimental or theoretical insights would provide useful context.

      We thank the reviewer for this helpful suggestion. We will expand the Introduction to briefly discuss the known molecular driving forces of LLPS in IDR-containing proteins. Specifically, we will discuss the role of π–π interactions between aromatic residues (Vernon et al. eLife 2018; 7:e31486), cation–π interactions between aromatic and positively charged residues such as tyrosine–arginine pairs, which have been experimentally demonstrated to drive condensate formation in proteins such as FUS (Qamar et al. Cell 2018; 173:720–734), and the broader sequence-encoded molecular grammar governing these interactions in prion-like RNA-binding proteins (Wang et al. Cell 2018; 174:688–699, Rekhi et al. Nat Chem 2024 16:1113–1124 ). We will discuss previous findings on how ions shape interactions in condensates (MacAinsh et al. eLife 2024; 13:RP100282). We will also note the contribution of electrostatic interactions arising from charge patterning within the IDR, and contextualize how these general principles apply to the specific sequence composition of MUT-16 FFR, motivating the simulation-based investigation presented in this work.

      (6) On page 18, the authors state: "MUT-16 FFR satisfies the length (172 residues), aromatic content (20.35%), and Arg enrichment (85.71%) criteria. Its charge content (10.47%) and charge balance (38.89% positive charge fraction) are slightly below the nominal thresholds." It would be very helpful to include a schematic representation of the protein sequence highlighting these features (aromatic residues, charge distribution, etc.) in the corresponding figure, to provide a more intuitive understanding.

      We thank the reviewer for this helpful suggestion. We will include a figure showing a schematic representation of the MUT-16 FFR sequence, with aromatic residues, charged residues (positive and negative), and arginine content highlighted.

      (7) A question regarding ion hydration: What is the coordination environment of the ions that bridge proteins? Are they still hydrated by water molecules, or does the reduced water content inside the condensate significantly affect their solvation. Typically, Na<sup>+</sup> and Cl<sup>-</sup> ions have coordination numbers around 5-6 in aqueous solution. Do protein interactions and reduced solvent conditions within the condensate alter this coordination? A brief analysis or discussion would be valuable.

      We will calculate the coordination numbers of Na⁺ and Cl⁻ ions that mediate residue–residue bridging interactions inside the condensate and compare them against ions in the bulk dilute phase. This will directly reveal the degree to which bridging ions retain or lose their hydration shell when engaging with protein residues, and whether the condensate environment meaningfully perturbs ion solvation. The results will be presented as an additional figure in the Supplementary Information.

      Reviewer 2:

      (1) The large amount of detail in the results section sometimes makes it difficult to identify the central take-home messages. I encourage the authors to more clearly highlight the principal findings and the physical insights that may generalize to other condensate-forming systems. The authors may also consider streamlining parts of the Results section to improve focus and readability.

      We thank the reviewer for this constructive feedback. We will revise the Results section by adding brief concluding remarks at the end of each subsection that explicitly state the key physical insight emerging from that analysis. We will consider which secondary findings can be moved to the Supplementary Information. We will also strengthen the Conclusion section to more clearly distil the principal findings of the study as a whole and highlight the broader insights that may generalize to other condensate-forming systems, ensuring the central take-home messages are clearly communicated to the reader.

      Reviewer 3:

      (1) In its current form, several technical issues need to be addressed before the main conclusions can be considered robust. Most importantly, the simulated sequence is 172 residues long, while the atomistic slab has box dimensions of only 12 nm in two directions. This length scale is comparable to the expected end-to-end distances of a disordered 172-residue chain. It is therefore not clear whether individual protein chains interact with their own periodic images, which could substantially affect overall chain dynamics and subsequently bias contact lifetimes, residue-residue interaction statistics, and the inferred condensate dynamics. The authors should check, for each chain, histograms of end-to-end distances. For chains for which more than ~2-3% of the end-to-end distances exceed ~11 nm, the authors should explicitly check for self-image interactions (for example, using "gmx mindist -pi") and report whether such interactions occur and for what fraction of the trajectory. Without this control, at least in the Supporting Information, I do not think the simulation-derived contact dynamics are sufficiently trustworthy.

      We thank the reviewer for raising this important point. Indeed the box size in x and y dimensions is only marginal, which may influence the dynamics in our simulations and could affect our conclusions. In response, we will perform a control simulation with a larger box, increasing the x and y dimensions to ~16 nm. We will compare the contact dynamics of the resulting trajectory with our original results. This control simulation is initiated from an independently assembled coarse-grained condensate (see our response to Question 6) and therefore also addresses the replica-independence concern raised there.

      (2) A second major concern is the treatment of ions. The manuscript makes important conclusions about Na<sup>+</sup> association and Na<sup>+</sup>-mediated bridging, but the atomistic ion model is not explicitly stated. This is a reproducibility problem and also affects interpretation - for example, standard Amber ions are known to bind too strongly to the oppositely charged residues. In their results, one acidic residue appears to interact on average with roughly two Na⁺ ions, which is not obviously expected from charge balance alone. The authors should state the exact Na<sup>+</sup>/Cl<sup>-</sup> parameters used, justify their compatibility with TIP4P-D and the protein force field, and explicitly interpret why such a strong Na<sup>+</sup> association with acidic residues is observed.

      We thank the reviewer for raising this important point. We will explicitly state in the Methods section how the Na<sup>+</sup> and Cl<sup>-</sup> ions, including the force field parameters of the ions, were modelled in our setup, and discuss its compatibility with TIP4P-D and the protein force field. In the presented simulations we have used the Joung and Cheatham parameters (Joung et al, J. Phys. Chem. B 2008, 112 (30), 9020–9041) with σ = 0.243934 nm and ε = 0.365846 (kJ mol<sup>-1</sup>) for Na<sup>+</sup> and σ = 0.447766 nm and ε = 0.148913 (kJ mol<sup>-1</sup>) for Cl<sup>-</sup>. While similar setups have been used, these ion parameters have not been optimized for TIP4P-D (originally developed for TIP3P water) and thus a lack of compatibility of the parameters could affect our conclusions.

      In response to the Reviewer and also in response to Reviewer 1 (Question 3), we will perform a sensitivity check by running an additional molecular dynamics simulation with scaled ion parameters as suggested by Reviewer 1 ( J. Phys. Chem. Lett. 2019, 10, 23, 7531-7536). In this way we will assess to what extent the degree of Na<sup>+</sup> association with acidic residues is sensitive to the choice of ion parameters and discuss the implications for our conclusions regarding Na⁺-mediated bridging interactions.

      (3) More generally, because the manuscript is centered on contact lifetimes, the choice of the atomistic force field needs stronger justification. Salt bridges, cation-pi contacts, pi-pi stacking, ion coordination, and water-mediated interactions are all force-field-sensitive. Since there is no direct experimental observable used here to validate the simulations, the authors should discuss the expected limitations of the chosen force field (while I do acknowledge that testing different force fields would be computationally too demanding).

      We thank the reviewer for this fair comment. We will add a short discussion justifying the choice of both TIP4P-D and Amber99sb-star-ILDN-q force field, discussing their performance for disordered proteins. We will explicitly acknowledge that absolute contact lifetime values should be interpreted with caution given the inherent force field sensitivities of salt bridges, cation-π, and π-π interactions, while relative trends and qualitative insights are expected to be more robust. We believe this transparent discussion will strengthen the manuscript and place our findings in the appropriate context for the reader.

      (4) I also find the sequence-comparison section somewhat confusing. The authors compare one specific IDR, MUT-16 FFR, with the average properties of human IDRs and then frame it as more representative than FUS LCD. It is not clear how informative this is because IDR behavior depends strongly on sequence-specific patterning, molecular connectivity, and the particular interaction network of each protein. Averages over human IDRs may provide a broad context, but they do not necessarily define what is physically or biologically representative for phase separation. In addition, FUS LCD is not intended to be a representative human IDR; it is an unusually low-complexity, phase-separating domain. Therefore, the "more representative than FUS" framing should be toned down. At most, this analysis shows that MUT-16 FFR is compositionally less extreme than FUS LCD.

      We thank the reviewer for this valid criticism. We agree that the framing of MUT-16 FFR as "more representative than FUS LCD" is an overstatement, and we will revise the text accordingly. The comparison against human IDR averages was intended to provide broad compositional context rather than make claims about functional or dynamical representativeness, and we will make this distinction explicit. We will reframe the statement to simply note that MUT-16 FFR is compositionally less extreme than FUS LCD, without implying broader representativeness, which as the reviewer correctly points out cannot be inferred from sequence composition alone given the strong dependence of IDR behavior on sequence-specific patterning and interaction networks.

      (5) The ion- and water-bridging analyses are also potentially overinterpreted. A distance-based simultaneous contact with two residues does not by itself establish functional mediation or regulation of condensate dynamics. The authors should either add appropriate controls, such as local-density-normalized baselines or randomized-contact expectations, or soften the language to describe these as geometrically defined co-contact events rather than mechanistic bridging interactions.

      We thank the reviewer for this valid point. We agree that distance-based co-contact events do not by themselves establish mechanistic bridging or functional regulation, and we will revise the manuscript language throughout to describe these observations as geometrically defined co-contact events rather than mechanistic bridging interactions. We will also explore appropriate controls such as local-density normalized baselines or randomized-contact expectations. In this respect we will also consider our results in light of a recent paper that showed that salt-bridges are overestimated in atomistic molecular dynamics simulations (Ivanović et al, JACS Au 2026, 6(3), 1900–1913). We will ensure the interpretation is appropriately cautious and does not overstate the mechanistic implications of these findings.

      (6) Finally, the independence of the atomistic replicas is unclear. The manuscript should state whether all ten all-atom simulations were initiated from the same coarse-grained condensate configuration or from distinct CG frames. If the starting structures came from one CG trajectory, the authors should report how far apart those frames were in simulation time and provide evidence that the initial atomistic configurations are structurally independent. If only velocities differ, the simulations should not be described as fully independent structural replicas.

      We thank the reviewer for this important clarification request. We confirm that all ten atomistic replicas were initiated from the same coarse-grained condensate configuration following backmapping, but were equilibrated independently using different random velocity seeds. Only the last 800 ns of each trajectory was used for analysis, discarding the initial 200 ns as equilibration. We will add these details explicitly to the Methods section and make clearer that these simulations are not fully independent structural replicas. We will report the overlap of residue–residue contact maps between replicas to provide an indication of how the contact statistics have decorrelated, given the shared starting structure.

      In response to this question and also question 1, we are initiating an all-atom simulation from an independently formed CG condensate (16 nm x 16 nm x 60 nm). This will provide a valuable check as to the conclusions from our ten initial simulation trajectories.

      References

      Blazquez S, Conde MM, Abascal JLF, Vega C. J. Chem. Phys. 2022;156(4):044505.

      Chakravarti A, Joseph JA. Protein Sci. 2025;34(10):e70284.

      Changiarath A, Flores-Solis D, Michels JJ, Herrera Rodriguez R, Hanson SM, Schmid F, Zweckstetter M, Padeken J, Stelzl LS. bioRxiv. 2024. doi:10.1101/2024.03.16.585180.

      Dignon GL, Zheng W, Kim YC, Mittal J. ACS Cent. Sci. 2019;5(5):821–830.

      Gaurav K, Busetto V, Páez-Moscoso DJ, Changiarath A, Hanson SM, Falk S, Ketting RF, Stelzl LS. Biophys. J. 2025;124:3987–4004.

      Ivanović MT, Holla A, Nüesch MF, von Roten V, Schuler B, Best RB. JACS Au. 2026;6(3):1900–1913.

      Joung IS, Cheatham TE III. J. Phys. Chem. B. 2008;112(30):9020–9041.

      Kirby BJ, Jungwirth P. J. Phys. Chem. Lett. 2019;10(23):7531–7536.

      MacAinsh M, Dey S, Zhou HX. eLife. 2024;13:RP100282.

      Panagiotopoulos AZ. J. Chem. Phys. 2020;153(1):010903.

      Qamar S, et al. Cell. 2018;173:720–734.

      Rekhi S, Garcia CG, Barai M, Rizuan A, Schuster BS, Kiick KL, Mittal J. Nat. Chem. 2024;16:1113–1124.

      Uebel CJ, Anderson DC, Mandarino LM, Manage KI, Aynaszyan S, et al. PLOS Genet. 2018;14(7):e1007542.

      Vernon RM, Chong PA, Tsang B, Kim TH, Bah A, Farber P, Lin H, Forman-Kay JD. eLife. 2018;7:e31486.

      Wang J, et al. Cell. 2018;174:688–699.

      Zeron IM, Abascal JLF, Vega C. J. Chem. Phys. 2019;151:134504.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors investigated the relationship between physical activity (PA) and both structural (MRI) and cognitive brain health in the LIFE-Adult Study, with total baseline recruitment of 2576. Hippocampal volume, an MRI-derived BrainAGE marker, and scores from the Trail Making Test were used as outcomes, with the majority of participants measured at baseline and subsets also measured in a follow-up session. The key findings were a lack of direct association between PA and outcomes, but longitudinal evidence for a higher BrainAge at baseline leading to lower physical capacity at follow-up. This supports a reverse-causation hypothesis in contrast to the prevailing understanding of the positive effects of physical activity on brain health.

      Strengths:

      The Life-Adult study is a rich and carefully acquired dataset, with multiple follow-up time points. The statistical analyses were conducted carefully with appropriate control for confounds and multiple testing. The study design enables an important assessment for reverse causality. The authors are scrupulous in their consideration of a number of factors that could potentially bias their results, performing an age-stratified analysis, and emphasising discrepancies in PA measurements (specifically, age-reporting bias) across the dataset and other limitations.

      Weaknesses:

      This is an observational study with inconsistent measures of physical activity. Previous studies have used physical activity interventions, and might be more strongly weighted when considering evidence for these effects (specific confounders involved in interventions notwithstanding).

      The model identifying potential reverse causality is relatively limited - it seems possible/likely that brainAge could reflect more general health status, which would expand the potential range of factors underlying this observation.

      The important quantitative actigraphy subset is small (n=227), as are the longitudinal subsets. Along with the discrepancy of physical activity/capacity at baseline and follow-up, and other complexities of the dataset, it is difficult to make firm conclusions. The authors point out that the actigraphy subset was quite inactive.

      We would like to thank the reviewer for their valuable feedback. We agree with the limitations mentioned, and we have extended the discussion section in order to address the drawbacks more effectively. In particular, we agree that the null findings of this study do not suggest that physical activity has no effect on the brain; for such a conclusion, an intervention study would be necessary.

      Furthermore, we agree that BrainAGE might reflect a more general health status. Although we excluded images of individuals with visible acquired brain injuries, we did not control for other medical conditions (e.g. hypertension or diabetes), which may have affected the results.

      Please see the revised discussion parts in the response below.

      Reviewer #2 (Public review):

      Summary:

      This population-based cohort study found no evidence that physical activity, whether self-reported or objectively measured, positively influenced brain structure (hippocampal volume or BrainAGE) or cognitive function (Trail Making Test scores). Notably, longitudinal analyses suggested the opposite temporal relationship: a higher BrainAGE at baseline predicted higher physical capacity at follow-up, more in line with reverse causation rather than a neuroprotective effect of physical activity.

      Strengths:

      The study's statistical approach is thorough and well-documented, and the inclusion of two measurements of physical activity (self-report questionnaire and objective accelerometer data) is a strength. The longitudinal aspect also represents a strength.

      Weaknesses:

      Several aspects of the measurement timing warrant consideration. Physical activity was assessed over 7-day periods, creating a potential mismatch with (commonly less dynamic) brain outcomes examined (hippocampal volume, BrainAGE), which may reflect cumulative exposures over longer timescales. Additionally, the asynchronous measurement protocol (cognitive testing preceding accelerometry, and the MRI occurring weeks after baseline visits) may introduce time lags that attenuate associations. The observed null associations may be influenced by timing misalignment rather than reflecting the absence of consistent effects of physical activity on brain health and cognition.

      Other measurement characteristics also warrant consideration when interpreting the null findings. Physical activity was assessed using short-form self-report questionnaires and averaged accelerometer MET/day values, both of which have limited reliability. Additionally, the modest accelerometer subsample size and low/insufficient variation in activity levels observed in this cohort increase the likelihood of missing effects. These factors collectively raise the possibility that true physical activity-brain health associations may have been obscured.

      The study's conclusions regarding brain health, structure, and cognitive functioning are broad despite the scope of the selection of outcomes examined. The analyses focus on hippocampal volume, BrainAGE (a global aging metric), and Trail Making Test performance (processing speed and executive function), while omitting other important neuroimaging markers such as cortical thickness, functional connectivity, or white matter microstructure. The null findings presented here cannot exclude positive effects of physical activity on broader constructs of brain health or cognitive functioning.

      While the authors appropriately note the use of different physical activity instruments across time points (IPAQ at baseline, VSAQ at follow-up) in the limitations section, the discussion should more explicitly address the interpretive challenges this creates. The observed association between higher baseline brain age gap and lower follow-up physical activity may reflect: (1) a true temporal relationship, (2) an artifact of switching from behavior-focused (IPAQ) to capacity-focused (VSAQ) measurement, or (3) some combination of both. This ambiguity substantially limits causal inference.

      Thank you for a thorough review of the manuscript. We appreciate the opportunity to consider the limitations in more detail. As you highlight, the null findings could be caused by a variety of reasons and changes may have occurred in white matter microstructure or functional connectivity that could not be observed using our chosen measures. We have expanded the discussion to address these and other issues (p. 12):

      “However, our results should be interpreted with caution, due to the limited sample size, potential attenuation of effects resulting from measurement error in the assessment of physical activity/capacity and the shift from an activity-based measure at baseline to a capacity-based measure at follow-up. This change limits the interpretability of longitudinal effects, as observed associations may reflect both changes in the underlying construct being measured and true changes in the relationship over time.

      Strengths and Limitations

      The results of this cross-sectional observational study may be affected by various factors, including bias in self-reported physical activity [48, 49], accelerometer measurement error [59, 60] and reverse causality, among others. Moreover, our results may also be affected by the general medical status of the participants, since we did not control for other diseases within the sample. In fact, BrainAGE may reflect overall health and the cumulative impact of various factors (including previous physical activity) on brain health over an extended period of time. Furthermore, our analysis focused on only a few cognitive and structural brain measures. While we did not observe any changes in hippocampal volume or BrainAGE, this does not exclude the possibility of changes in white matter integrity or functional connectivity. Another limitation of this observational study was the time lag between physical activity measurements and MRI scanning, which may have reduced the observed effects. Although the longitudinal design is a major strength of this study, attrition of participants at follow-up may have affected our estimates. Furthermore, the use of cross-lagged panel model design in the longitudinal setting has frequently been criticised for not distinguishing between within-person changes and between-person differences [61, 62], and our adapted design suffers from these limitations, as well as others arising from the use of different instruments to measure the construct related to physical activity at each time point (IPAQ and VSAQ). Nevertheless, compared to large volunteer-based cohorts such as the UK Biobank, the registry-based recruitment strategy of the LIFE-Adult Study may be less susceptible to healthy volunteer bias, although we cannot entirely eliminate the possibility of volunteer bias among the participants with accelerometry data in our case.

      Direct comparisons are limited in the absence of harmonised recruitment and assessment protocols.”

      Additionally, please note that in the Summary sentence ‘Notably, longitudinal analyses suggested the opposite temporal relationship: a higher BrainAGE at baseline predicted higher physical capacity at follow-up, more in line with reverse causation rather than a neuroprotective effect of physical activity’

      The opposite is actually true; higher BrainAGE at baseline predicted lower physical capacity at follow-up.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The analysis and discussion are somewhat limited. More detail and discussion of the demographic features of the study dataset, and perhaps a stronger concluding position regarding the potential impacts of PA on brain and cognition would be helpful - this might also be integrated into the abstract.

      Thank you for pointing this out. As both of the reviewers have highlighted that the discussion is limited, we appreciate the opportunity to revise it (see also the response above). We hope that it now gives a more comprehensive interpretation of the results.

      We have also updated our abstract to include a bit stronger concluding position:

      “Physical activity is believed to positively influence brain health and cognition and is considered a modifiable lifestyle factor that may protect against cognitive decline and neurodegeneration. In this observational study, we investigated the cross-sectional and longitudinal effects of self-reported total and moderate-to-vigorous physical activity on cognitive scores on the Trail Making Test (TMT-A and TMT-B), hippocampal volume, and BrainAGE, in a large population-based cohort from the LIFE-Adult Study (n = 2576). Furthermore, we examined the effect of objectively measured physical activity on brain structure in a subgroup with available accelerometry data (n = 227). Multiple linear regression analyses did not show any positive effects of self-reported or objectively measured physical activity on hippocampal volume or processing speed and executive function. Longitudinal path analyses suggested a potential for reverse causation, where a higher BrainAGE at baseline was associated with lower physical capacity at follow-up. Additionally, we observed an age-related bias in the self-reporting of physical activity, indicating that older individuals tend to overestimate their level of activity. Future interventions targeting middle-aged adults may be necessary to raise awareness of potential misperception and encourage increased physical activity.”

      There might be more careful inspection of alternative models and dissection of the impact of covariates (e.g. smoking, which is very prevalent in this cohort). For example, did PA show any benefit specifically in the "non-smoker" vs. "smoker" subgroups?

      Thank you for this suggestion. We decided against including analysis of various subgroups, as this would have shifted the focus of the manuscript. However, we do provide the results of the analysis with the interaction term here. There was no evidence that smoking status moderated the association between self-reported physical activity and BrainAGE at follow-up (p = .192). In both non-smokers and smokers, physical activity was not significantly associated with BrainAGE at follow-up (b = 0.038, SE = 0.033, p = .242 and b = −0.028, SE = 0.038, p = .471, respectively).

      The age-dependent reporting bias seems important and should be assessed and discussed in more detail - it could have important implications for other studies. Why might this occur?

      We appreciate your drawing more attention to this point. As you have mentioned, it can have important implications for other studies, suggesting that objective measures of activity should, if possible, be used alongside self-report questionnaires. We have expanded on this topic in more detail in the revised discussion (p. 11):

      “This age-dependent reporting bias was previously demonstrated by other studies, where higher age was associated with overreporting activity levels [51-54]. Overreporting could stem from worsening recall, socially desirable responses, and the subjective nature of self-report questionnaires, which also depend on a person’s physical fitness [51]. Future (observational) studies would greatly benefit from including both accelerometer and self-reported measures of physical activity.”

      The reverse causation result could also be discussed (and possibly analysed) in more detail - what might the neurobiological mechanisms underlying this be? Is general health a factor - were confounds like smoking/health assessed here?

      Thank you for raising this important point. We have not added confounds other than age here, as we did not have enough degrees of freedom to add additional parameters to the model. However, we agree with you that general health might have played an important role here, although we don’t have a specific measure for it. We have expanded upon these limitations in the revised discussion (p. 12):

      “Similarly to Hofman et al. [62] and Rodriguez-Ayllon et al. [63], who found a bidirectional association between physical activity and brain structure, with a more consistent pattern of brain structural measures affecting physical activity, the results of our path analysis partially supported the reverse causality explanation, indicating that baseline brain health influences follow-up physical capacity, rather than baseline physical activity affecting follow-up brain health. Possible mechanisms may involve decreasing health status with age-related mitochondrial dysfunction [64, 65] and potential low-grade inflammation, which could result in fatigue [66] and a possible decline in fitness and physical capacity. Further studies are necessary to investigate this in more detail.”

      Results should contain greater detail; in particular, they should summarise key results from tables and not rely on the reader to carefully look through all figures. Reporting relatively non-informative results (e.g. entirely unadjusted model results) does not add much.

      We appreciate your feedback. We have tried to summarise the results in greater detail, reporting statistical values within the text (e.g., main results, p. 7):

      “The results indicate no statistically significant effects of self-reported total PA on brain structure (β = -0.029, p = .137 and β = 0.035, p = .137 for hippocampal volume and BrainAGE, respectively). There is a statistically significant effect of total self-reported PA on cognitive function, indicating that higher levels of PA lead to higher time scores on TMT-B (β = 0.053, p = .042). Similar results can be observed for MVPA, however, the results do not survive the correction for multiple comparisons (β = 0.045, p = .072). The analysis of objectively measured PA indicated no statistically significant effects on brain structure (β = -0.036, p = .582 and β = -0.076, p = .393 for hippocampal volume and BrainAGE, respectively).”

      We report the results of the unadjusted model to provide greater transparency, in line with the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guidelines.

      Reviewer #2 (Recommendations for the authors):

      Please review the manuscript for consistent use of abbreviations and definitions (e.g., write out BDNF line 49). Specifically, definitions and language to describe BrainAGE, brain age, brain age gap, neuroimaging-derived biomarker of brain ageing, Brain Age Gap Estimate, BA, BrainAGE (BA), etc. would benefit from consistency.

      Thank you very much for this observation. We have revised the manuscript in the hope that it is now more consistent. Specifically, we spelled out the ‘brain-derived neurotrophic factor’ and replaced the abbreviation 'BA' in Figure 2 with BrainAGE. However, we acknowledge that there are already many naming inconsistencies within the field. We therefore tried to follow the different authors' notations, which were established within the field: we used 'brain age gap' when referring to James Cole's model (or the concept in general), and 'BrainAGE' for our own.

      We thank both reviewers for their helpful feedback, which has improved our manuscript.

    1. Author response:

      We thank the reviewers for their careful and constructive assessment. We are glad they found the theory well motivated, that they recognised it unifies the previously unexplained center–surround suppression and facilitation in MT, and that they appreciated the methodological innovation in mapping ideal-observer predictions onto neural responses. We will make the manuscript more self-contained and mathematically explicit.

      To clarify our central claim and the connection to neural data: Our model can account for both what people perceive and what neurons do. Specifically, we take a Bayesian causal-inference model that was built and fitted to human behaviour (Shivkumar et al., 2025), and use it to derive neural predictions for center–surround interactions in area MT during motion perception. We then compare these predictions to previously reported MT single-neuron responses – qualitatively, but without any further parameter fitting.

      The reviewers are correct that there are two types of latent variables in our model, which imply two different sets of neural predictions. While one might conjecture relationships to other neural properties like classic center-surround suppression, a separate determination of which latent variable a neuron corresponds to is a simple matter of model comparison after fitting its responses to both. If the responses of a recorded neuron correspond to one of these predictions (as many existing neurons appear to do as we show in our paper), then this constitutes evidence in favor of them representing the corresponding posterior in our model. On the other hand, if they do not, then this can be due to a number of factors: our generative model being wrong, the neural encoding assumption (sampling or LDC) being wrong, or the Bayesian brain hypothesis being wrong (also see Lengyel et al. 2023; Haefner et al. 2024).

      Furthermore, we’d like to also clarify that both types of latents support each of the perceptual states (including integration and segmentation) and that there is no 1-1 correspondence between them. Importantly, the related velocity latent represents the velocity in the inferred reference frame – which may be the surround, or the retinal, or an intermediate reference frame (see Fig. 5 in Shivkumar et al. 2025). As a result, the same neuron can show both suppressive and facilitatory effects (yellow and blue regions in the difference panels in Fig. 6 of our paper).

      Finally, while specifying both the likelihood and the prior constitutes our model definition, we have made reasonable assumptions about the shape of each. The physics of the world — for instance, that objects tend to be stationary or to move slowly — motivates a spike-and-slab prior (Knill & Richards, 1996), and the likelihood is well described by a unimodal form (Stocker & Simoncelli, 2006). Our qualitative predictions do not depend on the exact specification of the prior and likelihood; other unimodal likelihoods yield similar results.

      We will make corresponding edits throughout the text to clarify each of these points.

      References:

      Haefner, R. M., Beck, J., Savin, C., Salmasi, M., & Pitkow, X. (2024). How does the brain compute with probabilities? arXiv.

      Knill, D. C., & Richards, W. (Eds.). (1996). Perception as Bayesian inference. Cambridge University Press.

      Lengyel, G., Shivkumar, S., & Haefner, R. M. (2024). A general method for testing Bayesian models using neural data. In Proceedings of UniReps: The First Workshop on Unifying Representations in Neural Models (Proceedings of Machine Learning Research, Vol. 243).

      Shivkumar, S., DeAngelis, G. C., & Haefner, R. M. (2025). Hierarchical motion perception as causal inference. Nature Communications, 16, Article 3868.

      Stocker, A. A., & Simoncelli, E. P. (2006). Noise characteristics and prior expectations in human visual speed perception. Nature Neuroscience, 9(4), 578–585.

    1. Author response:

      We would like to express our sincere gratitude for your time and constructive feedback. We are highly encouraged by the positive assessment highlighting the solid evidence and convincing methods of our study. We also deeply value the insightful and constructive comments regarding our conceptual framing, the integration with established ecological theories, and the underlying dynamic mechanisms. We believe that incorporating these excellent suggestions will substantially enhance the conceptual clarity and theoretical depth of our manuscript. To achieve this, we are fully committed to conducting a comprehensive revision to address all the points raised. Below, we outline our main strategies for the forthcoming revision:

      (1) Structural Reorganization

      We fully agree with the reviewers and the Editor that the manuscript's structure requires improvement. We are especially grateful to Reviewer 1 for providing such a detailed and constructive roadmap for the revision. We will adopt all of the suggested changes. Specifically, in the revised manuscript, we will:

      (1.1) Rewrite the Abstract: provide a clearer introduction to the "Tragedy of the Commons" and a more accessible description of our modeling framework.

      (1.2) Establish a dedicated Methods section: move the core model equations, key assumptions, parameter choices, and the detailed explanation of our graph-theoretic framework (the Benefit Transfer Graph) from the Supplementary Information (SI) into the main text.

      (1.3) Restructure results and figures: We will reorganize the Results section to improve the logical flow. As suggested, we will split the current Figure 2, move critical diagrams from the SI into the main text, and expand our figure captions to ensure all data representations are immediately clear.

      (2) Reframing the Conceptual Framework and Terminology

      We thank Reviewer 1 for the insightful critique regarding the use of the term “cheating.” We have reflected on our previous phrasing and fully agree that "cheating" introduces an unnecessarily humanized judgment and conflates pure exploitation with metabolic generalism. To ensure mechanistic accuracy and alignment with recent ecological literature, we will systematically update our terminology throughout the text, from title to supplement:

      (2.1) Species strategies: "Pure-producers," "partial-producers," and "pure-cheaters" will be redefined as "single-receptor producers," "multi-receptor producers," and "non-producers," respectively.

      (2.2) Receptor types: "Cheating-receptors" will be renamed to "exogenous-receptors" (or foreign-receptors, exploitative-receptors) to objectively describe the uptake of siderophore types that are not produced by the focal microbe.

      (2.3) Updating the key parameter: To avoid ambiguity regarding synthesis versus uptake, we will rename "Cheating Breadth (CB)" to "Siderophore Exploitative Breadth (SEB)," defined strictly as the number of distinct exogenous-receptors expressed by a species.

      (2.3) Updating the core paradigm: We will reframe "The Paradox of Cheating" to "The Paradox of Siderophore Exploitation." We will clarify that the transition to high-diversity coexistence is not driven by "cheating", but by the topological connectivity of the mBTG.

      (3) Contextualizing within BQH and Hypercycles

      We sincerely thank the Editor and Reviewer 2 for highlighting the connections between our work, the Black Queen Hypothesis (BQH), and Hypercycle theory. We will dedicate a new section in the Discussion to thoroughly compare our siderophore-mediated network with these established frameworks.

      We will explicitly discuss the key similarities and differences. While the exploitation of siderophores in our model resembles the producer-beneficiary dependency described in BQH, the evolutionary drivers are distinct, in that BQH is primarily driven by the adaptive loss of costly genes (reductive evolution), whereas siderophore exploitation is driven by the acquisition of exogenous-receptors (e.g., via horizontal gene transfer). More importantly, the high diversity and lock-and-key specificity of siderophore-receptor interactions, renders each siderophore a "mixed good." This dynamic can actually drive the community into a Red Queen-like arms race, as suggested by the high probability of oscillatory dynamics observed in our simulations. Although we did not explicitly consider genetic mutations in the current ecological framework, unidirectional exploitation typically drives the involved species to extinction; consequently, the system naturally selects for communities where exploitation is reciprocated, organically giving rise to closed, distributed loops of benefit transfer.

      In the revised text, we will cite recent theoretical progress on structured and multi-goods BQH networks. We will also discuss how our topological loops link to Eigen's Hypercycle theory by illustrating how specific structures of exploitative interactions foster community diversity.

      (4) Addressing Siderophore Exploitative Breadth (SEB) Interpretations

      (4.1) The biological realism of the SEB range

      Both reviewers raised insightful questions regarding the settings and impacts of SEB (previously "CB"). While some of these questions will be addressed through new control simulations, we would like to immediately clarify the biological realism of the SEB parameter, particularly addressing Reviewer 1's concern about the simultaneous expression of multiple receptors.

      We completely agree that possessing a vast genomic repertoire of siderophore receptors does not mean a microbe expresses all of them simultaneously. Receptor expression in nature is a highly regulated and substrate-specific process. In Gram-negative bacteria like Pseudomonas, the expression of exogenous-receptors is tightly regulated by cell-surface signaling pathways (e.g., ECF sigma/anti-sigma factor systems). Under iron-limited conditions, a specific receptor is upregulated only when it detects its corresponding siderophore in the environment. Based on our literature review, while a bacterium may not express 30 receptors at once, expressing a substantial subset (e.g., 5–15) is biologically realistic. Therefore, in our model, SEB does not represent a static genomic capacity, but rather the number of active receptors that actually have corresponding siderophore producers present within the local community. We extended the SEB axis up to 30 in our initial figures primarily to capture the complete theoretical phase transition. However, following the reviewer's excellent suggestion, we will adjust the x-axis in our primary revised figures to highlight the more realistic regime (e.g., SEB 0–15) and add a dedicated paragraph detailing these biological regulatory mechanisms, with appropriate citations.

      (4.2) Disentangling the receptor allocation trade-off

      We highly appreciate Reviewer 2’s perceptive insight regarding the confounding effect: under a normalized allocation scheme, increasing SEB inevitably decreases the expression level of the self-receptor, thereby reducing self-reliance. We completely agree that explicitly addressing this trade-off is crucial.

      Biologically, this strong trade-off is realistic: receptor operations are energetically costly, and the initiation of their expression requires competing for a finite pool of RNA polymerase core enzymes. Therefore, investing in the capacity to exploit heterologous siderophores inherently incurs a cost to self-reliance. To rigorously test whether our central paradox is merely an artifact of this specific trade-off, we immediately initiated a series of control simulations. In these new models, we mathematically decoupled the variables by fixing the allocation fraction of the self-receptor as a constant.

      We are encouraged to report that our preliminary results support the core of the original paradox. Even when self-reliance is mathematically maintained, community-level extinction risk and the biodiversity of surviving communities remain positively linked. Interestingly, these controlled simulations exhibit an even clearer non-monotonic pattern, where both diversity and extinction risk peak at a biologically realistic SEB of approximately 5. This suggests that the paradox is fundamentally driven by network topology changes rather than the allocation trade-off alone:Viewed through our maximal Benefit Transfer Graph (mBTG) framework, a higher probability of non-self-directed edges in the mBTG forces the community to "gamble" between collapsing into a Sink Core or surviving in a high-diversity Cyclic Core. We are currently performing exhaustive simulations to gather detailed statistics on this decoupled model, particularly the non-monotonic behavior, which will be prominently featured in the revised manuscript.

      (5) Evolutionary Stability and Topological Resilience

      We also deeply appreciate Reviewer 2’s insightful critique regarding the evolutionary stability of our proposed cyclic networks, particularly their potential vulnerability to self-serving or short-circuit mutants that bypass intermediate species in a loop.

      To rigorously address this, we are currently conducting invasion simulations in which established communities are challenged by randomly generated mutant species. While the exhaustive computational analysis is ongoing, our preliminary results suggest the absence of a strict, static Evolutionarily Stable Strategy (ESS). Instead, the topological space fosters complex, intransitive competition. Intriguingly, these early data suggest that communities exhibiting oscillatory dynamics are actually more robust against invaders than those at a stable equilibrium. We intend to explore this phenomenon fully.

      Furthermore, we will expand our Discussion to address the implications of longer evolutionary timescales. When true structural mutations occur (e.g., the appearance of novel siderophore-receptor pairs to evade existing exploitation), the system will likely transition into a continuous Red Queen regime of ongoing molecular arms races. We will thoroughly discuss these evolutionary horizons and present our complete invasion simulation data in the revised manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      Amyotrophic lateral sclerosis (ALS) affects nerve cells in the brain and spinal cord. The authors' approach to use genetic code expansion to tag two ALS proteins associated with stress granules has value and should be useful in the ALS field. Parts of the work are well done, but there are concerns that the evidence is incomplete overall, and additional controls would strengthen the study.

      We thank the editors and reviewers for their thoughtful assessment and for highlighting the potential value of applying genetic code expansion (GCE) to study ALSassociated proteins involved in stress granule biology. Our goal in this work was to establish and validate a minimally perturbative labeling strategy using the noncanonical amino acid Anap to monitor the localization and stress-dependent behavior of TDP-43 and G3BP1.

      We agree that additional controls can further strengthen the conclusions. In the revised manuscript, we have clarified the experimental design and added essential controls to better support the reliability of the Anap labeling approach (Supplementary Fig. 1).

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors utilize genetic code expansion to tag TDP-43 and G3BP1, and evaluate this protein tagging system (ANAP) compared to antibodies, and evaluate protein trafficking and stress granule formation in response to stress with sodium arsenite treatment. They find similar staining to antibodies in HeLa cells, mouse embryonic stem cells, and primary mouse cortical neurons. This is a useful study that demonstrates the utility of ANAP tagging to evaluate ALS proteins.

      We sincerely thank the reviewer for the positive assessment of our work and for recognizing the utility of the Anap-based GCE system for studying ALS-associated proteins.

      Strengths:

      Rescue of cell survival by ANAP-tagged TDP-43 is compelling

      We appreciate the reviewer’s highlighting of this point. Demonstrating that TDP43-Anap can rescue cell survival was an important validation in our study, as it indicates that incorporation of the noncanonical amino acid does not substantially disrupt the biological function of TDP-43. Additionally, we also tested the RNA splicing function recovery potency of TDP-43-Anap. As shown in Fig. 1K and 1L, a recovery of expression of PFKP, a protein undergoing cryptic exon when TDP-43 lost its function [1], was observed when expressing TDP-43-Anap in TDP-43 knockout Hela cells.

      Weaknesses:

      While the ANAP-tagged proteins had similar distributions to antibody staining, there were some discrepancies that may be more explained by the technique than by novel findings, as the authors suggested. The inclusion of additional controls to evaluate this would be helpful.

      This is a helpful suggestion. To ensure that the fluorescence signal observed in our experiments was specifically derived from site-specific Anap incorporation rather than background fluorescence, we performed three control conditions. Specifically, we tested: (1) cells cultured with Anap supplement, (2) cells expressing the Anap incorporation system with the addition of Anap, and (3) cells expressing both the TAG-mutated protein plasmid and the Anap incorporation system but without the addition of Anap. These control experiments were performed for both TDP-43 and G3BP1, and no observable fluorescence signal was detected under any of these conditions (Supplementary Fig. 1). We have clarified this control experiment in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Chen and colleagues describe a novel means of labeling two RNAbinding proteins, G3BP1 and TDP-43, using genetic code expansion. Overexpressed constructs that incorporate the intrinsically fluorescent non-canonical amino acid Anap redistribute to cytoplasmic granules upon application of external stressors such as sodium arsenite. Similar labeling and redistribution of overexpressed G3BP1 and TDP43 were observed in cultures of mouse primary neurons.

      We are grateful for the reviewer’s accurate summary of our study and recognition of the value of GCE strategy for labeling the RNA-binding proteins G3BP1 and TDP-43.

      Strengths:

      Genetic code expansion and non-canonical amino acid labeling have quite a few advantages over traditional fusion proteins for tracking protein redistribution in living cells. The authors show that they are able to label exogenous G3BP1 and TDP-43 with the non-canonical amino acid Anap and follow labeled proteins in living cells with and without stress.

      We acknowledge the reviewer’s comment on the advantages of GCE-based noncanonical amino acid labeling for studying protein dynamics in living cells.

      Weaknesses:

      The authors do not convincingly leverage the advantages of genetic code expansion in the current study. There is no specific question posed by the authors that can be or is answered using this approach, and several of the experiments lack critical controls. This is also not the first example of TDP-43 labeling by genetic code expansion (see PMID: 38290242). As a result, the study as a whole adds little to our understanding of protein trafficking and behavior under stress.

      We thank the reviewer for raising these important points. Although as reviewer mentioned, genetic code expansion has previously been applied to TDP-43 [2], it mainly employed the photocaged lysine incorporation system to optogenetic control of TDP-43 translocation, and the protein was still labeled by mRubby. Our paper has totally different goal, to establish and validate a minimally perturbative labeling strategy using the intrinsically fluorescent noncanonical amino acid Anap to monitor the localization and stress-dependent behavior of both TDP-43 and G3BP1. And our work extends this approach in several important ways.

      First, we demonstrate that Anap incorporation enables visualization of stress-dependent redistribution of both TDP-43 and G3BP1, two key proteins involved in stress granule biology. Importantly, we validate this approach across multiple cellular systems, including HeLa cells, mouse embryonic stem cells, and primary mouse cortical neurons, which broadens the applicability of this labeling strategy.

      Second, we provide functional validation of the Anap-tagged protein, showing that TDP43-Anap rescues both cell survival and RNA splicing activity in TDP-43 knockout cells, including restoration of PFKP expression, a known cryptic exon target of TDP-43. These results support that Anap incorporation does not substantially disrupt protein function.

      We performed additional control experiments to ensure the specificity of the labeling system. Specifically, we tested three control conditions: (1) cells cultured with Anap supplement, (2) cells expressing the Anap incorporation system with the addition of Anap, and (3) cells expressing both the TAG-mutated protein plasmid and the Anap incorporation system but without the addition of Anap. These control experiments were performed for both TDP-43 and G3BP1, and no observable fluorescence signal was detected under any of these conditions (Supplementary Fig. 1).

      We agree that the manuscript would benefit from clearer articulation of the advantages of genetic code expansion in this context. Accordingly, we have revised the manuscript to more explicitly emphasize how Anap labeling provides a minimally perturbative alternative to large fluorescent protein fusions, which can alter the phase behavior and localization of stress granule proteins.

      “Conventional fluorescent protein tags have enabled visualization of TDP-43 and G3BP1 in living cells; however, these approaches can perturb the native biophysical properties of the proteins being studied. For example, GFP or other fluorescently tagged TDP-43 usually requires additional modifications, such as deletion of the nuclear localization signal (NLS) [3, 4], to induce cytoplasmic inclusion formation. Such manipulations introduce non-physiological conditions that may alter the native trafficking and aggregation behavior of TDP-43. As for G3BP1, tags like GFP may also cause unexpected effects on the phase separation or other dynamics of the protein. In contrast, Anap based GCE strategy allows the minimally perturbative labeling and visualization of protein localization and stress-induced redistribution while preserving native protein architecture and function of both proteins. Importantly, the approach provides a generalizable genetically encoded platform for quantitatively examining the behavior of ALS-associated proteins in living cells. By enabling faithful monitoring of protein trafficking and stressgranule dynamics without extensive protein engineering, Anap-based GCE can offer a powerful strategy for probing molecular-scale mechanisms underlying ALS-linked proteinopathies”.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 1A

      The authors report that the nuclear staining of G3BP1 by ANAP labeling shows the presence of nuclear pools of G3BP1 that aren't detected with antibody staining. However, unspecific nuclear staining by aminoacylated tRNAs bound to synthetases has been described. It would be important to have a control to evaluate for this possibility.

      This is an important point. We agree that the nuclear ANAP signal should be carefully controlled to exclude the possibility of nonspecific staining arising from the Anap incorporation machinery itself, such as aminoacylated tRNAs and/or synthetases.

      To address this concern, in methods and material part, we note that after DPBS washes to remove excess Anap, cells were incubated in fresh medium for 2 hours to allow sufficient time for the decay of unstable aminoacylated tRNAs, which are generally cleared within minutes to tens of munites [5].

      Also, we performed three control conditions for both TDP-43 and G3BP1: (1) cells cultured with Anap supplement, (2) cells expressing the Anap incorporation system with the addition of Anap, and (3) cells expressing both the TAG-mutated protein plasmid and the Anap incorporation system but without the addition of Anap. Under all three conditions, we observed no detectable fluorescence signal (Supplementary Fig. 1).

      In addition, as shown in Fig. 1I, the nuclear signal of G3BP1-Anap partially colocalizes with the nuclear signal of TIA-1 in several condensate-like structures. This observation further supports that the nuclear Anap signal reflects protein-associated localization rather than nonspecific fluorescence, as it overlaps with a known RNA-binding protein that can form nuclear condensates under certain conditions.

      (2) Figure 1A, 1B

      Anap labeling appears to stain fewer cytoplasmic structures compared to antibody staining for both G3BP1 and TDP-43 after sodium arsenite treatment. Quantification would be useful to address whether this is the case. If so, might this be due to unincorporated/truncated proteins competing with Anap-labeled proteins?

      We appreciate the reviewer’s helpful suggestion. To address this point, we performed quantitative colocalization analysis using Fiji/ImageJ, calculating the Pearson correlation coefficient (R) for regions of interest between the Anap signal and antibody staining. These analyses indicate a strong overall agreement between the two detection methods under stress conditions, supporting that Anap labeling reliably reports the localization of both G3BP1 and TDP-43 (see Fig1. A, B).

      Regarding the possibility that truncated or unincorporated proteins could influence the observed signal, we note that fluorescence from Anap depends on successful amber suppression and incorporation of Anap at the engineered TAG site. Proteins that fail to incorporate Anap, such as truncated products generated by premature termination, would not produce fluorescence, and therefore would not contribute to the Anap signal. Thus, the Anap fluorescence selectively reports the population of successfully labeled full-length proteins, whereas antibody staining detects both labeled and unlabeled protein pools. This difference may partially explain why antibody staining appears to label a larger number of cytoplasmic structures.

      (3) Figure 1F

      FRAP of G3BP1-GFP in stress granules is slower than in previous publications. The underlying reasons for this should also be addressed.

      We thank the reviewer for this important observation. Differences in FRAP recovery kinetics of G3BP1 in stress granules may arise from several experimental variables that are known to influence stress granule dynamics. These include differences in cell type, expression levels of G3BP1-GFP, and imaging or photobleaching parameters. In our experiments, FRAP measurements were performed under specific conditions optimized for our experimental system, which may lead to recovery kinetics that differ from those reported in previous studies.

      (4) Figure 1H

      A full-size Western blot would be useful to evaluate for amount of truncated protein for G3BP1 and TDP-43. Could truncated proteins be competing with and altering ANAPtagged G3BP1 and TDP-43 localization in response to stress? This should be addressed.

      We acknowledge this important point. Full-size Western blotting can provide information on the overall presence of truncated species in the transfected population; however, it represents a bulk measurement and does not capture cell-to-cell variability in amber suppression efficiency at the single-cell level. We therefore cannot exclude the possibility that truncated products are present at varying levels in individual cells and may contribute, directly or indirectly, to differences between antibody staining and Anap fluorescence.

      Importantly, we observe that cells with successful Anap incorporation consistently exhibit strong antibody staining for TDP-43 or G3BP1, indicating that full-length protein is the predominant species in these cells. Because Anap fluorescence depends on successful amber suppression, it selectively reports the full-length protein population, whereas truncated products are not detected in the imaging assay. The concordance between Anap fluorescence and antibody staining therefore argues against a major contribution of truncated species to the observed localization patterns (Supplementary Fig. 1).

      Accordingly, we interpret the Anap signal as reflecting the localization of successfully labeled full-length protein, while acknowledging that heterogeneity in suppression efficiency is an important limitation of the current approach.

      (5) Figure 3

      This is a well-designed diagram.

      We are grateful for the reviewer’s positive feedback on the diagram and are pleased that the schematic effectively illustrates the experimental design and the principles of the genetic code expansion strategy used in this study.

      Reviewer #2 (Recommendations for the authors):

      The authors present a one-sided viewpoint concerning the connection between stress granules and disease (lines 45-46). A more balanced discussion is recommended, including data arguing against a role for abnormal stress granules in neurodegeneration.

      This is an important suggestion. We agree that the relationship between stress granules and neurodegeneration remains an active area of investigation and that evidence both supporting and questioning a causal role of stress granules in disease has been reported. In the revised manuscript, we have modified the Introduction to provide a more balanced discussion of this topic.

      “Altered stress-granule dynamics have been associated with ALS/FTD [6, 7]; however, whether stress granules directly drive neurodegeneration remains debated, as several studies suggest that stress granules primarily function as protective stress responses [8].”

      (1) A central rationale for the study is missing. The authors state only that G3BP1 and TDP-43 'undergo dynamic stress-dependent redistribution, making them ideal candidates for minimally invasive, site-specific fluorescent labeling.' Is there a controversy or question that can be resolved using these approaches?

      We thank the reviewer for raising this important point. The central motivation of this study is that the dynamic behavior and phase separation properties of stressgranule proteins are highly sensitive to protein modifications and tagging strategies.

      “Conventional fluorescent protein tags have enabled visualization of TDP-43 and G3BP1 in living cells; however, these approaches can perturb the native biophysical properties of the proteins being studied. For example, GFP or other fluorescently tagged TDP-43 usually requires additional modifications, such as deletion of the nuclear localization signal (NLS) [3, 4], to induce cytoplasmic inclusion formation. Such manipulations introduce non-physiological conditions that may alter the native trafficking and aggregation behavior of TDP-43. As for G3BP1, tags like GFP may also cause unexpected effects on the phase separation or other dynamics of the protein.”

      (2) Related to this, there is little context for how or why genetic code expansion is utilized for these studies

      We agree that the rationale for using genetic code expansion should be more clearly explained. In this study, genetic code expansion was employed to enable sitespecific incorporation of the small fluorescent noncanonical amino acid Anap, allowing minimally perturbative labeling of proteins of interest.

      “Anap based GCE strategy allows the minimally perturbative labeling and visualization of protein localization and stress-induced redistribution while preserving native protein architecture and function of both proteins. Importantly, the approach provides a generalizable genetically encoded platform for quantitatively examining the behavior of ALS-associated proteins in living cells. By enabling faithful monitoring of protein trafficking and stress-granule dynamics without extensive protein engineering, Anapbased GCE can offer a powerful strategy for probing molecular-scale mechanisms underlying ALS-linked proteinopathies.”

      (3) The justification for the criteria for selecting the site for incorporation of non-canonical amino acids in G3BP1 or TDP-43 is missing.

      We acknowledge this important comment and agree that the rationale for selecting the incorporation sites should be stated more clearly.

      “For TDP-43, the incorporation site was selected to avoid the major functional domains involved in RNA binding, nuclear localization, and aggregation-related behavior, thereby reducing the likelihood that Anap incorporation would perturb its native trafficking or function. For G3BP1, the selected site was chosen to minimize interference with domains important for stress granule assembly, RNA binding, and protein-protein interactions. More generally, we aimed to place the ncAA at positions likely to be solventaccessible and tolerant of substitution, while avoiding highly conserved or functionally essential residues.”

      (4) Studies in Figures 1 and 2 lack essential controls, including background signal from Anap in non-transfected cells, or those transfected with plasmids lacking the tRNA or tRS.

      This is an important point, also raised by Reviewer 1. To evaluate potential background fluorescence arising from Anap or the labeling system, we performed several control experiments. Specifically, we examined three conditions: (1) cells cultured with Anap supplement, (2) cells expressing the Anap incorporation system with the addition of Anap, and (3) cells expressing both the TAG-mutated protein plasmid and the Anap incorporation system but without the addition of Anap. Under all three conditions, we observed no detectable fluorescence signal (Supplementary Fig. 1).

      (5) Another marker of stress granules should be used for confirming the identity of G3BP1-Anap (+) or TDP-43-Anap (+) structures, including TIA1, TAF15, or polyA RNA.

      We appreciate this helpful suggestion. To further confirm the identity of the stress granule structures observed in our experiments, we performed colocalization analysis with TIA-1, a well-established marker of stress granules. The results have been included in revised manuscript.

      “Additionally, we examined the colocalization of G3BP1-Anap with TIA-1, another established stress granule marker. Under stress conditions, G3BP1-Anap largely colocalized with TIA-1 within stress granules. Interestingly, under basal conditions, the nuclear signal of G3BP1-Anap, which was not detected by antibody staining, appeared to partially colocalize with TIA-1 in several condensate-like structures. (Fig. 1I).”

      (6) There is no information on the number of granules bleached or the number of cells selected for FRAP studies. There is no information on the shaded areas in Figure 1F or 1G, and no information on statistical comparisons between regressions in Figure 1F.

      We thank the reviewer for pointing out these omissions. We have revised the figure legends to clarify these details.

      “One granule from each of three independent cells was selected and photobleached for FRAP analysis.”

      “Here, error bars with filled area are used for better data presentation. FRAP recovery curves were compared using two-way ANOVA.”

      (7) Protein dynamics measured by FRAP are highly dependent on the concentration and/or expression level of each protein. Because of this, the authors need to control for expression level in all FRAP studies.

      We agree that protein concentration and expression level can influence FRAP recovery kinetics. Since Anap incorporation is based on amber suppression, and the suppression rate in each cell varies, so it is difficult to control the expression of Anap labeled proteins, however, to minimize this potential effect, we performed FRAP measurements on cells exhibiting comparable fluorescence intensities, which served as a proxy for similar expression levels of the labeled proteins. In addition, FRAP analyses were conducted on individual granules within cells expressing moderate levels of the protein, avoiding cells with unusually high fluorescence intensity that might reflect overexpression.

      Furthermore, fluorescence recovery was normalized to the pre-bleach intensity of the selected granules, which reduces variability arising from differences in overall expression levels between cells.

      (8) There is no point of reference for TDP-43-Anap FRAP results in Figure 1G. Additional studies using variants harboring a mutated NLS (mNLS) can be used in place of TDP43-YFP.

      This is a helpful suggestion. In response, we have performed additional FRAP experiments using TDP-43<sup>ΔNLS</sup>, a commonly used construct that promotes cytoplasmic localization and facilitates analysis of TDP-43 granules. The results from TDP-43<sup>ΔNLS</sup> have now been included as a reference for the FRAP measurements of TDP-43-Anap in the revised manuscript (Fig. 1D, 1G).

      “We then used YFP-tagged nuclear localization signal (NLS)-deleted TDP-43 (TDP43<sup>ΔNLS</sup>-YFP) as a reference and performed FRAP analysis to compare the mobility of TDP-43-Anap and TDP-43<sup>ΔNLS</sup>-YFP. Fluorescence recovery of TDP-43-Anap reached ~45% within 20 s after photobleaching, consistent with liquid-like dynamics. In contrast, TDP-43<sup>ΔNLS</sup>-YFP showed only ~22% recovery, suggesting more solid-like dynamics (Fig. 1D, 1G). These results are consistent with previous reports describing relatively immobile aggregates formed by TDP-43<sup>ΔNLS4</sup>and illustrate the advantage of Anap-based labeling, which preserves native protein properties and enables real-time assessment of protein dynamics without introducing disruptive mutations.”

      (9) There is no point of reference for comparing FRAP results from G3BP1-GFP to G3BP1-Anap. What is the 'gold standard'? Without this, it is difficult to conclude that "... Anap labeling better preserved the native mobility and biophysical properties of G3BP1 than the conventional GFP tag."

      We acknowledge this important point and agree that there is currently no definitive gold standard for measuring the native mobility of endogenous G3BP1 within stress granules in living cells. Our intention was not to claim that the Anap-labeled protein definitively represents the native state, but rather to compare the relative effects of different labeling strategies.

      Thus, we rewrite the sentence as “These results suggest that G3BP1-Anap displays higher mobility compared with G3BP1-GFP, indicating that Anap labeling may provide a less perturbative approach for monitoring G3BP1 dynamics.”

      (10) The WB in Figure 1H is overexposed, making it difficult to compare expression levels between WT and V100Anap-transfected cells. In addition, there is no similar assay for confirming G3BP1-Anap expression.

      Thank you for pointing this out. In the revised manuscript, we have replaced the image with a properly exposed Western blot to allow clearer comparison of protein expression levels.

      In addition, we have now included a corresponding western blot analysis to confirm the expression of G3BP1-Anap in G3BP knockout U2OS cell (Fig. 1H). These results verify that the Anap-labeled proteins are expressed at detectable levels and support the interpretation of the imaging and FRAP experiments.

      (11) Although survival studies in Figures 1I and J are promising, a more convincing demonstration of functional replacement of TDP-43 would involve an assessment of cryptic exon splicing, comparing WT to TDP-43 KO, V100Stop- and V100Anaptransfected cells.

      This is a valuable suggestion.

      “We also evaluated TDP-43-dependent RNA splicing activity by examining the expression of PFKP, a well-established target that undergoes cryptic exon inclusion upon loss of TDP-43 function17. As shown in Figures 1K and 1L, expression of TDP-43Anap in TDP-43 knockout HeLa cells restored PFKP expression, indicating that the Anap-labeled protein retains functional RNA splicing activity. These results demonstrate that TDP-43-Anap is capable of functionally compensating for endogenous TDP-43, supporting that the incorporation of Anap does not substantially disrupt the protein’s biological function.”

      (12) Tuj1 staining in Figure 2 is inconsistent and often fails to confirm neuronal identity.

      We thank the reviewer for this important comment. We acknowledge that Tuj1 staining in Figure 2 is variable and, in some cases, does not clearly delineate neuronal identity. Notably, the reduced Tuj1 signal is primarily observed in neurons that express Anap-labeled proteins under sodium arsenite treatment, which likely reflects the combined effects of transfection-associated stress and oxidative stress on neuronal morphology and cytoskeletal integrity.

      In addition, transfection efficiency in primary neurons is inherently low and variable, and cells that successfully express the constructs may represent a more stress-sensitive subpopulation, further contributing to variability in staining quality. Despite optimization efforts, these technical constraints limit the consistency of Tuj1 labeling under these experimental conditions.

      (13) Close-up images and correlation scatter plots in Figures 1 and 2 do not add very much information.

      We thank the reviewer for this comment. To address the reviewer’s concern, we have revised the figure legends to better clarify the purpose of these panels and how they support the quantitative analysis presented in the manuscript.

      For scatter plot, “Colocalization threshold analysis was performed in Fiji/ImageJ to calculate the Pearson correlation coefficient (R) for each region of interest (A, B, I, J). The X- and Y-axes represent the fluorescence intensity values of the red and green channels, respectively. When signals are colocalized, pixels with high intensity in one channel correspond to high intensity in the other, forming a diagonal distribution. In contrast, non-colocalized signals cluster along the axes. A higher R value indicates a greater degree of colocalization. Scale bar, 3 μm.”

      Same information was added to figure legend of figure 2.

      For the scheme, please see line 412-413 in the revised manuscript.

      Reference:

      (1) Rothstein, J.D. et al. Sporadic ALS induced pluripotent stem cell derived neurons reveal hallmarks of TDP-43 loss of function. Nature Communications 16, 7092 (2025).

      (2) Shadish, J.A. & Lee, J.C. Genetically encoded lysine photocage for spatiotemporal control of TDP-43 nuclear import. Biophys Chem 307, 107191 (2024).

      (3) Gasset-Rosa, F. et al. Cytoplasmic TDP-43 De-mixing Independent of Stress Granules Drives Inhibition of Nuclear Import, Loss of Nuclear TDP-43, and Cell Death. Neuron 102, 339–357.e337 (2019).

      (4) Yan, X. et al. Intra-condensate demixing of TDP-43 inside stress granules generates pathological aggregates. Cell 188, 4123–4140.e4118 (2025).

      (5) Walker, S.E. & Fredrick, K. Preparation and evaluation of acylated tRNAs. Methods 44, 81–86 (2008).

      (6) Kassouf, T. et al. Targeting the NEDP1 enzyme to ameliorate ALS phenotypes through stress granule disassembly. Science Advances 9, eabq7585 (2023).

      (7) Van Nerom, M. et al. C9orf72-linked arginine-rich dipeptide repeats aggravate pathological phase separation of G3BP1. Proceedings of the National Academy of Sciences 121, e2402847121 (2024).

      (8) Wolozin, B. & Ivanov, P. Stress granules and neurodegeneration. Nat Rev Neurosci 20, 649–666 (2019).

    1. Author response:

      eLife Assessment

      This important study provides evidence that plateau pikas, at moderate densities, can facilitate yak nutrition by suppressing a poisonous plant, offering a helpful perspective on reciprocal interactions between small mammal ecosystem engineers and large herbivores. The evidence is solid, supported by a manipulative field experiment and appropriate measurements of intermediary ecological processes, although some claims about density dependence, competition, and stress-gradient mechanisms are not fully supported by the experimental design. The work will be of interest to ecologists, conservation biologists, and rangeland managers, particularly those studying grassland herbivore interactions and livestock management on the Qinghai-Tibetan Plateau.

      Thank you very much for these positive assessments of our work, below we provided the point-by-point responses to the comments from the 2 peer reviewers, and we hope these revisions are satisfied.

      Reviewer #1 (Public review):

      Summary:

      This is important and significant work because it helps describe the complexity of interactions between system components where two herbivores interact with vegetation. Whereas other studies have shown that the larger ungulate (yaks, Bos grunniens, in this case) can facilitate the abundance and population growth of the smaller (the semi-fossorial lagomorph, Ochotona curzoniae, plateau pika hereafter), this study flips the tables and shows that, at least under some conditions, moderate densities of the plateau facilitate the nutritional condition of yaks.

      The study was not designed to investigate the reasons that pikas clip Stellera chamaejasme. That said, based on other studies and general knowledge of the ecology of these pikas, it is likely that they clip (although do not eat) this plant because its relatively large size hinders predator detection. This species of pika does better where vegetation height is low than where it is higher.

      Strengths:

      Notably, the strong inference the authors can claim for their results is supported by the careful experimental design. A weaker paper would have simply noted correlations between pika burrow density and yak feeding efficiency without experimental removal. This paper, to its credit, not only used experimental removals but also documented the various intermediary results that support the ultimate conclusions. The statistical approaches used appear to be appropriate. (Readers are encouraged to read the full Materials and Methods, which are available in the Supplementary Materials section.)

      We appreciate these positive comments on our work.

      Weaknesses:

      Although the study was well designed and executed, and its conclusions appear strongly supported, readers interested in the management implications of the Qinghai-Tibetan Plateau should be mindful of its limitations. First, the study site, at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera chamaejasme becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. Thus, it would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace S. chamaejasme as the problematic plant for pastoralists.

      Agree! We will acknowledge this limitation in the Discussion, by adding the paragraph below (see the Third point):

      “Despite of these, several questions remain deserve further investigation. First, our study examined pika–yak interactions only during the summer period, when food resources are most abundant. Whether such facilitative effects weaken or even shift toward competition under more stressful conditions—for example, when forage becomes limited during autumn or winter—remains to be tested. Second, if pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition (Yang et al., 2026)? Third, our study site located at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. It would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace Stellera as the problematic plants for pastoralists (Lu et al., 2012; Li and Zhao, 2025). Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      Second, the authors make no mention of wild ungulates, so it is unclear what, if any, role they may have played in this system. At least one study in Qinghai Province, albeit at a slightly higher elevation, showed that not only pikas, but also Tibetan gazelles (Procapra picticaudata), which were commonly observed on grazed pastures, grazed more frequently on some dicots avoided by domestic sheep than did the livestock themselves (Harris et al. 2015). Citation:

      Harris RB, Wang, WY, Badinqiuying , Smith AT, Bedunah DJ (2015) Herbivory and Competition of Tibetan Steppe Vegetation in Winter Pasture: Effects of Livestock Exclosure and Plateau Pika Reduction. PLoS ONE 10(7): e0132897.

      doi:10.1371/journal.pone.0132897

      Agree! We will add more details about the study site, particularly regarding wild ungulates, in the Methods section. Specifically, we will include the following sentence: “Wild ungulates, such as Tibetan gazelles (Procapra picticaudata) (Harris et al., 2015), and other small mammals such as rabbits and zokors, occur rarely in the area.” This key reference will also be cited in this section.

      It would also be instructive to learn if similar facilitation as observed here applied to the other principal livestock species in the area, domestic sheep (which are often herded together with smaller numbers of domestic goats).

      Agree! The same as mentioned above. We will acknowledge this limitation in the Discussion, by adding the paragraph below (see the Final point):

      “Despite of these, several questions remain deserve further investigation. First, our study examined pika–yak interactions only during the summer period, when food resources are most abundant. Whether such facilitative effects weaken or even shift toward competition under more stressful conditions—for example, when forage becomes limited during autumn or winter—remains to be tested. Second, if pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition (Yang et al., 2026)? Third, our study site located at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. It would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace Stellera as the problematic plants for pastoralists (Lu et al., 2012; Li and Zhao, 2025). Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      Finally, as suggested by this study, the interactions between all components of the system are complex and interactive. If pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition?

      Agree! The same as mentioned above. We will acknowledge this limitation in the Discussion, by adding the paragraph below (see the Second point):

      “Despite of these, several questions remain deserve further investigation. First, our study examined pika–yak interactions only during the summer period, when food resources are most abundant. Whether such facilitative effects weaken or even shift toward competition under more stressful conditions—for example, when forage becomes limited during autumn or winter—remains to be tested. Second, if pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition (Yang et al., 2026)? Third, our study site located at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. It would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace Stellera as the problematic plants for pastoralists (Lu et al., 2012; Li and Zhao, 2025). Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      Reviewer #2 (Public review):

      Summary:

      This study uses a combination of field sampling and manipulative experiments to test for facilitative impacts of pikas on yaks via suppression of a poisonous forb. The authors found that, when Stellera forbs were present, yak weight increases over the growing season were greater in the presence of pikas compared to in their absence. This occurred because, although pikas do not consume Stellera, they clip it and use it in nest/burrow construction, thereby decreasing its relative abundance in the plant community. Thus, overall, the study contributes to our understanding of how herbivores of different size classes indirectly affect each other via the use of shared resources.

      Strengths:

      It is well known that large herbivores on grasslands impact smaller animals, but the reciprocal interaction is rarely tested. Thus, this study asks a valuable question, and the experiment is well-designed to test it. The authors also do a good job of demonstrating the potential conservation impacts of their research.

      We appreciate these positive comments on our work.

      Weaknesses:

      What the authors tested is really cool, but their claims go far beyond what they can say based on their experimental design. For example, the authors claim to show that pika impacts on yaks display density-dependent transitions from competition to facilitation. However, their experiment only looked at the presence (at moderate densities) and absence of pikas, and they only tested for facilitation, not competition. The paper would also benefit from changes to the framing in the introduction and discussion. For example, the authors pitch the work as a test of the stress-gradient hypothesis. However, there is no abiotic stress gradient in the study, which is an essential component of the SGH. They also pitch the work in terms of density dependence, but there is no significant variation in population densities beyond the presence-absence binary. The paper would be stronger if they focused their framing around the literature on facilitative interactions across mammals of different size classes, especially indirect facilitation via use of shared resources, which is what this paper is really about.

      We agree that our work had explored only the facilitative effects of pikas on yaks, rather than the density-dependent balance between competition and facilitation, and the Stress Gradient Hypothesis (SGH).

      We plan to make the major revisions below to address this important concern.

      (1) We will revise the title as “Moderate density of small mammalian herbivores facilitates livestock growth in grasslands ”.

      (2) We will delete all the statements about density-dependent transition of facilitation and competition and the SGH in the Abstract, Introduction, Discussion, and the References sections.

      Finally, the paper has significant weaknesses in the experimental and statistical methodology. Most importantly, there are inconsistencies in what is visualized in the figures compared to the model results. For example, the results section in several places notes a lack of significant interaction terms in the model but shows interactions in the p-values on the figures.

      In the Results section, there are only two locations where we discussed non-significant interactions: Line 148–149 “Pikas and Stellera had no interactive effects on abundance of sedges, forbs, and neutral detergent fiber (NDF) of total forage for yaks (Fig. 3F,I, fig. S1, table S3,5).” and Line 161–162 “Pikas and Stellera had no interactive effects on yaks’ foraging efficiency on forbs (fig. S2, table S7).”.

      We have cross-checked both the manuscript as submitted and the website, and in every instance we are consistent in not reporting interactions as non-significant when the model output shows significance.

      We will confirm these details in the revised version as “Pikas and Stellera had no interactive effects on abundance of sedges, forbs, and neutral detergent fiber (NDF) of total forage for yaks (Fig. 3F,I, Fig. S1, Table S5, S8). ”; and “Pikas and Stellera had no interactive effects on yaks’ foraging efficiency on forbs (Fig. S2, Table S10).” in the Results section.

      The authors also plot smoothed lines rather than their model results and then draw interpretations from those lines that cannot be tested in the models that they used.

      Agree! There are only two figures in which we used generalized additive models (GAMs) to plot smoothed lines: Figure 2C and Figure 3C.

      For Figure 2C, the supplementary table for the GAMM associated with the smoothed line was not originally included, but we will add it as Table S4 in the revised version. For Figure 3C, we explicitly fit a GAMM corresponding to the plotted line, and the model results will be reported in the Table S7 in the revised version.

      There are also missing details that are important for model interpretation, including the distributions used and the sample sizes.

      Agree! We will provide the Table S13 to summarize all statistical models used in the study, including the distributions used and the sample sizes in the Supplementary Materials. We will also add a sentence of “A summary of all statistical models used in the study is available in table S13.” in the Statistical analyses section to indicate this information.

      Another major concern with experimental design is in the forage nutrient analyses. The authors picked plants along a grazing trail, then measured nutrient content without standardizing based on plant species, so any differences across treatments could be because of what they happened to grab rather than overall forage quality.

      We will revise this section to provide more details on how forage samples were collected and their quality were analyzed. Specifically, five forage samples were collected per grazing plot, focusing on the two dominant plant species —one sedge and one grass—that were most frequently grazed by yaks. To ensure comparability across plots and treatments, we mixed the two species at equal dry mass (5 g). We will revise this section as below.

      “To assess forage quality, five forage samples were collected from each grazing plot to quantify their nutritive values. To obtain samples that reflect the forage actually consumed by yaks, we tracked the animals along their grazing paths and collected the plant tissues of the two most frequently consumed species: the dominant sedge Kobresia humilis and the dominant grass Elymus nutans (Fig. 2B; Pan et al., 2019). The collected tissues of each species were dried in a forced-air oven at 60 °C for 48 h, then ground through a 1-mm mesh. Subsequently, 5 g of each dried and ground species were combined in a 1:1 dry mass ratio, and the resulting mixture was stored in plastic bags for subsequent analyses.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This manuscript used deep learning to highlight the role of inhibition in shaping selectivity in primary and higher visual cortex. The findings hint at hitherto unknown axes of structured inhibition operating in cortical networks with a potentially key role in object recognition.

      The multi-species approach of testing the model in macaque and mouse is excellent, as it improves the chances that the observed findings are a general property of mammalian visual cortex. However, it would be useful to delineate any notable differences between these species, which are to be expected given their lifestyle.

      The overall performance of the model appears to be excellent in V1, with over 80% performance, but it falls substantially in V4. It would be important to consider the implications of this finding; for example, in the context of studying temporal lobe structures that are central to recognizing objects. Would one expect that model performance decreases further here, and what measures could be taken to avoid this? Or is this type of model better restricted to V1 or even LGN?

      While the manuscript delineates novel axes of inhibitory interactions, it remains unclear what exactly these axes are and how they arise. What are the steps that need to be taken to make progress along these lines?

      Reviewer #2 (Public review):

      The classic view of sensory coding states that (excitatory) neurons are active to some preferred stimuli and otherwise silent. In contrast, inhibitory neurons are considered broadly tuned. Due to the gigantic potential image space, it is hard to comprehensively map the tuning of individual neurons. In this tour de force study, Franke et al. combine electrophysiological recordings in macaque (V1, V4) and mouse (V1, LM, LI) visual cortex with large-scale screens based on digital twin models, as well as beautiful systems identification (most/least activating stimuli). Based on these digital twins, they discover dual-feature selectivity (which they validate both in macaques and mice). Dual-feature selectivity involves a bidirectional modulation of firing rates around an elevated baseline. Neurons are excited by specific preferred features and systematically suppressed by distinct, non-preferred features. This tuning was identified by excellently combining advances in AI & high-throughput ephys.

      The study is comprehensive and convincing. Overall, this work showcases how in silico experiments can generate concrete hypotheses about neuronal coding that are difficult to discover experimentally, but that can be experimentally validated! I think this work is of substantial interest to the neuroscience community. I'm sure it will motivate many future experimental and computational studies. In particular, it will be of great interest to understand when and how the brain leverages dual-feature selectivity. The discussion of the article is already an interesting starting point for these considerations.

      Strengths:

      (1) Using computational models to predict neuronal responses allowed them to go through millions of images, which may not be possible in vivo.

      (2) The cross-species and cross-area consistency of the results is another major strength. Pointing out that the results may be a fundamental strategy of mammalian cortical processing.

      (3) They show that the feature causing peak excitation in one neuron often drives suppression in another. This may be an efficient coding scheme where the population covers the visual manifold. I'd like to understand better why the authors believe that this shows that there are low-dimensional subspaces based on preferred and non-preferred stimulus features (vs. many more, but some axes are stronger).

      We thank the reviewers for their constructive and helpful feedback on our manuscript. We are delighted that they found the study to be “comprehensive and convincing” and a “tour de force” in its combination of electrophysiological recordings with large-scale digital twin screening. We appreciate that the reviewers highlighted the strengths of our multi-species approach and the “cross-species and cross-area consistency” of the results, noting that the work showcases how in silico experiments can generate concrete, experimentally validatable hypotheses. Overall, we agree with the assessment of the reviewers. We have performed the following changes to the text to clarify and strengthen the manuscript, without introducing new analyses or altering the conclusions. 

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Page 3: The authors state that RFs were mapped using sparse noise, with the goal to ensure that the RFs align with the visual stimulus, but no data appear to be shown regarding this alignment. It would be important to provide a full analysis of the sparse noise-mapped RFs for both V1 and V4. Also, is it correct that the V4 data analyzed here came from a single animal? This could potentially be problematic and would need to be addressed, for example, by performing analyses also in V1 for participant animals separately. Please elaborate.

      We have added a sentence to the Results section clarifying the sparse noise RF mapping procedure, noting that probe insertions were targeted orthogonal to the cortical surface so that neurons sampled along the probe depth share overlapping receptive fields, allowing a single stimulus configuration to adequately drive the entire recorded population. We have also corrected the text to clarify that V4 data were collected from 2 animals (not 3 as previously stated in an earlier draft), consistent with the Methods section.

      (2) Page 4: Only half the neurons in V4 are "high confidence" in terms of test image performance, which seems a little low and probably significantly lower than the corresponding value for V1 of 84%. It is unclear how to interpret this confidence, but it seems to suggest that half of the V4 neurons are not well captured by the model. If true, this fraction appears large enough to cast doubt on the validity of the V4 results. Please elaborate.

      We have expanded the text to explicitly discuss the lower proportion of high-confidence in-silico neurons in V4 relative to V1. We attribute this to the greater complexity of V4 tuning compared to V1, as well as missing contextual information such as image surrounds and sequential image context—factors that likely limit model performance in higher visual areas. We note that our restriction of analyses to high-confidence neurons provides resilience against these limitations, and that the goal was not to maximize predictive performance per se but to identify response patterns—dual-feature selectivity—that are robust across neurons, areas, and species.

      (3) Page 5: It seems that identical L2 norms are valid for discounting contrast variations, particularly if the neural responses are linear, since the L2 norm is computed on the entire RF. It might be judicious to attenuate the claim that contrast variation has no effect.

      We have softened the claim that contrast variation has no effect. The revised text now states that L2 normalization controls for root-mean-squared contrast but does not fully equate effective contrast in nonlinear cells, whose responses depend on the spatial structure of the stimulus beyond its total energy. We note that residual contrast dependent effects, particularly in the suppressive regime, cannot be entirely excluded.

      (4) Page 6: The authors acknowledge that, at least for simple cells, a phase shift in the grating and concomitant ON-OFF overlap is an inhibitory axis, which is correct. It does not really become clear what other axes were found, and whether any of these represent a novel discovery about V1.

      We have clarified the description of inhibitory axes in V1, noting that while phase-shifted stimuli represent a well-established suppressive axis for simple cells reflecting linear On-Off subfield structure, and complex cells exhibit no coherent suppressive pattern due to phase pooling, neither model class accounts for the multidimensional suppressive structure we observe. We have made explicit that our unbiased approach reveals suppressive structure spanning simultaneous changes across orientation, spatial frequency, phase, and texture, exceeding what any single known suppressive mechanism predicts.

      (5) Page 7: Dreamsim is based on human similarity judgements, whereas the data is from macaques. Is there any evidence suggesting that macaque similarity judgements might be similar to those of humans?

      We have added a paragraph to the Discussion acknowledging that DreamSim was trained on human perceptual similarity judgments while our neuronal data are from macaques. We note that this cross-species application is supported by the deep homology between primate ventral visual streams, and that natural-image similarity judgments have been found to be highly consistent across macaques and humans. Importantly, we clarify that we deploy DreamSim not as a model of macaque perception but as an image feature embedding to test whether stimuli that cluster in perceptual space evoke similar neuronal responses—a use that is robust to the precise calibration of the metric. We also note that we are developing custom macaque-specific embeddings for future work.

      (6) Page 7: How many images were in the test set?

      We have added the number of test images to the relevant text (n=75 for V1, n=150 for V4) and to the Figure 1 caption.

      (7) Page 8: As mentioned above, performing the analysis on V1 data of individual subjects and demonstrating similar digital twins might be an additional way to confirm the models' accuracy.

      We have added text noting that for V4, 1digital twin models were fit independently per neuron without sharing information across animals, and that extreme image sets identified by the model elicited correspondingly extreme responses in neurons from the other animal, confirming that identified selectivity patterns are not idiosyncratic to individual subjects.

      (8) Page 11: The mouse data is presented very briefly only, and the authors seem to imply that there is a high degree of coding similarity between this rodent species and macaques and, by extension, humans. Were there any notable differences between the mouse and macaque data?

      We have added text explicitly noting that while macaque and mouse visual cortex differ substantially in their functional organization and the complexity of neuronal selectivity, the broader principle—that non-sparse neurons are jointly defined by distinct excitatory and suppressive feature sets—generalizes across mammalian visual systems. We clarify that this does not imply that mouse and macaque visual cortex share similar functional organization or equivalent complexity of neuronal selectivity; rather, within the representational regime of each area, neurons are organized such that excitatory and suppressive feature sets are jointly structured and distinct.

      (9) Page 13: One main finding of the study is that inhibition appears to operate along additional dimensions that had not been previously recognized, but what is the nature of these dimensions, how do they arise and relate to known inhibitory effects in V1 such as centre-surround effects? The fact that suppression is tuned in response to natural images or other complex objects is not a new finding, and there is plenty of published work along these lines; the authors may want to cite Tamura et al 10.1152/jn.01267.2003. I am not sure introducing the term "dual feature selectivity" is really a major conceptual advance.

      We have added a citation to Tamura et al. (2004) in the Discussion, alongside other prior work documenting suppression by non-optimal stimuli. We have also expanded the Discussion to more carefully position our findings relative to existing work on feature-selective suppression, noting that while prior work has established that inhibition can be structured and feature-selective, our results suggest a broader organizing principle: within each visual area, there exists a set of feature combinations from which individual neurons draw both their excitatory and suppressive preferences.

      (10) Page 14: The authors enumerate a number of technical limitations, which is to be commended. It would be useful for them to comment on the particular advantages of the digital twin model, compared to a more traditional analysis of the responses to the thousands of natural images that were experimentally obtained. It seems likely that the main finding, i.e. tuned inhibition, is also evident directly in this population (?). While the digital twin is to some degree validated by the test images, its responses to the much larger set of images studied are not validated, and one must trust that the ResNet50 indeed captures V4 selectivity. It would be useful to discuss some of these points, and highlight a potential way that digital twins (maybe as a shared model between laboratories) can learn from a large number of animals and datasets, and maybe even be used to generate novel visual stimuli suitable to test emergent hypotheses.

      We have added a paragraph to the Discussion explicitly contrasting the advantages of digital twin models with direct analysis of experimentally recorded responses, noting that digital twins enable screening of more than one million images per neuron in silico, gradient-based synthesis of stimuli precisely optimized to drive or suppress individual neurons, and cross-model verification of identified selectivity patterns—a test that has no analog when working with fixed experimental image sets.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) Call out Figure 1/b in the main text. 

      We have added a callout to Figure 1b in the main text

      (2) Can you make a supplementary figure illustrating more examples with skewness around the middle (e.g. 1.5, 2, 2.5)? Namely, you state that 2 is a good threshold for deciding if it is non-sparse, but you only present clear-cut cases in Figure 2 (with <0.75 and >3.5). I am wondering if 2 is a good threshold?

      We have revised the text to clarify that the skewness threshold of 2.0 is adopted purely for analytical convenience to focus subsequent analyses on neurons with sufficiently graded response distributions, and that the key findings are not dependent on the exact threshold chosen. We explicitly note that the underlying distribution of sparsity is continuous, consistent with recent findings (Gondur et al., 2025).

      (3) The reference "A tale of two tails: Preferred and anti-preferred natural stimuli in visual cortex." Has no authors. I know it's anonymous, but maybe put that for now? I also congratulate including a paper that is anonymously under review at ICLR 2026. I don't find Unk, 2025 in the list of references. Perhaps related?

      We have updated the reference “A tale of two tails” to include the authors (Gondur et al., 2025) and ensured it appears consistently in the reference list. We have also resolved the missing “Unk, 2025” citation, which now correctly refers to this same work.

      (4) Why do you use a different model for the analysis in Figure 8?

      We have added text to the Methods and Results clarifying why a distinct architecture was used for the V4 evaluator model in Figure 8. Specifically, the V4 generator model uses a fixed, pretrained ResNet50 backbone whose weights are deterministic; any re-trained model sharing this backbone would not constitute a genuinely independent evaluation. By contrast, for V1, the ConvNeXt core is fine-tuned from different random initializations, producing architecturally equivalent but computationally independent models. A truly independent V4 evaluator therefore required a fundamentally different architecture.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      The study by Raiola et al. conducted a quantitative analysis of tissue deformation during the formation of the primitive heart tube from the cardiac crescent in mouse embryos. Using the tools developed to analyze growth, anisotropy, strain, and cell fate from timelapse imaging data of mouse embryos, the authors elucidated the compartmentalization of tissue deformation during heart tube formation and ventricular expansion. This paper describes how each region of the cardiac tissue changes to form the heart tube and ventricular chamber, contributing to our understanding of the earliest stages of cardiac development.

      Strengths:

      In order to understand tissue deformation in cardiac formation, it is commendable that the authors effectively utilized time-lapse imaging data, a data pipeline, and in silico fate mapping.

      The study clarifies the compartmentalization of tissue deformation by integrating growth, anisotropy, and strain patterns in each region of the heart.

      Weaknesses:

      The significance of the compartmentalization of tissue deformation for the heart tube formation remains unclear.

      While it is obvious that the patterns of deformation should be relevant to model the cardiac crescent into the primitive cardiac tube, we do not provide direct evidence that changing these patterns affects heart tube formation. In this sense, the Reviewer is correct and this is a limitation of the study.

      Reviewer #1 (Recommendations for the authors):

      (1) It is interesting that growth rate and anisotropy are anticorrelated. However, the functional significance of this anticorrelation in heart formation remains unclear. It may be worthwhile to analyse the importance of the relationship between the two by adding inhibitors to cultured embryos or using mutant mouse models.

      We appreciate this thoughtful suggestion and agree that such experimental approaches, involving inhibitors or mutant mouse models, could provide powerful validation of the proposed relationship. However, generating the appropriate lines and performing the necessary quantifications would represent a substantial effort that extends beyond the scope of the current study. Our focus here is to establish the correlation and its potential implications, leaving these more in-depth mechanistic investigations for future work.

      (2) The authors claim to have analysed tissue deformation at the cellular level. Although cell labelling of specific regions using Tat-Cre and DiI injection and tracking of their fate have been performed, this still gives the impression of tissue-level analysis. An analysis "at the cellular level" would be expected to describe morphology, proliferation, polarity, etc., at the single-cell to multi-cell level.

      We thank the reviewer for the comment. Our analysis does not involve single-cell characterization (e.g., morphology, proliferation, polarity) but focuses on quantifying tissue motion. The motion extracted from the images achieves cellular-level precision, as demonstrated by testing the registration algorithm and validating it against cellular tracking experiments. The accuracy of the method is therefore at the cellular scale. The goal of our study is to describe tissue dynamics during heart development, not to perform detailed cellular analyses. The novelty of our approach is that it enables tissuescale quantification in developmental mouse heart imaging, where cell density and image resolution make automated single-cell tracking unfeasible. By using fluorescence labelling as markers, we obtain cellular-level accuracy in tissue motion quantification.

      (3) It is stated that cardiomyocytes, cardiac mesodermal cells, and SHF cells were labeled with Nkx2.5-GFP/Nkx2.5-Cre, Mesp1-Cre, and Islet1-Cre, respectively; however, the results of the labeling using these mice are not presented, and the reason for using different mouse strains is not apparent. Information on these mouse strains is missing in the Materials & Methods section. In particular, attention must be paid to mice of the same name but different strains. Islet1-Cre mice are not SHF-specific and exhibit activity in part of the left ventricular progenitors. Sparse labeling induced by low-dose tamoxifen administration is also unclear regarding the timing and concentration of tamoxifen administration. The authors should provide data on labeling efficiency and region, and also discuss the usefulness of analyses using different mouse strains.

      We thank the reviewer for raising these important points. In this study, the use of different mouse strains was not driven by a biological comparison or lineage-specific analysis, but by the availability of high-quality cardiac imaging datasets generated in the laboratory. The primary goal of this work is methodological: to demonstrate that developmental cardiac imaging data can be reused within an engineering framework to quantify tissue deformation. For this purpose, we do not track individual cells but instead use fluorescence labelling as a versatile strategy to follow tissue motion without requiring a strain- or lineage-specific labelling.

      We acknowledge that Islet1-Cre mice are not SHF-specific and exhibit activity in part of the left ventricular progenitors. However, this limitation does not affect our analysis, as the specificity of the labelled cells is not used in the image processing or deformation quantification pipeline.

      Regarding tamoxifen, we clarified the dosage and administration in the revised Experimental workflow section. Importantly, tamoxifen treatment does not influence the proposed image analysis framework, since labelled cells are employed solely as fiducial references to provide ground-truth validation of the tissue motion estimated from image registration.

      We made these points clearer in the Results section and in Materials and Methods.

      (4) It is noteworthy that the authors have utilized many new analytical methods that they have developed. In the analysis presented in this paper, it is understandable that the methods described in another paper by the authors (Raiola M et al., 2025) are utilized; however, it is important to note that this causes some overlap. It is necessary to clearly distinguish and describe whether the novelty of the methods is based on those developed in this paper or those described in the paper by Raiola M et al. (2025).

      We thank the reviewer for this important observation. We agree that it is essential to clearly separate the methodological developments reported in Raiola M et al. (2025) from the present work. As described in Raiola M et al. (2025), the methods have already been fully developed and validated. In this paper, our focus is to apply these approaches to cardiac development in order to generate and describe the new biological insights. In the revised version, we made this distinction more explicit in the Results and Discussion, highlighting the methodological continuity with our previous work and the biological contribution of the present study.

      Minor points

      (1) In Figure 4, the labels and legends for A, A', B, and B' are reversed. Similar colours are used for C through F, making it difficult to distinguish between them.

      We thank the reviewer for noticing this error. We have corrected it.

      (2) In Figure 5, the labels start with B.

      We thank the reviewer for noticing this error. We have corrected the labels in Figure 5.

      Reviewer #2 (Public review):

      The authors address an important challenge in developmental biology: the quantitative description of tissue deformation during organogenesis. They have developed a new pipeline to quantify early heart tube morphogenesis in the mouse, with cellular resolution. They adopt an elegant approach by integrating multiple 3D time-lapse datasets into a dynamic atlas of cardiac morphogenesis in order to compute spatio-temporal deformation patterns. The main findings highlight a strong compartmentalization of cell behaviors, with tissue growth and anisotropy exhibiting complementary and spatially segregated patterns. Using these data, the authors developed an in-silico fate mapping tool to interrogate cell displacement within the myocardium. This virtual model provides new mechanistic insights into how the bilateral cardiac primordia converge and transform into a three-dimensional heart tube. The authors identify "belt-like" constraints at the arterial and venous poles that prevent tissue expansion and thus shape the ventricular barrel morphology.

      The computational framework is highly innovative and impressive, providing an unprecedented 3D model of tissue deformation during heart morphogenesis. It also opens avenues for testing hypotheses regarding tissue growth and the forces that cause cell motion. However, the proposed model of ventricular chamber formation with the two constraining belts remains hypothetical, lacking biological validation and requiring strengthening or modulation.

      Overall, this carefully performed study provides a new model for exploring tissue deformation during organogenesis and will be of broad interest to computational and developmental biologists.

      We agree with the Reviewer on the limitations of the proposed model due to limited experimental validation. In the revised version of the manuscript we provide further experimental evidence that strengthens the biological validation of the proposed barrel model with two transversal “belts” generating the barrel shape of the primitive ventricle.

      Reviewer #2 (Recommendations for the authors):

      (1) The study proposes a new model of heart morphogenesis by identifying two regions of tissue contraction at the arterial and venous poles. Although the fate map tool has been validated using two ex vivo approaches (DyeI microinjection and TAT-Cre genetic labelling), the conclusions regarding the two belts still need to be demonstrated using in vivo/ex vivo experiments and quantification of cell movements.

      We thank the reviewer for this important suggestion. We agree that experimental validation of the two contraction belts is essential to strengthen the conclusions of the study. In the revised manuscript, we have addressed this point by adding new experimental data directly supporting the existence and dynamics of both D1 and D2 contraction boundaries.

      Specifically, we performed microinjection experiments in which four anchor points, two along D1 and two along D2, were labeled in living embryos and tracked after 10–14 hours (Figure 5C–E, Table S2). For D2, the Euclidean distance between the two anchor points was computed from multiphoton microscopy images acquired at t0 and tfinal (voxel size: 0.57 × 0.57 × 2.5–6.0 µm). In all three embryos analyzed, the D2 anchor points converged over time, with the segment retaining on average 0.27 ± 0.14 of its initial length (range: 0.13–0.40), confirming the lateral compression predicted by the model. For D1, the in-plane geodesic distance between anchor points was measured at t0 and after 10–14 hours. Given the difficulty of imaging the arterial pole at high resolution by whole-mount microscopy, cryosections were used for these measurements (pixel size: 0.65 × 0.65 × 0.042 µm). The D1 segment similarly underwent contraction, retaining on average 0.50 ± 0.22 of its initial length (range: 0.23–0.77). Together, these results provide direct experimental evidence that both boundaries undergo compression during heart tube formation, consistent with the contraction dynamics predicted by the virtual model and supporting the existence of the two belts described in the study.

      We acknowledge that the quantitative values show variability across embryos, which reflects two main sources of uncertainty: (i) the exact position of microinjection along D1 and D2 could not be perfectly standardized; (ii) embryos were not staged at exactly stage 2 at t0 nor did they all reach exactly stage 8 at tend, introducing stage-dependent variability. The primary goal of this experiment was therefore not to precisely quantify compression rates, but to demonstrate that tissue contraction along both boundaries occurs in vivo, consistent with the barrel model predictions. The fact that contraction was observed in all six embryos analyzed, despite the inherent variability of the experimental setup, supports the robustness of this conclusion. These points have been discussed in the revised manuscript.

      (2) The region labelled as OFT appears to correspond instead to the right ventricle primordium, as demonstrated previously by cell labelling of the anterior heart field (Zaffran et al., 2004, PMID: 15217909). The nomenclature should be corrected in the figures and the text. Alternatively, the term "arterial pole" may be useful.

      We thank the reviewer for this observation. We aligned our nomenclature with the literature, correcting the labelling in figures and text.

      (3) The integration of 12 different time-lapses into the model is very impressive. However, while the early stages (2 to 5) are very well covered, the number of replicates for the later stages is much lower. Figure S4 highlights variability between some of the samples, but this is not commented on in the results or the discussion. How does this impact the averaging of tissue deformation patterns and the subsequent model predictions? We thank the reviewer for this comment. We acknowledge that the number of specimens is lower and more variable at later stages. This limitation primarily arises from technical constraints associated with long time-lapse imaging. Because embryo positioning could not be actively tracked during growth, manual repositioning was required, and since embryo development proceeded overnight, maintaining perfect alignment throughout the acquisition was challenging. As a result, several embryos gradually drifted out of the imaging volume and had to be excluded due to incomplete coverage. In addition, at later stages the onset of uncoordinated and subsequently coordinated cardiomyocyte contractions introduces motion-related blurring, which further limits image quality at the acquisition frequency used. These technical limitations were already discussed in the context of the imaging methodology and Limitation and Future Directions section in Raiola et al. (2025).

      As shown in Figure S4, variability between embryos is present and reflects natural biological diversity. Figure S4 also indicates that this variability is highly localized, whereas the regions identified as anticorrelated growth and anisotropy zones remain consistently preserved across embryos. The variability observed in Figure S4, we note that while inter-embryo variability is present, it mainly affects the magnitude of tissue deformation rather than the spatial pattern of deformation. As shown in the additional analyses presented in Figures S5 and S6, the overall organization of deformation, both in terms of growth and anisotropy, is consistently preserved among embryos within the same stage group, within the expected range of natural intra-embryonic variability.

      Finally, regarding the in-silico fate map, our model was not constructed as a statistical average but as a descriptive framework obtained from the concatenation of selected representative embryos. Constructing a statistical model was not feasible due to the limited number of embryos at later stages and the frequent occurrence of incomplete datasets (e.g., randomly missing inflow or arterial pole regions). Under such conditions, only the left ventricular primordium and the inner curvature would have been consistently preserved, thereby limiting the analysis to a very restricted and less informative region. We emphasized these points more clearly in the revised Result section.

      (4) Since the growth rate appears to be highly regionalized, could the authors provide a molecular mechanism for one of these growth patterns?

      We thank the reviewer for this insightful suggestion. Although correlating growth patterns with specific molecular mechanisms would greatly enhance the study, such an effort necessitates extensive additional experimentation, including spatial transcriptomics and detailed molecular analyses. As this falls outside the scope of the present work, we have chosen not to incorporate molecular mechanism data in this manuscript, reserving it for future research.

      (5) Could the model be used to predict new experimental outcomes? For example, could the author simulate a perturbation and validate it through in vivo experiments using mouse mutants?

      We thank the reviewer for this interesting suggestion. At this stage, the model cannot be used to predict new experimental outcomes, as it was designed as a descriptive rather than a statistical or predictive framework. The predictive potential of the model, including the simulation of perturbations, was discussed in detail in Raiola et al. (2025), where this aspect was indicated as a direction for future work.

      We clarified this more explicitly in the revised Results and Discussion sections.

      Minor points

      (1) The readership of eLife is diverse. The methodology and figures could be further annotated, and the axes (A/P, D/V, L/R) could be labelled in all figure panels.

      We thank the reviewer for this helpful suggestion. We revised the figures to include clearer annotations and ensure that the axes (A/P, D/V, L/R) are consistently labelled across all panels.

      (2) It is sometimes difficult to follow without reference to the companion paper. For example, machine learning is mentioned in the summary but is not described in this paper.

      We thank the reviewer for this comment. We clarified in the revised manuscript that the staging system is machine learning-based, using morphometric features to align specimens over time, and indicate that full methodological details are provided in Raiola et al. (2025). This will help readers understand the approach while keeping the focus on the biological findings.

      (3) The authors state the versatility of the model in the introduction, but this is not really addressed in the manuscript; please modulate.

      We thank the reviewer for this feedback. We agree that the versatility of the model was not sufficiently demonstrated throughout the manuscript. In the revised version, we rephrased the Summary to ensure that our claims are aligned with the descriptive scope shown in the current study.

      (4) The authors describe a rightward rotation of the ventricle in stage 9, which they relate to the arterial pole rotation described by Le Garrec et al., 2017. However, this event was reported to occur at E8.5f (which would be equivalent to stage 7). Please modulate or modify.

      We thank the reviewer for this observation. Heart tube rotation is a gradual process that begins at earlier stages, including stage 7, depending on embryo developmental variability. In our study, using the Atlas-based framework described by Esteban et al. (2022), this rotation becomes clearly detectable and morphologically prominent at stage 9, as illustrated in Figure 6d of Esteban et al. At this stage, rightward rotation of the ventricle emerges as the dominant feature in terms of tissue deformation and associated growth patterns, providing a robust reference point to describe and quantify the process. Thus, the description of stage 9 does not indicate the initiation of ventricular rotation, but rather the stage at which the process is most evident and measurable. We moderated it into the revised manuscript to avoid potential ambiguity.

      (5) Some rationales are missing. Why aren't all of the initial 16 time-lapses used for the cumulative deformation pattern analysis? Please explain the impact on the virtual fate mapping of using either labelling of cell clusters or cell continuums. Explain how the Strain Agreement Index neighborhood size (6-7 cells) was chosen, and whether the results are robust at other scales.

      We thank the reviewer for raising these important points. We agree that this section requires clarification and will expand it in the revised Results and Discussion. Not all 16 time-lapses could be included in the cumulative deformation analysis, as this approach relies on concatenating individual embryos into the Atlas framework while preserving the largest possible overlap of tissue. A technical limitation of our recordings was the nonsystematic loss of cardiac tube extremities (inflow tract or arterial pole) due to embryo drift during acquisition. Consequently, several time-lapses provided incomplete tissue coverage and were excluded to avoid an inconsistent assessment of cumulative deformation. In fact, some regions of the tissue would have reflected the contribution of multiple embryos, whereas others would not. Moreover, the registration required to align anatomical regions across stages and embryos would have yielded inaccurate correspondences. For these reasons, we decided to exclude such cases. We commented on this point in more detail in the revised manuscript. For the Strain Agreement Index, the choice of a 6–7 cell neighbourhood size represented a balance between local resolution and robustness. This scale was small enough to allow the tissue to be computationally flattened, while larger neighbourhoods would have included folded regions and created artefacts during the flattening step. Conversely, smaller neighbourhoods would have produced fragmented, “salt-and-pepper” patterns lacking generalization. We commented on this point in more detail in the revised manuscript.

      (6) Figure 5: The panels are mislabelled (B-C versus A-B).

      We thank the reviewer for noticing this mistake. We have corrected the panel labels in Figure 5 to ensure consistency.

      (7) Figure 5C: The red region in stage 2 within the IFT is missing.

      We thank the reviewer for this observation. We have corrected Figure 5C accordingly.

      (8) Typo in Figure 1 legend (p.5): "Our dataset includes multiple specimens raging from E7.75 to E8.25" - should be "ranging".

      We thank the reviewer for pointing this out. We have corrected the typo in the Figure 1 legend.

      (10) Figure S3 legend should state: "Deformation analysis for stage 7, stage 8, and stage 9."

      We thank the reviewer for pointing this out. We have revised the Figure S3 legend accordingly.

      Reviewer #3 (Public review):

      Summary:

      The manuscript by Raiola and colleagues entitled "Quantitative computerized analysis demonstrates strongly compartmentalized tissue deformation patterns underlying mammalian heart tube formation" takes a highly quantitative approach to interrogating the earliest stages of cardiogenesis (12 hours, from early cardiac crescent to early heart tube) in a new and innovative way. The paper presents a new computational framework to help identify both regional and temporal patterns of tissue deformation at cellular resolution. The method is applied to live embryo imaging data (newly generated and from the group's previous pioneering work). In the initial setup, the new model was applied directly to raw time-lapse data, and the results were compared to actual cell tracks identified manually, showing close correlations of the model with the manual tracking. Next, they integrated spatial and temporal information from different embryos to generate a new model for tissue movement, driven by parameters such as tissue growth and anisotropy. Key findings from their model suggest that there are distinct compartments of tissue deformation patterns as the bilateral cardiac crescent develops into the linear heart tube, and that the ventricular chamber forms by a defined expansion pattern, as a 'hemi-barrel shape', with the aterial and venous poles (IFT and OFT) acting as the harnessing belts constraining the expansion of the chamber further. Lastly, the model is tested for its ability to predict future residence of cardiac crescent cells in the heart tube, which it seems to be able to do successfully based on fate tracking validation experiments.

      Strengths:

      The manuscript provides an exceptionally careful analysis of a critical stage during heart development - that of the earliest stages of morphogenesis, when the heart forms its first tube and chamber structures. While numerous studies have interrogated this stage of heart development, few studies have performed time-lapse imaging, and, to my knowledge, no other report has performed such in in-depth quantitative analysis and modeling of this complex process. The computational model applied to normal heart development of the myocardium (labelled by Nkx2-5) has revealed multiple new and interesting concepts, such as the distinct compartments of tissue deformation patterns and the growth trajectories of the emerging ventricle. The fact that the model operates at cellular resolution and over a nearly continuous time period of approximately 12 hours allows for unprecedented depth of the analysis in a largely unbiased manner. Going forward, one can imagine such models revealing additional information on these processes, performing analyses of subpopulations that form the heart, and maybe most importantly, applying the model to various perturbation models (genetic or otherwise). The manuscript is very well written, and the data display is accessible and transparent.

      Weaknesses:

      No major weaknesses are noted with the study. It would have been very exciting to see the model applied to any kind of perturbation, for example, a left-right defect model, or a model with compromised cardiac progenitor populations. However, the amount of live imaging required for such analyses renders this out of scope for the current study.

      We agree with the Reviewer on the relevance of applying this pipeline to mutant conditions. We are engaged on those experiments but they represent a major effort beyond the scope of this manuscript, as also indicated by the Reviewer.

      Reviewer #3 (Recommendations for the authors):

      (1) Application of the model to defective heart development:

      While including perturbation models seems out of scope for the present work, some discussion on how the model might benefit our understanding of early cardiac defects, or any currently unknown mechanisms acting at this stage of development, could be included in the discussion of the manuscript. This would help highlight the enormous power that this new model could bring to understanding these critical steps during heart development, in a quantitative and unbiased manner.

      We thank the reviewer for this insightful comment. Our approach is a deterministic, descriptive framework that integrates individual tissue motion into a common spatiotemporal Atlas, providing a quantitative description of early HT morphogenesis. The primary goal of this framework is to establish a robust baseline of normal HT development under wild-type conditions.

      This baseline is essential for studying heart defects, as deviations from normal tissue motion and deformation patterns can reveal developmental defects like altered growth or aberrant morphogenetic trajectories. Currently, the limited number of embryos per developmental stage (typically 2-4) does not allow the construction of statistically robust inferential models. Nevertheless, by mapping all embryos into a unified reference system and providing quantitative descriptors of tissue motion, our framework already enables meaningful comparisons between normal and abnormal development.

      We have clarified this point in the Discussion section.

      (2) Confusion with Raiola et al., 2025:

      The manuscript frequently references an accompanying manuscript, which is currently a preprint on bioRxiv. The relationship of these two papers is not clear from the description. Not only is the majority of the data shared between the reports, but some figures seem to overlap quite substantially. The methods state that "the computational workflow is detailed in Raiola et al 2025". Any clarification on this would be helpful.

      We thank the reviewer for raising this important point and we appreciate the opportunity to clarify the relationship between the two manuscripts. The two studies indeed rely on the same underlying dataset; however, their aims and scope are fundamentally different. Raiola et al. (2025) is a purely methodological study, whose sole objective is to describe, validate, and benchmark a computational framework for spatiotemporal alignment, motion integration, and in-silico fate mapping. That work deliberately avoids biological interpretation, as the proposed approach is designed to be general and transferable to other organs or developmental systems.

      In contrast, the present manuscript represents the biological application of this validated framework. Here, the computational model is used as a tool to extract, characterize, and interpret biologically meaningful information about early heart morphogenesis, including myocardial motion patterns, regional growth and anisotropy, and fate relationships, supported by experimental validation.

      To avoid any ambiguity, we revised the Introduction and Materials and Methods to explicitly state this distinction and clarify why the methodological details are provided in Raiola et al. (2025), while the current manuscript focuses on biological insight rather than computational development.

      (3) Additional point:

      Concerning overlap with the authors' related manuscript in Bioarchive on the computational workflow: the number of specimens analysed should be noted without referral to the second manuscript (as currently mentioned in the figure legends). Is the "b" necessary when referring to the second manuscript?

      We thank the reviewer for this suggestion. We included the number of specimens analysed directly in the revised manuscript to improve clarity for the reader. Regarding the citation format, the "b" in Raiola et al. (2025) is used to distinguish between two manuscripts from the same group published in the same year.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This study presents valuable findings on the differential effects of RNA on the phase separation, aggregation dynamics, and bioactivity of PSMα3 and LL-37. The authors provide solid evidence from complementary biophysical and cell-based experiments that RNA influences peptide assembly and associated in vitro activities. The study is of interest for understanding interactions between amyloidogenic peptides and nucleic acids, although the physiological significance and some aspects of the mechanistic interpretation would benefit from further clarification.

      We are grateful for the positive assessment. The two outstanding concerns about physiological significance and mechanistic interpretation are addressed in detail below through Reviewer #2's comments. We have made targeted revisions throughout the manuscript, and have been careful to distinguish genuine clarifications from reframing that would misrepresent what the data show.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Rayan et al. aims to elucidate the role of RNA as a contextdependent modulator of liquid-liquid phase separation (LLPS), aggregation, and bioactivity of the amyloidogenic peptides PSMα3 and LL-37, motivated by their structural and functional similarities.

      Strengths:

      The authors combine extensive biophysical characterization with cell-based assays to investigate how RNA differentially regulates peptide aggregation states and associated cytotoxic and antimicrobial functions.

      Weaknesses:

      While the study addresses an interesting and timely question with potentially broad implications for host-pathogen interactions and amyloid biology, some aspects of the experimental design and data analysis require further clarification and strengthening.

      We thank Reviewer #1 for the positive assessment. Previous revision round incorporated all major quantitative additions requested:

      Quantitative EMSA binding analysis with Kd values and Hill coefficients (Fig. S1)

      Quantitative FRAP recovery curves with mobile fractions and half-times (Figs. S4, S8, S12)

      Colocalization metrics — Pearson's correlation coefficient and Manders' overlap coefficients (Fig. S5)

      Quantification of AmyTracker630 amyloid signal intensity (Fig. S6)

      Explicit acknowledgment of limitations regarding phase diagram boundaries and csat

      Revised interpretation clarifying nucleolar localization as phenomenological, not causal

      Reviewer #2 (Public review):

      In this paper, Rayan et al. report that RNA influences cytotoxic activity of the staphylococcal secreted peptide cytolysin PSMalpha3 versus human cells and E. coli by impacting its aggregation. The authors used sophisticated methods of structural analysis and describe the associated liquidliquid phase separation. They also compare to the influence of RNA on aggregation and activity of LL-37, which shows differences to that on PSMalpha3.

      That RNA impacts PSM cytotoxicity when co-incubated in vitro becomes clear. However, I have two major problems with this study:

      The premise, as stated in the introduction and elsewhere, that PSMalpha3 amyloids are biologically functional, is highly debatable and has never been conclusively substantiated. The property that matters most for the present study, cytotoxicity, is generally attributed to PSM monomers, not amyloids. The likely erroneous notion that PSM amyloids are the predominant cytotoxic form is derived from an earlier study by the authors that has described a specific amyloid structure of aggregated PSMalpha3. Other authors have later produced evidence that, quite unsurprisingly, indicated that aggregation into amyloids decreases, rather than increases, PSM cytotoxicity. Unfortunately, yet other groups have in the meantime published in-vitro studies on "functional amyloids" by PSMs without critically challenging the concept of PSM amyloid "functionality". Of note, the authors' own data in the present study that show strongly decreased cytotoxicity of PSMalpha3 after prolonged incubation are in agreement with monomerassociated cytotoxicity as they can be easily explained by the removal of biologically active monomers from the solution.

      In their revision and in the rebuttal, the authors have further described their concept regarding what they call "functionality" of PSMalpha3 amyloids. They now admit that monomers are the active cytolytic form, like other researchers have stressed, whereas amyloids are not. This represents a considerable difference to earlier papers in which they ascribed functionality, i.e. cytolytic capacity, to PSMalpha3 amyloids, a claim that has raised considerable controversy. Now, they use the term "functional " to describe that PSMalpha3 amyloids, while not cytolytic, can be reversed to a cytolytic monomeric state, calling them a "dynamic reservoir". There is no evidence that such a reservoir is necessary for the cytolytic activity of the monomers to be established; also, there is no evidence that in a biological system, such an amyloid reservoir exists. To continue calling PSMalpha3 amyloids "functional" based on this - considerably changed - concept of the authors appears inappropriate, given the finally admitted absence of cytolytic activity of the PSM amyloids in addition to the continuing complete lack of evidence of any biological relevance of PSM amyloid formation.

      That RNA may interfere with PSM aggregation and influence activity is not very surprising, given that PSM attachment to nucleic acids - while not studied in as much detail as here - has been described. Importantly, it does not become clear whether this effect has biologically significant consequences beyond influencing, again not surprisingly, cytotoxicity in vitro. The authors do show in nice microscopic analyses that labeled PSMalpha3 attaches to nuclei when incubated with HeLa cells. However, given that the cells are killed rapidly by membrane perturbation by the applied PSM concentrations, it remains unclear and untested whether the attachment to nucleic acids in dying cells makes any contribution to PSM-induced cell death or has any other biological significance. Overall, the findings can be explained in a much more straightforward way with the common concept of cytotoxicity being due to monomeric PSMs, and the impact of nucleic acids on cytotoxicity being due to lowering of the concentration of that active form by RNA attachment. Further limiting the significance of the findings, whether this interaction has any biological significance on the physiology or infectivity of the PSM producer remains largely unexplored.

      We thank the reviewer for the detailed comments. We appreciate the opportunity to further clarify our interpretation of the relationship between PSMα3 assembly, cytotoxicity, and RNA-mediated regulation. In the revised manuscript, and building on the previous revision round, we substantially expanded and refined the Discussion and Introduction to more clearly distinguish between mature fibrils, transient assembly intermediates, and broader assembly state-dependent mechanisms. We also incorporated additional literature representing different perspectives from the field. The revised manuscript presents a model in which biological activity is governed by dynamic assembly pathways and membrane-associated intermediates whose formation, persistence, and structural organization are modulated by environmental conditions, including RNA.

      A central point raised by the reviewer is the suggestion that the RNA effects observed here can be explained simply by sequestration of active monomeric PSMα3. We respectfully disagree that this interpretation can account for the data. A monomer-depletion model makes a clear experimental prediction: conditions that promote aggregation should proportionally reduce activity by reducing the free monomer pool. However, our data show the opposite behavior. RNA promotes PSMα3 aggregation, induces liquid–liquid phase separation, and reshapes fibril morphology into distinct polymorphic assemblies, yet preserves cytotoxic and antimicrobial activity over incubation periods during which peptide alone progressively loses activity. Thus, activity does not correlate with suppression of aggregation or maintenance of soluble peptide. Instead, the data indicate that assembly trajectory and supramolecular organization are functionally relevant parameters. We state this point explicitly in the section “RNA preserves PSMα3 bioactivity,” where we added text clarifying that RNA does not prevent aggregation but redirects the assembly pathway toward structurally and functionally distinct states.

      To further clarify our interpretation, we substantially revised the section “PSMα3 cytotoxicity arises from dynamic assembly intermediates.” This section now integrates multiple independent lines of evidence supporting an assembly-state-dependent model. Together, these observations argue against a simple binary model in which either monomers alone or mature fibrils alone determine activity. Instead, they support a framework in which transient intermediates formed along the assembly pathway contribute to membrane disruption and cytotoxicity. Consistent with this interpretation, our confocal and super-resolution microscopy experiments directly show PSMα3 accumulation and aggregation at bacterial and cellular membranes (Figs. 5, 6C, S10), supporting a model in which assembly occurs in direct association with membrane interfaces rather than exclusively in bulk solution prior to membrane contact. We expanded the Discussion accordingly.

      We acknowledge the reviewer’s alternative interpretation that the nucleolar/nucleic-acid association observed in HeLa cells may reflect post-lysis binding following membrane permeabilization. We agree that this is a valid consideration at the cytotoxic concentrations used here, where membrane disruption is rapid (Figs. 5–6, Movies S1–S2). The Discussion therefore clarifies that nucleolar localization under these conditions is unlikely to represent a distinct intracellular toxic mechanism, but instead reflects the intrinsic nucleic-acid binding capacity of PSMα3 after cellular entry. We accordingly do not claim that intracellular nucleic-acid interactions contribute causally to cell death in these experiments. The potential biological relevance of PSMα3–nucleic acid interactions at sub-cytotoxic concentrations, where membrane disruption does not dominate, remains an important question for future investigation.

      We additionally revised the manuscript to clarify the significance of the EGCG comparison. We agree with the reviewer that the EGCG data alone do not demonstrate “amyloid-mediated cytotoxicity,” and we do not make that claim. Rather, the comparison between EGCG and RNA provides evidence that different assembly trajectories produce different functional outcomes. EGCG redirects PSMα3 into amorphous, non-fibrillar assemblies that lose activity, whereas RNA promotes aggregation while preserving activity and generating distinct supramolecular morphologies. If activity depended solely on monomer concentration, both conditions would be expected to reduce activity similarly through sequestration. Instead, the divergent outcomes support the conclusion that assembly architecture and assembly pathway are functionally important.

      In response to the reviewer’s concern that the manuscript overstates the concept of “functional amyloid,” we explicitly distinguish between mature fibrils and dynamic assembly processes, and we avoid wording that could be interpreted as implying that mature fibrils themselves are the active cytotoxic entities. At the same time, we note that the broader concept of functional amyloid-like assembly pathways is widely used in biology to describe assemblies whose formation regulates storage, localization, stabilization, or timing of bioactive states, including hormone-storage amyloids, RNA-binding protein assemblies, and bacterial curli systems. Within this framework, our interpretation is that PSMα3 assembly dynamics modulate the availability and lifetime of bioactive species rather than that mature fibrils themselves are directly toxic.

      Importantly, we also broadened the manuscript substantially by incorporating independent studies from multiple unrelated systems supporting the principle that supramolecular organization influences biological function. These additions include: studies showing that structured fibrillar assemblies of LL-37 are required for specific antibacterial activities; work demonstrating that the nanoscale organization of β-defensin–nucleic acid complexes governs immunostimulatory potency; studies correlating α-helical solid-state conformations with cytotoxicity across fibril-forming antimicrobial peptides; salt-induced PSMα3 polymorphism studies showing distinct toxicities for amorphous versus fibrillar assemblies; and real-time AFM work demonstrating that membrane-associated protofibrillar intermediates are more disruptive than mature fibrils. We also added discussion of recent cryo-EM structures showing that RNA acts as a structural cofactor shaping tau fibril polymorphism at atomic resolution, as well as two-dimensional infrared spectroscopy studies demonstrating coexistence of cross-α and cross-β PSMα3 polymorphs. Together, these orthogonal observations from multiple systems support the broader principle that assembly architecture is a major determinant of biological behavior.

      We also addressed the reviewer’s concern regarding biological relevance. We agree that direct in vivo validation remains an important future direction and state this explicitly in the revised Discussion. However, we respectfully submit that establishing the mechanistic principle that RNA regulates PSMα3 assembly state and functional output is itself a meaningful contribution independent of immediate in vivo confirmation. To better contextualize potential physiological relevance, we expanded the “Biological and therapeutic implications” section to discuss biologically plausible extracellular environments in which PSMα3 may encounter nucleic acids, including biofilms enriched in extracellular RNA, extracellular vesicles, damaged host tissues, inflammatory milieus, and host-derived extracellular RNA released as DAMPs.

      Overall, the revised manuscript reflects a substantially expanded discussion of PSMα3 assemblystate-dependent activity, the role of RNA in modulating assembly trajectories, and the broader conceptual implications for membrane-active peptide assemblies.

      Further remarks:

      (1) Circumstantial evidence based on the "amyloid inhibitor", EGCG: The results with EGCG, which has been shown to have a moderate amyloid-reducing effect on PSMalpha 1 and PSMalpha4, should not be taken as evidence for amyloid-based cytotoxicity. While increased concentrations of EGCG reduced the cytotoxic effect of PSMalpha3, it is not convincingly shown that this is due to a lower concentration of amyloid vs. monomeric PSM.

      We agree that the EGCG data alone should not be interpreted as evidence that mature amyloid fibrils are the directly cytotoxic species. Our interpretation is more limited and focuses on the effect of assembly redirection. Specifically, EGCG redirects PSMα3 into amorphous, non-fibrillar assemblies that lose activity, whereas RNA promotes aggregation while preserving activity and producing structurally distinct assemblies. The key conclusion is therefore that functional outcome depends on the nature and trajectory of assembly rather than on aggregation versus non-aggregation alone. We clarified this distinction in the revised Discussion section addressing RNA- versus EGCG-mediated modulation of PSMα3 assembly.

      (2) It is appreciated that the authors refrain from presenting the unsubstantiated concept of "functional" PSM amyloids in the discussion. However, wording in that direction must also be removed from other parts of the manuscript (e.g. "bioactive fibrillar polymorphs". "The formation of cross-alpha amyloids has been correlated with toxic activity", etc.), generally refraining from uncritically implying that amyloid formation underlies PSM biological activity, and rather discussing that the much more likely explanation of the findings is a lowering of cytolytically active, monomeric PSM concentration.

      In the Introduction, the phrasing 'may enable dynamic switching' has been used to soften the mechanistic claim regarding cross-α assemblies. The phrase 'bioactive fibrillar polymorphs' was revised in the previous round. At the same time, statements such as “cross-α amyloid formation has been correlated with toxic activity” are retained because they describe experimental observations reported in multiple studies (including Tayeb-Fligelman et al., 2017, 2020; Malishev 2018), without implying direct causality (correlation is not causation). We now explicitly frame these observations within a broader discussion of transient assembly intermediates and assembly-state-dependent toxicity.

      (3) Discussion: "PSM alpha3 interaction with nucleic acids within human cells ...supports a comparable mechanism...". Delete. Unsubstantiated.

      This sentence was removed in the previous revision round and remains absent from the current manuscript.

      (4) The authors should cite papers that have argued against their hypothesis and not only their own manuscripts.

      We appreciate this suggestion and agree that alternative interpretations should be represented explicitly. In the revised manuscript, we added and discussed studies including Zheng et al. (2018) and Yao et al. (2019) (already cited in both earlier versions), which support models in which advanced amyloid formation reduces cytotoxicity and active species are prefibrillar. These studies are now discussed substantively in both the Introduction and Discussion alongside our own work and that of others.

      More broadly, we revised the manuscript to present the current understanding of PSMα3 toxicity as an actively debated question in the field rather than as a settled model. At the same time, we note that citing our prior studies remains necessary where the present work directly builds upon previously reported structural, biophysical, and mechanistic observations.

      If the reviewer has additional specific references in mind, we welcome them and will incorporate them.

    1. Author response:

      The following is the authors’ response to the original reviews.

      In preparation for release of the analysis code used in the paper, we made many analyses more parallel to one another in their exact preprocessing. This resulted in very slight changes to many panels, but these changes are nearly invisible and conclusions did not change. In one case, though, we realized that the way we were presenting data was potentially misleading (the timing plot in Figure 3A). The original plot was of the distribution of pixel values from the spatially smoothed map instead of distributions over individual neurons. We have now swapped it out for better interpretability and changed the accompanying text accordingly.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Here, the authors address the organization of reach-related activity in layer 2/3 across a broad swath of anterodorsal neocortex that included large subregions of M1, M2, and S1. In mice performing a novel variant water-reaching task, the authors measured activity using two-photon fluorescence imaging of a GECI expressed in excitatory projection neurons. The authors found a substantial diversity of response patterns using a number of metrics they developed for characterizing the PETHs of neurons across reach conditions (target locations). By mapping single-neuron properties across the cortex, the authors found substantial spatial variation, only some of which aligned with traditional boundaries between cortical regions. Using Gaussian mixture models, the authors found evidence of distinct response types in each region, with several types prominent in multiple cortical regions. Aggregating across regions, four primary subpopulations were apparent, each distinct in its average response properties. Strikingly, each subpopulation was observed in multiple regions, but subpopulation members from different regions exhibited largely similar response properties.

      Strengths:

      The work addresses a fundamental question in the field that has not previously been addressed at cellular resolution across such a broad cortical extent. I see this as truly foundational work that will support future investigation of how the rodent brain drives and controls reaching.

      The quantification is thoughtful and rigorous. It is great that the authors provide an explanation for and intuition behind their response metrics, rather than burying everything in the Methods.

      The Discussion and general contextualization of the results are thorough, thoughtful, and strong. It is great that the authors avoid the common over-interpretation of classical observations regarding cortical organization that are endemic in the field.

      All things considered, this is the best paper regarding spatial structure in the motor system I have ever read. The breadth of cellular resolution activity measurement, the rigor of the quantification, and the clear and open-minded interrogation of the data collectively have produced a very special piece of work.

      Thank you! We really, really appreciate this!

      Weaknesses:

      The behavioral task is very impressive and an important contribution to the field in its own right. However, given that it appears substantially different from the one used in the previous paper, the characterization of the behavior provided in the Results is too brief. More illustration of the behavior would be helpful. For example, it is rather deep into the paper when the authors reveal that the mice can whisk to help localize the target location. That should be expressed at the outset when the behavior is first described. Other suggestions for elaborating the behavior description are included below.

      Thank you. Although the task will be treated in greater detail in the next paper (where we more closely relate neural activity to the kinematics), we have added more exposition of the task here. In particular, we now include a figure with a characterization of the trial-to-trial variability across reaches to the same target versus across reaches to different targets (Figure 2-figure supplement 1B). This supports the idea that the mice aimed their reaches. We have also expanded that text.

      Regarding whisking, we have now revised that text to make clear that we do not know how the mice localize the spout. The original work by Galinanes and Huber argued that they find the spout by sniffing the water; they may do the same here, or may find it via whisking. It is also possible that the whisking they do is simply because the spout moves in and they are excited, or startled, or do it by reflex. We simply have no evidence one way or another. We have therefore revised the text to make it clearer that whisking-related activation could have occurred for a variety of reasons.

      Statistical support for key claims is lacking. For example, "The five areas of interest varied in the fraction of neurons that were modulated: M2 had 14%, M1 had 23%, S1-fl had 30%, S1-hl had 25%, and S1-tr had 27%" - I cannot locate the statistical tests showing that these values are actually different. Another example is Figure 7, where a key observation is that distributions of PETH features are distinct across regions. It is clear that at least some distributions are not overlapping, but a clearer statistical basis for this key claim should be provided.

      Good idea. For the proportions, we have now added first a Chi-square test for homogeneity to show that there is variation in the proportions, then shown the results of pairwise two-proportion Z tests (Bonferroni-corrected for multiple comparisons) as a binary matrix in Figure 3-figure supplement 1B. For the area distributions in the t-SNE space (Figure 7), we have added a 2-dimensional Kolmogorov-Smirnov test, again corrected for multiple comparisons, with p-values quoted in the text.

      I understand that the authors are planning a follow-up study that addresses the relation between activity patterns and kinematics. One question about interpreting the results here though, is how much the activity variation across target locations may relate to the kinematic differences across these different conditions, as opposed to true higher-order movement features like reach direction.

      We agree this is a very important question. However, having done many of the analyses to examine the question for the next paper in the series, we do not know of a shortcut to the right answer. This question requires thorough treatment, and so we leave it to be covered in subsequent work. Instead, after our speculation about how responses suggest function, we are now explicit that these hypotheses needs testing:

      “In each of these cases, determining the relationships of the observed activity patterns to function will require specific attempts to link the activity to kinematics, target location, sensory feedback, and more; these relationships will be addressed in future work.”

      Reviewer #2 (Public review):

      Summary:

      The functional parcellation of cortical areas is a critical question in neuroscience. This is particularly true in frontal areas in mice. While sensory areas are relatively well characterized by their tuning to sensory stimuli, the situation is much less clear for motor areas. This has become even more ambiguous since recent studies using large-scale neuronal recordings consistently report mixed sensory and motor-related activity throughout the brain, and motor mapping studies have shown that movements evoked by cortical stimulation are by no means limited to motor areas alone. Here, the authors use a correlation approach combining large-scale functional imaging at cellular resolution with movement-tracking in mice executing a reaching task. Across multiple recording sessions in the same animals, the authors have imaged a large portion of the sensorimotor cortex at cellular resolution in mice performing a reaching task, recording the activity of nearly 40,000 neurons. By aligning the calcium signal of each neuron to three task events-the Go cue triggering the reach, the onset of paw lift, and the contact between the paw and the target-for different target positions, the authors identified different response patterns distributed differently across cortical areas. They defined a set of features that describe the neurons' response pattern, representing the temporal dynamics and tuning properties for the different target positions. These features were used to construct cortical maps, and the authors show that, interestingly, gradient maps obtained from the first derivative of the feature maps reveal sharp discontinuities at the boundaries between anatomically defined cortical areas. Using dimensionality reduction of the neuronal response features, the authors found that, despite clear differences in their average response properties, individual neurons from the same cortical areas do not form distinct clusters in the reduced-dimensional space. In fact, most areas contain heterogeneous neuronal populations, and most neuronal populations are present in multiple areas, albeit in different proportions. Interestingly, the authors identified four neuronal subpopulations based on the distance between the components of the Gaussian mixture model used to model the distribution of neurons within each area. One of these subpopulations is almost exclusively represented in the anterior M2 cortex, while another is broadly distributed across the different areas.

      Strengths:

      This article is based on an impressive dataset of nearly 40,000 neurons covering a large portion of the sensorimotor cortex and on innovative analytical approaches. This study is likely the first to clearly demonstrate boundaries between cortical areas defined based on the responses of individual neurons. This innovative approach to functional mapping of cortical areas potentially opens up new perspectives for higher-resolution mapping of frontal cortical areas, using a broader repertoire of sensory and motor evoked responses.

      Thank you!

      Weaknesses:

      The second part of the article, which presents multimodal responses in the cortical areas, seems to be a perhaps overly complicated way of showing what has already been demonstrated in numerous recent publications, but these new analyses expand upon these previous observations by revealing an interesting functional organization of the sensorimotor cortex, highlighting interesting similarities and differences between certain areas.

      We understand the concern: a number of recent papers have also noted different neuron response characteristics distributed throughout the motor system. We compare and contrast in greater detail following the more specific comments on this below, but we briefly summarize here. The way previous work handled the data – for example, starting with PCA – mixes what neurons are tuned for and when they are tuned for it with what we refer to as the “response format”: properties like tuning sharpness, response duration, etc. We focused primarily on this response format, and designed our features to be mostly independent of tuning preferences or peak response timing. We therefore pick up on different properties of neurons’ responses than those prior works. In addition, no previous work we know of examined these properties across large swathes of cortex at single-cell resolution in the context of forelimb control. Together, these aspects of our work allowed us to produce high-resolution mapping of response properties in a way we have not seen in any prior work.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      In addition to addressing the weaknesses stated above, I suggest the authors also consider the following.

      The one big question left unresolved here is whether we should be thinking about these four subpopulations as distinct types with a biological basis and importance, or just reflections of activity pattern heterogeneity. The authors say that "we did not observe tight clusters in feature space separated by gaps," but their discussion here is light and a bit unclear, and their engagement with the issue of types versus heterogeneity, in my view, could be improved. We do not need "gaps" where the density goes to zero in parameter space, but we do need reproducible troughs between peaks. The authors should clarify if there are substantial and reproducible troughs in the parameter space between their four subpopulations.

      This is a great idea, and we have added three analyses and additional text to address it. We break this concern down into two more specific questions, based on the next comment by this reviewer.

      (1) Are the clusters well separated / do they have troughs between them? (Note that even with troughs, clustering might not be stable if the clustering algorithm is poorly matched to the shapes of the clusters.)

      (2) Is the clustering stable? (It can be stable even without troughs, if, for example, the distribution has a long tail and a GMM needs one Gaussian for the body of the distribution and a second for the tail.)

      First, to directly address the presence or absence of troughs between clusters, we have added Figure 9-figure supplement 2A and 2B. For each pair of subpopulations, we trained a logistic regression classifier to separate the 5D feature vectors of the neurons in one subpopulation from the feature vectors of the neurons in the other subpopulation, then projected the feature vectors onto this axis. Note that because the subpopulations are defined by GMMs, which have nonlinear boundaries, the (linear) logistic classifier does not typically produce perfect classification. Nevertheless, this analysis provides a window onto how well separated each cluster is from each other cluster in feature space. In 5 of the 6 pairwise comparisons, it is obvious that the distributions are different and have at least some dip in the distribution density at the boundary. The one pair of clusters without a trough between them were the forelimb somatomotor and hindlimb somatomotor subpopulations. This was surprising to us, given that their likelihood maps are so strongly distinct, but this presumably reflects trying to capture a nonlinear classifier boundary with a linear one (see below). Overall, this analysis argues that the clusters do have fuzzy edges that blend into one another, but reflect concentration of mass near the centers of the clusters we identified.

      Second, to address the same question with a different nonlinear method, we have added a version of the t-SNE plot from Figure 7 that is instead colored and contoured by subpopulation identity instead of area (Figure 9-figure supplement 2B). Agreement with the GMMs is not a given here either, because t-SNE is a fundamentally different and independent nonlinear transform from that performed by the GMM classification. Nevertheless, the subpopulations were again nicely separated – though not with troughs, possibly thanks to the inherent difficulty of interpreting point density with t-SNE. Interestingly, here the hindlimb somatomotor subpopulation was the best separated from the other subpopulations, supporting the idea that the lack of separation we observed above with the logistic projections was indeed due to a nonlinear boundary. This analysis again argues that neurons are more likely to have features that lie near the center of a cluster, but that the edges of the clusters run into one another. Additionally, this analysis makes clear that treating the hindlimb somatomotor subpopulation as a second cluster can be supported by other analyses, even if not by the logistic regression projection.

      Third, to address the question of cluster stability, we have performed random splits of our data, GMM clustered the two halves independently, applied the GMM from one half to the other, and asked how similar the clusterings are using the Adjusted Rand Index. This produced a value of 0.856, which for this sensitive measure argues that the clustering is rather stable (at least for the three clusters that can be found with all data together, which does not include the smaller-in-size Anterior subpopulation). Note that we did not perform this analysis on the more complicated version where we fit a GMM to each area separately then cluster those; in our main analysis, the hierarchical clustering agreed with what we found by eye, but determining the number of clusters for hierarchical clustering is in general very unstable and so we did not have an objective way to determine the “true” number of clusters.

      In addition to these new analyses, we note that three analyses we had already included bore strongly on this issue. Regarding separability of the clusters, the fact that our likelihood maps (Figure 9C-F) were quite distinct for different subpopulations argues that we picked up on ‘real’ differences. Second, Figure 9B found that when clustering non-overlapping data – different cells from different areas – we obtained clusters that were nearly identical in their feature distributions. Third, Figure 10E used the clusterings from different areas’ data to create likelihood maps, and found that they were extremely similar. These analyses together argue strongly that we are finding ‘types’ in a meaningful sense; given that we know the areas do have different distributions of properties, if there weren’t types then clustering would yield different clusters for different areas. Given the importance of the question, however, we are grateful that the reviewer encouraged us to find additional ways to make this point!

      The original t-SNE plot is beautiful and quasi-fractalic, but it does not show clear signs of four cell types. The single-neuron activity profiles are clearly heterogeneous in very interesting ways, but heterogeneous does not imply a strong or reproducible multimodality that would indicate meaningful cell types. Clustering algorithms will always spit out an answer. If you just have elements uniformly distributed across a parameter space, plus some noise, when you ask for X clusters, you will get X clusters that have different centroids. When you ask an algorithm to cluster without defining the number of clusters, noise can lead the algorithm to produce a particular number of clusters that again will have distinct centroids. The salient question, though, is whether in the present case there is a parameter space in which the clusters are substantially and or reproducibly distinct. Distinct here would mean that peaks in the density across some parameter space are separated by troughs - again, we don't need true gaps. The more substantial the differences between clusters are (again, not the differences between centroids but the prominence of the density troughs between them), the more biologically meaningful the clustering is likely to be. Reproducibility here could be addressed with resampling methods (e.g., how often do two separate halves of the cells produce the same clusters?).

      Please see the reply above, which includes our addressing of this concern.

      The Introduction is generally good, but it could further develop existing ideas about how function is distributed across cell types and regions. We would like to be able to imagine different answers to the question of how activity patterns are organized that might have divergent implications for how the circuit works. I understand we have very little to go on in terms of data, but I think it would be helpful for readers to be given more of a sense of what *could* be important.

      Good idea. We have added such a paragraph to the Introduction:

      “To frame possible outcomes, consider that single neuron responses can vary along many dimensions. Cells could differ according to which movements or time periods they are recruited for (tuning), what movement parameters their activities reflect (encoding), or how their responses are structured across different movements (e.g., nonlinear encoding structure). Further, differences in these response properties across cells could be distributed over the cortical sheet in a variety of ways. Cells could form distinct “categories” or clusters that are spatially well-aligned to the boundaries of anatomically defined regions. Or, categories of neurons might span area boundaries in spatial footprints that do not relate obviously to area boundaries, and that either abut or overlap. At a fine-grained scale, cells with similar responses could be physically located near one another as in primate and feline visual cortex, or similarly-responsive neurons might be salt-and-pepper intermingled as seen in rodent visual cortex or in primate motor cortices during reaching behaviors.”

      It should be clarified in the Results how the cue relates to the target location. Most would assume a different cue for each location, but this does not appear to be the case. The authors should clarify whether there was some amount of searching for the precise target location after the reach, or else how the block structure or other sensory information allowed mice to learn where exactly the target would be. In the absence of target-specific cues, some sense of how the mice achieved target-specific reach trajectories should be offered.

      Related to this, in Figure 1, it would be good to see some individual trajectories, as they all overlap near the target in the current plot. Clearly, the reaches were targeted, but it is unclear how targeted. Some of the adjustments at the end may reflect searching or palpation to resolve the precise spout location. It is very much ok if the mice were not reaching with micron precision each time to each of 15 different targets, but it would be good to provide the reader a better sense of what the mice were doing.

      These are important points. First, to clarify, the Cue is just a Go cue, and was the same for all targets. It is now described in the Results as “non-target-specific”. For additional explanation about supplemental analyses to assess “aiming”, see replies to Reviewer #1 Public Review comments above. Finally, regarding how the mice locate the target: we just don’t know. As discussed above, Galiñanes and Huber found evidence for the mice using stereo sniffing, but whisking, listening to the motors, or some other strategy are also conceivable. We simply don’t have data to weigh in on this. We now make this limitation clear where we describe the task.

      In Figure 1A, CFA does not look well aligned with Tennant et al. (2011). CFA should only extend to +1 AP. The overlap of CFA and RFO seems strange. RFO also does not totally align with the injection coordinates used in An et al, biorxiv 2022.

      Thank you for your attention to these points. Our designation of the name CFA to the red dashed outline in Figure 1A was consistent with an earlier version of our previous work (Grier et al 2026) wherein we referred to the anatomical outline “MOp-ul” from Munoz-Casteneda et al 2021 as CFA. We have since revised that nomenclature to now refer to the outline as M1-fl, or the forelimb representation of primary motor cortex.

      Our placement of RFO was obtained by aligning the Allen CCF from Figure 1K of An et al 2022 to our version of the Allen CCF and outlining the hotspot of RFO with a circle. We have slightly adjusted the location of RFO posterior and medial to more closely align with the injection coordinates reported in the methods of An et al. 2022 of “1.5-1.88 mm anterior from Bregma, 2.25-2.63 mm lateral from the midline.” Because (as far as we understand) the injection coordinates and the map are not perfectly in register, we show a compromise between the two.

      We stress that the Figure 1A map is meant to be descriptive in its illustration of the variety of organizational zones that have been identified across mouse sensorimotor cortex.

      Discrepancies in the alignment procedure, animal strain, and mapping modality all introduce heterogeneity across mapping attempts that we do not aim to reconcile or resolve here.

      Related to this, aspects of the results do seem consistent with the distinction between RFA and CFA, but this is not acknowledged or discussed. For example, the barriers in Figure 6H that lie along the M1/M2 border - these seem consistent with the gap between RFA and CFA. The same could be said for the dim trough along the M1/M2 boundary that appears to separate RFA and CFA in Figure 3B. A slightly more rostral and lateral location of CFA compared to Tennant's definition or the regions backlabeled from cervical spinal injections (see Wang, Maunze et al. J Nsci, 2018) could be expected if flattening the brain under the coverslip for imaging effectively stretches the ML axis, and Bregma (notoriously hard to define reliably at this spatial scale) was defined a bit more caudally here than in other studies. Related to this, it would be better for the field if people described their method for defining Bregma in the Methods. I suggest the authors do this here.

      We appreciate the suggestion and have acknowledged the suggested correspondence in the discussion. Given the difference in our approach from those that originally characterized RFA (through ICMS and deep layer projection tracing) we have avoided making overly strong conclusions about this correspondence in our data. See the quoted text below.

      “The spatial distribution of modulated cells in Figure 3 suggests a distinction between the caudal forelimb area (CFA, involving M1 and S1-fl) and the rostral forelimb area (RFA) in M2, while the feature gradient boundaries suggest a distinction between M1 and M2 more generally. The absence of a clearly delineated RFA was surprising, given its distinct projection patterns (Carmona et al. 2024; Hira et al. 2013b; Wang et al 2018) and functional differences from CFA (Kristl et al. 2025; Morandell and Huber 2017; Saiki-Ishikawa et al. 2025), but our results might suggest that the activity in layer 2/3 of RFA does not differ markedly from other nearby subregions of M2.”

      Regarding bregma, we did not use it for atlas alignment here. Alignment was accomplished through a combination of paw vibration mapping and the location of the central sinus. Bregma’s location was only relevant for our injection of tdTomato labeling, and that labeling was used here only to stabilize the image plane. We include an estimate of it on the map solely in an attempt to be helpful, but we cannot claim we have the most reliable method for defining it.

      The authors focus on activity aligned to cue timing. This is sensible, but it could be meaningful to know how this choice affects the definition of organization. If response clustering is largely different across time, it would seem important. I understand that addressing this question may be beyond the scope of this paper. I just wanted to raise the issue with the authors for their future consideration.

      We agree that this is important to address directly. There are two aspects to this comment: (1) does it matter if activity from approximately the same time period is aligned to the paw lift or contact instead of the cue? (2) What changes if we use data from a different period of time?

      Regarding the first question (alignment), if we switch to aligning our data based on lift or contact, we have more statistically modulated neurons (see Figure 3C), but everything else is qualitatively similar with one exception: the GMM optimization doesn’t separate out the Anterior subpopulation from the Forelimb Motor subpopulation. The Anterior subpopulation only has a relatively small number of members, and they mostly exhibit the strongest peaks in their PETHs when Cue-aligned, so this makes sense. We now show the modulation maps for all of the locking events (Figure 3-figure supplement 1).

      The issue of the time window is a little more complicated. There are many choices we made in this work, of course, not least of which are the task we used and the features we chose based on hand-inspection of thousands of PETHs. As we noted in the Discussion, different tasks or different features would likely distinguish more subpopulations from one another. We think of the time window as a feature choice, albeit an implicit one. We chose not to include later time points because this begins to strongly include reward signals, which are known to be large (Levy et al 2020) and can dominate other aspects of the responses. The largest differences we noted when trying time windows that extended later are that mouth-related areas are separated out in the subpopulation analyses, perhaps because of later licking/consummatory responses, but we have not explored fully enough to speak confidently on this point without much more work and another 10 figures. To keep the scope of the paper manageable, we now call out this choice explicitly (see text below). We thank the reviewer for raising these important points.

      “Crafting additional PETH features, or using end-to-end neural network approaches to discover other features, might enable the discovery of additional structure (Minderer et al. 2019; Wang et al. 2023b). For example, our PETH features were chosen to be invariant to the onset time of activity, but these onset times were markedly later in lateral M1 than in adjacent M2 or S1-fl. Including onset times, using a wider window of time that includes more of the reward/licking period, aligning data to other behavioral events, or adding other PETH features would presumably result in finer subdivisions of sensorimotor cortex.”

      The map in Figure 4 is very cool, and the spatial structure is quite striking. In terms of the actual values of the onset times in each region, I am a little concerned with a dependence on the level of reach-related activity modulation, especially relative to the level of background activity (potentially related to posture). Less reach-related activity and more background activity, which we might expect for trunk and hindlimb regions, could seemingly skew the onset times earlier. We could be getting the right answer, or an answer that makes intuitive sense, for the wrong reason. Can this potential confound be excluded with some sort of control analysis?

      The previous text wasn’t clear. We have now clarified what we meant, very much in line with the reviewer’s thoughts. In addition, note that our change to what is displayed in the histogram (now neurons, previously pixel values) makes clearer that there is a multi-peaked distribution of onset times and it is mostly the prevalence of each peak in each area that varies. The text now reads:

      “These distributions over neurons revealed clear differences in the overall profile of activation: early onsets were more prevalent in S1 trunk and hindlimb regions, perhaps due to activity related to the animal stabilizing itself even if the neurons became more active later; then M2, and finally S1-fl and M1. Nevertheless, each area contained neurons activated at any given time in the trial.”

      The "Peak time variation" metric could potentially vary with activity level, with lower, noisier activity levels making cells appear less persistent. Perhaps a control analysis, based on SNR or some reasonable assumptions of the linkage between calcium signals and spiking, could be performed to measure the extent to which this could be creating differences between regions.

      Good idea. We have now performed this analysis, and the reviewer was correct: the correlation between peak time variation and a simple metric of SNR (assessed as range of PETH / max s.e.m.) was substantial: ⍴=-0.53. We now report this correlation and describe in the Results that this metric is driven by both true peak time variation and trial-by-trial variation. Thank you for this!

      “Peak time variation. To quantify whether a neuron’s firing peaked at the same time for every target or varied by target, we found the peak firing rate of the response to each target, then computed the standard deviation of these peak times across targets. This value is therefore higher if the peak time varied and nearly zero if the timing was consistent. Notably, this measure correlated substantially with overall signal-to-noise ratio of a neuron’s PETH (Spearman’s ⍴=-0.53; Methods), and thus partly measures trial-to-trial variability, not just true peak timing variability. This metric was quite low in M1, indicating highly consistent timing of the activity peak (and reliable responses), and was highest in the posteromedial part of M2 (presumably corresponding to the hindlimb representation) and the posterior tip of S1-hl (Figure 5B).”

      One could argue that the likelihood calculations illustrated in Figure 8 are biased higher for neurons within each region since they were used for defining the likelihood for that region. I think these likelihood calculations should be done for separate neurons other than the ones used to compute the mixture model for each region.

      We agree with the point about bias: the by-area GMM in Figure 8 is biased toward cells within the area, though the effect is probably quite mild given the large numbers of neurons and modest number of parameters. However, this model was intended to make the point that even if you give an area an unfair advantage, you still can’t cleanly isolate it. This was intended to help motivate the following analysis of subpopulations, and we have now made this logic clearer. Doing it this way has the advantage that the GMM components are identical between Figures 8 and 9, while if we held out the test neurons it would not be possible to make them the same without some complicated version of bagging on the GMM components. The reviewer is right that we should make this bias explicit, though, and we have now done so:

      “This mapping approach is explicitly biased toward finding feature differences between areas, allowing for a direct test of the hypothesis that response profile distributions are area-specific.”

      To me, the last Results section (Spatial overlaps between subpopulations indicate intermingled members) does two things: it shows you get the same results when you map each cell to a subpopulation independently of its area, and it shows that defining the subpopulations with cells from each area gives you essentially the same results, arguing against spatial variation of properties within subpopulations. I worry that these two points are getting merged together or not made clearly enough here, especially the first one. In general, the logic of this section does not seem well conveyed.

      Thank you for the feedback. In particular, your first point is made by Figure 9-figure supplement 4 when we fit an area-agnostic GMM to all modulated cells in the five main areas. However, your second point is one of the two main goals of the last Results section, along with the demonstration of the spatial distributions of cells after hard-clustering them by subpopulations. We have tried to clarify these main points further through substantial edits of the results section for Figure 10.

      One set of ideas that is highly relevant and should be raised concerns an ethological organization of the motor cortex. Since the observations of Graziano, there has been a steady stream of results describing ethological organization in rodents as well. This literature is briefly reviewed in Kristl et al., Nature Communications, 2025. For example, because of the potential for a differential involvement of grasping movements across different target locations, some of the variation in neuronal tuning described in the present manuscript may stem from a region preferentially involved in grasping.

      We agree that the Graziano literature, and the substantial literature in rodent that was inspired by Graziano’s work, is highly relevant to understanding the organization of motor areas. Kristl 2025 handles these issues very thoughtfully. The challenge here is that there are many possible different reconciliations of the stimulation results with ours, and some seriously unresolved challenges in doing so. To name a few:

      Our subpopulations and high-gradient boundaries both give quite different pictures than microstimulation does in rodent motor and sensory cortices. In particular, microstim produces more subregions that evoke different movements than we identify, and the borders don’t generally line up. This implies that the mapping between the two approaches is probably complicated.

      There is a completely alternative possibility to explaining the Graziano-like results: microstimulation is thought to preferentially hit axons, and some of these projections reach the medullary motor regions. Given that the medullary motor regions have known topography in the movements they evoke (Yang et al 2023) – but may or may not be driving the movements during flexible behavior – the two approaches may not be reconcilable. Or, it may require a much deeper understanding of medulla as driving the primary movement and cortex acting as a residual controller. This is an exciting set of ideas, but as yet very underdeveloped in our understanding.

      We don’t know if the subpopulation structure exists at all in L5, or in the PT cells, and if it does whether it differs. This is crucial given the frequent targeting of deep layers by ICMS stimulation protocols.

      As we caution in the Discussion, it is possible that our subpopulation findings are at least partly specific to the task we used.

      Although it is beyond the scope of this paper and will be addressed thoroughly in separate work, we have spent significant time with encoding models for joint angles and high-level target encoding in these same data. Given those results, we are fairly confident that the reviewer’s reasonable guess, of tuning variation due to intersections between body parts, does not seem to be the main driver of the subpopulation structure we find.

      After careful thought and discussion amongst the authors, we did not think that including this discussion in the paper was likely to improve interpretability of the present results for most readers. We very much agree with the point, though, and when we can narrow down the possible explanations in the future (likely in our next paper on this topic, which will address encoding) we plan to address it. We thank the reviewer for encouraging us to think through this.

      Minor:

      (1) Page 3: "densely shared" - perhaps "broadly shared"? Dense implies most/all the neurons get the same signals, which may not be true.

      Changed to “widely”.

      (2) Page 4: "data-driven approaches" - could be more specific - isn't everything we do data-driven?

      Changed to “bottom-up”.

      (3) Page 4: "spanned areas" - perhaps "spanned multiple cortical areas", since everything spans an area.

      Changed to “spanned multiple areas” (we mention cortex just a few words earlier).

      (4) Page 5: "intervals were generally fast" - awkward, "short" perhaps.

      Agreed, changed.

      (5) Page 5: "which asks whether the activity for a neuron changes over time consistently in relation to any target" - Rephrase to disambiguate between consistent temporal variation in firing for all targets and variation across targets in the firing patterns. In other words, are we talking about cells that are just modulated during reaching, or cells whose firing patterns differ across targets?

      Changed ending to “to any given target”. The ZETA measure really does simply ask whether there is a change in firing rate over time that is consistent across trials, for each target independently. A neuron that exhibits an identical bump for all targets would register as modulated. We chose this measure in part because of the number of temporally-modulated but untuned cells. This wasn’t very clear as we had written the text, so we now note this explicitly in the Methods. Thank you for pointing out that this wasn’t clear.

      “For all analyses, only neurons modulated by the relevant locking event were included. Note that this measure looks for modulation over time to any target; it is indifferent to whether the neuron exhibits tuning across targets.”

      (6) Figure 1: It seems like some of the abbreviations used in 1A have not been defined yet in the paper.

      Yes. It’s a long list, and we wanted to put the citations for the description of each area together with the definition of the acronym. Moreover, we wanted all this info together with the description of how we aligned these area descriptions from others’ work with one another on the Allen atlas. This was impractical in the caption, and would be a long digression for what is intended as a simple point in the Results, which is why we refer to the Methods here.

      (7) Page 8: "Given that these areas have known spatial organization within them and structure was apparent by eye in the spatial scatterplot of modulated neurons (Fig. 3A)," - it is not clear what spatial structure we are supposed to see in 3A.

      Good point. We have changed the parenthetical to: “(for example, the less modulated band along the M1/M2 border in Fig. 3A)”

      (8) Page 8-10: The region-wide onset analysis breaks up the flow from PETHs to the metrics used to quantify them. I suggest moving this section (Onset of neural activity varied with somatotopy and subregion) to later in the manuscript.

      We appreciate the reviewer’s input on organization. We went back and forth many times in how to organize the many results in this paper. The reviewer is right that this analysis breaks the flow, but the reason we included it where we did was threefold. First, it uses an easily-understood metric to introduce the reader to how we made maps from single-neuron features. Second, it easily introduces the power of making such maps. Finally, it makes clear that if we are not careful with how we handle time in the feature design, timing will dominate.

      All these things said, this has helped inspire us to add a result in which we re-examine timing broken down by subpopulation (Figure 9-figure supplement 2C). It shows that subpopulations timing distributions appear more distinct than distributions for areas, but there is still substantial heterogeneity in timing that is explained by location in cortex and not subpopulation membership alone.

      (9) Page 12, Target tuning linearity: This metric should be clarified in the Results. It is not clear how the 2D of targets is turned into 1D. Also, the plot in the figure has correlation on the y-axis, and it is not clear how each target location gets its own correlation value. The phrase "optimized anchor target" is unclear.

      Agreed this needed to be clearer. The text in the Results now reads: “To quantify how linearly a neuron’s activity related to target location in physical space, we correlated the 15D vector of mean activity of the neuron for each target with the 15D vector of the targets’ ordinal distances from the neuron’s preferred target (Methods).” In agreement with your suggestion, we have dropped use of the phrase “anchor target” in favor of “preferred target”, which should be clearer. We have also revised the Methods text accordingly to clarify.

      To directly answer your question, we turn the targets from 3D positions into 1D by computing the ordinal distance of each target from a preferred target. (Note that the preferred target is actually the one that maximizes the resulting correlation; this is detailed in the Methods). There therefore aren’t 15 correlations; we’re correlating two 15D vectors, where each has one element per target and the “ordinal distance” vector has a zero for the preferred target. Hopefully the new description makes this clearer.

      The figure schematic was unclear, thank you for catching that. We have updated the Y axis to read “mean activity” and the X axis now reads “dist. to pref. target.”

      (10) Page 12, paragraph beginning "We also compared our metric maps simply using the top 20 PCs." - This paragraph is unclear, since both sentences refer to using the metrics. I would guess the authors mean that the metric maps were compared with and without PCA and basis rotation, but this is not clearly stated.

      Thank you, this was unclear as written. We have changed it to:

      “We also compared our metric maps with maps generated from the top 20 PCs of the PETHs (Methods), rotated using VARIMAX to identify a sparser basis (Musall et al. 2019).”

      (11) Page 18: "These results make clear that the working hypothesis - of areas with well-separated feature distributions - is incorrect." This is the clearest statement of the impact of the results. The authors could consider including this in the Abstract or Introduction.

      Thank you for pointing this out. We agree, and have added a similar phrase to the Abstract.

      (12) Figure 9: It would be great to also just see the average PETHs for each of the four clusters to get a better sense of how their time series differ.

      Good idea. The feature computations are a many-to-one mapping, so it’s not possible to literally generate a PETH from the mean of the cluster, but we have added PETHs from well-modulated neurons that are near the means of their subpopulations (Figure 9-figure supplement 1).

      (13) Figure 9B: Colorbar has no label.

      Fixed, thanks.

      (14) Figure 9C: Need a colorbar - need to see the difference in density for locations.

      The color map is the same Figure 8B, which is now noted in the caption for Figure 9C. The scaling of likelihoods is almost totally uninformative; they’re not well-behaved like probability distributions, so you’ll note that even on Figure 8B the labels are simply “max likelihood” and “min likelihood”. The important pieces of information here are that these are log likelihoods (noted in the Figure 8 caption), and the visualization of the color map itself (from the color bar). Given these considerations, we have elected to keep the maps themselves a little larger by not trying to squeeze in a minimally-informative colorbar to all of the plots, but thank you for noting that the reference to 8B was needed.

      (15) Page 22: "additional spatial structure could be present" - The nature of the additional spatial structure here is a bit opaque. The authors could clarify what additional structure may be present.

      Good idea. This paragraph now reads:

      “The overlaps in the subpopulation likelihood maps above imply that members of different subpopulations are spatially intermingled, but it is less clear whether each subpopulation has homogeneous response profiles across space. In particular, the use of likelihoods mixes two properties: the fraction of neurons in a given neighborhood that are members of each subpopulation, and the heterogeneity of response profiles amongst members of that subpopulation. These properties could vary systematically with respect to one another, and the spatial structure shown by the likelihood map does not disentangle them.”

      (16) Figure 10E, legend: "GMM component" - I think this should be "GMM subpopulation" to avoid confusion with the previous use of "component" above, referring to the components of the GMM models for each region.

      Thank you – good catch. Changed to “Likelihood map”.

      (17) Page 24: "Note that this consistency also validates the use of clustering to combine components and identify the subpopulations in the first place." - I don't totally get this, and how this result validates the method of combining components, as opposed to just clustering all the cells from all regions at once. Perhaps the implied opposing strategy is not clear here.

      We have changed this sentence to:

      “Note that this consistency mirrors the low Bhattcharyya distances between corresponding GMM components in Figure 9B, and further validates the use of clustering to combine components from different areas.”

      Regarding the reviewer’s larger point, we have three thoughts. First, we do also show the result of fitting the GMM to all cells together (Figure 9-figure supplement 4).The result is similar, but the Anterior subpopulation is lost because its membership is low and so the ICL criterion can’t justify a fourth cluster. Second, because we imaged more neurons in some areas than others, fitting the GMMs to each area separately put their representations on a more equal footing. Finally, doing the analysis this way allowed us to most directly compare our two hypotheses, as illustrated in Figures 8A and 9A.

      (18) Page 25: "in the zones where different subpopulations overlapped" - I would omit this, since "intermingled" seems to mean exactly this.

      We included this phrase to prevent quickly-skimming readers from incorrectly concluding that the subpopulations overlapped entirely and were therefore intermingled everywhere. The reviewer is right that it’s unnecessary for a careful reader, but we aimed to prevent misinterpretation by readers that might skip to the Discussion for a results summary.

      (19) Page 25: "content of the activity, but also its format" - the difference between content and format is not entirely clear. Metaphor not quite metaphoring here. Agreed. We have added examples to clarify.

      “This makes clear that there are potentially important differences not just in the content of the activity (e.g., encoding target vs. movement commands (Grier et al. 2026)), but also its format (e.g., linear encoding vs. nonlinear, persistent vs. brief responses).”

      (20) Page 30, bottom: In the description of the behavior, more details should be provided, especially since the paradigm is new. For example, it says the block size was reduced - what was the ultimate block size?

      Targets were cued randomly in the behavior performed during neural recordings. Blocked trials were used during training and were phased out incrementally as performance improved. This and various other details have been added. Please let us know if there are other specific details you would like to see in the final version.

      (21) Page 39, citation of An, Mulcahey et al.: There is a biorxiv version with a different author list that could be cited.

      This was an error with our citation manager, and has been corrected. Thanks for catching it.

      Reviewer #2 (Recommendations for the authors):

      Overall, this is a remarkable study with well-designed in-depth analyses, and I only have some minor suggestions that could help improve the clarity of the paper.

      Thank you!

      General:

      It is not immediately clear to me why the GMM approach used in this study is more interesting than a clustering approach based on single-neuron response patterns (See Esmaeli et al., Neuron 2021 or Oryshchuk et al., Cell Report 2024). But my impression is that it led to the same observation that most clusters are widely distributed across cortical areas, with different proportions, but a few clusters are quite specific to a few areas. A noticeable difference perhaps is the number of clusters - or response profile - that seems particularly low (only 4) in the current study. Could the authors clarify and comment on that, maybe?

      The reviewer brings up an interesting point: at heart, these works ask related questions, albeit about different effectors, tasks, recording modalities, and types of information encoded. Those differences probably mean that results cannot be directly compared, but we can certainly discuss the methodological tradeoffs. The two papers mentioned take a more traditional first step, using PCA on the vectorized PETHs to reduce dimensionality, then layer on a spectral approach to improve clusterability. These are good methods; we use something similar as our alternate method, applying VARIMAX to the PCs instead of spectral methods to preserve linearity of transforms. For the kinds of responses both they and we have, PCA will tend to most strongly pick up two aspects of the responses: tuning and timing. This is because vectorized PETHs will have large values in the rows corresponding to the target/condition and time points where the high activity is, and the alignment of these profiles with those of the other neurons will capture a large fraction of the variance. For data like either theirs or ours, this would tend to cluster apart left-tuned cells from right-tuned, and (more importantly here for revealing spatial structure) early-response cells from later response cells. That intuition is consistent with what those papers report, and examining our VARIMAX’ed PC plots closely (which have sharpened in the latest version thanks to improved normalization), we can see that they break apart sub-regions largely based on timing. In our feature approach, we intentionally chose our features to be largely invariant to both tuning preferences and timing. Instead, we chose our features to pick up on what we call the single cell “response format”: response duration; peak time variation (but not absolute timing); and tuning sharpness, persistence, and linearity. These different methods pick up on different aspects of responses.

      To double-check that the PCA-then-spectral approach reveals similar structure to our use of VARIMAX on the PCs, we tried applying the suggested method to our data. We applied spectral clustering to the N x 20 PETH PC feature matrix, then fit an area-agnostic GMM to the spectral features. We plot the likelihood map for the components of a GMM with 10 modes. The GMM components did not display clear spatial structure beyond that observed in the VARIMAX’ed PCs (Figure 5-figure supplement 1) and were less interpretable than those identified by area-agnostic clustering of our response features (Author response image 1). As noted, the number of subpopulations identified by the clustering of our hand-engineered features is lower than what would be obtained from clustering the PCs of the PETHs. This is likely the result of the substantial heterogeneity in activity onset and preferred target that is preserved by PCA. Because our central approach is largely agnostic to these two sources of variation, the number of identified clusters reflects the dominant patterns of variation beyond these two sources.

      Author response image 1.

      GMM fit to spectrally transformed PETH PCs, agnostic to anatomical areas. One GMM was fit to the spectrally-embedded PC feature vectors of cells from all 5 main areas. Each component of a 10 component model is shown.

      Also, I think it would greatly help the reader to return to PETHs at some point, if possible, to show the response profiles of each identified neuronal subgroup (page 20). To what extent are they similar or different across the cortical areas (for the same neuronal subgroup)?

      This is a good idea. We have added a figure to address this question and the related question by R1 (Figure 9-figure supplement 1). In short, given the wide variety of PETHs we observed, there is of course still substantial variation within subpopulation, and some mild but systematic differences in the distribution of what we observe across areas. We now discuss the conclusions from this plot in the Results:

      “As a qualitative depiction of the response profiles identified with each subpopulation, we plotted the two highest-likelihood cells for each area/subpopulation combination (Figure 9-figure supplement 1). These examples reveal stereotypy in the subpopulation responses across areas, but also show variation across areas, especially for the two somatomotor subpopulations.”

      Specific:

      (1) Figure 2B and M&M: the 3D spatial organization of the target locations is not immediately clear. What is the spacing between target locations? What is the 'final azimuthal spacing'?

      Added, thanks. The pairwise horizontal distances between targets were between 1.72 and 6 mm apart and the vertical spacing within a column was 1 mm. “Final azimuthal spacing” just referred to the targets being closer together during training and our gradually spacing them apart to their final locations. We have also added some relevant details about the training.

      (2) Figure 2C: It would help to have a scale bar (mm).

      Added, thanks.

      (3) Figure 2C: It would be easier to appreciate the variability of the trajectories across trials to plot an overlay of trajectories to one target only (could be a Supplementary Figure).

      The reviewer has a good point: the variability and accuracy of aiming was hard to ascertain from the plot. We experimented with a few options for making this clearer most effectively. We have now added Figure 2-figure supplement 1 that shows in the third subpanel of panel A the finger centroid trajectories for one of the 15 targets highlighted for the mouse shown in Figure 2C, mouse 3. The centroid trajectories for all other mice are shown as well to illustrate similarities and differences across animals as well as the overall variability. As noted elsewhere we have also included an analysis of the variability of the centroid trajectories, showing that reaches to a given target were more similar than reaches to different targets. We think this provides a fuller picture of the behavior and intend to provide still more detail in future work. Thank you for suggesting additional detail here!

      (4) Figure 4: It would be nice to also show the amplitude-normalized grand-average PETHs for the different areas.

      This is an interesting suggestion. After careful consideration, we think that this analysis is not as effective for depicting overall timing and modulation profiles as the current ones, given the strong amount of target selectivity and response time heterogeneity (now better visible in the revised Figure 4A). When computing the grand mean of all cells within each area, the dominant features distinguishing areas are onset time and response duration. The differences across areas in these two features are better supported by the analyses of Figures 4 and 5 due to the large amount of heterogeneity in responses within each area. We thank the reviewer for encouraging this exploration; more complicated spin-offs will likely inform additional timing analysis in the next paper on these data.

      (5) Figure 7C: figure legend - although it is quite self-explanatory, please explicitly indicate which pattern corresponds to the 'Three contour levels (98%, 95%, 90%)'.

      We have now added this as a legend on the figure panel itself (here and on similar plots). Thanks for pointing this out.

      (6) Figure 8: Is there also an interesting asymmetry between sensory are motor areas, with neurons in sensory areas being more likely associated with motor areas (B and C), whereas neurons in motor regions are less likely to arise from the distribution of sensory areas (dark blue color in frontal regions in D, E, and F)?

      This is an interesting observation, but we understand it to be an artifact of colormap scaling. As mentioned above, likelihoods are not well-behaved like probability distributions are: for example, they are not bounded at 1, and their sums over a dataset can have any positive value. The only things that can be interpreted are their relative values. This makes their scaling functionally arbitrary – you’ll notice we used “min likelihood” and “max likelihood” instead of numbers, which would be nearly meaningless – and therefore presents a problem for scaling the colormaps. We don’t know of a principled way around this problem. To deal with it, we simply put the ends of our colormap at the extreme pixel values. It so happens that both the M1 and M2 maps had a handful of neurons in a less-sampled spot at the bottom of M2 that were very low-likelihood, which results in what you noticed. We debated removing those neurons for this purpose, but we had no basis on which to do that kind of manipulation, so we left it as the most honest representation of the data we could produce.

      To clarify this, we now mention in the caption “The ends of the colormap were set to the maximum and minimum likelihood values for each map.”

      (7) Figure 9B: there are two-time 'S1-hl: 1' indicated at the two bottom rows of the distance matrix. I suppose one of them should be 'S1-tr: 1' instead?

      Fixed, thanks for catching it.

      (8) Page 20: 'This hinted at a second hypothesis: that some of the 'modes' (groups of neurons) discovered separately in each area might correspond.' ???

      We had meant “mode” as in “multimodal”, but it was very unclear. We have rewritten the sentence:

      “This hinted at a second hypothesis: that a peak in the multimodal distribution from one area might correspond to a peak in the multimodal distribution of a different area.”

      (9) Figure 9S2: Please indicate for which area each map is computed.

      The caption was not clear enough about what we were doing here: we fit the GMM on all neurons together, ignoring which area they came from. We have now clarified it in the caption:

      “One GMM was fit to the feature vectors of cells from all 5 main areas. Each map plots the likelihood for all cells to each of the three components of this area-agnostic GMM.”

      (10) M&M, Subjects and surgical procedures: 'ambient temperature of 71.5 {degree sign}F', please use international units.

      Done.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It was nice to see that the authors could distinguish differences between the OFC signals that they observed and those in the visual regions based on changes through the session. However, the linkage between these brain activations and a functional role in generating behavior was left unexplored. Without further exploration, it is hard to tell exactly what role the signals might be playing, if any, in the behavior of interest.

      To link the behavioral with the fMRI data, we now correlated fMRI decoding accuracy with behavioral performance. We studied behavioral performance in two ways: the difference in high versus low noise environment learning rates, and mean accuracy (i.e., absolute prediction error). We correlated both measures with the decodability of the environment in the central OFC. Each correlation was calculated either in the full experiment, or only the second half. However, none of these correlations were significant (all p > .1). Given the difficulty of interpreting this result, and our lack of statistical power for doing individual difference analyses, we decided not to report these analyses in the final paper.

      Reviewer 2 (public review):

      (1) The authors make the distinction between meta-learned "global" learning rates and within environment learning rate adaptation in response to "local" fluctuations/observations. Though the experimental paradigm is novel, there are certainly links to prior work - for instance, though change point structures don't entail revisiting unique environments, they do require meta-learning from environmental statistics that is distinct from transient local adaptation to prediction errors. This tendency to increase one's learning rate after large prediction errors is appropriate in change point environments, though, as is true in this study, the amount of increase should be dependent on. This represents a similar kind of slower-timescale learning or reuse of more "global" parameters, and can be seen to different extents in prior work. It might benefit readers if the authors were to link the current work to previous research more explicitly to draw clearer connections between the approaches and findings.

      We thank the reviewer for their very helpful literature suggestions and now contextualize and discuss our findings in light of relevant literature.

      (2) Throughout much of the paper, the authors refer to the distinctions between environments primarily as differences in "initial learning rates" or "environment-specific learning rates." This is particularly prominent when discussing fMRI results. Though the optimal initial learning rate did differ across environments, this was the result of differences in underlying task statistics. It will be important to clarify this throughout the text, because of the confounds between task statistics and initial learning rate (and to some extent, the position on the screen), it is not possible to separate the impact of these specific variables. This is also relevant to understanding the justification for using methods like RSA to test whether brain regions represent task states similarly. If the main hypothesis is that neural activity reflects the (initial) learning rate itself, then a univariate analysis approach would seem more natural.

      We agree that task statistics are not the same as differences in learning rates. However, we do not consider this as a confound: The point of the differences in task statistics is exactly to generate differences in learning rates. With our paradigm, we deliberately tried to dissociate variations in learning rate that were induced by learned environmental differences versus local task statistics. We tried to make this dissociation more clear, especially when discussing the fMRI results.

      (3) For the neuroimaging results in particular, the specificity of some of the results (e.g. ventral striatum showing an effect of prediction error only in the low noise condition in the second half of task experience, only on the first trial) is a bit surprising. Additional justification of or context for these results would be useful to help readers gauge how expected or surprising these findings are.

      We agree some of these findings were unexpected. We now also highlight that while we expected the ventral striatum to be involved in prediction error processing, we had no strong a priori expectations regarding these further modulations by time and environment. We also tried to contextualize these interactions more.

      (4) There are some methodological details that are unclear (e.g., how were the positions of the crabs selected relative to the location they emerged from? Looking at Figure 1C, it looks like the crabs spread out unevenly, and that the single position they emerge from is not necessarily at the center of the crab locations.) Additional detail and clarity would help address some unanswered questions (more details below).

      We clarified the experimental procedure at several places, and now added a video that helps illustrate the trial timeline better.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) With regards to the primary weakness mentioned above, it would be nice to have some link between the brain signals of interest and upcoming behavior. For example, can you read something out of OFC that enables you to better predict what the participant will do next? Or even better, do so beyond any behavioral variability that is explained by the computational model?

      To link the behavioral with the fMRI data, we now correlated fMRI decoding accuracy with behavioral performance. We studied behavioral performance in two ways: the difference in high versus low noise environment learning rates, and mean accuracy (i.e., absolute prediction error). We correlated both measures with the decodability of the environment in the central OFC. Each correlation was calculated either in the full experiment, or only the second half. However, none of these correlations were significant (all p > .1; see plots in Author response image 1). Given the difficulty of interpreting this result, and our lack of statistical power for doing individual difference analyses, we decided not to report this analysis in the paper.

      Author response image 1.

      (2) A number of the learning analyses are based on splitting the session into halves. As a first pass, this seems like a reasonable thing to do, but I certainly wonder what the dynamics of the meta-learning actually look like, and it seems like the data collected would be sufficient to gain some insight into those dynamics through some sort of sliding window analysis.

      We thank the reviewer for this interesting suggestion, which was also raised by Reviewer 2. We now calculated the learning rate in a sliding window of 20 trials (i.e., trial x to x + 19), and provide revised figures for each experiment separately (Fig. 2E and Fig. 4E, respectively).

      (3) The model selection procedures described make sense, but it would still be useful if the authors justified them by showing that they work in synthetic data (ie, generate a confusion matrix). I may be confused about what delta-SE is, but I'm confused about why two models with very different fits have the same value (211) for that metric.

      We report model recovery on synthetic data, which yielded model recovery rates of 100%, and added these to our Methods section. To clarify the Reviewer’s second point, ∆SE is the standard error of the difference between a model’s LOOIC and the top ranked model’s LOOIC. There is no one-to-one mapping between the ∆SE and a model’s LOOIC.

      (4) Was the central OFC anatomical ROI overlapping with the cluster surviving in the whole brain analysis? I didn't see this mentioned in the text, and it certainly would be important for interpreting the two results together.

      The central OFC indeed overlapped with the cluster surviving whole brain analysis, which we report on page 17-18.

      (5) The authors found regions that reflected learning rate at the "island presentation" phase of the task - it could be distinguishing this analysis and its meaning from other work that has focused on representations of learning rate at the time of feedback.

      We agree that this is an important distinction worth emphasizing. Therefore, we added the following lines to our discussion paragraph: 

      “Importantly, previous studies examined neural correlates of learning rates during outcome evaluation, where learning rates may be adjusted online as a function of locally experienced prediction errors (e.g., (Behrens et al., 2007; Browning et al., 2015; Nassar et al., 2012). In contrast, our RSA analysis targeted neural activity at island presentation, before any outcome information was available. At this moment, learning rates cannot be updated based on current feedback and instead reflected the retrieval of a previously learned, environment-specific learning-rate settings. This difference reflects our hypothesis that the OFC represents the latent states in a cognitive map of the task (Knudsen & Wallis, 2022; Moneta et al., 2024; Schuck et al., 2018; Wilson et al., 2014), which are expected to activate as soon as the agents can infer which task state it is in. Several studies have identified such “partially observable” task states in the medial OFC (Bradfield et al., 2015; Schuck et al., 2016; Tan et al., 2025; Wimmer & Büchel, 2019), in line with the region identified here (but see e.g., (Ongur & Price, 2000), for important anatomical distinctions between medial and lateral OFC and (Tan et al., 2025) for an example of related functions in lateral OFC). Our finding extends this notion by suggesting a link between OFC and meta learning, wherein meta-learned information becomes encapsulated in task states (Hattori et al., 2023; Moneta et al., 2024).”

      (6) "Specifically, it showed a more negative response to larger (location) prediction errors, which is consistent with its documented role in showing a more positive response to more positive reward prediction errors (Calderon et al., 2021) - keeping in mind that being closer to the centre of where the crabs appeared (i.e., smaller location prediction errors) is less negatively or more positively surprising (i.e. smaller negative or larger positive reward prediction errors)."

      I found this sentence very hard to parse. Do PE responses in the high noise environment get "compressed" in their representation over time (ie, it takes a larger error to get the same BOLD response)? If so, this relates to claims made in Diederen 2016... but see also Mah 2024 Cell Reports, who fails to see learning rate encoded in DA system in striatum of rodents that appear to adjust their learning rates.

      Thank you for pointing to this. We agree that this sentence was hard to parse, and so we now split it in three revised sentences. We also agree with the Reviewer’s interpretation, and would like to thank the Reviewer for their useful literature suggestions which we now added to our discussion. 

      (7) Figure 7 should use a different color scheme because many of the activations just appear black, and I can't tell whether they are positive or negative. It was also notable in Figure 7A that regions are not visible, including ACC, which is typically thought to encode prediction errors in such paradigms. It would probably be useful for the authors to include a table of all clusters exceeding multiple comparisons correction and to on differences to other work examining absolute prediction errors. ACC does appear on the second trial, which made me wonder whether there were changes in the prediction error coding from first to subsequent trials. 

      Thank you for pointing this out. We now revised our color scheme which we agree makes it much clearer now. Although the ACC is frequently implicated in prediction error–related signals (e.g., Behrens et al., 2007), models suggest that ACC responses more strongly reflect unsigned prediction errors, surprise, or the need for control and model updating (Alexander & Brown, 2019; Hayden et al., 2011; Silvetti et al., 2018). In our task, ACC activity only emerged on the second trial, when participants had formed an initial estimate and prediction errors could meaningfully signal the need to update internal models or control settings. We now added a to the Discussion highlighting this distinction and relating our findings to this prior work emphasizing prediction errors and control-related signals in ACC.

      (8) The authors suggest that fast learning would presumably occur in a neural activation space, whereas slow learning would occur through weight adjustments. This makes sense, but activity-based dynamics have been suggested to do rapid adjustments by encoding a "latent state" though (Razmi 2022 j neurosci) -- and such a latent state has been shown in OFC (Schuck etc)... but here OFC is more implicated in the slow learning. I am curious about whether authors could on this a bit in the discussion. 

      Thank you for bringing up this interesting question. We can only speculate but a crucial factor is on which level of resolution tasks states operate. On the one hand “detailed” trial-level states are needed that map a specific sensory input onto a specific latent state and its value. Such states would change quickly, possibly through activation dynamics, and are in line with how they have been operationalized in Razmi or Schuck etc. On the other hand, successful task performance also needs “higher level” states that describe entire task phases or full tasks, as in the present experiment. Due to the different speeds of learning, it appears plausible that these would be learned with synaptic changes. We expand on this in the discussion as follows: 

      “Our finding extends this notion by suggesting a link between OFC and meta learning, wherein meta-learned information becomes encapsulated in task states (Hattori et al., 2023; Moneta et al., 2024). Consistently, OFC has been shown to represent task states (Moneta et al., 2024; Stalnaker et al., 2015; Wilson et al., 2014). While earlier evidence shows that the OFC represents concrete aspects of task states, such as task-relevant stimulus features (Schuck et al., 2016), we hypothesized that the OFC also represents more abstract aspects, such as learned, environment-specific learning rates. Indeed, we showed that the central OFC gradually came to represent these environment-specific learning rates (or the environment-specific statistics that drive them). While previous work speculated that these different levels could have different neural underpinnings (Sharpe et al., 2019), our findings indicate OFC might signal states on multiple levels. This does not imply identical learning dynamics; fast-changing trial-specific states might be learned through activity dynamics, while higher-level contextual states could involve synaptic plasticity.”

      (1.9) Also, as a more minor point in the same section, the sentence about blocking synaptic plasticity in OFC sounded interesting, but should have a reference.

      Thank you for noticing, we now added the reference (Hattori et al., 2023).

      Reviewer #2 (Recommendations for the authors):

      (1) Additional links to prior literature: In terms of prior work in which there is something akin to more "global" adaptation, some examples of potentially relevant prior work include:

      McGuire, Nassar, Gold, & Kable (2014) Neuron 

      D'Acremont & Bossaerts (2016) Cerebral Cortex 

      Lee, Gold, & Kable (2020) Decision 

      Bakst & McGuire (2021) JEP: General 

      Bakst & McGuire (2023) Cognition

      We would like to thank the reviewer for pointing us to these different literature suggestions which we agree help us contextualize and discuss some of our findings better. We now refer to McGuire et al. (2014) when discussing the fMRI results, and d'Acremont & Bossaerts (2016) when discussing potential alternative strategies in the high noise environment (the Reviewer’s last point). Finally, we integrated the clearly relevant works of Bakst & McGuire (2021; 2023) and Lee et al. (2020) in our discussion of meta-learning different adaptive strategies. 

      (2) Individual differences: Though not always the focus of work on predictive inference, one common finding has been that there are pronounced individual differences in behavior (see, e.g., coefficients in Figure 2 in Nassar et al. 2019 eLife, or Figure 2 McGuire et al. 2014 Neuron, or Bakst & McGuire 2023 Cognition). There appears to be substantial variability between individuals in your data as well (i.e., Figure 2B, 4B, and the modeling figures). It would be interesting to see some direct exploration of this variability: baseline learning rate appears to differ between participants to a large extent, does their rate of adaptation (across trials within a block) also differ? Does their metalearning occur at different rates (in fact, do some participants not show evidence of appropriate meta-learning at all)? 

      Relatedly, your computational modeling approach fits the six candidate models hierarchically, and therefore the reported results show the overall best fit for the group. It might be worthwhile to determine whether individuals have different best-fitting models. This could be another way to characterize the variability between individuals. 

      In concert with this, it could be a useful complement to determine whether either the strength of the OFC neural similarity results or their time course reflects aspects of behavior. Put another way, is it the case that not only does OFC activity and behavior both come to reflect task structure, but that these changes happen to a similar extent and over a similar time course across individuals?

      We agree it would be highly interesting to investigate meaningful individual differences in both fast and slow adaptations in learning rate. However, our sample was not set up and is underpowered to conduct such analyses. In response to a similar by Reviewer 1, we did run correlational analyses between differences in learning rate, performance accuracy, and the responsiveness of the OFC. However, none of these analyses yielded a significant effect. We decided to not include these results in the paper, for reasons of statistical power, but we report them in Author response image 1.

      (3) fMRI:

      (3a) The primary finding in OFC is restricted to the central OFC. The manuscript would benefit from additional explanation regarding this specific subregion. 

      Thank you for bringing up this important distinction. In the discussion we now clarify as follows: 

      “This difference reflects our hypothesis that the OFC represents the latent states in a cognitive map of the task (Wilson et al., 2014; Schuck et al. 2018; Knudsen & Wallis, 2022; Moneta et al, 2023), which are expected to activate as soon as the agents can infer which task state it is in. Several studies have identified such “partially observable” task states in the medial OFC (Schuck et al., 2016; Bradfield et al., 2015; Wimmer et al., 2019; Tan et al., 2025), in line with the region identified here (but see e.g., Öngur & Price, 2000, for important anatomical distinctions between medial and lateral OFC and Tan et al., 2025, for an example of related functions in lateral OFC).”

      (3b) Though the main clusters visible in Figure 6 are the occipital and OFC clusters, there appear to be others. Did other clusters indeed rise to statistical significance in the whole-brain analysis? If so, is there a reason they aren't included or discussed? 

      All clusters visible in Figure 6C survived FDR correction. However, we refrained from interpreting these other clusters, because we had no prior hypotheses about them like we did for the OFC.

      (3c) Why do you posit that the ventral striatum becomes less sensitive to RPE on the second trial over time? And why is the ventral striatum only sensitive to RPE in the low noise environment generally?

      We reasoned the ventral striatum should be more responsive to more positive reward prediction errors. While we further assumed this response could be modulated by both time and environment, we would like to emphasize that we had no specific hypotheses about the direction of this modulation. We now also make this clearer in the manuscript. This being said, we believe both the pattern that its responsiveness to the second trial decreases over time, and the pattern that it was most sensitive to the low noise environment, can be considered fitting with its broader involvement in coding behaviorally relevant reward prediction errors. Namely: 

      First, we believe that as the participants learn more about the global reward structure of the task, they should obtain a better understanding of the fact that, per round, all crabs always center around a fixed mean. Therefore, the first RPE is most behaviorally relevant, and every later RPE has an exponentially decreasing relevance. As participants obtain more experience with this aspect of the task over time, the VS should show a lower responsiveness to the second RPE over time.

      Second, as participants learn more about the local differences between the three different environments, they should learn that especially in the low noise environment, RPEs are most behaviorally informative. That is, in this environment it makes most sense to have a high learning rate and thus let the RPEs substantially inform the placement of the cage on the next trial. Accordingly, participants showed that the ventral striatum was most responsive to RPEs in these environments.

      (4) Methods

      (4a) This section could generally benefit from some proofreading. 

      We now proofread the method section. 

      (4b) The main results text states that 49 participants performed Experiment 1, while the methods section reports 50 participants. Which is correct? 

      (4c) Following this, on page 8, statistical results are reported with a df = 49 (which would be appropriate only if n=50). 

      The correct sample size was actually 50, we adjusted the text and degrees of freedom where incorrect accordingly (note: only text is in track changes, but degrees of freedom were also changed accordingly). 

      (4d) Additionally, I am a bit surprised by the Experiment 1 findings that learning rates on the second trial were significantly different between low and high noise conditions, in that the effect size found using all trials was stronger than both the first half of trials (no significant effect) and the second half (significant but weaker than all trials). Are these all the same type of statistical test? Double-checking the statistics might be worthwhile. 

      It is not the effect size that is larger across the full experiment, but the t-statistic. This is possible because a t-statistic depends on both effect size and noise estimate, and the latter is smaller with more data. 

      (4e) The methods and results both state that the five crabs always emerged from one position in the sand. How were the locations of the crabs selected relative to this position? Looking at Figure 1C, it looks like the crabs spread out unevenly, and that the single position they emerge from is not necessarily at the center of the crab locations. 

      The crabs did indeed spread out evenly. However, we can see how the graphic in Figure 1C can be confusing, as two crabs are shown to be caught, which breaks the symmetry of the dispersion (because some crabs can run away after the even spreading phase, see Methods). We emphasized the even spreading more clearly in the new version of the paper. We think the flow of events will be much clearer with our newly added animation (Video 1).

      (4f) The methods section states that the crabs "spread out to cover the same proportion of the screen width as the cage (18.75%)" (page 23). The corresponding visual in Figure 1C appears to show something different. 

      This looks different because the graphic illustrates the last 500 msec, where crabs can run away (see also response to 4e, and the novel animation that was added).

      (4g) Information on the timing of the trials would be useful to include in Figure 1C or similar. 

      The reader can find this information in the Methods section. We chose not to include it in the caption to avoid information overload.

      (4h) The methods section specifies that there was a 3-7s ITI after the first and second trials of each block. How was the ITI selected for each trial? Were there ITIs between the other trials? If so, what were they? 

      The ITIs were selected from a truncated exponential distribution. This selection was not random, but rather a distribution was carefully constructed for each environment (and event of interest: boat presentation, first trial of each block, second trial of each block) separately to ensure that enough longer ITIs were selected for each environment (and event of interest). Of course, the order in which the ITIs were used across blocks, was random. The same approach was used to determine the duration of the presentation of the boat at the start of each block. There were no ITIs after later trials.

      (4i) Please provide a link to the data and analysis materials on OSF in the text. 

      We now provide a link to the data and analysis materials in our methods section.

      (4j) In the methods section, there are some references to information provided "below" (page 26: "The two approaches resulted in different posterior densities (see below) for estimate uncertainties, but in similar posterior densities (see below) for learning rates..."). Where in the paper is this referencing? 

      We indeed did not detail this further as we considered it not further relevant to our main study, and now removed the references to “below”.

      (4k) The methods section specifies using uniform priors between the lower and upper bounds of the relevant parameters. This seems likely to be 0 and 1, but should be listed explicitly. 

      Thank you for noticing. We now added this to our manuscript.

      (4l) For parameter recovery, correlations are provided to indicate effective recovery. These correlations are indeed high and suggest excellent recovery, but correlations wouldn't reveal if there was systematic over- or underestimation occurring. It might be useful to provide some visualizations of the parameters and their estimates to speak to this potential issue. 

      We now visualize the parameter recovery results in Author response image 2, which show that, indeed, there was a slight underestimation of the decay rates, but not the learning rates. Importantly, our main analyses and results all pertain to the learning rates, and we never made hypotheses or conclusions about the decay rates.

      Author response image 2.

      (4m) The methods section ends with a reference to a reward localizer (page 32). This localizer doesn't appear to be mentioned/used elsewhere. 

      Indeed. We implemented the localizer because we wanted to independently identify reward processing areas. However, this localizer did not succeed in localizing a reward area (no significant results), possibly due to the fact that (1) it was performed by the end of the experiment when participants may have been fatigued, and (2) there was no learning component in this localizer task. For these reasons, we did not use it after all.

      (5) Analysis: 

      (5a) Did you consider fitting a Bai model that only allowed for environment-specific initial learning rates (with a non-environment-specific decay rate)? Given that the data (e.g., Figure 2, Figure 4) seems to support differences in initial learning rate but not necessarily a difference in the rate of change, it might be worthwhile to see whether a model like that fits best. 

      We now fitted this extra model, which we called the semi-environment-specific Bai model. See Author response tables 1 and 2 for result in experiments 1 and 2, respectively) for the results. This new model has the best (in Experiment 2) and second-to-best (in Experiment 1) LOOIC. In a way, this is not surprising, because the model formulation is entirely based on the data. We think that we can draw the same substantive conclusions with or without this extra model, so for simplicity we did not include this new model in the paper itself.

      Author response table 1.

      Note. Models are ranked in descending order according to how well they fit the data. LOOIC refers to a model’s approximated expected log pointwise predictive density. Higher values indicate higher out-of-sample predictive fit. SE refers to the standard error of a model’s LOOIC. ∆LOOIC refers to the difference between a model’s LOOIC and the top ranked model’s LOOIC. ∆SE refers to the standard error of the difference between a model’s LOOIC and the top ranked model’s LOOIC.

      Author response table 2.

      Note. Models are ranked in descending order according to how well they fit the data. LOOIC refers to a model’s approximated expected log pointwise predictive density. Higher values indicate higher out-of-sample predictive fit. SE refers to the standard error of a model’s LOOIC. ∆LOOIC refers to the difference between a model’s LOOIC and the top ranked model’s LOOIC. ∆SE refers to the standard error of the difference between a model’s LOOIC and the top ranked model’s LOOIC.

      (5b) If part of the goal is to investigate whether there is a distinct local change in LR between conditions (dependent on prediction errors), then there might be more direct ways of doing so as a complement to the modeling approach. One potential way could be to visualize the LR or change in LR as a function of PE. 

      We agree that it’s beneficial to use a direct (model-free) approach to represent learning rate as a function of condition; that is also part of our approach. For example, see Figures 2, 4, which shows learning rate as a function of condition, but in a model-free manner. We think learning rate as a function of prediction error is less informative, because the idea is that prediction error can (in Kalman-filter terminology) be indicative of either noise variance or process variance, and participants are able to distinguish between them. This is also why we constructed the conditions in such a way that on the very first trial, prediction errors were on average the same across conditions. The fact that participants did respond appropriately to prediction errors on the very first trial (i.e., larger updates or learning rates in the low noise condition), suggested they are able to assign the prediction error to process variance (in the low noise condition) versus noise variance (in the high noise condition).

      (5c) In addition to looking at the evolution of LR across trials within a block separated by task epoch (i.e., Figure 2C-D & Figure 4C-F), the structure of the task would lend itself very nicely to visualizing the evolution of the second trial LR on its own across instances. This could provide additional insight into the meta-learning process.

      We thank the reviewer for this interesting suggestion, which was also raised by Reviewer 1. We now calculated the learning rate in a sliding window of 20 trials (i.e., trial x to x + 19), and provide revised figures for each experiment separately (Fig. 2 and 4, respectively).

      (6) The environment-specific Bai model appeared to become less good at capturing participant behavior with increased environmental noise. Why do you think this is?

      We thank the reviewer for raising this point. In this environment, individual outcomes are considerably less indicative of the latent mean, which may reduce the usefulness of the trial-by-trial, prediction-error–driven learning-rate adjustments that we see in the other environments. Under such extreme conditions of variability, people may rely less on delta-rule updating and more on alternative strategies (D'Acremont & Bossaerts, 2016; Reynders et al., 2026), such as exploratory adjustments or heuristics that are not explicitly captured by the Bai model but also outside the scope of the present paper.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      Hüppe and colleagues characterized the network of neurons in the central nervous system of Antarctic krill that contained pigment-dispersing hormone (PDH), an important output factor in the circadian clock of insects. These neurons in the brain are putative clock neurons since a subset also expressed the clock genes period and cryptochrome 2. As one of the ocean's major contributors to biomass, krill is an ecologically important marine species that experiences challenging daily and seasonal environmental fluctuations in its high-latitude habitat. A comprehensive study of krill's internal clock may help to understand the extent of its resilience to the rapidly changing climate.

      The authors used antibody staining against PDH across the whole central nervous system and additional in situ hybridization for cry2 and per mRNA, with a focus on the supraesophageal ganglion. There, they identified the major neuropils in the eye stalks and central brain of Antarctic krill. The resulting staining pattern aligns with the identified circadian clock network in insects and PDH-expressing networks in other crustaceans, making these neurons highly likely candidates for krill clock neurons.

      Strengths:

      (1) This study provides the first clues about the circadian clock architecture in a non-model organism in chronobiology, Antarctic krill, with a clear 3D reconstruction of the putative clock network.

      (2) The authors effectively place their results within the extensive body of literature on arthropod circadian clock networks to argue that the neurons they describe are likely the circadian clock in krill.

      Weaknesses:  

      (1) The data presented here are not sufficient to support the claim that the described network is the circadian clock because functional evidence is missing.

      (2) Additionally, the study falls short of identifying any elements of the positive limb of the canonical circadian clock transcriptional-translational feedback loop, e.g., clk or cyc, in the PDH-expressing neurons.

      (3) No sample sizes are reported, making it difficult for readers to assess the generalizability of the presented data.

      We thank the reviewer for recognizing the contribution of this study to advancing our understanding of clock systems in non-traditional model organisms. We acknowledge that definitive functional evidence would require the generation of null mutants of core clock components, which is currently not feasible in this species. In a revised version, we will adjust our claims to more precisely reflect the evidence presented and include sample sizes to allow the reader to better assess the representativeness of the results.

      Reviewer #2 (Public review):

      Summary:

      This study advances our understanding of the neuronal basis of the circadian clock in pancrustaceans. It extends our knowledge on the pigment-dispersing hormone system and provides links to information on the expression of core clock components, cryptochrome 2, and period. The data are sound and well-documented.

      Comments:

      The neuronal components of the arthropod circadian clock system have been analysed extensively in insects. Much less information on this system is available on malacostraca crustacea crustaceans. However, considering that malacostracan crustaceans and insects go back to a common pancrustacean ancestor and considering that we know that the brain architecture in these two groups shares many commonalities (see, e. g., extensive reviews by N. J. Strausfeld), we have to expect that crustaceans and insects share many of the characteristics of the circadian system. This is the case, e. g., for the network of pigment-dispersing hormone-positive neurons. The authors cite these studies, although late in the paper (discussion, line 339ff), and I suggest to move this info into the introduction: "339 ff: The arborization pattern of the PDH-network has been described in various malacostracan crustaceans, including Carcinus maenas (Alexander et al., 2020; Mangerich & Keller, 1988; Mangerich et al., 1987), Cancer productus (Hsu et al., 2008), Orconectes limosus (de Kleijn et al., 1993; Mangerich & Keller, 1988; Mangerich et al., 1987), Homarus americanus (Harzsch etal., 2009), Cherax destructor, Procambarus clarkii (Sullivan et al., 2009), and Procambarus virginalis (Luna et al., 2010)."

      The strength of this paper is that it extends our knowledge on the PDH system and brings together neuroanatomical information on PDH-positive neurons with information on the expression of core clock components, cryptochrome 2, and period. That way, it advances our understanding of the neuronal basis of the circadian clock in pancrustaceans. The data are sound and well documented, and the authors are to be applauded for the superb dissection presented in Figure 1.

      Below, please find some essential suggestions on how to further improve the paper.

      (1) Framing of the study:

      I know that krill is a key element of the Southern Ocean's food webs, but my sense is that discussing the current findings in a context of resilience of this species to global ocean change means largely overselling this study:

      Lines 47, 48: "and the resilience of this key species in a rapidly changing Southern Ocean."

      Lines 70 ff: "Hence, understanding the mechanisms of adaptation, including biological clocks, is crucial for predicting how species, populations, and whole ecosystems will respond to climate change."

      154 ff: "The Southern Ocean environment experiences rapid change (Abram et al., 2025; Meredith et al., 2019; Thomalla et al., 2023). To assess krill's resilience to environmental changes, understanding the mechanisms that govern daily and seasonal timing in krill is essential."

      325 ff: "The rhythmic adaptation of krill to its high-latitude environment is key to its success in the Southern Ocean, which in turn represents a cornerstone for the well-being of the whole krill centred ecosystem. To predict krill's resilience to rapid environmental changes, it is essential to understand the mechanisms that govern daily and seasonal timing in krill."

      597 ff: "A detailed mechanistic understanding of the flexibility of clock-based processes is therefore essential to predict krill resilience in a changing Southern Ocean."

      My understanding is that duration of day length is one of the most predictable environmental drivers, and - despite the seasonal changes of day length - nevertheless a very stable one compared to fluctuations of environmental drivers such as temperature or salinity (see, e.g. this recent review on environmental driver fluctuations on nervous system functioning in crustaceans: Stein W, Harzsch S (2021) The Neurobiology of Ocean Change - insights from decapod crustaceans. Zoology: 125887. https://www.sciencedirect.com/science/article/pii/S094420062030146X).

      I do not see how global ocean change may significantly change day length, and what this study has to do with understanding this species' resilience against ocean change. I suggest that you explain in more detail why the light day length will change in the future or strongly tone this aspect. Statements such as Line 76 ff: "Due to their disproportionate importance for ecosystem function, understanding the resilience of ecological key species is essential in assessing the fate of ecosystems in the future." are completely out of focus here and, again, trying to oversell the current study.

      (2) Uncited essential studies of crustacean neuroanatomy, missing connection to contemporary crustacean neurobiology:

      Line 157: "despite the ecological importance of E. superba, only very little is known about its neurobiology".

      Line 329: "However, so far, little was known about the neurobiology of krill in general."

      I agree that this species' brain is understudied, but this makes it even more important to cite the little information that IS available. Please consider this essential reading for any crustacean neurobiologist: "Sandeman, D.C., Scholtz, G., Sandeman, R.E., 1993. Brain evolution in decapod crustacea. J. Exp. Zool. 265, 112-133." to find information on the basic brain anatomy in E. superba.

      The manuscript in many places seems to reinvent the wheel and raises the impression that our knowledge of crustacean brain morphology is close to zero. The authors in places seem to operate in a vacuum, and I find it disturbing that in a study on the crustacean brain, very few references are provided to studies on crustacean brain anatomy, such as the following essential book chapter: "Schmidt, M., 2016. Malacostraca. In: Schmidt-Rhaesa, A., Harzsch, S., Purschke, G. (Eds.), Structure & Evolution of Invertebrate Nervous Systems. Oxford University Press, Oxford, pp. 529-582. https://www.researchgate.net/publication/315366157"

      In terms of brain anatomy, I would like to know if the authors have a hypothesis on whether and how their target species' brain structure may be similar or different to the brains of other "shrimps" as described, e. g., in the following studies. If so, please elaborate in the introduction:

      Krieger J, Hörnig MK, Sandeman RE, Sandeman DC, Harzsch S (2020), Masters of communication: The brain of the banded cleaner shrimp Stenopus hispidus (Olivier, 1811) with an emphasis on sensory processing areas. Journal of Comparative Neurology 528(9): 1561-1587.

      Meth R, Wittfoth C, Harzsch S (2017) Brain architecture of the Pacific White Shrimp Penaeus vannamei Boone, 1931 (Malacostraca, Dendrobranchiata): correspondence of brain structure and sensory input? Cell and Tissue Research 369(2): 255-271.

      (3) Lacking rigor and command of crustacean brain nomenclature

      I suggest that for their brain nomenclature, the authors should rigorously stick to that laid out by Sandeman et al. 1992 (not yet cited in the ms): Sandeman, D.C., Sandeman, R.E., Derby, C.D., Schmidt, M., 1992. Morphology of the brain of crayfish, crabs, and spiny lobsters: a common nomenclature for homologous structures. Biol. Bull. 183, 304-326.

      More specifically, in lines 41, 163, 199, 204, 207, and throughout the paper, the authors use the terms "Optic lobes" or "optic lobe neuropils". To the best of my knowledge, "optic lobe" is not a term used in crustacean neuroanatomy at all (as opposed to insects). Lamina, medulla, and lobula are collectively referred to as "visual neuropils" (see Krieger, J., Hörnig, M. K., Sandeman, R. E., Sandeman, D. C., & Harzsch, S. (2020). Masters of communication: The brain of the banded cleaner shrimp Stenopus hispidus (Olivier, 1811) with an emphasis on sensory processing areas. Journal of Comparative Neurology, 528(9), 1561-1587. https://doi.org/10.1002/CNE.24831). The medulla terminalis and mushroom bodies are referred to as "lateral protocerebrum". All afore-mentioned neuropils are summarized as "eyestalk neuropils" (compare nomenclature in Schmidt 2016 as referenced above).

      Line 170, 172, 175 ff, and Figure 1. "abdomen", "abdominal ganglia": Contra the book chapter by Siegel 2016 "Introducing Antarctic Krill Euphausia superba Dana, 1850", his Fig. 1.2, the "tail" of crustaceans in most books on crustacean anatomy is not called "abdomen" but instead "pleon"; hence the name "pleopods" for the appendages of the pleon (instead of "abdomipods"). What is more, I suggest using the terms "pleon ganglia" instead of "abdominal ganglia", following the terminology suggested in "Harzsch S, Sandeman D, Chaigneau J (2012) Morphology and development of the central nervous system. In: Forest J and von Vaupel Klein JC (Eds.). Treatise on Zoology - Anatomy, Taxonomy, Biology. The Crustacea Vol. 3. Brill, Leiden pp. 9-236."

      Line 174: "thoracic ganglia". In Figure 1, there is a labelling mistake as these ganglia are named "thoracaic ganglia".

      Line 176, and throughout the paper: "supraesophageal ganglion". Following the standard nomenclature for crustaceans (see, e. g., Schmidt, M., 2016. Malacostraca. In: Schmidt-Rhaesa, A., Harzsch, S., Purschke, G. (Eds.), Structure & Evolution of Invertebrate Nervous Systems. Oxford University Press, Oxford, pp. 529-582. https://www.researchgate.net/publication/315366157", this structure (as in insects) is typically called a "brain". For terminology, also consult the following nomenclature paper: "Richter, S., Loesel, R., Purschke, G., Schmidt-Rhaesa, A., Scholtz, G., Stach, T., Vogt, L., Wanninger, A., Brenneis, G., Döring, C., Faller, S., Fritsch, M., Grobe, P., Heuer, C. M., Kaul, S., Møller, O. S., Müller, C. H. G., Rieger, V., Rothe, B. H., Stegner, M., Harzsch, S. (2010). Invertebrate neurophylogeny: Suggested terms and definitions for a neuroanatomical glossary. Frontiers in Zoology, 7. https://doi.org/10.1186/1742-9994-7-29".

      Line 212, and throughout the paper - hemielliposoid body: please refer to Harzsch Krieger 2011 and the numerous references to studies by Strausfeld cited therein in crustaceans. Strausfeld has provided compelling evidence that the crustacean hemiellipsoid body is equivalent to the insect mushroom body, so this term should be replaced. Harzsch, S., & Krieger, J. (2021). Genealogical relationships of mushroom bodies, hemiellipsoid bodies, and their afferent pathways in the brains of Pancrustacea: Recent progress and open questions. Arthropod Structure & Development, 65, 101100. HYPERLINK "https://doi.org/10.1016/J.ASD.2021.101100" https://doi.org/10.1016/J.ASD.2021.101100.

      Legend, figure 2, and others, and throughout the paper: "The olfactory neuropiles comprise the lateral antennal neuropile (LAN, ochre), the olfactory lobes (OL, yellow), and the antennal neuropile (AnN, green)." This is a strange terminological mix that you should urgently revise according to the standard terminology by Sandeman et al. 1992 (as referenced above). The LAN is the lateral antenna 1 neuropil. The AnN is the antenna 2 neuropil. The AnN is NOT deutocerebral but tritocerebral.  

      We thank the reviewer for acknowledging this paper's contribution to our understanding of the neuronal basis of the circadian clock in Pancrustaceans, as well as for the positive evaluation of the data documentation and presentation.

      We would like to clarify that we are aware of the existing body of literature on crustacean neuroanatomy and did not intend to present our data as a first in this field. This study intersects multiple communities (e.g., chronobiology, crustacean neurobiology, krill ecology), and the current focus of the manuscript arose from an attempt to make the paper as accessible to these communities as possible. We acknowledge, however, that the current version falls short in its engagement with the existing literature on crustacean brain anatomy. We therefore thank the reviewer for the input on crustacean neuroanatomy and its nomenclature, which will help us improve the manuscript in these respects. In a revised version, we plan to adjust the framing of the study to more precisely reflect the data presented. This will include better situating the present findings within the existing literature on crustacean neuroanatomy and its specific nomenclature, while toning down the emphasis on ecological importance and implications.

      Reviewer #3 (Public review):

      Summary:  

      A solid and very descriptive study of gene expression of three factors in krill, PDH, per, and cry2 that are important for circadian rhythms in insects. The results reveal optic areas in which PDH colocalises with each or per and cry2, and central brain areas where it does not. The authors speculate on the functional implications of their results for biological rhythms.  

      Comments:

      This manuscript describes a detailed anatomical study of the brain of krill in a circadian gene expression context. The results are well described, and the work is well done considering the obvious technical/practical difficulties of working with this species. Having stated that, the authors in their Methods write that the animals, after being caught, were placed in constant darkness. Is there any idea at all of when in ZT these brains were processed? Are the representations of gene expression taken at random around the clock? Perhaps the authors might make this explicit somewhere in the ms as it is an important point.

      The manuscript focuses mostly on PDH and its overlap or not with per or cry2. I found Figures 5 and 6 particularly confusing. The panels show PDH colocalising (or not-filled or unfilled arrows) with cry2 or with per. What they do not show (to me) is that per and cry2 colocalise. Now, of course, they probably do, but Figure 5 does not show this - or am I misinterpreting it? In Figure 6 again, I cannot see any panels with per and cry2 overlaid. Seems different sections were used for each probe? Is that what 'Areas with high per/ cry2-expression are marked by white arrowheads' means? I see that lines 493 and 494 confirm my suspicions that per/cry were not shown to be colocalised. Perhaps the authors could make this clearer up front than halfway through the Discussion, and clarify this in their legends, which are a little misleading in this respect?

      We thank the reviewer for his positive evaluation of our work, acknowledging the difficulty when working with this organism, and for the constructive comments. In a revised version of the manuscript, we will clarify the sampling time in the Methods. We will also state upfront — and in the figure legends — that per and cry2 were assessed on separate sections and their direct co-localization was therefore not demonstrated. However, as both components were independently shown to co-localize with PDH, their spatial overlap is nevertheless suggested by the shared co-localization with PDH. We will make this reasoning explicit earlier in the manuscript to avoid any misleading implications.

    1. Author response:

      We agree that the manuscript would benefit from a more clearly articulated conceptual framing, stronger model validation, more explicit statistical and ERP comparisons, and improved transparency regarding task design, sample inclusion, and preregistration. In the revised manuscript, we plan to address these points through substantial revision of the Introduction and Discussion, along with additional robustness and validation analyses, and more cautious interpretation of the main findings.

      Reviewer #1 raised important points about the framing of the cooperation task, the interpretation of betrayal, the standard statistical analyses, the modelling, and the ERP analyses. In response, we plan to clarify that the present task captures betrayal-related social risk or anticipated partner defection, rather than betrayal in its full interpersonal and emotional sense, and to better motivate this operationalization with reference to the betrayal-aversion and trust-game literature. We will moderate our claims regarding “emotional costs,” incorporate a more explicit task overview and accompanying schematic into the main text, and frame individual differences as a key avenue for future research. In addition, we will streamline the standard behavioral analyses, make the expected-value structure of the task explicit, add EV-based analyses of choice and reaction time, strengthen the ERP analyses, clarify that the study was not preregistered, and provide a complete report of data-quality checks. For the modelling section, a central revision will be to simplify the model structure and refit the models using a Bayesian hierarchical approach.

      Reviewer #2 emphasized the need for stronger theoretical framing and more specific distinctions between related constructs. In the revised manuscript, we will substantially revise the Introduction to better situate the present task in relation to the Trust Game literature and prior work comparing social and non-social decision-making under matched payoff structures. We will also define risk aversion, loss aversion, anticipated partner defection, and betrayal-related aversion more explicitly, and clarify that the distinction between betrayal-related aversion and loss aversion is inferred through computational modelling rather than directly manipulated as separate experimental factors. We also plan to introduce the computational model earlier in the manuscript, clarify how the key constructs are operationalized, replace unclear wording such as “impersonal losses,” strengthen the rationale for our hypotheses, and acknowledge the lack of preregistration more clearly.

      Reviewer #3 highlighted the need to align our conclusions more closely with the current evidence. In the revised manuscript, we will moderate the interpretation of the betrayal-related parameter, acknowledging that the cooperation task differs from the non-social risk task not only in social versus non-social uncertainty, but also in partner outcome, intentionality, and potential inequity structure. We therefore plan to avoid treating this parameter as a pure betrayal-specific construct and to describe it more cautiously as capturing betrayal-related social risk or aversion to anticipated partner defection. We also plan to report robustness analyses excluding participants who expressed doubts about the social interaction, as well as participants with poor catch-trial performance or otherwise low-quality data, and to clarify the sample sizes and exclusion criteria used for behavioral, modelling, and ERP analyses. Finally, we will strengthen model validation and ERP reporting, including broader validation analyses and more cautious interpretation if the evidence for temporal dissociation between betrayal-related aversion and loss aversion proves weaker than currently stated.

      Across these revisions, we also intend to simplify the model structure and use Bayesian hierarchical fitting to strengthen model validation, while avoiding overly strong claims if the additional analyses provide only modest support for a single preferred model.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Demeshkina and Ferré-D'Amaré showed that extrachromosomal circular DNA (eccDNA) and chromatin-associated proteins are present in stress granules, based on proteomic and sequencing analyses. Using HCR-FISH combined with imaging, the authors showed the colocalization of eccDNA with stress granule proteins. Furthermore, they found that CRISPR machinery targeting the eccDNA component of stress granules disrupts stress granule assembly, and that this effect is largely independent of Cas9 endonuclease activity. Notably, expression of cytoplasmic chromatin factors restores stress granule formation in the presence of CRISPR machinery in yeasts. This also rescues the growth defect caused by hypoxic stress, which correlates with impaired stress granule formation. Together, this manuscript provides insight into the presence of eccDNA in cytoplasmic membraneless organelles, specifically stress granules, and suggests a functional role for eccDNA within these structures under stress conditions.

      Strengths:

      The authors used a panel of ribonucleases to demonstrate that stress granule cores isolated from yeast and HEK293 cells are resistant to plasmid-safe DNase, an enzyme that does not degrade circular double-stranded DNA. To further support the presence of extrachromosomal circular DNA (eccDNA) in stress granules, they performed Circle-Seq on stress granule cores. The gel electrophoresis and sequencing experiments complement each other well, providing consistent evidence for eccDNA within these granules. Overall, this study provides insight into potential cytoplasmic roles for eccDNA, an area that remains largely unexplored.

      Weaknesses:

      (1) Figure 1F suggests that stress granule cores are susceptible to DNase I but not to plasmid-safe DNase (psDNase). However, its smearing pattern in the psDNase condition appears similar to that in the DNase I treatment shown in Figure 1E, although psDNase produces more discrete bands. The authors should comment on these differences between Figures 1E and 1F, or consider revising Figure 1F to improve consistency with Figures 1E and 1D.

      We suggest that the appropriate comparisons are between the DNase I and psDNase treatments within each figure panel, and not between panels (e.g., Figures 1E vs. 1F). The electrophoretic gels in the different panels were run for different lengths of time, and therefore the comparison between gels would be spurious. In Figure 1E, electrophoresis after DNase I treatment results in a characteristic smear, while after psDNase treatment yields discrete bands (lanes 2–3 vs. 4–5). Electrophoretic conditions for this figure were optimized to minimize diffusion and allow quantitative evaluation. The electrophoresis shown in Figure 1F, which compares yeast and mammalian stress granule core nucleic acids, was run for a longer period — as evidenced by the greater migration distance from the loading wells — yet still clearly shows the same qualitative difference between DNase I (smear, lane 3) and psDNase (discrete bands, lanes 1–2) treatments for the yeast samples. The apparent discrepancy noted by the referee therefore simply reflects the difference in electrophoretic conditions between the gels shown in the two separate figure panels.

      (2) The authors should clearly define "colocalization". Does it refer to complete spatial overlap between two signals (i.e., VCP and T30), or partial overlap (i.e., AHNAK DNA and G3BP)? Figure 3 and the associated text are descriptive. Quantitative analysis would strengthen the conclusions. For example, the authors could analyze the fraction of molecules localized to stress granules or provide Pearson's correlation coefficient or similar measurements.

      In our considered opinion, categorizing colocalization as either "partial" or "complete" implies a level of molecular precision that is physically unattainable at the resolution limits of any current light microscopy modality, and would therefore be misleading. Our approach employs super-resolution confocal laser scanning microscopy (Airyscan) with hybridization chain reaction fluorescence in situ hybridization (HCR-FISH) or with immunofluorescence. The detection method used offers higher spatial resolution and signal-to-noise ratio than single-point detector/physical pinhole confocal (or widefield epifluorescence) microscopy used in most prior stress granule studies. Despite these enhancements, the system retains inherent diffraction-imposed limits: a lateral (XY) resolution of ~130 nm and an axial (Z) resolution of ~350–400 nm, defining the minimum separable distance between two fluorescent signals. Structures smaller than these thresholds remain unresolved within a single point spread function (PSF) maximum – a volume sufficiently large to simultaneously accommodate multiple stress granule cores or tens of thousands of individual proteins (such as G3BP) and dozens of nucleic acid molecules several thousand nucleotides in length. Consequently, any detected fluorescence signal may represent the superimposition of a large and indeterminate number of individual molecules or particles. True molecular interaction analysis remains for future studies using technologies with angstrom resolution (e.g., cryo-electron tomography, cryo-EM, X-ray crystallography, smFRET, EPR, NMR, etc.). Metrics such as Pearson's correlation coefficient report solely on the degree of signal overlap at the PSF scale (hundreds of nanometers) and would not provide any insight beyond what is already conveyed by our data.

      (3) The authors used a CRISPR-based approach to target the Ty1 LTR retrotransposon, an abundant stress granule eccDNA, and they observed a loss of stress granule formation. However, this phenotype may be specific to Ty1 eccDNA rather than representative of all eccDNA species present in granules. In particular, the title "Cytoplasmic circular DNA is a key constituent of stress granules" implies a broader role. To support this claim, the authors should consider approaches that more globally deplete eccDNA rather than targeting a single eccDNA.

      We respectfully disagree with the referee that further depletion of eccDNA would alter our conclusions. A central finding of our study is that stress granules can be abrogated cytoplasmically by co-expressing a Cas9 endonuclease, active or inactivated by point mutations (D10A /H840A), and a gRNA (which is itself a fusion of the crRNA and trcrRNA, natively separate RNAs in the source bacterium). We show in Figure 4 that when the gRNA targets the Ty1 sequences, endonucleolytically active holoenzyme co-expression in the cytoplasm results in loss of the corresponding eccDNAs, as assayed by sequencing of the relevant cytoplasmic fractions. Critically, when a catalytically inactive Cas9 protein (dCas9) is co-expressed with the gRNA instead of the wild-type endonuclease, depletion of the eccDNAs containing Ty1 sequences no longer takes place (Figures 4D and 4E), but stress granule formation is still abrogated (Figure 4C).

      In our manuscript, we indicated (as "data not shown”) that co-expression with Cas9 of a gRNA "targeting" a sequence that is absent from the S. cerevisiae genome still results in abrogation of stress granule formation. These data are shown in Author response image 1. The gRNA is targeted to the sequence 5’-agaatcgatgcattt, which is absent in the genome of the yeast strain used.

      Author response image 1.

      It follows from our experiments that stress granule abrogation (1) is not a result of the catalytically active Cas9 endonuclease; (2) is not a result of the presence of a gRNA-directed but catalytically inactive Cas9 holoenzyme, but (3) is the result of the presence of a CRISPR holoenzyme (as defined above) in the cytoplasm.

      To reiterate, abrogation of stress granules occurs when a Cas9-gRNA complex is present in the cytoplasm, regardless of whether the nuclease activity exists, or the gRNA targets a sequence that is present in the genome. Importantly, the holoenzyme is required for this phenomenon: presence of the endonuclease or the gRNA alone does not abrogate stress granule formation (Figures S5).

      It is because of this unexpected observation that we next hypothesized that activities of the Cas9-gRNA complex other than sequence-specific gRNA-targeted endonucleolytic activity is driving the suppression of stress granule formation. The best documented such activity is DNA sequence sampling (1-dimensional diffusion). We think that 1-dimensional diffusion of the Cas9-gRNA holoenzyme is displacing from the cytoplasmic eccDNA interactors whose association with the DNA is required to drive stress granule assembly. The fact that the stress-granule suppressive effect of cytoplasmic Cas9-gRNA expression can itself be suppressed by two completely unrelated proteins whose only shared feature is action on chromatin (CHD1 and GCN5) strongly supports this hypothesis (Figures 4G, 4H and S6; also response to point 4, below), in addition to confirming that cytoplasmic eccDNA is packaged by histones in a conformation that CHD1 and GCN5 can both recognize.

      (4) The authors should provide additional experimental evidence to support the claim that eccDNA is packaged in a chromatin-like state. The rescue of stress granule formation by ectopic expression of modified chromatin-associated proteins (CHD1NES and GCN5NES) following CRISPR treatment does not necessarily demonstrate that eccDNA is packaged like chromatin under basal conditions.

      We would like to reiterate the temporal order in our experimental design (detailed in full in Methods and summarized in Results). Cas9<sub>NES</sub>-gRNA and CHD1<sub>NES</sub> (or GCN5<sub>NES</sub>) were expressed simultaneously (not sequentially) in the cytoplasm. This was intentional, so as to give each player ample opportunity to engage its preferred substrate under non-stress conditions, prior to the brief oxidative stress. The referee appears to believe that cytoplasmic eccDNA was pre-exposed to Cas9<sub>NES</sub>-gRNA, and then the bound endonuclease challenged with chromatin-modifying enzymes.

      Our experimental design accounts for the contrasting substrate specificities of CRISPR and chromatin-modifying enzymes. Cas9-gRNA (holoenzyme) binds to nucleosome-free DNA with sub-nanomolar dissociation constant (Kd 0.1–1 nM) but its association with chromatinized DNA is impeded 5- to 100-fold (Isaac et al., 2016; Yarrington et al., 2018; Strohkendl et al., 2021). In contrast, whereas CHD1 binding to DNA is strictly nucleosome-dependent — its chromodomains actively block engagement with protein-free DNA (Hauk et al., 2010), and its productive binding (Kd 10–200 nM) relies on obligate multivalent contacts with the histone octamer, H4 tail, and wrapped DNA (Farnung et al., 2017; Sundaramoorthy et al., 2018).

      Our observation that stress granule formation was unperturbed following oxidative stress is most parsimoniously interpreted as CHD1<sub>NES</sub> outcompeting the CRISPR machinery for cytoplasmic binding to eccDNA by virtue of the latter existing in a histone-bound state that is recognized as chromatin by CHD1 –simultaneously favoring CHD1<sub>NES</sub> engagement and impeding Cas9 access. Thus, our experiment in effect employs stress granule formation as a readout for differential binding to chromatin or chromatin-like eccDNA.

      Farnung, L., Vos, S.M., Wigge, C., and Cramer, P. (2017). Nucleosome-Chd1 structure and implications for chromatin remodelling. Nature, 550(7677), 539–542.

      Hauk, G., McKnight, J.N., Nodelman, I.M., and Bharat, T.A.M. (2010). The chromodomains of the Chd1 chromatin remodeler regulate DNA access to the ATPase motor. Mol Cell, 39(5), 711–723.

      Isaac, R.S., Jiang, F., Doudna, J.A., Lim, W.A., Narlikar, G.J., and Bhatt, D.L. (2016). Nucleosome breathing and remodeling constrain CRISPR-Cas9 function. Nature Struct Mol Biol, 23(12), 1097–1103.

      Strohkendl, I., Saifuddin, F.A., Gibson, B.A., Bhatt, D.L., Russell, R., and Bharat, T.A.M. (2021). Inhibition of CRISPR-Cas9 by bacteriophage-encoded proteins. Mol Cell, 81(8), 1665–1679.

      Sundaramoorthy, R., Hughes, A.L., Singh, V., Wiechens, N., Ryan, D.P., El-Mkami, H., Petoukhov, M., Svergun, D.I., Treutlein, B., Sproll, P., and Owen-Hughes, T. (2018). Structural reorganization of the chromatin remodeling enzyme Chd1 upon engagement with nucleosomes. eLife, 7, e35720.

      Yarrington, R.M., Verma, S., Schwartz, S., Trautman, J.K., and Carroll, D. (2018). Nucleosomes inhibit target cleavage by CRISPR-Cas9 in vivo.PNAS, 115(38), 9450–9455.

      Reviewer #2 (Public review):

      Summary:

      The authors report the presence of extrachromosomal circular DNAs (eccDNAs) within the core of stress granules purified from both yeast and mammalian cells.

      Strengths:

      This study is important for understanding the molecular mechanisms underlying stress granules containing eccDNAs and is likely to have a major impact on future research. A major strength of the study is the extensive experimental validation performed in yeast cells. In particular, cytoplasmic CRISPR-mediated targeting of eccDNAs suppresses stress granule formation and impairs recovery from hypoxic stress in yeast cells.

      Weaknesses:

      The conclusions would be further strengthened by validating the functional findings in an additional model system, such as mammalian cells.

      Comments:

      (1) Section: "Stress granule cores contain eccDNA"

      (a) The presence of eccDNAs would be more convincingly demonstrated using an orthogonal validation approach, such as DNA FISH targeting MYC and Centromere 8 (CEN8) on metaphase spreads from HEK293T cells (as performed in PMID: 34819668).

      The relationship between eccDNA dynamics and stress granule assembly across distinct cell cycle phases remains an important and poorly explored question. To our knowledge, no published data currently describe how stress response mechanisms are regulated during mitotic division, particularly in metaphase. Our identification of eccDNA as a component of stress granule cores can provide a first tractable framework to investigate this relationship. However, a systematic and in-depth characterization of this phenomenon warrants a dedicated future investigation.

      (b) The study would also benefit from assessing the presence of eccDNAs in the extracellular medium. For example, DNA could be extracted from conditioned media and analyzed by PCR using primers spanning eccDNA breakpoint junctions (as performed in PMID: 40074906; PMID: 36123406).

      We agree with the referee that eccDNA biology represents a fascinating and rapidly evolving area of research, particularly given the emerging role of eccDNA in oncogenesis. In this context, our identification of eccDNA as a core structural component of stress granules opens a novel avenue for exploring the connection between stress-dependent translational regulation and disease-associated eccDNA dynamics. While we acknowledge the importance of this direction, a rigorous investigation of this relationship requires extensive multifaceted experimentation that falls beyond the scope of the current study.

      (2) Section: "eccDNA-CRISPR abrogates stress granules"

      These findings should be further validated under additional stress conditions, such as drug-induced stress (like methotrexate) or nutrient deprivation in the cell medium. In addition, the same set of experiments should be performed in HEK293T cells to support the broader relevance of the observations.

      We agree with the referee that the composition and dynamics of stress granules arising from different stressors is an important endeavor. However, given the range of stressors documented to result in stress granule formation, those studies fall well beyond the scope of this manuscript. We will note however that the presence of eccDNA in stress granules of yeast and human cells is strong evidence for conservation of function(s). We think that exploration of the role of eccDNA in stress granule formation across the kingdoms of life (stress granules were first observed in heat-shocked tomato plants), cell cycle stages, stressors, etc. will be important research programs for the future.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figures 3D and 3I: The use of magenta and red makes it difficult to distinguish between the two labeled signals. Consider using more contrasting colors to improve visual clarity.

      We appreciate the comment regarding color choices in the figures. In our view, magenta and red are sufficiently distinguishable as nucleic acid labels, particularly when combined with the green signal representing G3BP in these panels.

      (2) Figures 3F and 3G: Do the authors have an explanation for why AHNAK or MAPT DNA (white) does not colocalize with the anti-DNA immunofluorescence signal?

      Immunofluorescence (IF) is standard for detecting protein antigens but has limitations when the target is a non-protein molecule such as DNA, owing to its compacted chromatinized state. Anti-DNA antibodies can miss a significant fraction of their targets because the DNA backbone remains largely inaccessible, a limitation that DNA-FISH overcomes by directly hybridizing probes to denatured DNA sequences with high specificity. The fixation step required for both IF and FISH imaging can introduce additional steric barriers that disproportionately restrict antibody access compared to small nucleic acid probes. Even under optimized conditions, the IF signal with anti-DNA antibodies is inherently reflective of a subset of the total cellular DNA content.

      (3) Adding a subtitle on page 12 ("The abundant histones in purified stress granule...") would improve the overall structure and readability of the manuscript.

      We think that an additional subtitle would not substantially improve the readability of what is, admittedly, a very dense manuscript that employs a diversity of experimental approaches.

      (4) It would strengthen the analysis if statistical significance were included for the different time points in Figure 5C.

      We appreciate the reviewer’s suggestion. Figure 5C shows the largest difference at 40–45 hours after stress recovery, which is statistically significant between Cas9NES-gRNA (or dCas9NES-gRNA) and Cas9NES or gRNA only (two-tailed Student’s t-test, *, p ≤ 0.05). All primary experimental data are publicly available (FigShare) so further analyses can be performed by interested future parties.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      Summary:

      Eroglu and Hobert demonstrate that injecting CRISPR guides and repair constructs to target three genes at a time, tagging each with a different fluorescent protein, and selecting which gene to tag with which fluorophore based on genes' expression levels, can improve efficiency of gene tagging.

      Strengths:

      This manuscript demonstrates that three genes can be targeted efficiently with three different fluorophores. It also presents some practical considerations, like using the fluorophore least complicated by agar/worm autofluorescence for genes with low expression levels, and cost calculations if the same methods were used on all genes.

      Weaknesses:

      Eroglu has demonstrated in a previous publication that single-stranded DNA injection can increase efficiency of CRISPR in C. elegans, while inserting two fluorescent proteins and a co-CRISPR marker into three loci, and Paix et al 2015 demonstrated simultaneous insertion of two fluorescent tags. The current work is valuable and incremental advance. In general, I applaud the authors' willingness to strategize about how whole proteome tagging might be accomplished. I predict that the advance here will be one of many small advances that will get the field to that goal. The title oversells the advance presented, in my view, since seems like one among many key advances, and the first sentence of the Discussion seems a more apt summary of the key advance here.

      Some injections targeted genes on the same chromosome together, which will create unnecessary issues when doing crossing that will be useful for some future experiments. This made me wonder if injecting 3 together really is helpful vs targeting each gene separately, since only 5 worms need to be injected. It cuts time down by 2/3, but perhaps avoiding targeting the same chromosome with two tags would be useful.

      The limited utility of current blue fluorescent proteins makes me wonder if it's worth using at this stage, before there are better blue fluorescent proteins, or better yet, far red, to avoid issues with live imaging under phototoxic UV or near-UV illumination.

      These comments are a repeat of the original comments, and we refer the reader to our response to the original comments.

      Reviewer #2 (Public review):

      Original Review:

      The manuscript by Eroglu and Hobert presents a set of strains each harboring up to three fluorescently tagged endogenous proteins. While there is technically nothing wrong with the method and the images are beautiful, we struggled to appreciate the advance of this work - who is this paper for?

      As a technical method, the advance is minimal since the first author had already demonstrated that three mutations (fluorophore insertion and co-CRISPR marker) could be introduced simultaneously.

      As a pilot for creating genome-scale resources, it is not clear whether three different fluorophores in one animal, while elegantly designed and implemented, will be desired by the broader community.

      Finally, the interpretation of the patterns observed in the created lines leaves much to be desired. A Table with all the observations must be included and can replace the tedious (and often wrong) descriptions of the observations with the different lines. It would be too much to point out every mistaken expectation of protein expression. Two examples include:

      The expectation that ACDH-10 is enriched in the intestine and epidermal tissues (hypodermis) is naïve - there are multiple paralogs of this protein (look at WormPaths or WormFlux) that may share functions in different tissues. There is also no reason to assume that fatty acid metabolism does not occur in other tissues (including the germline). Finally, there are no published studies about this enzyme, so we really don't know for sure what it's doing.

      The expectation that HXK-1 is ubiquitously expressed is similarly naïve. There are three paralogous enzymes that are all associated with the same reaction, and we have shown that these three function redundantly in vivo, perhaps in different tissues (PMID: 40011787). Moreover, single cell RNA-seq data (PMID: 38816550) also shows enrichment of hxk-1 in gonadal sheath cells.

      The table should have at least the following information: gene/protein name - Wormbase ID - TPM levels of single cell data assigned to tissues for L2, L4 and adult (all published) - tissues in which expression is observed in the lines presented by the authors.

      Other points:

      (1) We would encourage the authors to provide systematic validation of the reported insertions. The manuscript reports that 24 of 30 tags were isolated and visible but does not clearly state whether each isolated line was confirmed by sequence‑level validation to be correctly in‑frame and free of unintended mutations at the target locus.

      (2) The manuscript presents aggregated success counts (e.g., 8/10 mTagBFP2 tags, 9/10 mStayGold, 7/10 mScarlet3) and useful narrative descriptions of injection outcomes. We suggest also to include per‑locus success rates.

      (3) For pools that required re‑injection after initial failures, we would like to see a description of the specific changes that were made to the injection mixes or procedures (e.g., new repair template prep, different Cas9 reagent lot, guide redesign). This will be useful troubleshooting information for others.

      (4) The authors states that the fluorophore sequences are codon-optimized for C. elegans. We suggest they provide the exact donor/tag sequences used specifically state whether the fluorophore sequences contain any synthetic/artificial introns or other sequence modifications (e.g., silent PAM‑disrupting mutations) were included in the donor templates.

      (5) Page 3: Include a reference for "The C. elegans genome encodes around 20,000 genes"

      We hope these comments are useful.

      Comments on Revised Version:

      Overall, we found the responses to be quite recalcitrant.

      We have one remaining composite concern about the comparison between observed expression patterns with the new strains versus published data.

      First, the authors only report patterns for one stage while it should be not too much effort to image the different life stages. However, since this is a revision, we are not formally requesting they do this.

      Second, in the now provided Table (thank you) 'observed expression' (last column) is lacking for 9 of the 30 proteins, and for 6 of these the procedure was not successful. Why not report patterns for the other three? It is confusing also because on page 5, the authors say that "overall, 24 of 30 tags ...all of which were visible with fluorescence stereomicroscopy" - are we missing something? Also, they then said that they "obtained 6/9 of the originally failed tags"; why are the corresponding patterns not included in table 1, and are 9 proteins still labeled as "no" in the "success?" Column?

      We appreciate the chance to clarify this matter: There are only 6 “no” in the “success” column. In two cases, HAT-1 and CBP-1, expression was dim at F1 but still sufficient to pick positive worms and quantify success rate at the locus. We noted these as “dim” on the table to indicate that if expression was lower, we likely would not have been able to isolate them at F1. In one case, COX-6B, expression was too dim at F1 to be isolated but was sufficient at F2 to be visualized and isolated from parents that were positive for the other two tags. We now clarified this distinction in the table and accompanying text: “Fluorescent signals of HAT-1::mScarlet3 and CBP-1::mScarlet3 in F1 progeny were dim but still sufficiently visible for quantification of knock-in efficiency, indicating that they are at the lower end of detectability for mScarlet3.”

      We imaged worms that had multiple tags as proof of principle and are happy to provide strains to those who would like to image/study them. At this point we are not convinced that imaging more worms would add to the conceptual framework.

      Third, we strongly feel that the response to our comments about expression patterns is not adequate. On page 5 the authors say that "all proteins were expected to be ubiquitously expressed" and that "scRNA-seq indicated that transcript abundance was ubiquitous and without strong tissue-specific enrichment with few exceptions". However, in their rebuttal, the authors now argue for tissue-specific expression for proteins with paralogs, turning around their own argument! Moreover, their Table indicates that many genes show tissue-enriched expression by RNA-seq while many of their tagged proteins exhibit ubiquitous expression.

      We respectfully disagree that there is contradiction. In our response, the discussion on paralogs was added as a clarification in response to the referee’s original comments (e.g., regarding ACDH-10): “There is also no reason to assume that fatty acid metabolism does not occur in other tissues (including the germline).” We wanted to make it clear that we were not concluding fatty acid metabolism (or other processes) does not occur in other tissues.

      We wish to stress that we never argued that paralogs could not fulfil the same essential function across tissues. The proteins were selected because their biological functions (e.g., glycolysis, fatty acid β-oxidation, translation) are broadly required, and that scRNA seq generally predicted broad expression with few exceptions as detailed in the text. Paralogs with similar activities (e.g., hxk-1, -2, -3) may overlap broadly in expression, or individual paralogs may carry out the process in different tissues provided one carries out the reaction in each tissue. For acdh-10 and hxk-1 specifically, both appear broadly expressed across tissues by scRNA-seq, with no consistent enrichment or depletion across datasets. So, our central point is that: for a specific gene involved in an essential process, transcript data alone are not sufficient to accurately predict tissue specific enrichment. Not that the processes do not occur in tissues where one paralog is absent. The possibility that a paralog may compensate for lack of expression is in no way contradictory with our conclusion.

      The table does not generally show tissue-enriched expression: it simply lists three tissues with the highest quantitative value in the respective dataset. For instance, taking the first gene from the list (Y82E9BR.3) and looking at the Ghaddar dataset, the top 3 tissues (log2(TPM)) are: pharyngeal muscle (13.4), gonadal sheath (12.9), marginal cells (12.9). The next 3 tissues are: body wall muscle (12.9), pharyngeal epithelium (12.8), and intestine (12.3). Even when there were apparent enrichments among the top 3 tissues, there were significant disagreements between datasets, and beyond top 3 even greater disagreements (the datasets agreed on the top tissue only 4 times over the 30 genes). These indicate that much of the variation is attributable to experimental noise rather than true predicted enrichment. The referee points to HXK-1 being correctly gonadal sheath enriched in one scRNA dataset; however, the other two datasets actually show different sites as being highest, and the same dataset misses effects in other cases. This is precisely why protein level data is needed.

      We further clarified this issue in the text: “We thus selected 30 genes across a variety of bulk transcript expression ranges which are generally predicted to be broadly expressed based on molecular function or, where molecular function was unknown (e.g., ZK632.9), single cell RNA sequencing (scRNA-seq) data (Table 1, Fig. 2A, B) (Gao et al., 2024; Ghaddar et al., 2023; Taylor et al., 2021).”

      Overall, this indicates that both the overall accomplishment of generating tagged protein strains and analyzing their expression is oversold.

      We have tried to make clear that our contribution is not a handful of new tagged strains added to the many that already exist. Rather, as stated in the abstract and elsewhere, we propose a strategy and provide proof-of-concept for scaling up tagging efforts. We believe the importance of this cannot be oversold.

      Reviewer #3 (Public review):

      Summary:

      The authors argue that establishing the expression pattern and sub-cellular localisation of an animal's proteome will highlight hypotheses for further study. This claim is probably accepted by many in the community. This manuscript seeks to confirm the feasibility of establishing such a resource, by using current transgenic methods to knock in DNA encoding different colored fluorescent tags into C. elegans genes.

      Strengths:

      The authors make the points above. For example, they provide evidence that the C. elegans germline harbors two populations of mitochondria that differ qualitatively in the proteins they express. They also confirm that labelling the whole proteome is an achievable goal with relatively limited resources and time.

      Weaknesses:

      The work is somewhat incremental in that it uses existing transgenic technology. Cell biology in C. elegans is challenging because of the small size of many of its cells, notably neurons. This can make establishing the sub-cellular localisation of a fluorescently tagged protein, or co-localizing it with another protein, tricky. The authors point out in their introduction that advances in light microscopy such as diSPIM, STED and ISM (a close relative of SIM), have increased the resolution of light microscopy. They also point out that recent advances in expansion microscopy can similarly help overcome the resolution limit. However, they do not use these technologies to characterize their transgenic strains.

      Reviewer #4 (Public review):

      Summary:

      Tagging the entire proteome of a metazoan would be a landmark achievement, providing a powerful complement and extension to existing "omic" catalogs in model systems. Here, Eroglu and Hobert argue that efficiently tagging multiple loci in a single "batch" would make the community-based achievement of this goal realistic. They provide rigorous evidence that such an approach is indeed feasible, exploring issues related to efficiency, design and screening strategies, disruption of gene function, and the potential for endogenously tagged alleles to reveal unexpected aspects of protein expression and localization. While the work has some minor gaps that are important to rigorously assess the feasibility of the proposed effort, the detailed and valuable insights that emerge should provide impetus to the community to coordinate efforts to make this ambitious goal a reality.

      Strengths:

      The work has numerous strengths. The authors provide compelling evidence that:

      Three distinct loci can be efficiently targeted with three distinct fluorescent tags in a single injection.

      Thoughtful targeting design can reduce the likelihood of disruption of function by the tag.

      Systematic design principles based on expression level and predicted localization/function can be used to optimize tagging strategies.

      The resulting tags can provide unexpected insight into patterns of protein production and subcellular localization.

      Not all of these advances are novel in themselves, but taken together, they represent an important technical and conceptual advance. The most important strength comes from the exceptionally high value of the goal itself, in that the work is that it has the potential to spur a community-wide effort toward achieving the ambitious goal of proteome-wide tagging.

      We appreciate the referee’s enthusiasm and hope that this will engage members of the community in a collective effort.

      Weaknesses:

      The work's shortcomings are minor.

      One concern has to do with the feasibility of the proposed screening strategies. The experimental design cleverly coinjects tags for three loci in different gene expression 'zones'; this expression level determines which tag will be used. As the authors allude to, there is an important distinction between genes with the same overall FKPM value between those that are expressed broadly and those focally expressed in a specific tissue. The proposed strategy claims that there are a sufficient number of highly expressed genes "to be used as visible markers" for recovering successfully edited animals. It would be useful for the authors to discuss the issue of broad vs focused expression among this set of genes a bit more thoroughly, with an eye toward the issue of how likely it is that these genes could indeed consistently be used as visible markers, particularly for those at the low end of this limit.

      To give two examples, this principle aided us with screening F54C8.1 and HAT-1. We added additional discussion on this to the first paragraph of the discussion: “For instance, we could clearly visualize F54C8.1::mScarlet3 in adult sperm by fluorescence stereomicroscopy despite a bulk FPKM of 16. Similarly, nuclear localized proteins will likely be easier to detect even at low expression levels, given the concentration of signal in small subcellular compartments. Indeed, this helped us detect HAT-1::mScarlet3 (56 bulk FPKM), which may have been too dim if distributed more broadly within cells.”

      What fraction of the proteome (on a per-gene basis) is secreted proteins? How difficult will it be to screen these for successful tags? Are there specific tags that would be more optimal for secreted proteins? (The authors mention the use of an SL2 or T2A cassette to label the cells in which these proteins are expressed but note that there are technical challenges associated with doing this at scale.)

      We added some of these points to the discussion: “Moreover, around 17% of the C. elegans genome (3,484 genes) may encode for secreted proteins (Suh and Hutter, 2012). Endogenous tagging of a substantial fraction of these proteins could reveal spatial patterns of secretion, distinguishing components that remain near their cell of origin from those that disperse to distal sites (Keeley et al., 2020). Tagging secreted proteins can also reveal sites of secretion – such as apical or basolateral membranes, or neurites – as has been observed for specific insulins (Sural et al., 2025) and for neuropeptides that localize selectively to synaptic regions (Toker et al., 2025).”

      Various tags have been used for secreted proteins including Venus, TagRFP, and mNeonGreen. The pH of secretory vesicles is ~5.0-5.5, so chosen FPs should have a pKa below this range to avoid denaturation. All 3 fluorophores used here (mStayGold, mScarlet3 and mTagBFP2) have pKa’s below this range and would likely be fluorescent within secretory vesicles.

      For secreted and/or weakly expressed genes, it would be useful for the authors to estimate for what fraction of these would successful insertions need to be screened by PCR, and what resources (time and money) this would likely entail. 

      We think that the bulk of ECM proteins would likely be visualizable without PCR due to their broad and stable expression, and as mentioned a good portion of these have been already tagged. However, it is likely that most of the secreted small peptides will have to be screened by PCR. We use homemade Taq, which makes material cost of the reagents minimal. A pair of genotyping primers costs ~$8 (~$27,872 for all secreted genes).

      Hands on time for lysis of 48-96 worms is approximately 20-30 minutes, with time to set up PCR around 5-10 minutes per target, and time to load a gel of 10 mins. In a given pool, 2/3 could be a putative secreted protein; thus, the same lysed population would enable screening for two targets at once. Collectively, around 40-60 mins of hands-on time would be required for two genes (around 20-30 mins per gene). Given 18 targets are injected per day, if 12 are screened by PCR, the screening could be done in 6 hours per day without affecting throughput. Most of the time spent on PCR would be replacing fluorescence screening time and would not overlap with the rate limiting injection step, performed by a separate specialist.

      For how many genes would a single tag not capture all predicted isoforms?

      Around 25% of C. elegans genes are thought to undergo alternative splicing (PMID: 21177968), with on average, ~2 isoforms per transcript. Among our selected genes, we only had one case where a single tag would not capture all isoforms (flad-1). We examined an additional 30 random genes and found no more examples by chance. So, in our view, this will be rare though we recognize in some cases a practical decision will need to be made, which could involve consideration of expression levels of each terminal exon.

      Finally, some readers might object to the authors' assertion in the abstract that this work is "a first step in this direction" (presumably referring to designing a strategy for whole-proteome tagging). There is no concern that the authors are disregarding the extensive work of other groups, as they explicitly mention the contributions of other groups to the foundation that enables the present work. However, the spirit of the abstract could be misinterpreted by a well-intentioned reader.

      We appreciate the referee’s perspective and have reworded this phrase in the abstract to: “As proof-of-principle for scalable pooled tagging, we undertook a pilot study in the nematode C. elegans, in which we set out to tag 30 different genetic loci with three different fluorophores, with 3 tags being introduced at a time.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study uses the yeast two-hybrid assay to identify proteins that may interact with yeast Set1 and other subunits of COMPASS/Set1C, the histone H3K4 methyltransferase, providing also some evidence for Set1 sumoylation and a role of SET1C methylating other factors in vitro. The results are valuable, and they should contribute to understanding the functions of the conserved SET1C complex, as they suggest potential functional connections with RNA biogenesis, chromatin remodeling, and non-histone methylation, whose implications would yet need to be explored. Nevertheless, apart from the fact that only a small subset of the Y2H interactions is further examined, the validating experiments are only partial or inconclusive, the strength of evidence being at this point incomplete.

      We present a systematic SET1C interaction map that provides a structured resource for generating and testing new hypotheses on SET1C function. We emphasise that these interactions represent a hypothesis generating resource rather than a set of validated protein–protein interactions. To reflect this, the manuscript has been carefully revised to distinguish clearly between observation and interpretation, and to avoid overstatement of the data. Accordingly, we have revised the title and the abstract. Selected examples are explored further to illustrate how candidates from the dataset can be followed up, but the primary contribution of this work is to provide a structured framework and resource that can guide future mechanistic studies of SET1C function.

      We thank the reviewers for their thoughtful comments. We have followed their recommendations by modifying the structure of the manuscript, removing distracting results and relocating some figures to the supplementary materials to improve the readability of the manuscript. At the same time, the reviewers acknowledge that the dataset is extensive and that aspects of the validation work are valuable.

      The changes made to the manuscript's structure in accordance with the reviewers' recommendations are as follows:

      (1) Figure 1 is accompanied by a table (Table S2) with the raw data describing all the interactions from the ten 2H screens. This table also lists common interactors found in the independent screens. I'm afraid Table S2 was omitted from the initial submission of the manuscript

      (2) Figure 2 has been modified to include an AlphaFold modeling of a seven-subunit Set1C complex (Set1– Bre2–Sdc1<sub>2</sub>–Swd1–Swd3–Spp1) together with Kap104. Figure 2D has been moved to a new Figure S2

      (3) The initial figure S2, which was problematic, has been removed, along with the accompanying text.

      (4) Figure 3 of the original paper has been moved to the supplementary material and is now shown as a new Figure S3.

      (5) Figure 5 in the original paper becomes Figure 3 in the revised version

      (6) Figure S3 (Co-IP between Set1 and Prp22), which serves as validation data, has been moved to the main figures and is now presented as Figure 4.

      (7) Figure 6 in the original paper becomes Figure 5 in the revised version

      (8) Figure 4 from the original paper has been repositioned as the first figure (new Figure 6) of the biochemical characterization of the interaction between Snf2 and Set1C.

      (9) Figure 7 has been removed from the manuscript. We have kept the original Figure 7E as a new Figure S6.

      (10) Figures 8, 9, 10 become Figures 7, 8, 9.

      Public Reviews:

      Reviewer #1 (Public review):

      We thank Reviewer 1 for the careful and thoughtful evaluation of our manuscript. We fully agree that yeast two hybrid screening provides candidate interactions that require cautious interpretation, and we recognise that our original version did not always make this sufficiently explicit.

      In the revised manuscript, we have made substantial changes to address this central concern. All Y2H interactions are now consistently presented as candidate or potential interactions, and speculative statements have been either removed or explicitly framed as hypotheses. Our intention is that the reader can clearly separate the dataset itself from any proposed biological implications.

      Second, we have refocused the manuscript to better reflect its primary contribution. We now present the Y2H screens as a comprehensive resource that defines a set of candidate interactions for SET1C, rather than as a set of validated functional relationships. In line with this, we have reduced the emphasis on speculative models and removed sections where the connection to experimental evidence was not sufficiently strong. This includes the removal of Fig. S2 and Fig. 7 and the associated text, as well as the relocation of several figures to the supplementary material. Where appropriate, we have added statements highlighting the limitations of the approaches used and the need for future work to establish physiological relevance.

      More generally, we agree with the reviewer that the value of Y2H data lies in generating testable hypotheses rather than establishing conclusions. We have therefore revised the manuscript throughout to ensure that the interpretation remains proportionate to the strength of the evidence.

      We hope that these changes address the reviewer’s concerns and result in a clearer and more appropriately balanced presentation of the data.

      The manuscript by Luciano et al is a collection of experiments about the yeast histone 3 lysine 4 methyltransferase, Set1, starting with 10 yeast two-hybrid screens (Y2H). Y2H screens were briefly popular 20+ years ago, but the persistently unfavourable false-to-true positive ratios limited their utility, and the conclusion emerged that Y2H is an unreliable approach for gathering protein-protein interaction data. Y2H outcomes are candidate interaction lists at best, strongly contaminated by false positives. Here, the authors employed a company (Hybridomics) to perform the Y2H screens.

      The primary data is not presented, and the outcomes are summarized using the Hybridomics in-house quality scoring system in Figure 1A. It is not possible to evaluate these data, and the manuscript presents cartoon summaries that the reader must accept as valuable.

      Hybrigenics brings extensive experience from conducting numerous screens, enabling the team to recognize recurring false positives that commonly arise in screening assays. In their detailed analysis, Hybrigenics reports the number of clones recovered and the extent of overlap among interaction regions, both of which contribute to the confidence scores they assign. Table S2, provided in the revised version, more accurately reflects the raw data obtained by Hybrigenics. Nevertheless, we agree that false positives contaminate the list of potential interactors. Some interactions may also be indirect through a common interactor and do not reflect a physiological interaction.

      (1) Based on the extensive knowledge about Set1C/COMPASS acquired from genetics and biochemistry by many labs (including the Geli lab), the results presented here from the 10 Y2H screens are notably patchy. Of the 7 subunits of this complex, only one (Spp1) was identified using Set1 as bait. Conversely, as baits, Swd2, Spp1, Shg1, captured Set1, and the Bre2-Sdc1 interaction was reciprocally identified. These interactions were scored at the highest confidence level, which lends some confidence to the screens. However, the missing interactions, even at the third confidence level, indicate that any Y2H conclusions using these data must be qualified with caution. The authors do not appear to be cautious in their lengthy evaluations of these candidate interactions, which are illustrated with cartoons in Figures 2 and 3, with some support from the literature but almost without additional evidence. Snf2 is a particularly interesting candidate, which the authors support with pull-down experiments after mixing the two proteins in vitro (Figure 4). After Y2H, this is the least convincing evidence for a protein-protein interaction, and no further, more reliable evidence is supplied.

      We thank the reviewer for raising this important point regarding the strength of the evidence supporting the Set1– Snf2 interaction. We agree that the current data do not establish a definitive physiological interaction. In the discussion, we explicitly note the limitations of the current data.

      For Figure 2, as recommended by referee 2, we performed AlphaFold modeling of a seven-subunit Set1C complex (Set1–Bre2–Sdc1<sub>2</sub>–Swd1–Swd3–Spp1) together with Kap104. Consistent with the Y2H data, the model recapitulates binding of the Kap104 SID to the PY-NLS region of Set1 (residues 40–90).

      We have moved Figure 3 in the supplementary materials.

      (2) Figure 5 continues the cartoon summary of extrapolations from the Y2H screens, again without supporting evidence, except that the authors state.

      Figure 5 is now Figure 3. We have added the statement in the text: “It is not feasible to validate all of these interactions within the limits of this manuscript, and their validity should therefore be interpreted with caution. Nonetheless, these findings provide a useful basis for future research”.

      "We have refined the interaction region between Set1, Prp8 and Prp22, showing that Prp8 and Prp22 interact strongly with Set1-F4 (n-SET). Prp22 interacts in addition with Set1-F1 (Figure S2)." However, Figure S2 does not show this evidence and is incoherent.

      When we say that we have refined the interaction region between Set1, Prp8, and Prp22, we mean that we have restricted the interaction regions according to Y2H criteria. Indeed, we have not shown the spots illustrating the results. This statement has been deleted as well as Fig. S2

      The figure legends for Figure S2B and C do not correspond to the figure.

      (B) Expression of the F1-F5 fragments in yeast cells. Fusion proteins were detected with an anti-GAL4 monoclonal antibody. TOTO yeast cells (Hybrigenics) were transformed with the different pB66-Set1-F1 to F5 plasmids and subsequently with either P6, pP6-Snf2 762-968, pP6-Prp8 37-250, or pP6-Prp22 379-763 that were identified in the Y2H screens. Transformed cells were incubated 3 days at 30{degree sign}C on SD-LEU-TRP and then restreaked on SD-LEU-TRP-HIS with 3AT. Cell growth was monitored after 2 days at 30{degree sign}C.

      (C) Solid and dotted arrows indicate that transformed TOTO cells transformed with pB66-Set1-F1 to F5 and the indicated prey (Snf2, Prp8, and Prp22) are growing in the presence of 20 mM and 5 mM of AT, respectively.

      Figure S2D is two almost featureless dark grey panels accompanied by the figure legend D) Control experiment showing that TOTO cells transformed with p6 and pB66-Set1-F4 are not gowing (sic) in the presence of 5 mM or 20 mM AT.

      We agree that the legend for Figure S2 was unclear and does not accurately describe the panels shown in the figure. Fig; S2 has been deleted in the revised version. The results shown in the original Fig. S2 add limited information and may detract from the clarity of the main points.

      In the revised version, we have moved the CoIP analysis demonstrating the interaction between Set1 and Prp22 (previously shown in Figure S3) into the main figures (now Figure 4) to further support and validate the two-hybrid screening results presented there.

      Line 343. Interestingly, the two-hybrid screens reveal that Set1 1-754 interacted with Gag capsid-like proteins of Ty1 (Figure S5), raising the possibility that Set1 binding to Ty1 mRNA is linked to the interaction of Set1 1-754 with Gag.

      This is another example of the primary mistake repeatedly made by the authors -Y2H interactions are candidate results and not conclusive evidence.

      This statement is supported by our previous findings showing that Set1 binds Ty1 mRNA independently of its dRRM domain and represses Ty1 mobility at a post-transcriptional stage (Luciano et al., Cell Discovery, 2017; PMID: 29071121). One possible explanation for Set1 association with Ty1 mRNA is its interaction with the Gag capsidlike protein. In this context, the observed interaction between Set1(1–754) and Gag capsid-like proteins is consistent with this model.

      To further illustrate this point, the authors highlight the candidate interaction between Nis1 and 3 Set1C subunits.

      While we agree that the Nis1-Set1C interaction has not been demonstrated beyond doubt, we feel that our Y2H and in vitro binding experiments provide reasonable evidence that the interactions may be relevant. It is important to consider that any interaction assay can provide negative (and false positive) results, this includes Y2H, in vitro binding and mass-spec analysis of purified complexes from cells. We feel that it is not appropriate to only trust protein interactions that are strong and stable enough to be demonstrated via purified complexes. It is clear that some protein interactions do occur in transient and weak manner and therefore are not compatible with biochemical purification approach. This indeed is the strength of alternative methods like Y2H and in vitro binding assays, that interactions can be identified and tested even if the physiological context of the interaction may be more complex.

      (3) After multiple speculations based on the Y2H candidates, the authors changed to focus on sumoylation of Set1, which has previously reported to be sumoylated. Evidence identifying two sumoylation sites in Set1, in the N-SET and SET domains, is valuable and adds important progress to the role of sumoylation in the regulation of H3K4 methyltransferase, relevant for all eukaryotes. This illuminating part of the manuscript is only tenuously connected to the preceding Y2H screens and concomitant speculations.

      We thank Referee 1 for their comment. While it is true that there is only a modest connection between Set1 interactors involved in direct or indirect sumoylation and the characterization of Set1 SUMOylation sites, we believe that this does not constitute a weakness of the manuscript.

      (4) The manuscript then describes a red herring exercise involving Set1 methylation of Nrm1. In an already speculative and difficult manuscript, it is exasperating to read a paragraph about a failed idea. Apart from panel E, Figure 7 is a distraction, and I believe it should not be shared.

      (5) However, despite the failure with Nrm1, Line 443 - The H3K4-like domain in Nrm1 raised our attention to other yeast proteins that carry such sequences.

      This line of thinking is even less connected to the Y2H screens than the sumoylation work.

      However, the authors present a reasonable evaluation of the yeast proteome screened for six amino acids similar to the known H3K4 motif ARTKQT (Figure 7e).

      (6) However, this evaluation goes nowhere and has no connection with the next section of the manuscript, which is entirely speculation about the regulation of metabolism and stress responses based on the Y2H results and selected evidence from the literature.

      In response to comments 4 and 5, we have removed Fig. 7 and the paragraph titled “The transcriptional corepressor Nrm1 interacts with SET1C.” Part of this paragraph and the section describing the screen of the yeast proteome for six–amino acid sequences resembling the H3K4 motif (ARTKQT) has been kept as Fig. S6.

      In the abstract, we have removed the sentence: We demonstrate that the transcriptional corepressor Nrm1 is methylated by SET1C in vitro suggesting that H3K4-like domains may represent a class of non-histone substrates for SET1C.

      At the end of the introduction, we have deleted “the transcriptional corepressor Nrm1” in the sentence: In addition, we demonstrate that the transcriptional corepressor Nrm1 and the Snf2 AT-hook are both methylated by SET1C in vitro

      (7) The manuscript then describes more failed experiments regarding lysine methylation of Snf2 by Set1C, which unexpectedly reports arginine methylation rather than lysine. The manuscript does not currently meet the standard expected for this type of paper - the composition is somewhat incoherent and there are no previous reports of arginine methylation by SET domain proteins.

      We have integrated extensive in vitro reconstruction experiments with complementary in vivo studies, all conducted according to the rigorous standards expected by leading journals. These approaches have allowed us to reach the conclusions presented in this manuscript. While some of these findings are unexpected, they are supported by the data. We have carefully discussed the results and their limitations to provide a comprehensive interpretation.

      The manuscript presents a very experienced grasp of the literature and a sophisticated appreciation of the forefront issues, but a surprising failure to eliminate uninformative failures and peripheral distractions. The over interpretation of Y2H results is a dominating failure. There are some valuable parts within this manuscript, and hopefully, the authors can reformat to eliminate the defects and appropriately qualify the candidate data.

      We thank Referee 1 for these insightful comments. In the revised version, we have followed the advice to remove non-informative failures and peripheral distractions. Additionally, we exercise greater caution to avoid over-interpreting the Y2H results.

      Reviewer #2 (Public review):

      Summary:

      This paper starts with a large-scale yeast two-hybrid (Y2H) screen using Set1 (full-length and smaller parts) and other Set1C/COMPASS subunits as bait. There are hundreds of possible interactions identified, but only a small number are given any follow-up. While it's useful to document all the possible interactions, the unfocused and preliminary nature of the results makes the paper feel scattered and incomplete.

      Strengths:

      The Y2H screen was very comprehensive, producing lots of interesting possible leads for further experiments.

      Weaknesses:

      The results are useful but incomplete because only a small subset of the Y2H interactions is further examined. Even in the case of those that were further tested, the validating experiments are only partial or inconclusive.

      Referee 2’s comments align in some respects with those of Referee 1. In the revised version, we have followed the detailed Referee 2 suggestions to reduce the scattered nature of the manuscript. In addition, we include an AlphaFold model of the interaction between the Set1 N-term 1-754 with the SID domain of Kap104 that involves the proposed Set1 PY-NLS sequence.

      Reviewer #3 (Public review):

      The SET1C/COMPASS complex is the histone H3K4 methyltransferase in Saccharomyces cerevisiae, where it plays pivotal roles in transcriptional regulation, DNA repair, and chromatin dynamics. While its canonical function in histone methylation is well-established, its full interactome remains poorly defined. Moreover, whether SET1C methylates non-histone substrates has been an open question. In this study, Luciano et al. employ systematic yeast two-hybrid (Y2H) screening to uncover novel interactors and functions of SET1C. Their findings reveal potential functional connections to RNA biogenesis, chromatin remodeling, and non-histone methylation.

      The authors performed multiple Y2H screens using Set1 (full-length, N-terminal, and C-terminal fragments) and each of its seven subunits as baits. They identified high-confidence interactors that link SET1C to diverse cellular processes, including chromatin regulation (e.g., the SWI/SNF complex via Snf2), DNA replication (e.g., Mcm2, Orc6), RNA biogenesis (e.g., spliceosome components Prp8 and Prp22; polyadenylation factors Pta1 and Ref2), tRNA processing (e.g., Trm1, Trm732), and nuclear import/export (e.g., importins Kap104 and Kap123). Some of these interactions were further validated by immunoprecipitation or in vitro assays.

      Given the interaction of Set1 with Slx5 and Wss1 - proteins involved in SUMO-dependent processes - the authors investigated and convincingly demonstrated that Set1 is sumoylated. This modification may influence the function and regulation of the SET1C complex.

      Finally, the authors provide evidence that SET1C methylates proteins beyond histone H3K4, notably Nrm1, a transcriptional corepressor, and Snf2, the catalytic subunit of the SWI/SNF chromatin remodeling complex. Although Nrm1 contains a domain resembling the H3K4-methylated sequence (H3K4-like domain), this region does not appear to be required for its methylation. The search for other proteins containing similar domains as potential methylation candidates (p.12, first paragraph) seems less justified, given the lack of evidence supporting the requirement for the H3K4-like domain in methylation.

      This study offers valuable insights into the interactome of SET1C, suggesting potential links between the complex and a wide range of cellular processes. However, the functional implications of the Y2H interactions remain to be explored further. Additionally, the study provides intriguing information on the possible regulation of Set1 by sumoylation. The discovery of Nrm1 and Snf2 as methylation substrates could significantly expand the known targets and functions of SET1C.

      The results are supported by high-quality data.

      We thank referee 3 for their positive comments

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Restructure the manuscript into at least two papers.

      We thank the reviewer for this suggestion. In the revised manuscript, we have addressed this concern by substantially restructuring and streamlining the presentation. We consider the dataset, validation experiments, and functional observations to be closely integrated, and we believe that presenting them together provides the most coherent and impactful account of the work.

      Minor points

      There are several basic flaws in the manuscript that I feel indicate the co-authors have not proofread the manuscript sufficiently - 4 examples from early in the manuscript are listed below.

      (1) The reference for Hybridomics is (73) - obviously from an earlier version that used a different referencing system that has not been corrected.

      Thank you. This has been corrected.

      (2) Line 194 - 197. These screens have proven their power and effectiveness. In particular, they identified ...... the CTD of Rpb1 as an interactor of the N-terminal region of Set1 (Bae et al, 2020) (Figure S1). Rbp1 interaction is not identified in the screens presented here, and Figure S1 is a cartoon and not primary evidence.

      The interaction between the CTD of Rpb1 (Rpo21) and Set1 is reported in Table S2. The detailed characterization presented in Bae et al. (2020) was subsequently carried out as a direct follow-up to this screen.

      (3) Line 205-211. The highly confident interactors of the seven SET1C subunits are shown in Figure 1C-E. We found that Spp1, Shg1 and Swd2 interact alone with Set1 (Figure 1C). The minimum Set1 region for which an interaction is found for each of these 3 subunits is shown in Figure 1C. The high confidence interactors of the seven SET1C subunits are shown in Figure 1C-E. We found that Spp1, Shg1 and Swd2 display Y2H interactions with Set1 (Figure 1C). The high confidence interactors of Spp1, Shg1 and Swd2 are indicated in Figure 1D (see also Table S2).

      It is possible that Table S2 was omitted from the original submission, as it was requested during the production stage.

      (4) Line 335. We have classified all Set1 and subunit interactors according to these SET1C roles (Figure S5). However, this refers to Figure S4 - many further references to Figure S5 are also to Figure S4.

      Thank you. This has been corrected.

      Reviewer #2 (Recommendations for the authors):

      General recommendations:

      (1) Figures 1, 2, 3, and 5 and their associated main text are essentially just lists of interactors, put in graphic form and grouped to allow speculation about possible biological functions for the interactions. But almost none of the ideas are tested, so these sections take much more space than warranted. Having so much preliminary Y2H data actually distracts attention from the follow-up experiments that are shown. I would move most or all of this to the supplement, consolidating the Y2H results into fewer figures (or even just the Table).

      As mentioned earlier, the manuscript has been reorganized and Table S2 is provided.

      (2) The Snf2 interaction gets the most follow-up, so separating Figure 4 from Figures 8-10 broke the flow of that story. I would group these figures together since all are related to the Snf2 AT hook story.

      This was done accordingly.

      (3) I understand that it's impossible to validate all the possible interactions, particularly if resources are limited. However, at least for the interactions that get further attention, it could be very useful to try some AlphaFold multimer predictions. A high confidence AlphaFold score would provide a second orthogonal piece of evidence to support the Y2H results.

      We generated an AlphaFold model (Figure 2C) that recapitulates the key predictions for the Set1-Kap104 Y2H interaction.

      Comments on specific sections:

      (1) Y2H results. The text says Figure 1 shows all the high-confidence interactors. But the Set1 NTD interaction with the Rpb1 CTD is not shown here (it's in the supplement).

      In Table S2, an interaction is observed between full-length Set1 and the Rpb1-CTD (14 repeats), where Rpb1 is referred to as Rpo21.

      Figure 2 shows additional high-confidence interactors that do not appear in Figure 1, while others (like the Shg1Mog1 interaction) are shown in both Figures 1 and 2. It's confusing to scatter the data like this, which is why I recommend consolidating into a single figure or table.

      In Figure 2, the high-confidence interactors of Set1 (1–754) are highlighted in red and green (Snf2, Gbp2, and Kap104), and all are also present in Figure 1. Dbp1, identified as a high-confidence interactor of Spp1, likewise appears in Figure 1. Table S2 summarizes all of these interactions.

      (2) Line 219. How does a "high confidence" Set1-Kap104 Y2H interaction suggest the interaction is direct? Couldn't an indirect interaction also be tight and reproducible? This is an example where it would be worth seeing if AlphaFold also predicts an interaction and, if so, whether it involves the proposed NLS sequences.

      Y2H screening indicated that Kap104 binds to the N-terminal region (aa 1–754) of Set1 via its Set1 interaction domain (SID). To validate this, we used AlphaFold to model the seven-subunit Set1C complex (Set1-Bre2-Sdc1(x2) Swd1-Swd3-Spp1) with Kap104. The resulting model showed borderline confidence for the overall fold (pTM = 0.53) and low confidence in subunit positioning (ipTM = 0.5). Visualization in PyMOL confirmed Kap104 SID binding to Set1(1–754), consistent with Y2H results. The structure highlights Kap104 SID interaction with Set1’s PY-NLS at residues 40–90; the second PY-NLS is neither visible nor engaged in this model.

      (3) In the discussion of nuclear import interactors, what does it mean to say the Shg1-Mog1 interaction is "along the same line" as Set1-Kap104?

      We meant that the interaction between Shg1 and Mog1 represents another example of an interaction between a Set1C subunit and a protein involved in nuclear import. Along the same line has been deleted in the revised version.

      (4) To follow up on the Swd1-Nrm1 Y2H interaction, the paper shows that Nrm1 is methylated by Set1 in vitro (Figure 7), but it's not clear whether this has any biological significance. Without any in vivo follow-up, this figure is probably more appropriate for the Supplement.

      As noted above, Figure 7 has been removed, only panel E of Figure 7 is retained in the revised version.

      (5) Figures 6 and S8 show that Set1 is SUMOylated. Although it's not clear what this does to Set1 function or which E3 is responsible, the modification data looks convincing. The legend to Figures 6A and B says the Elutes samples are purified on nickel columns. Why are the Myc-Set1 and GB-Set1 proteins without the his-SUMO modification also binding to the nickel column? That's not happening in panels C and D. In the blots on the right for his-SUMO, is there any way to show that one of those bands is Set1? Maybe IP for MYC and then probe for the His tag?

      We thank the reviewer for this observation. His-SUMO purification using Nickel beads was used to purify HisSUMOylated proteins. Purified proteins were analyzed by Western blot using anti-MYC or anti-GAL4 antibodies to detect SET1-His-SUMO, as well as anti-His antibodies to confirm the presence of purified His-SUMOylated proteins. As mentioned by the reviewer, we detected unmodified MYC-Set1 and GAL4-Set1 in both the (-) and (+) His-SUMO eluates. This phenomenon is most likely due to the stickiness of unmodified Set1 to the beads. This is a commonly observed phenomenon in this type of biochemical assay, particularly when analyzing large proteins such as Set1 (124 kDa). This stickiness behavior has been observed in similar SUMOylation assays, e.g., for Hpr1 (88 kDa) (Bretes H, 2014. PMID: 24500206), Nup1 (114 kDa), and Nup2 (78 kDa) (Folz H, 2019. PMID: 30837289). This stickiness was not observed when using Set1 fragments (panels C and D), most likely because the fragments lost the stickiness to the beads, a characteristic belonging only to the full-length Set1. We mention this point in the legend of the new figure 5.

      (6) The Snf2 interaction gets the most follow-up. The GST pulldown validation of Set1 interaction with Snf2 AThook looks pretty good. However, the RGG repeats are necessary for the Set1 interaction with recombinant Snf2 proteins, but not for the co-IP of in vivo material. Again, AlphaFold could lend further support here.

      Thank you for this helpful suggestion. We agree that structural modelling could, in principle, provide an additional and orthogonal line of support for the Set1-Snf2 interaction. We did explore this using AlphaFold. However, both Set1 and Snf2 contain extensive intrinsically disordered regions, including the regions implicated in the interaction, and none of the models we obtained provided interpretable structural insight into the interaction interface. In particular, the predicted complexes showed low confidence in relative domain positioning, which limits their usefulness for supporting or refining the interaction model. One possible explanation is that additional components are required to stabilise a meaningful interaction in silico. While we modelled Set1 within a seven-subunit Set1C complex, Snf2 was necessarily included in isolation from its native context. Given that Snf2 functions as part of multiple, heterogeneous chromatin remodelling complexes, the absence of its physiological binding partners may prevent AlphaFold from resolving a relevant interaction interface. In light of these limitations, we have not included the AlphaFold models in the manuscript, as we felt they would not provide reliable or informative support. Instead, we have focused on the experimental evidence presented. We have clarified this point in the revised discussion to acknowledge both the potential and the current limitations of structural prediction approaches in this context.

      (7) The Snf2 methylation by Set1 is less convincing, and its biological significance is still unclear. I think it's pretty unlikely that Set1 could methylate arginine. The mass spectrometry is used for in vivo validation (mass spec), but mutating the lysines (Figure S11, S12) or Set1 deletion (Figure S14) doesn't seem to affect the signal. Could there be quantitative differences? Is there any way to quantitate the mass spec data to estimate the modified/unmodified ratio?

      We thank the reviewer for highlighting the unexpected nature of the methylation results. We agree that the observation of arginine methylation in this context is surprising, particularly given that SET domain proteins are classically associated with lysine methylation. This is why we performed multiple in vitro and in vivo experiments, and careful interpretation data that were clear led us to conclude that Set1C methylates the arginines within the ARTSTRGR motif of the AT-hook. We agree that the biological significance of this modification remains unclear. We obtained data showing that deletion of the SID domain of Snf2 impairs yeast growth on lactate, whereas this mutant grows normally on glucose and galactose, in contrast to the Snf2Δ mutant, which exhibits poor growth on both glucose and galactose. In comparison, deletion of the RG motif of Snf2 does not affect growth on lactate. These results provide insight into the interaction between Set1 and Snf2 but do not shed light on the potential importance of methylation of the RG motif. We therefore chose not to include them. In the discussion, we acknowledge the limitations of the current evidence. Our intention is to retain these findings as potentially interesting observations while ensuring that their interpretation remains appropriately cautious.

      Minor comments:

      (1) Lines 153 and 163: Stress response is listed twice, but with different references. Maybe these need to be further defined or else combined?

      We have deleted stress response line 163 and moved the references “Deshpande et al, 2022 and Nadal-Ribelles et al, 2015” line 153.

      (2) Line 193: better to say the proteins were fused to the C- or N-terminus (rather than upstream/downstream). It would be worth mentioning if there was a reason why Swd2 was fused to the N-terminus, unlike all the others.

      This has been done accordingly. In our hands, C-terminal fusions of Swd2 are not functional.

      (3) Is the scoring scheme (highest, high, good) that produces the colors in Figure 1 shown in the table? It doesn't say what the tan color (two of the Bre2 interactors) means.

      It is a mistake, Tea1 should be blue and Swi1 should not appear here. This has been fixed.

      (4) Line 206. It's not clear what it means to say that three of the subunits "interact alone with Set1". It can't mean they only interact with Set1, since other interactors are shown in Figure 1B. If it meant to say the interactions don't require other COMPASS subunits? I don't see how you can tell that from the Y2H assay. Please clarify.

      It means that these 3 subunits interact directly with Set1 without the need of another subunit, unlike of the other subunits.

      (5) Line 252. While discussing the Set1 - Snf2 interaction, the paper cites Hirschhorn et al. That paper talks about Swi-Snf, but doesn't mention Set1 anywhere. Maybe the authors meant to cite a different paper?

      We agree, this reference is not appropriated. It has been deleted.

      (6) Figure S2 panels A and C are redundant and could easily be combined.

      Figure S2 has been deleted.

      (7) Figure S4: Should the green category also include transcription? Ssl1 is a TFIIH subunit, which could be involved in either transcription initiation or NER. Sen1 and Nrd1 are transcription termination factors, although Sen1 may also function in R-loop resolution.

      We agree but it is already complicated as it is.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors identify and investigate a specific population of PVNOT neurons (oxytocin neurons of the paraventricular hypothalamus) that seem to be involved in both behavioral and autonomic thermoregulation. These cells are activated by social thermoregulatory behaviors, but can influence thermoregulation in both social and nonsocial contexts, specifically during transitions and when mice are at low core body temperature (Tb).

      Strengths:

      The manuscript has many strengths.

      This is a novel study, with a clear question that is addressed using an array of well-designed experiments employing integrative methods. Most of the figures are well-developed, and the analysis is generally rigorous and well-detailed. The authors are clearly very experienced in this field, and indeed, their scholarly introduction and discussion sections are to their credit.

      We are grateful for the reviewer’s careful reading and positive assessment, including their remarks on the clarity of the question, experimental design, and analysis.

      The link between thermoregulation and the oxytocin system is well established, as is the link between social behavior and the same broad system. However, the link between these three things is novel, if it can be well substantiated. I am not persuaded that was achieved here, but I do think this manuscript has many novel and useful offerings.

      We thank the reviewer for this thoughtful comment and for recognizing the novelty of the study. We wish to clarify the central goal of the manuscript: while social thermoregulation provided the initial influence for studying PVNOT neurons, our principal finding is that PVNOT activity during rest-to-arousal transitions is independent of social context. As stated in the manuscript, "To our surprise, these peaks were observed in both social and non-social contexts." Thus, our study demonstrates a broader role for PVNOT neurons in state-dependent thermoregulatory transitions—one that includes, but is not limited to, social contexts. We have revised the text to make this emphasis clearer throughout.

      We also added a short piece to the Discussion on this point. This is the fourth and final paragraph of the Discussion section called “State-dependent PVNOT activity during thermo-behavioral transitions.”

      The authors use a cooling floor, and only go down to 10 degrees Celsius. This is fine, but I would like to see the effects using ambient temperature also. This is not a crucial issue, as it is not necessary for the authors' interpretations, but it could improve measurement sensitivity.

      Both Reviewer 1 and Reviewer 2 raise important and related points: manipulating floor temperature provides a thermal stimulus that is distinct from manipulating whole-chamber ambient air temperature, and these modalities could engage partially different sensory pathways and circuits. (Note this response is copy-pasted to other relevant comments).

      We intentionally used floor cooling/heating because it provides a reliable, well-controlled stimulus that elicits thermoregulatory behaviors while keeping the experimental environment stable (e.g., avoiding changes in airflow/humidity that can accompany ambient cooling). To prevent conflation of these modalities, we revised the manuscript to consistently describe the manipulation as “floor temperature” (and not “ambient temperature”), and we added to the Discussion acknowledging that conductive floor temperature changes may differentially recruit peripheral thermoreceptors compared to ambient air temperature.

      While extending these experiments to whole-chamber ambient temperature changes could be informative in future work, it is not required for the central interpretations here, which focus on PVNOT activity dynamics during thermoregulatory behavior under controlled thermal conditions.

      Through an elegant behavioral experiment in Figure 1, the authors identify c-Fos patterns in the PVN that are activated by active social huddling, and they show that at the RNA level these cells overlap with oxytocin, indicating that they are oxytocin-producing cells. But this is not well discussed or indeed quantified.

      We thank the reviewer for catching this; Reviewer 2 made a similar comment. A typo in the figure legend led to this confusion. Figure 1I is in fact a quantification of the percent Oxytocin:Fos colocalized cells (not Fos:DAPI, as was written) in dorsal and ventral subregions of the PVN during active huddling and quiescent huddling. We have corrected the legend and clarified the quantification in the revised manuscript.

      The authors engage in a deep analysis of fiber photometry experiments, first by observing PVNOT neuron overall activity during a variety of different behaviors in the context of three different temperatures. Activity was associated with nesting, quiescence, and both types of huddling (when social opportunities exist). Social situations did not strongly affect this, nor did temperature conditions. These analyses indicate that the PVNOT neurons are involved in mediating specific behavioral outputs.

      With more detailed analysis, the authors investigated how PVNOT neuronal activity relates to behavioral state transition. They found that the probability of peak PVNOT neural activity strongly predicts the offset of quiescence or quiescent huddling, and therefore can be argued to signal an increase in physical activity, and as such, increased metabolism. However, the opposite pattern was observed for huddling and nesting (onset being associated with PVNOT activity), again arguing for increased thermogenesis as a function.

      What is particularly compelling is that these peaks of activity tend to occur during low Tb, again arguing for the function in increasing body warmth.

      The authors then employ an impressive setup where they image brown adipose tissue (BAT) in tandem with DeepLabCut (DLC) based animal tracking. Crucially, BAT activity and surface temperature correlated with the calcium peak of PVNOT neurons.

      Lastly, optogenetic activation of PVNOT neurons increased Tb when it was in the lower range, but not when in the higher range. It also affected BAT and rump temperature, again at low Tb. However, there is no real effect on behavior, except a trend in activity.

      The authors do some interesting tracing work at the end, though this is not functionally explored. That is not a criticism, as it does seem like this would be a whole follow-up study.

      Weaknesses:

      While novel and valuable, the manuscript feels incomplete in its current form.

      The main evidence lacking is a loss of function of the experiment. Ideally, the authors would chronically and/or acutely inhibit PVNOT neurons to establish their necessity. I know this seems obvious, but I think it is important.

      We agree with the reviewer that loss-of-function experiments are a valuable component of circuit mapping and we appreciate this suggestion. For transparency, we did attempt a chronic chemogenetic inhibition experiment using DREADDs in PVNOT neurons. However, the results were inconclusive, primarily owing to the confounding effects of pharmacological injections: both drug and vehicle-treated animals exhibited stress-induced hyperthermia following injection, and because inhibition could not be delivered while animals were asleep/resting the experimental conditions did not recapitulate the low-Tb quiescent state during which PVNOT peaks naturally occur. Given these confounds, we do not believe these data meet the standard required for inclusion in this manuscript.

      We did consider acute optogenetic inhibition. However, a clear prediction about inhibition was not as apparent in our model. Our photometry data identified a, testable hypothesis for activation: PVNOT peaks precede the exit from quiescence, therefore activation during quiescence should increase the transition, which it did (Figures 5 and 6).

      That said, new analyses of our data, driven by these reviews, have now uncovered what might be inhibition of PVNOT neurons during the approximate 60 seconds prior to entry to resting states (i.e., quiescence and quiescent huddling); see the new Fig. S3I-L. This raises the possibility that an appropriately timed photoinhibition of PVNOT neurons could facilitate the establishment of resting states. We believe that, in light of our chemogenetic and optogenetic activation experiments, for an inhibition experiment to be done appropriately would require a real-time, closed loop setup that is currently not available in our laboratory.

      We have added a caveat to the Discussion acknowledging the lack of LOF data as a limitation and have identified this as an important direction for future investigation.

      The relative lack of behavioral analysis following optogenetic activation of PVNOT neurons is puzzling. The authors must surely want to study what this intervention does to behavioral state transitions. I feel that the current level of analysis limits the overall conclusions of this study to a large extent.

      We appreciate this concern and wish to clarify two points.

      First, our decision to perform optogenetic activation in isolated (solo-housed) animals was driven by our initial finding that PVNOT activity profiles are mostly social-context independent during the transition from rest to arousal (Figures 2 and 3). By studying isolated animals, we could test the fundamental relationship between PVNOT activation and the rest-to-active transition without confounding social feedback. Additionally, we encountered technical challenges when using the SGBS thermographic model in paired contexts: the high thermal intensity at the point of contact between huddling mice created a thermal merging artifact that prevented accurate segmentation of individual body regions (BAT vs. rump).

      Second, we did examine the post-stimulation behaviors of solo-housed animals (Fig. S5B). While PVNOT activation significantly increased the probability of exiting quiescence, it did not trigger a singular, stereotyped behavioral output. Instead, it facilitated a generalized transition to an active state, within which animals engaged in various context-appropriate actions (nesting, grooming, locomotion). We note in the discussion that “Analysis of manually- annotated behaviors suggested that PVNOT stimulation did not activate a specific motor pattern output but instead resulted in combined increases in the time spent in nesting (linear mixed model estimate coefficient of ChR2+ stimulation: +38.0 sec), locomotion (+54.0 sec), and grooming (+14.5 sec), but not in eating/drinking (-0.4 sec) (Fig. S4B).”

      That photostimulation had relatively larger effects on nesting and locomotion is consistent with our model.

      Last, in the Discussion we acknowledge that future experiments should seek to disentangle the effects of PVNOT light simulation in the non-social vs social context (last paragraph of the Discussion section called “State-dependent PVNOT activity during thermo-behavioral transitions”).

      A broader criticism is that the social dimension of this manuscript seems overplayed. Naturally, oxytocin signaling can be implicated in social behavior based on a large literature. However, the focus on social thermogenesis seems like a crude integration of social behavior and thermogenesis. Given that the authors see their effects in both social and nonsocial cases of thermoregulation, I am not sure the attempts at integrating social functions and thermogenic functions of PVNOT neurons are warranted. That is, unless the authors have further experiments or analysis that can convincingly justify this link.

      We thank the reviewer for this comment. We understand the concern and wish to reframe our position. We argue that the equivalence of PVNOT signals across social and non-social contexts is itself a central finding. While the oxytocin system is widely regarded as a mediator of social bonding, and therefore a candidate mechanism underlying huddling, our data demonstrate that PVNOT neurons provide a signal for state-dependent thermoregulatory transitions that is unbiased by social context. Rather than overplaying the social dimension, we believe our study contextualizes the social function within a broader homeostatic role: PVNOT neurons facilitate transitions from rest to thermogenesis and arousal regardless of whether the resting state involves social huddling or solitary quiescence.

      While the thermoregulatory transitions are present in both contexts, we note that social context appears to modestly enhance some PVNOT downstream effects. Specifically, peak probability and frequency were slightly higher in the paired compared to solo context (Fig. 3F-I, Fig. S2D), and peaks were associated with a somewhat stronger increase in physical activity when a cagemate was present (Fig. 3B-E). Additionally, quiescent huddling (paired) bouts were associated with stronger body temperature regulation compared to solo quiescence (Fig. S3Q-V). This nuance supports that the social dimension is not overplayed but rather situated within a broader homeostatic function.

      We have revised the manuscript to ensure that this framing is consistent and clear. We emphasize that our goal was to uncover neural mechanisms underlying physiological transitions across behavioral and arousal states, using our social thermoregulation assay as a starting point (based on our previous publication). Counter to our initial hypothesis, the PVNOT signals generalized beyond the social setting.

      In addition, the analysis of virgin females and lactating mothers seems out of place in Figure 4.

      This point was echoed by Reviewers 1 and 3, and one we have taken several actions to address this. (Note this response is copy-pasted to the other reviewers).

      We agree with the reviewers that the rationale for the lactation data should be made more explicit. The primary purpose of this experiment was to validate the identity of oxytocinergic neurons of the PVN.

      Our efforts to use IHC to validate the identity of AAV-transfected cells were inconclusive, and we have now added new data to illustrate this point. We have added Fig. S4 that includes quantitative data on expression specificity. We observed significant variability in co-staining (OT+/GCaMP+) across brain slices, likely reflecting the dynamic nature of oxytocin peptide synthesis and storage, particularly with respect to processes lining the third ventricle. This finding is in accordance with other studies that are now cited in the text.

      We now emphasize that, because IHC provided variable co-localization, we employed the lactation model as an independent physiological validation of the identity of the recorded neurons.

      It is well established that PVNOT neurons undergo dramatic changes in firing dynamics and synchrony during lactation to support milk ejection (Yaguchi et al., 2023; Yukinaga et al., 2022). Conversely, AVP and CRF cell populations in the PVN do not appear to display synchronized pulsatile bursting during lactation (see response to Reviewer-2 comment-2 in ‘Recommendation for authors’ and our updated Discussion). Observing these characteristic changes in our recorded population provides high-confidence functional evidence that we are targeting oxytocin neurons. We have revised the text to clarify that Figure 4 serves primarily as a functional verification of genetic targeting.

      We also acknowledge in the Discussion the possibility that our Cre-line may capture a small percentage of nonoxytocinergic neurons, while noting that the dramatic shift in calcium dynamics during lactation (Figure 4I–L) strongly suggests the recorded population is dominated by oxytocin neurons.

      The c-Fos/oxytocin overlap needs to be quantified.

      We thank the reviewer for catching this; Reviewer 2 made a similar comment. A typo in the figure legend led to this confusion. Figure 1I is in fact a quantification of the percent Oxytocin:Fos colocalized cells (not Fos:DAPI, as was written) in dorsal and ventral subregions of the PVN during active huddling and quiescent huddling. We have corrected the legend and clarified the quantification in the revised manuscript. (Note this response is copy-pasted to other relevant comments).

      The methods section could be improved by explaining how the authors exclude animals that exhibit both types of huddling, if they occur within a 90-minute time window. This seems like it could cause significant confounds.

      We have clarified in the Methods that animals were not excluded if they exhibited both active and quiescent huddling during the recording session. Importantly, a prerequisite for inclusion in the FOS study was that animals had to be continually engaged in the target behavior for a minimum of 15 consecutive minutes from behavior onset, an established approach for behavior-driven immediate early gene mapping. The 90-minute window was then counted from that same onset for FOS IHC. Because active huddling frequently transitions directly into quiescent huddling (and vice versa), excluding such animals would have eliminated the majority of recordings. The heterogeneity of behavioral states within the FOS integration windows is precisely why we turned to fiber photometry, a technique with the temporal resolution necessary to dissociate neural signals associated with each behavioral state.

      The computer vision model is not well-explained. The authors need to be far more explicit here about how it was validated.

      We thank the reviewer for this comment and agree that the original manuscript did not sufficiently detail the validation framework. We have revised both the Methods and Results to explicitly detail how SGBS was evaluated.

      First, we now clearly describe model validation on a held-out dataset (20% of manually annotated images not used for training), reporting standard segmentation metrics (per-class IoU and Dice/F1) and directly comparing SGBS to an unmodified Mask R-CNN trained under identical conditions (same backbone initialization, dataset split, and training schedule). As shown in Fig. 5D, the skeleton-guided model converged more rapidly and achieved a lower final loss than the baseline network, demonstrating improved segmentation performance in occlusion-rich thermographic recordings.

      Second, we more explicitly describe an independent physiological validation step. SGBS-derived surface temperature trajectories were temporally aligned with simultaneously recorded implanted thermologger measurements, which were not used during model training. As shown in Fig. 5E, SGBS-derived signals strongly corresponded with core body temperature dynamics and reproduced expected thermophysiological relationships (e.g., BAT warming preceding core temperature rise). This establishes external validity beyond pixel-level segmentation metrics.

      The authors should cite and consider this preprint: https://www.biorxiv.org/content/10.1101/2024.09.17.613378v1

      We have cited this preprint (Raam et al., 2024) in the revised manuscript and integrated relevant findings into the Discussion, in the section called “Limitations and caveats”.

      Reviewer #2 (Public review):

      Summary:

      This is a very interesting study from Vandendoren and colleagues examining the role of PVN oxytocin neurons during thermoregulatory behaviors, in particular during thermoregulatory huddling. The findings are important and compelling, and have implications for the thermoregulation field as well as the social/naturalistic behavior field.

      Strengths:

      The study is very creative and tackles a challenging task to examine how natural and social behavior influences neural circuits for a homeostatic system such as thermoregulation. The authors use a combination of state-of-the-art tools (photometry, optogenetics, automated behavior tracking, thermal imaging, and core body temperature measurement), often in combination with each other, to produce a rigorous and high-dimensional dataset. Carrying out tightly temperature-controlled experiments and examining natural behavior, neural activity, and body physiology simultaneously is quite a feat. I applaud the authors for taking this on in a rigorous and detailed manner. This paper will be valuable for both the thermoregulation field as well as for researchers interested in naturalistic social behaviors. The conclusions are supported by the data.

      We appreciate the reviewer’s careful read and positive assessment of our integrated behavioral, neural, and physiological measurements and their relevance to both thermoregulation and social behavior.

      Weaknesses:

      I have a number of questions and suggestions for clarification that would help improve the interpretation of the findings.

      (1) Figure 1D-F: It would be helpful to include representative images of cFos expression in the PVN, LS, and DMH during both quiescent and solo huddling conditions, to better illustrate the reported differences.

      We have now addressed this in the revised manuscript. We had originally shown active huddle FOS expression in Fig. 1D-F and quiescent huddle in Fig. S1A-C. We have now added solo groom FOS expression to Fig. S1D-F.

      (2) Figure 1C: The data suggest a general suppression of neural activity during sleep-associated quiescent huddling, which somewhat complicates the interpretation of what specifically the active huddling cells are responding to. A more informative control might have been a comparison between huddling and a more generic form of social engagement (e.g., dyadic sniffing) to assess whether huddling-responsive neurons are broadly tuned to social stimuli. While it may not be feasible to add this experimentally at this time, a brief discussion of this limitation in the main text would be valuable.

      We thank the reviewer for this thoughtful suggestion. We agree that comparing huddling-responsive neurons with a more generic social engagement is an important consideration.

      We first note that the FOS study required animals to be continuously engaged in the target behavior for a minimum of 15 consecutive minutes, ensuring that FOS expression reflects sustained behavioral engagement rather than brief social contact. Furthermore, we believe the FOS association with active huddling in Figure 1C is likely driven by preceding bouts of quiescent huddling. Because these experiments were conducted during the light phase, active huddling bouts were almost always preceded by bouts of quiescent huddling.

      Given that FOS protein often integrates neural activity over ~60-90 minutes, the FOS signal during active huddling may reflect cumulative PVNOT activity during the quiescent to active transition, rather than active huddling by itself. This interpretation aligns with our fiber photometry data, which show that PVNOT peaks are concentrated at the offset of quiescent states and the onset of active states. Moreover, a broad-scale analysis of calcium data driven by these reviews, now shows there is a local minimum of PVNOT neurons during the transition into quiescent states and a local maximum of calcium activity during the offset of resting states and the onset of nesting and active huddling (Fig. S3I-L).

      To directly address whether PVNOT neurons are broadly tuned to social engagement or specifically associated with thermoregulatory state transitions, we examined neural activity during "Contact Initiated" (ConI) and "Contact Received" (ConR) events—brief social interactions (e.g., dyadic sniffing) that occur outside the context of huddling. These interactions, which typically last less than one second, did not trigger the large-amplitude calcium peaks observed during rest-to-arousal transitions. Specifically, there was no significant association between ConI or ConR events and PVNOT peak frequency or amplitude (Fig. S2H; Table S1; p = 0.505, p = 0.575, respectively). This reinforces our conclusion that PVNOT peaks are not a generic response to social stimuli but are specifically aligned with the coordinated autonomic and behavioral transitions required to exit a low-temperature quiescent state. We have added a clarifying paragraph to the Discussion.

      (3) Figure 2H-J vs. Figure 1: The fiber photometry data suggest increased PVN activity during quiescent huddling vs active huddling, which appears to contrast with the cFos results from Figure 1. It would be helpful for the authors to comment on possible reasons for this discrepancy-e.g., methodological differences, temporal resolution, or cell-type specificity.

      We agree that this apparent contrast deserves explicit discussion. The difference arises from the dramatically different temporal resolutions of the two techniques. Fiber photometry captures real-time neural dynamics at subsecond resolution, revealing that PVNOT neurons exhibit high-amplitude bursts primarily during the offset of quiescence (and to a lesser extent the onset of post-quiescence behaviors) (Figs. 3 and 5). Because these peaks occur while the animal is categorized as "quiescent," they appear as quiescence-associated activity in the photometry ethogram.

      Conversely, FOS integrates neural activity over ~30–90 minutes. In retrospect, and in light of our photometry data, an animal categorized as "Active Huddling" in the FOS study is one that has likely experienced PVNOT bursts and subsequently transitioned to an active state. The higher FOS signal in active animals therefore likely represents the cumulative activity of the transition itself and sustained activity in the active state.

      We have added a clarifying statement to the Discussion section, in the section called “State-dependent PVNOT activity during thermo-behavioral transitions”.

      (4) Figure 2O: A comparable linear regression for active huddling would be informative to assess whether the observed relationships extend across behavioral states.

      We agree. We have added linear regression analyses for active huddling and nesting to Fig. S2K-N including rsquared values, to complement the resting analyses in Figure 2O and 2L.

      This analysis shows that active huddling peak counts are also positively correlated with active huddle duration (but not nesting duration). The text has been updated accordingly.

      (5) Temperature manipulation: The use of floor temperature changes presents a distinct physiological and sensory experience from, for example, manipulation of ambient temperature. A discussion of how this choice may affect neural circuit engagement or interpretation of thermoregulatory responses would be beneficial.

      Both Reviewer 1 and Reviewer 2 raise important and related points: manipulating floor temperature provides a thermal stimulus that is distinct from manipulating whole-chamber ambient air temperature, and these modalities could engage partially different sensory pathways and circuits. (Note this response is copy-pasted to other relevant comments).

      We intentionally used floor cooling/heating because it provides a reliable, well-controlled stimulus that elicits thermoregulatory behaviors while keeping the experimental environment stable (e.g., avoiding changes in airflow/humidity that can accompany ambient cooling). To prevent conflation of these modalities, we revised the manuscript to consistently describe the manipulation as “floor temperature” (and not “ambient temperature”), and we added Discussion acknowledging that conductive floor temperature changes may differentially recruit peripheral thermoreceptors compared to ambient air temperature.

      While extending these experiments to whole-chamber ambient temperature changes could be informative in future work, it is not required for the central interpretations here, which focus on PVNOT activity dynamics during thermoregulatory behavior under controlled thermal conditions.

      (6) Correlations with behavior: Across the manuscript, it would be informative to see correlations between huddle duration and neural activity (e.g., cFos expression, calcium signal magnitude). Similarly, do longer huddles produce greater thermogenic effects?

      This is a great suggestion. The first point about huddle duration and neural activity echoes the Reviewer’s comment (4) above. For this point, we now show that the duration of active huddling is positively correlated with PVNOT peak count (Fig. S2K), which is similar to what we had shown for quiescence and quiescent huddling (Fig. 2K-P).

      Next, the point about huddle duration and thermogenic effects is also helpful. We have now added new analysis and panels to address this (Fig. S3M-R). We find that the duration of quiescent huddle bouts is negatively correlated with Tb (Fig. S3V). The other behaviors examined did not show correlations between duration and Tb. This finding supports our previous demonstration that quiescent huddling is an energy saving state in mice (Landen et al., 2024).

      Finally, we note that longitudinal correlations between bout length and peak counts are already reported in Fig. S3A-H.

      (7) Lactating vs. virgin mothers: The inclusion of maternal data is intriguing but feels somewhat disconnected from the central huddling-thermoregulation narrative. If these experiments are to remain, additional explanation of their rationale and how they fit into the broader story would help clarify their relevance.

      This point was echoed by Reviewers 1 and 3, and one we have taken several actions to address this.

      We agree with the reviewers that the rationale for the lactation data should be made more explicit. The primary purpose of this experiment was to validate the identity of oxytocinergic neurons of the PVN.

      Our efforts to use IHC to validate the identity of AAV-transfected cells were inconclusive, and we have now added new data to illustrate this point. We have added Fig. S4 that includes quantitative data on expression specificity. We observed significant variability in co-staining (OT+/GCaMP+) across brain slices, likely reflecting the dynamic nature of oxytocin peptide synthesis and storage, particularly with respect to processes lining the third ventricle. This finding is in accordance with other studies that are now cited in the text.

      We now emphasize that, because IHC provided variable co-localization, we employed the lactation model as an independent physiological validation of the identity of the recorded neurons.

      It is well established that PVNOT neurons undergo dramatic changes in firing dynamics and synchrony during lactation to support milk ejection (Yaguchi et al., 2023; Yukinaga et al., 2022). Conversely, AVP and CRF cell populations in the PVN do not appear to display synchronized pulsatile bursting during lactation (see response to Reviewer-2 comment-2 in ‘Recommendation for authors’ and our updated Discussion). Observing these characteristic changes in our recorded population provides high-confidence functional evidence that we are targeting oxytocin neurons. We have revised the text to clarify that Figure 4 serves primarily as a functional verification of genetic targeting.

      We also acknowledge in the Discussion the possibility that our Cre-line may capture a small percentage of non-oxytocinergic neurons, while noting that the dramatic shift in calcium dynamics during lactation (Figure 4I–L) strongly suggests the recorded population is dominated by oxytocin neurons.

      (8) Optogenetic manipulation: Have the authors tested the effect of PVN OT neuron stimulation or inhibition during huddling? Even a negative result would be of interest to the field. If these data exist (main or supplementary), I apologize for missing them. If not, the authors might consider including them or commenting briefly on any attempts or challenges in carrying out these experiments.

      We thank the reviewer for this question. We have not performed optogenetic manipulation during huddling. Our decision to perform optogenetic activation in solo-housed animals was driven by our fiber photometry finding that PVNOT activity profiles during the rest-to-arousal transition are social-context independent (Figures 2 and 3). Had the GCaMP data suggested that PVNOT peaks were specific to social huddling, optogenetic manipulation during huddling would have been the natural next experiment. However, because peaks aligned with thermoregulation broadly, rather than social behavior specifically, we designed our functional experiments to test the circuit's role in driving the autonomic and behavioral arousal transition.

      We also note that our experience with chemogenetic manipulation suggests that pharmacological approaches to study the rest-arousal transitions during huddling are not currently feasible. As described to our response to Reviewer 1, our DREADD inhibition experiments were confounded by stress-induced hyperthermia following injection, and because drug delivery could not occur while animals were asleep and resting, the experimental conditions failed to recapitulate the low-Tb quiescent state during which PVNOT peaks naturally occur. We share this experience because we believe it will be informative for others in the field considering similar approaches.

      Additionally, as described above (Reviewer 1, #5), the SGBS thermographic model encounters artifacts in paired contexts due to thermal merging between huddling mice. We have added a note in the Discussion addressing this, in the section called “Limitations and caveats”.

      Reviewer #3 (Public review):

      Summary:

      The authors aimed to elucidate the relationship between physiological state (i.e., behavioral status and thermogenic sympathetic activity) and the activity of hypothalamic paraventricular oxytocin (PVNOT) neurons in female mice. They studied this by combining automated classification of mouse behavior via video-based analysis with calcium imaging of PVNOT neuron activity. Sympathetic thermogenesis was inferred from surface temperature changes captured by infrared thermography, and the authors provided their custom analysis scripts in the manuscript. Notably, they found that a strong, pulsatile activation of PVNOT neurons was "occasionally" observed immediately before the animals transitioned from a resting to an active state. This pulsatile activity was observed in both pair-housed and individually housed animals. While PVNOT neurons are often associated with social behaviors, this finding suggests that the oxytocinergic system is also engaged during naturalistic behaviors, even in the absence of social interactions. If experiments were more convincingly performed and presented, the results would point to a broader physiological role of central oxytocin, including in the regulation of fundamental brain states and homeostatic processes, and offer a new perspective on the functional significance of central oxytocin signaling.

      Strengths:

      The oxytocinergic neural system is believed to subserve a wide range of physiological functions, and elucidating these roles requires monitoring PVNOT neuronal activity under various behavioral contexts, as well as manipulating this activity to establish causal links. In the present study, the authors show a technically sound experimental framework that integrates behavioral tracking in both individually and group-housed mice with the observation and manipulation of PVNOT neuron activity. This experimental setup represents a valuable methodological resource for researchers investigating the physiological functions of oxytocin.

      We thank the reviewer for the thoughtful review and for recognizing the value of our integrated framework for monitoring and manipulating PVNOT neuronal activity across behavioral contexts.

      Weaknesses:

      While this study successfully established a new experimental setup for simultaneous analyses of behavior and PVNOT neuronal activity, there are several concerns regarding the interpretation of the results and the robustness of the conclusions, which should be more thoroughly addressed.

      (1) The study relies on the assumption that calcium imaging and optogenetic manipulation were restricted only to PVNOT neurons. However, the specificity of AAV-mediated gene expression was not verified quantitatively. A fair number of cell bodies in the PVN expressed GCaMP8s, but not OT, indicating potential off-target expression (see Figure S2A, B). The lack of quantitative validation weakens confidence in the causal interpretation of the results.

      This point was echoed by Reviewers 1 and 3, and one we have taken several actions to address this.

      We agree with the reviewers that the rationale for the lactation data should be made more explicit. The primary purpose of this experiment was to validate the identity of oxytocinergic neurons of the PVN.

      Our efforts to use IHC to validate the identity of AAV-transfected cells were inconclusive, and we have now added new data to illustrate this point. We have added Fig. S4 that includes quantitative data on expression specificity. We observed significant variability in co-staining (OT+/GCaMP+) across brain slices, likely reflecting the dynamic nature of oxytocin peptide synthesis and storage, particularly with respect to processes lining the third ventricle. This finding is in accordance with other studies that are now cited in the text.

      We now emphasize that, because IHC provided variable co-localization, we employed the lactation model as an independent physiological validation of the identity of the recorded neurons.

      It is well established that PVNOT neurons undergo dramatic changes in firing dynamics and synchrony during lactation to support milk ejection (Yaguchi et al., 2023; Yukinaga et al., 2022). Conversely, AVP and CRF cell populations in the PVN do not appear to display synchronized pulsatile bursting during lactation (see response to Reviewer-2 comment-2 in ‘Recommendation for authors’ and our updated Discussion). Observing these characteristic changes in our recorded population provides high-confidence functional evidence that we are targeting oxytocin neurons. We have revised the text to clarify that Figure 4 serves primarily as a functional verification of genetic targeting.

      We also acknowledge in the Discussion the possibility that our Cre-line may capture a small percentage of nonoxytocinergic neurons, while noting that the dramatic shift in calcium dynamics during lactation (Figure 4I–L) strongly suggests the recorded population is dominated by oxytocin neurons.

      (Note, we have updated Figure S2A,B to more accurately reflect the extent of co-localization in this image).

      (2) The study focuses on the transition from rest to active states following pulsatile activity of PVNOT neurons. However, the physiological significance of this pulsatile activity remains unclear. According to the authors, pulsatile activity occurred with an approximately 20% probability within 100 seconds prior to the end of the resting state. This implies that, in the remaining 80% of rest-to-active transitions, pulsatile PVNOT activity did not occur, suggesting that it is not essential for initiating the transition. A comparative analysis of behavioral and thermogenic changes between transitions with and without pulsatile PVNOT activity would help to further clarify the functional relevance of this phenomenon and strengthen the authors' interpretation of the findings.

      These are excellent points, and here we address them separately.

      (1) probability of transitions.

      We agree that our wording could be misread and we have revised the text for clarity. The “~20%” value is not the fraction of rest-to-active transitions that exhibit pulsatile PVNOT activity within a 100-s window. Instead, Fig. 3F,H report an instantaneous (per-second) probability of observing a calcium peak as a function of time-to-bout offset (logistic regression). In other words, the probability of a peak increases sharply as the animal approaches rest offset (e.g., from ~2–3%/s near onset to ~14%/s for quiescence and ~25%/s for quiescent huddling near offset), indicating a strong state-dependent increase in peak likelihood rather than an all-or-none trigger.

      We further clarify in the Discussion that we do not claim PVNOT peaks are essential for initiating every transition; rather, PVNOT activity biases or enhances the probability of transition toward thermogenesis and behavioral arousal (added to section called “State-dependent PVNOT activity during thermo-behavioral transitions”).

      (2) the effect of peaks on transitions

      This is a very helpful suggestion and we agree that directly comparing transitions with vs. without pre-offset pulsatile PVNOT activity could strengthen interpretation of the functional relevance of these events. We have therefore added a new transition-aligned analysis of thermogenic dynamics at rest-to-active transitions (new Fig. 3P&S; and corresponding text in the Results and Statistics sections).

      Briefly, we extracted peri-transition body temperature (Tb) traces (−300 to +300 s) aligned to the offset of quiescence and quiescent-huddling bouts and classified each transition as Peak+ if it contained one or more calcium peaks in the 100 s preceding bout offset, and Peak− otherwise. To account for inter-individual differences in “balance point,” Tb was z-scored within mouse. We then quantified the post-offset thermogenic rise for each transition as the change in scaled temperature from a pre-offset baseline (−60 to 0 s) to the post-offset interval (0 to 300 s) and tested Peak+ vs Peak− differences using linear mixed-effects models. This revealed that Peak+ transitions exhibited significantly larger post-offset increases in scaled Tb than Peak− transitions for both quiescence offsets and quiescent-huddling offsets.

      Together, these results indicate that while pulsatile PVNOT activity is not present prior to every rest-to-active transition, when it occurs it is associated with a stronger thermogenic rise, consistent with a probabilistic modulatory role in promoting the transition rather than being strictly required to initiate it.

      We are grateful for this suggestion as this new data is very informative in the context of our model.

      (3) The study identifies a correlation between pulsatile activity of PVNOT neurons and rest-to-active transitions, and tests for a causal relationship using optogenetic stimulation. However, since PVNOT neurons are known to co-release other neurotransmitters such as glutamate, it remains unclear whether the observed effects are mediated specifically through oxytocin receptor signaling. To address this question, functional intervention experiments using oxytocin receptor antagonists or receptor knockout mice are necessary.

      We agree with the reviewer that PVNOT neurons co-release glutamate and that isolating the specific contribution of oxytocin signaling versus co-transmitted signals is an important question. However, our study was designed to identify the functional role of the PVNOT cell type during thermoregulatory state transitions, not to dissect the molecular mechanism of signaling at downstream targets. By demonstrating that the endogenous activity of this specific population aligns with the rest-arousal window and that their activation is sufficient to drive the phenotype, we provide an anatomical and functional framework for future mechanistic investigations.

      We also note that we provide anatomical evidence supporting a possible peptidergic mechanism: PVNOT neuron projections to the rostral medullary raphe (rMR), a key thermogenic control site, alongside oxytocin receptor mRNA expression in this region (Fig. S5). This anatomical link suggests a plausible pathway for oxytocinergic modulation of thermogenesis, but of course does not rule in/out glutamatergic signaling. We acknowledge this limitation in the Discussion and frame pharmacological and receptor knockout studies as important next steps.

      We address these points in the Discussion, in the section called “Limitations and caveats.”

      (4) The authors attempted to detect BAT thermogenesis and skin vasomotion using infrared thermography. This technique measures only skin hair temperatures (since the skin was not shaved), but does not measure "BAT temperature" or "vasomotor tone". As seen in Figure 5E, the temperatures of the body surface areas ("BAT", "Rump", and "Dorsal surface") mostly changed in parallel, indicating that these temperatures are strongly affected by body core temperature. Therefore, the thermographic measurements in this study did not provide convincing information on BAT thermogenesis or skin vasomotion. To avoid misleading reports, the authors need to use other techniques to directly measure temperatures, such as telemetry.

      We agree that infrared thermography measures surface radiance rather than internal tissue temperature. We have revised the manuscript to use more precise language (e.g., "surface temperature over the interscapular BAT region" rather than "BAT temperature"). However, surface measurements are not merely passive reflections of core temperature. Here we add background and explanation about our thermography data:

      Background on our approach

      Infrared thermography provides a non-invasive readout of heat emission over the interscapular region and has been validated as reporting UCP1-dependent BAT thermogenesis in mice under adrenergic stimulation (Crane et al., 2014). That said, there are known confounds (insulation/adiposity, blood flow, protocol variability) and standardized protocols are needed (Law et al., 2018). Direct telemetry or implanted thermocouples offer superior precision for measuring BAT temperature, so long as the probe is sutured to BAT itself or to Sulzer’s vein–a technical challenge because probes tend to drift over time (e.g., (Dodson et al., 2024)).

      Our BAT findings in context:

      Using SGBS, we demonstrate that the interscapular BAT region is significantly warmer than the adjacent rump surface (Fig. 5C). If surface temperature were purely a reflection of uniform core temperature, this consistent regional hotspot would not be observed.

      Our cross-correlation analysis from the photometry (Fig. 5E) shows the rise in BAT surface temperature precedes changes in other body regions by approximately 90 seconds, suggesting that BAT acts as a primary heat source during rest-to-arousal transitions rather than passively following core temperature. This finding is consistent with another study, using telemetric probes placed in BAT, finding that episodic onset of BAT temperature started to increase 3 minutes before body temperature (Ootsuka et al., 2009).

      Based on this Reviewer’s comment here and the subsequent one (5), we have now added a new analysis of the temporal patterning of arousal and thermogenesis in the optogenetic cohort of animals; see below for details.

      Vasomotor tone

      We agree that infrared thermography does not directly measure vasomotor tone. We have revised the text to remove language implying that our measurements directly quantify vasomotor tone, vasodilation or vasoconstriction.

      We note that the established approach for non-invasive assessment of vasomotion uses glabrous skin of the tail and ears (Garami et al., 2011; Meyer et al., 2017; Škop et al., 2020). Rump surface temperature measured over hairy, non-glabrous skin correlates more closely with core body temperature than with cutaneous vasomotor tone (Meyer et al., 2017; Škop et al., 2020) and is used in the literature as a reference point for calculating BAT thermogenesis.

      In our data, rump surface temperature decreased following PVNOT calcium peaks while BAT and dorsal surface temperatures increased (Fig. 5L-M). This pattern is consistent with sympathetically-driven thermogenesis in which peripheral heat loss is reduced while BAT drives core temperature upwards. We now acknowledge that our rump measurements do not isolate vasomotor contributions. We have revised the manuscript accordingly, replacing references to rump vasoconstriction with language describing the observed thermal pattern while avoiding attribution to a specific thermoeffector mechanism.

      Finally, we note that telemetry would strengthen deep-body temperature interpretation, but telemetry does not itself quantify vasomotor tone; the same distal heat-loss readouts described above would be required regardless of core Tb methodology.

      In sum, infrared thermography enables non-invasive, simultaneous tracking of multiple thermal features in freely moving, undisturbed animals—a requirement for studying the naturalistic state transitions central to this study. We have added a section to the Discussion acknowledging the limitations of surface infrared thermography.

      (5) Photostimulation of PVNOT neurons increased Tb after 400 sec (6.6 min) (Figure 5). This latency is too long to conclude that the neuronal stimulation elicited BAT thermogenesis. A more reasonable explanation is that the increase in Tb was caused by the induction of physical activity (Figure S4C), which slowly generates heat and contributes to the elevation of Tb. However, this view contradicts the authors' claim. To address this concern, the authors should directly measure BAT thermogenesis and compare it with the rate of Tb elevation. If BAT thermogenesis occurs, the rate at which the BAT temperature increases must exceed the rate at which Tb rises.

      We thank the reviewer for this thoughtful critique. With this response we first provide additional context about the timeline of temperature increases, and second add a new analysis addressing the relative contributions of activity and BAT-surface to Tb changes.

      (1) Additional context on the temporal progression

      First, the observed timescale does not, per se, rule out a contribution of BAT thermogenesis. While the kinetics of BAT activation and associated Tb increases can operate on a fast timescale in anesthetized animals, in vivo activation of BAT thermogenesis pathways can take several minutes to yield a statistically detectable difference. For example, activation of DMH→rMR glutamatergic signaling, a canonical thermogenic command pathway, takes several minutes to produce a significant increase in both Tb and BAT using telemetric temperature probes (Kataoka et al., 2014).

      This timescale could also be consistent with peptidergic neuromodulation by PVNOT neurons, which are more likely to be modulators (and not drivers) of the canonical thermogenic pathway. Oxytocin is known to act via volume transmission and metabotropic receptor signaling, which operate on slower timescales than ionotropic neurotransmission (Ludwig and Leng, 2006). Downstream recruitment of sympathetic outflow and BAT thermogenesis is likewise a multistep autonomic process, not an immediate synaptic event.

      Next, the thermal dynamics reported in Figure 5 and Figure S4 are not consistent with activity-induced heat production alone. Specifically:

      - Thermal increases were spatially localized to interscapular/dorsal regions corresponding to BAT depots before generalized surface warming.

      - Importantly, photostimulation-induced warming was observed even during behavioral states characterized by low baseline activity, suggesting that thermogenic activation was not simply a byproduct of movement.

      While we did not directly measure BAT sympathetic nerve activity, our surface thermography approach was designed specifically to resolve regional temperature dynamics over the interscapular BAT area. The spatial specificity and temporal profile of the warming are consistent with BAT thermogenesis rather than uniform musclegenerated heat.

      We acknowledge that direct measurement of BAT sympathetic activity or oxygen consumption would provide additional mechanistic resolution. However, given (i) the known role of PVN oxytocin neurons in autonomic regulation, (ii) the spatially localized dorsal temperature increase, and (iii) the temporal dissociation between stimulation onset and gradual systemic Tb rise, we conclude that BAT thermogenesis remains the most parsimonious explanation.

      We have revised the Discussion to more explicitly acknowledge these temporal dynamics by clarifying that photostimulation likely follows the timescales of peptidergic neuromodulation.

      (2) New analysis

      We have added a new analysis to address the relationship between Tb and BAT-surface temperature and locomotion the optogenetic cohort. In short, we show that across all mice changes in BAT typically precede changes in Tb, and that the effect of optogenetic stimulation on core Tb can’t be explained by physical activity (nor can it be explained by BAT-surface temperature).

      First, cross-correlation of derivatives suggested BAT surface temperature changes typically precede changes in dTb/dt across mice, whereas physical activity changes did not consistently precede dTb/dt. This result, now shown in Fig. S5G, is consistent with our cross-correlation analysis of the fiber-photometry cohort.

      Next, we used a lagged regression analysis to test whether photostimulation-evoked increases in core temperature are fully mediated by physical activity. Specifically, we modeled the derivative of core Tb (dTb/dt) using an impulseresponse representation of photostimulation, while controlling for distributed lags (0–120 s) of physical activity and BAT surface temperature derivative, with random effects for mouse and trial. Photostimulation remained a significant predictor of dTb/dt while controlling for activity and BAT-surface (likelihood ratio test, χ<sup>2</sup>=7.66, p=0.0056), indicating that the relationship between stimulation and Tb is not fully explained by activity.

      Recommendations for the authors:

      Editors note:

      We suggest including key statistical support for the claims in the main text (e.g., results or figure legends).

      We have added statistical support for key claims in the main text results. We have also added references to Table S1 where appropriate (e.g., where there is a long list of statistical results); we hope this aids the readability of the report.

      Reviewer #1 (Recommendations for the authors):

      See above - the authors should decide what to prioritize, but I only mention significant concerns above. The manuscript could be improved to 'Convincing' or even 'Compelling' with sufficient effort.

      Thanks for the careful reading of the manuscript. We’ve addressed many of these points, and feel the manuscript has been strengthened as a result.

      There were also some text errors here and there.

      Several text errors were identified and fixed. Thank you.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1I: The quantification shown here is a bit unclear from the figure and legend - are the authors reporting the percentage of cFos+ cells within the OXT+ population, or within the general DAPI+ population? If the latter, including a co-localization analysis to estimate the proportion of OXT+ cells activated would strengthen the interpretation.

      We thank the reviewer for catching this; Reviewer 1 made a similar comment. A typo in the figure legend led to this confusion. Figure 1I is in fact a quantification of the percent Oxytocin:Fos colocalized cells (not Fos:DAPI, as was written) in dorsal and ventral subregions of the PVN during active huddling and quiescent huddling. We have corrected the legend and clarified the quantification in the revised manuscript. (Note this response is copy-pasted to other relevant comments).

      (2) PVN cell types: It would be useful to briefly discuss the potential involvement of other PVN populations (e.g., CRF, AVP neurons) in huddling, given their known roles in social behavior, stress, and thermoregulation.

      Thank you for the insightful comment. We address these points in two parts.

      (1) PVN cell types and huddling

      Regarding the specific connection between these cell types and huddling: to our knowledge, no study has directly tested the effect of PVN CRF or PVN AVP neuron manipulation on huddling behavior. The most relevant data come from Bendesky et al. (Bendesky et al., 2017), who found that intracerebroventricular administration of AVP in Peromyscus inhibited nest building but had no effect on huddling, licking, or pup retrieval (though this pharmacological approach does not isolate PVN AVP neurons specifically). Their chemogenetic manipulation of PVN AVP neurons in Mus musculus confirmed the nest-building effect but did not assess huddling. For CRF, the available evidence suggests an opposing role to OT in social care contexts: chemogenetic activation of PVN CRF neurons impairs maternal behavior in postpartum mice (Melón et al., 2018), and intracerebroventricular CRF administration suppresses maternal care and can induce pup-killing in virgin rats (Pedersen et al., 1991).

      That said, PVN AVP neurons do promote wakefulness via lateral hypothalamic orexin neurons (Islam et al., 2022) and a recent preprint has implicated PVN AVP neurons in temperature-dependent maternal thermoregulatory behaviors, including co-nesting and shepherding, via projections to the central amygdala (Adahman et al., 2025). Notably, while that study focused on AVP neurons, their c-Fos data also revealed significant temperature-dependent modulation of PVNOT neurons (Fig. 3B), with suppressed activity at thermoneutrality relative to cooler conditions, a pattern suggesting that OT neurons are active under conditions where thermoregulatory effort is required. This data is consistent with our findings on PVNOT neuron involvement in rest-to-arousal transitions driven by thermoregulatory need.

      Additionally, Inada et al. (Inada et al., 2025) used an elegant series of viral-genetic experiments to demonstrate that PVN AVP neurons facilitate paternal caregiving behaviors via AVP to oxytocin receptor crosstalk in the preoptic area. Critically, their fiber photometry and circuit mapping data showed that chemogenetic activation of PVN AVP neurons did not recruit PVN OT neurons (Fig. 4), indicating that these populations operate independently in this context. We believe this finding is consistent with our interpretation that the thermoregulatory signals we observe reflect a cell-type specific property of PVNOT neurons. Future work examining how PVNOT, AVP, and CRF population interact during thermoregulatory state transitions would be valuable.

      (2) PVN cell types and stress and thermoregulation

      PVN CRF and AVP neurons have established roles in stress responses and social behavior, and future studies examining their involvement in huddling would be valuable. However, their direct roles in thermoregulation are limited. PVN CRF neurons are primarily stress-axis regulators whose thermoregulatory influence is mediated indirectly through downstream targets such as the DMH (reviewed in (Morrison and Nakamura, 2019)). AVP's thermoregulatory role is principally as an endogenous antipyretic acting via preoptic area neurons (Tabarean, 2021), rather than through PVN magnocellular AVP neurons.

      Importantly, the synchronized pulsatile bursting pattern that is characteristic of OT neurons during lactation (which serves as a key validation benchmark for our PVNOT calcium peaks), appears to be specific to OT neurons and does not generalize to other PVN populations. One study (Popescu et al., 2019) directly demonstrated that lactation-induced IPSC burst upregulation occurs selectively in OT magnocellular neurons, with no change in VP neurons within the same nucleus. VP neurons do exhibit phasic bursting, but these patterns are asynchronous, of longer duration, and serve antidiuretic rather than neuroendocrine-pulsatile functions (De Mota et al., 2004; Poulain et al., 1977; Wakerley et al., 1978). To our knowledge, no studies have reported synchronized burst activity in PVN CRF neurons during lactation or at rest. We have added a brief discussion of these points to the manuscript.

      (3) Figure 2B: Several behavioral abbreviations (e.g., LMA) are not intuitive and are missing from the legend. Spelling them out or including schematic illustrations would improve clarity.

      We have expanded the figure legends to define all behavioral abbreviations: LMA (Locomotor Activity), EaDr (Eating or Drinking), Groom (Grooming), Nest (Nesting or Nest Building), Quies (Quiescence), Sta (Stationary), ConI (Contact Initiated), ConR (Contact Received), AHud (Active Huddle), QHud (Quiescent Huddle).

      Reviewer #3 (Recommendations for the authors):

      (1) Figures 1D-F and S1A-C: The current magnification is insufficient to clearly resolve the distribution of FOS signals. FOS fluorescence is generally expected to be localized within cell nuclei. However, particularly in Figure 1F, the signals exhibit punctate or fibrous staining in addition to nuclear localization.

      This raises concerns about the quality of the tissue staining and the reliability of subsequent analyses. Including higher-magnification images would strengthen the credibility of the data presented.

      Thanks for the careful observation. We used a well-validated FOS protocol (see Methods; c-Fos (9F6) Rabbit mAb, Cell Signaling, 14609, 1:1000 dilution in block solution).

      To address this issue, in Figure 1 we have included better images of the regions of interest (DMH, LS, and PVN). We also show an inset with DAPI and the FOS IHC. These inset images show that the FOS signal does co-localize with nuclei.

      The reviewer notes that there is a fibrous staining in the PVN. We too noted this type of staining, due to clusters of bright dots in the PVN but not in other regions. This pattern was reproducible across several histological experiments. Fortunately, these bright dots were easily removed in our image processing routine using a selective median filter (pixel radius < 2.0 and and pixel intensity > 50).

      (2) Figures 2A, 4C, and 6A: As mentioned in the Public Review, the specificity of AAV-mediated gene expression is critical for the strength of the conclusions. Quantitative data demonstrating the expression specificity should be included.

      This point was echoed by Reviewers 1 and 3, and one we have taken several actions to address this. (Note this response is copy-pasted to the other reviewers).

      We agree with the reviewers that the rationale for the lactation data should be made more explicit. The primary purpose of this experiment was to validate the identity of oxytocinergic neurons of the PVN.

      Our efforts to use IHC to validate the identity of AAV-transfected cells were inconclusive, and we have now added new data to illustrate this point. We have added Fig. S4 that includes quantitative data on expression specificity. We observed significant variability in co-staining (OT+/GCaMP+) across brain slices, likely reflecting the dynamic nature of oxytocin peptide synthesis and storage, particularly with respect to processes lining the third ventricle. This finding is in accordance with other studies that are now cited in the text.

      We now emphasize that, because IHC provided variable co-localization, we employed the lactation model as an independent physiological validation of the identity of the recorded neurons.

      It is well established that PVNOT neurons undergo dramatic changes in firing dynamics and synchrony during lactation to support milk ejection (Yaguchi et al., 2023; Yukinaga et al., 2022). Conversely, AVP and CRF cell populations in the PVN do not appear to display synchronized pulsatile bursting during lactation (see response to Reviewer-2 comment-2 in ‘Recommendation for authors’ and our updated Discussion). Observing these characteristic changes in our recorded population provides high-confidence functional evidence that we are targeting oxytocin neurons. We have revised the text to clarify that Figure 4 serves primarily as a functional verification of genetic targeting.

      We also acknowledge in the Discussion the possibility that our Cre-line may capture a small percentage of non-oxytocinergic neurons, while noting that the dramatic shift in calcium dynamics during lactation (Figure 4I–L) strongly suggests the recorded population is dominated by oxytocin neurons.

      (3) Figure 2D: The authors should show an expanded view of a representative "PVNOT peak" from the spikes presented.

      We have added a representative peak to Fig. 2D.

      (4) Figure 2E-J: All the abbreviations of the behavioral states must be defined in the figure or legend.

      We added these abbreviations to the legend, and a text box reading “See legend for abbreviations” to the schematic.

      (5) Figure 2F, G, I, and J: The units on the y-axis should be indicated to facilitate interpretation.

      We have added these units. Thanks.

      (6) Figure 3A: Three large PVNOT peaks occurred between 01:30 and 02:00. However, these peaks did not cause an obvious transition in behavioral states or an increase in Tb within several minutes. Therefore, statements such as "PVNOT neurons predict transitions towards thermogenesis and behavioral arousal" in the text and subheading (pages 7 and 9) are questionable.

      We thank the reviewer for this careful observation. The three peaks between 01:30 and 02:00 that do not immediately lead to a behavioral transition illustrate a key aspect of our findings: the relationship between PVNOT activity and state transitions is probabilistic and state-dependent, not deterministic. Our logistic regression analysis (Fig. 3F, H, J, L) demonstrates that peaks increase the probability of a transition (up to ~20% per second) rather than acting as an obligatory "on switch." While individual variability exists in any single trace, the group-level analysis reveals a statistically significant increase in physical activity following PVNOT peaks (Fig. 3B–E).

      We therefore use ‘predict’ in a probabilistic sense: PVNOT peaks increase the conditional probability of impending state transitions in a manner that depends on behavioral context, rather than acting as an obligate trigger in every instance. We have taken care to not claim that PVNOT neurons are a necessary causal factor for transitions towards thermogenesis and arousal.

      We have updated the figure legend to clarify that Figure 3A shows an individual example trace, and revised the subheading on page 7 to more accurately reflect the probabilistic nature of this relationship: "PVNOT neurons predict increased likelihood of transitions towards thermogenesis and behavioral arousal in social and non-social contexts".

      We qualified the word “predicts” with “probabilistically” in the third paragraph of this section.

      Finally, this comment is related to the Reviewer’s comment-2 in the Public Reviews. To address that comment, we added a new analysis (now Fig. 3P&S) which shows that the presence of a peak in a bout of rest increases the thermogenic trajectory compared to bouts without a peak.

      (7) Figure 3F and H: If PVNOT peaks contribute to the initiation of transitions into the active state, the probability of peak occurrence should reach its maximum prior to the quiescence offset. However, the figures do not present the probability trajectory after the offset, which limits the ability to evaluate the authors' interpretation. Reanalysis extending to 150 seconds post-offset would be needed to clarify this issue.

      Thank you for this suggestion. We agree that examining PVNOT dynamics around the period following quiescence (and quiescent huddling) offset can further inform how PVNOT activity relates to rest-to-active transitions, and this has led to new insights within the manuscript.

      For background, in the original analysis (Fig. 3F,H), we used logistic regression to quantify how peak probability differs between bout onset versus near bout offset. We focused these analyses on the the timeframe of the bouts themselves (plus a small margin) because, in freely behaving animals, the pre-onset and post-offset period is heterogeneously composed of multiple potential subsequent behaviors (e.g., brief re-entry into quiescence, nesting, active huddling, locomotion, etc), which would confound a single post-offset probability trajectory (unless offsets are stratified by the identity of the subsequent behavioral state–beyond the scope of this paper).

      To address this concern, we now expand our peri-event baseline calcium analysis to include three minutes before and three minutes after both bout onset and bout offset for all four behaviors (new Fig. S3I–L). These extended traces show that for the two resting states (quiescence and quiescent huddling), baseline PVNOT calcium reaches a minimum near bout onset and a maximum near bout offset, whereas for the two active states (nesting and active huddling) baseline calcium shows the opposite pattern (maximum near onset, minimum near offset). Thus, the expanded post-offset analyses provide a more complete view of PVNOT calcium dynamics across the requested post-offset epoch and further support the conclusion that PVNOT activity is aligned with (and elevated around) behavioral transitions in a state-dependent manner. We have updated the Results text accordingly and now explicitly reference these new extended peri-event baseline analyses.

      (8) Figures 4H and I: Figure 4H shows that the waveform in the PPD2-7 group has a narrower FWHM than the Virgin group, which is the opposite of the group data in Figure 4I. Presenting scaled waveforms in parallel would allow for a clearer comparison across groups.

      Thank you for pointing out the inconsistency between the representative waveform in Fig. 4H and the group summary in Fig. 4I. You were correct: the PPD2–7 and Virgin waveforms in Fig. 4H had been mislabeled. We have corrected the labeling. (We verified that the underlying data are correct).

      As suggested, to enable visual comparison of waveform width across groups independent of amplitude differences, we derived peak-normalized average waveforms using a normalization procedure for every peak prior to averaging. Specifically, for each peak we (1) baseline-subtracted the trace by subtracting the mean fluorescence in a pre-peak baseline window, and then (2) divided the baseline-subtracted waveform by its own maximum value to scale the event amplitude to 1. We then computed the mean ± SEM of these peak-normalized waveforms across events within each group.

      We believe these changes resolve the discrepancy and improve the clarity of the figure, consistent with your suggestion.

      (9) Figure 5: In studies of thermoregulatory processes, tail blood flow or temperature is commonly used as an indicator of vasomotor responses. Is it feasible to track tail temperature using the SGBS system? If not, it may be helpful to acknowledge this as a technical limitation.

      We agree that tail temperature is a commonly used indicator of vasomotor responses. While SGBS could in principle be trained to segment the tail, the current model was optimized for dorsal body regions viewed from an overhead perspective. Reliable tail tracking presents substantial technical challenges in our configuration of homecage recordings. The tail’s thin geometry and rapid, multidirectional movement frequently result in partial or complete occlusion (e.g., beneath bedding or the animal’s body). In addition, during vasoconstriction the tail temperature approaches ambient floor temperature, reducing thermal contrast and making segmentation unreliable with the current thermal resolution limited by our camera. We have acknowledged this as a technical limitation in the Discussion, in the section called “Thermal tracking and validation of PVNOT recording specificity”.

      (10) Figure S5: Please describe the reason and histological background for the intravenous injection of FluoroGold.

      Intravenous injection of FluoroGold (FG) was used to histologically differentiate between magnocellular and parvicellular oxytocin neurons in the PVN. Because the posterior pituitary is located outside the blood-brain barrier,

      i.v. FG is selectively taken up by terminals of magnocellular neurons and retrogradely transported to their cell bodies. This allows us to infer the neuroanatomical identity (magno- vs. parvicellular) of the PVNOT neurons of interest. We have updated the Methods with a detailed description of the FG injection protocol as follows:

      “To distinguish between peripheral-projecting magnocellular and central-projecting parvicellular neurons, mice received 15 uL intravenous injection of 4% Fluoro-Gold (Fluorochrome) diluted in 100 uL of sterile saline. Prior to injection, mice were given an analgesic dose of carprofen (20 mg/kg, s.c.). Mice were briefly restrained using a modified 50 mL conical tube, in which holes were drilled to allow for proper air flow and respiration. Mouse tails were interposed between two heating pads to enhance visibility of the tail vein. Tails were wiped down with 70% ethanol and FG was administered via either right or left lateral tail vein using a 0.5 mL 28G syringe. Mice were sacrificed 24- 48 hours post-FG administration.”

      The following are minor points.

      (11) Figure 2E-G, Figure 3F,G, Figure S2G,I, Figure S3A: "quiesence" > "quiescence". This typo may appear elsewhere in the manuscript as well.

      Thanks. These edits have been made.

      (12) Page 7, line 14: Peaks were NOT significantly increased at 29{degree sign}C in Figure 2N.

      Thanks for the very careful read. By way of explanation: this difference had been significant in an earlier draft; however, when we added more replicates, the difference went away. We have corrected this sentence.

      (13) There are mislabeled figure numbers in the main text. The authors should carefully check this throughout the manuscript.

      We found mislabeled figure numbers and have corrected them.

      (14) Page 13, lines 1- 2: To make the description clearer, it might be better to rephrase the part that says, "some blue light stimulations occurred." As it stands, it could give the impression that the stimulations happened spontaneously. Using a phrase like "were delivered" would more clearly indicate that these were intentional, experimenter-controlled events.

      Agreed. Thanks. The edit has been made.

      Additional comments:

      The oxytocin system is thought to support a wide range of physiological and behavioral functions, and the circuits involving oxytocin neurons are likely to be regulated in complex and dynamic ways. As oxytocin research continues to expand, the growing body of evidence not only deepens our understanding but also highlights the system's complexity. In this context, the development of an approach that enables the observation of oxytocinergic neuron activity in parallel with naturalistic behavior represents a promising methodological contribution. It is likely that similar experimental frameworks will become increasingly common in future studies. While reading this manuscript, as a reader rather than a reviewer, I was wondering how OXT neurons detect or define the "rest balance-point," and how they might contribute to shifting the brain toward an "awake balance-point" (Figure 7). Given that eLife allows authors to include an "Ideas and Speculation" subsection within the Discussion, it would be appreciated - though not essential - if the authors could briefly share their perspective on this point. I believe such mechanistic insight would make the manuscript more intellectually stimulating.

      This is a great suggestion. We have added a new “Ideas and Speculation” section of the Discussion.

      References

      Adahman Z, Ooyama R, Gashi DB, Medik ZZ, Hollosi HK, Sahoo B, Akowuah ND, Riceberg JS, Carcea I. 2025. Hypothalamic Vasopressin Neurons Enable Maternal Thermoregulatory Behaviors. DOI: https://doi.org/10.1101/2025.01.23.634569

      Bendesky A, Kwon Y-M, Lassance J-M, Lewarch CL, Yao S, Peterson BK, He MX, Dulac C, Hoekstra HE. 2017. The genetic basis of parental care evolution in monogamous mice. Nature 544:434–439. DOI: https://doi.org/10.1038/nature22074

      Crane JD, Mottillo EP, Farncombe TH, Morrison KM, Steinberg GR. 2014. A standardized infrared imaging technique that specifically detects UCP1-mediated thermogenesis in vivo. Molecular Metabolism 3:490– 494. DOI: https://doi.org/10.1016/j.molmet.2014.04.007

      De Mota N, Reaux-Le Goazigo A, El Messari S, Chartrel N, Roesch D, Dujardin C, Kordon C, Vaudry H, Moos F, Llorens-Cortes C. 2004. Apelin, a potent diuretic neuropeptide counteracting vasopressin actions through inhibition of vasopressin neuron activity and vasopressin release. Proceedings of the National Academy of Sciences 101:10464–10469. DOI: https://doi.org/10.1073/pnas.0403518101

      Dodson AD, Herbertson AJ, Honeycutt MK, Vered R, Slattery JD, Goldberg M, Tsui E, Wolden-Hanson T, Graham JL, Wietecha TA, O’Brien KD, Havel PJ, Sikkema CL, Peskind ER, Mundinger TO, Taborsky GJ, Blevins JE. 2024. Sympathetic Innervation of Interscapular Brown Adipose Tissue Is Not a Predominant Mediator of Oxytocin-Induced Brown Adipose Tissue Thermogenesis in Female High Fat Diet-Fed Rats. Current Issues in Molecular Biology 46:11394–11424. DOI: https://doi.org/10.3390/cimb46100679

      Garami A, Pakai E, Oliveira DL, Steiner AA, Wanner SP, Almeida MC, Lesnikov VA, Gavva NR, Romanovsky AA. 2011. Thermoregulatory Phenotype of the Trpv1 Knockout Mouse: Thermoeffector Dysbalance with Hyperkinesis. The Journal of Neuroscience 31:1721–1733. DOI: https://doi.org/10.1523/JNEUROSCI.4671-10.2011

      Inada K, Hagihara M, Yaguchi K, Irie S, Inoue YU, Inoue T, Miyamichi K. 2025. Vasopressin-to-oxytocin receptor crosstalk in the preoptic area underlying parental behaviors in male mice. Nature Communications 16:10844. DOI: https://doi.org/10.1038/s41467-025-66908-0

      Islam MT, Rumpf F, Tsuno Y, Kodani S, Sakurai T, Matsui A, Maejima T, Mieda M. 2022. Vasopressin neurons in the paraventricular hypothalamus promote wakefulness via lateral hypothalamic orexin neurons. Current Biology 32:3871-3885.e4. DOI: https://doi.org/10.1016/j.cub.2022.07.020

      Kataoka N, Hioki H, Kaneko T, Nakamura K. 2014. Psychological Stress Activates a Dorsomedial HypothalamusMedullary Raphe Circuit Driving Brown Adipose Tissue Thermogenesis and Hyperthermia. Cell Metabolism 20:346–358. DOI: https://doi.org/10.1016/j.cmet.2014.05.018

      Landen JG, Vandendoren M, Killmer S, Bedford NL, Nelson AC. 2024. Huddling substates in mice facilitate dynamic changes in body temperature and are modulated by Shank3b and Trpm8 mutation. Communications Biology 7:1186. DOI: https://doi.org/10.1038/s42003-024-06781-7

      Law J, Chalmers J, Morris DE, Robinson L, Budge H, Symonds ME. 2018. The use of infrared thermography in the measurement and characterization of brown adipose tissue activation. Temperature 5:147–161. DOI: https://doi.org/10.1080/23328940.2017.1397085

      Ludwig M, Leng G. 2006. Dendritic peptide release and peptide-dependent behaviours. Nature Reviews Neuroscience 7:126–136. DOI: https://doi.org/10.1038/nrn1845

      Melón LC, Hooper A, Yang X, Moss SJ, Maguire J. 2018. Inability to suppress the stress-induced activation of the HPA axis during the peripartum period engenders deficits in postpartum behaviors in mice. Psychoneuroendocrinology 90:182–193. DOI: https://doi.org/10.1016/j.psyneuen.2017.12.003

      Meyer CW, Ootsuka Y, Romanovsky AA. 2017. Body Temperature Measurements for Metabolic Phenotyping in Mice. Frontiers in Physiology 8:520. DOI: https://doi.org/10.3389/fphys.2017.00520

      Morrison SF, Nakamura K. 2019. Central Mechanisms for Thermoregulation. Annual Review of Physiology 81:285– 308. DOI: https://doi.org/10.1146/annurev-physiol-020518-114546

      Ootsuka Y, de Menezes RC, Zaretsky DV, Alimoradian A, Hunt J, Stefanidis A, Oldfield BJ, Blessing WW. 2009. Brown adipose tissue thermogenesis heats brain and body as part of the brain-coordinated ultradian basic rest-activity cycle. Neuroscience 164:849–861. DOI: https://doi.org/10.1016/j.neuroscience.2009.08.013

      Pedersen CA, Caldwell JD, McGuire M, Evans DL. 1991. Corticotronpin-releasing hormone inhibits maternal behavior and induces pup-killing. Life Sciences 48:1537–1546. DOI: https://doi.org/10.1016/00243205(91)90278-J

      Popescu IR, Buraei Z, Haam J, Weng F, Tasker JG. 2019. Lactation induces increased IPSC bursting in oxytocinergic neurons. Physiological Reports 7:e14047. DOI: https://doi.org/10.14814/phy2.14047

      Poulain DA, Wakerley JB, Dyball REJ. 1977. Electrophysiological differentiation of oxytocin-and vasopressinsecreting neurones. Proceedings of the Royal Society of London. Series B. Biological Sciences 196:367– 384. DOI: https://doi.org/10.1098/rspb.1977.0046

      Škop V, Guo J, Liu N, Xiao C, Hall KD, Gavrilova O, Reitman ML. 2020. Mouse Thermoregulation: Introducing the Concept of the Thermoneutral Point. Cell Reports 31:107501. DOI: https://doi.org/10.1016/j.celrep.2020.03.065

      Tabarean IV. 2021. Activation of Preoptic Arginine Vasopressin Neurons Induces Hyperthermia in Male Mice. Endocrinology 162:bqaa217. DOI: https://doi.org/10.1210/endocr/bqaa217

      Wakerley JB, Poulain DA, Brown D. 1978. Comparison of firing patterns in oxytocin- and vasopressin-releasing neurones during progressive dehydration. Brain Research 148:425–440. DOI: https://doi.org/10.1016/00068993(78)90730-8

      Yaguchi K, Hagihara M, Konno A, Hirai H, Yukinaga H, Miyamichi K. 2023. Dynamic modulation of pulsatile activities of oxytocin neurons in lactating wild-type mice. PLOS ONE 18:e0285589. DOI: https://doi.org/10.1371/journal.pone.0285589

      Yukinaga H, Hagihara M, Tsujimoto K, Chiang H-L, Kato S, Kobayashi K, Miyamichi K. 2022. Recording and manipulation of the maternal oxytocin neural activities in mice. Current Biology 32:3821-3829.e6. DOI: https://doi.org/10.1016/j.cub.2022.06.083

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the submitted manuscript, Steinbach et al describe the formation of a detergent-resistant "cloud" around the Legionella-containing vacuole (LCV) that functions as a protective barrier. The authors show that formation of the "cloud" barrier is contingent upon the phosphoribosyl-ubiquitination activity of the SidE/SdeABC effector family, and is temporally regulated, with the assembly and subsequent disassembly of the "cloud" coinciding with replication and vacuolar expansion. The authors postulate a model of "cloud" barrier formation that relies upon a wave of initial ubiquitination by the SidC effector family, after which the SidE/SdeABC family expands the ubiquitination and forms cross-links that render the ubiquitin cloud resistant to harsh detergents. Additionally, Steinbach et al. also demonstrate that Rab5 is recruited to the LCV and remains associated for a considerable period.

      Strengths:

      This manuscript is very well written, with clear justification provided for experiments that make it very easy to follow along with the experimental logic. The figures have clearly been designed with much thought and are easy to interpret. Steinbach et al have also done a commendable job of addressing the previous reviewers' comments, even though some may suggest that some of these comments could be viewed as slightly unreasonable. This work would be of interest to both the Legionella and ubiquitin fields. Legionella researchers would potentially be interested to explore the proposed barrier model as the function for the ubiquitin "cloud," whereas ubiquitin researchers may be interested in exploring the mechanisms underlying SidE's crosslinking ability.

      Weaknesses:

      While the work is important and describes the physical nature of the ubiquitin cloud on the Legionella vacuole, it is somewhat descriptive in nature and does not dig deeply into what purpose this cloud serves. This is a complicated topic that will certainly stimulate additional research in this area.

      We thank Reviewer #1 for positive assessment of our work. We acknowledge that our study leaves many mechanistic questions open and, as suggested by Reviewer #1, we hope that our data is thought-provoking for researchers studying Legionella, ubiquitin signaling, or both. We are greatly looking forward to the results of future experimentation on the role of the “cloud” surrounding the bacterial vacuole.

      Reviewer #2 (Public review):

      Summary:

      The manuscript "Canonical and phosphoribosyl ubiquitination coordinate to stabilize a proteinaceous structure surrounding the Legionella-containing vacuole" by Steinbach et al. is well written and presents strong evidence that satisfactorily supports the main hypothesis and research objectives. The authors have clearly demonstrated the presence of cloud-like, detergent-resistant GTPase Rab5 surrounding the LCV, and formation of the structure is dependent on the SidE family of effectors. The study provides insights into the relevant (associated with described phenotype) ubiquitination pathways. The findings advance our understanding of Legionella pneumophila vacuole remodeling during intracellular infection and open directions for future research to establish broader implications of this structure on Legionella pathogenesis.

      Strengths:

      The manuscript convincingly demonstrates the presence of a cloud-like, detergent-resistant GTPase Rab5 surrounding the LCV through elegant microscopy. The experimental evidence about the dependence of the observed phenotype on the SidE family of effectors is compelling and presented with strong scientific rigor. The introduction is well-written, and the discussion is thorough and satisfactory. The article is thought-provoking and shows preliminary evidence for ubiquitin-mediated protection and spatial organization of the LCV.

      Weaknesses:

      The manuscript is well-organized and detailed, and it is hard to find weaknesses under the set goals of the research. A few weaknesses are that the molecular determinants or the regulatory mechanisms that drive selective versus non-selective incorporation of host proteins into this structure are unclear, and, as the authors mentioned, further work is required to establish the precise biophysical basis of the detergent resistance and expansive morphology of the ubiquitinated GTPase "cloud". Currently, the function or purpose of the structure is completely speculative. The effects or importance of the structure on bacterial replication is also not established in the current study. Figure 2D, right panel, Western blot results, the authors suggested the signal present in all four lanes between 37 and 25 kDa is 'nonspecific', which is probably a 'too intense' signal to be called so. Mass spec analysis would be interesting in order to identify sources of such intense signals. With these few limitations, the research presented in this manuscript is experimentally rigorous and opens avenues for future research.

      We thank Reviewer #2 for their positive assessment and constructive criticism of our study. We agree that the degree of selectivity of incorporation of proteins into the “cloud” is of great interest, as are the molecular details of the cloud structure, and we expect that future experimentation in this area will provide insight into these key questions.

      Reviewer #2 rightly points out that our study did not address the role of the LCV associated “cloud” in supporting bacterial replication. We note that previous studies have reported growth defects for knockout strains lacking SidC/SdcA (PMID 24483784) and the SidE family (PMID 27049943). However, given the multiple roles that these effector families appear to play during infection, we cannot ascribe these defects in bacterial growth solely to the absence of the LCV associated “cloud”.

      As for the band present in the four lanes in Fig 2D, we suggest that this band is non-specific (most likely detection of the light chain of the antibody used for immunoprecipitation) because we do not observe this band in the input lanes, and we also see this band in the IP samples in Fig 2C (uninfected samples), including the vector control in which no PR-ubiquitination is observed. In Fig 2C, the non-specific bands in the IP samples appear lower intensity because the HA signal is relatively intense in comparison to the infection experiment in 2D, as overexpression of SidE family effectors results in far more PR-ubiquitination than in infection.

      Reviewer #3 (Public review):

      Summary:

      This manuscript by Mukherjee and colleagues extended earlier studies on the coordination of the SidC and SidE effector families on the generation of a unique ubiquitin layer on the surface of the vacuoles containing the bacterial pathogen Legionella pneumophila (LCV).

      Strengths:

      The main strength of the manuscript is the identification of the small GTPase Rab5 as a major "carrier" of these differently modified ubiquitin and ubiquitin chains, which was nicely quantified.

      Weaknesses:

      (1) The results are mostly descriptive, based on mechanistic studies from earlier works.

      (2) The majority of the work was dedicated to the characterization of the unique ubiquitin layer on the LCV. One important question was ignored: what is the role of Rab5 in this process? Is the GTPase activity of Rab5 required for its ubiquitination by SidC and SidE? The authors should create a Rab5 KO cell line, complement the line with different mutants of Rab5, and examine their ubiquitination and association with the LCV.

      (3) The finding that Rab5 is associated with the LCV supports the notion that the LCV has characteristics of endo- or/late endosomes. The positioning of the LCV in the endocytic pathway should be discussed in the context of earlier studies (e.g.,PMID: 38739652; PMID: 11067875; PMID: 11067875).

      We thank Reviewer #3 for their constructive criticism of our work. While we appreciate this reviewer’s interest in Rab5, our data is not consistent with Rab5 being a primary “carrier” of ubiquitin species; many more LCVs are ubiquitin-positive than Rab5-positive during early infection, and in our live imaging experiments we observe many ubiquitin-positive, Rab5-negative LCVs. We used Rab5 as a model substrate in this study because it allowed us to compare modification at the LCV membrane between the WT and avirulent dotA strain. Our data is more consistent with a model in which Rab5 is one of many small GTPases, and likely other host proteins, caught in a crosslinked mesh around the LCV. However, we agree that discussing the interaction of the LCV with the endolysosomal system is relevant; while this is discussed at length in our previous publication (PMID 38117589), we have expanded the discussion in this study to include new publications and contextualize our latest findings.

      We agree with Reviewer #3 that assessing the role of nucleotide binding state in Rab5 ubiquitination is of interest. While creating a Rab5 KO cell line was not feasible given time and technical constraints, we conducted overexpression experiments with nucleotide binding mutants that exhibit dominant phenotypes (Q79L and S34N) and find that these mutants are still recruited to the LCV and ubiquitinated during infection (see new figure S1).

      Recommendations for the authors:

      Reviewing Editor Comments:

      There are suggestions from the reviewers to further address the role of Rab5 in LCV-associated ubiquitination, including whether its GTPase activity is required for modification by SidC and SidE, and the mechanism underlying the dissolution of the ubiquitin cloud during vacuolar expansion.

      Reviewer #1 (Recommendations for the authors):

      To improve upon the manuscript and its impact, the authors could consider the following:

      Major concern:

      The temporal regulation of the ubiquitin cloud is fascinating. The authors nicely demonstrate that SidE- and SidC-type ligases cooperate to form the cloud, but how is it dissolved during vacuolar expansion? They demonstrate that ectopic expression of DopA can do this, but do DopA and DopB regulate this process natively?

      We thank Reviewer #1 for their suggestions. While we agree that this line of experimentation is absolutely of interest, it is not feasible for our lab to carry out these experiments on a reasonable timeline to include in the current work.

      Minor concern:

      A syntax error on line 267 of the manuscript should be addressed.

      This has been corrected.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public Review):

      Summary:

      This study presents valuable data on diurnal patterns in aphid (Rhopalosiphum padi) feeding behavior and transcriptome profiles. The authors measured honeydew production by the aphids on plants and artificial diet during the day and night and conducted a comprehensive feeding behavior study using EPG with many biological replicates at 6 time-points in 24 hours. They also conducted transcriptome analyses of three samples of each 30 aphids at these time points. Differentially expressed transcripts were grouped into four clusters with distinct expression patterns. The expression of two genes found to be diurnally rhythmic was knocked down with RNAi and these aphids did less well, especially at night. They also analyzed the differential expression of candidate effector genes and found rhythmic ones to be enriched for more expression in aphid heads versus bodies - this pattern is expected given that effectors are most likely expressed in the salivary glands. Knockdown of a known effector (C002) that is diurnally rhythmic, and a novel effector gene, was found to alter aphid feeding dynamics and performance.

      Thank you for your thoughtful review and summary of our study. We would like to clarify one aspect of your summary regarding our clustering analysis. We did not cluster differentially expressed transcripts. Instead, our clustering analysis was performed exclusively on transcripts that were significantly diurnally rhythmic, as identified in our time-course transcriptomic analysis. This approach allowed us to reveal patterns of gene expression that exhibit robust rhythmicity over the 24-hour cycle, rather than grouping transcripts based solely on differential expression at individual time points.

      Strengths:

      The manuscript was highly accessible, with clear writing, and the figures provided were both comprehensive and of good quality. The datasets generated from this research are valuable to the research field, especially the findings for honeydew secretion, EPG analysis, and transcriptome experiments.

      The datasets generated in this study will be useful to scientists working on aphids and aphidplant interactions and will inform similar studies on other insect species.

      Weaknesses:

      The weaknesses mainly relate to the (depth of) analyses and interpretation of the data. Also, some methods require more explanation, as follows:

      In Figure 1, data show that aphids produce more honeydew at night than during the day. This suggests that the aphids ingest more phloem (E2 phase). However, in Figure 1d the duration of the E2 phase does not show obvious differences among the time points in the 24 hours. The authors contribute the explanation that the aphids may osmoregulate more during the night, leading to more honeydew secretion at night. This may be the case, but there could be other explanations. For example, the physiology, including regulation of water transport, of plants is known to change during night/day. The authors may focus this section more on the differences in the E1 phase, as this involves the delivery of aphid saliva and effectors into the plant phloem.

      Thank you for your constructive feedback. As noted, aphids excreted more honeydew at night, although the duration of the E2 phase did not differ significantly across time points. We agree that host plant physiology, particularly the composition of the phloem and its osmotic quality, also influences the observed osmoregulatory patterns in R. padi. However, a similar diurnal pattern of honeydew excretion was also observed on artificial diets (Fig. 1b), in which host-derived cues have been eliminated. This strongly suggests that increased nighttime honeydew excretion is primarily driven by enhanced aphid osmoregulation rather than plant factors alone. Nonetheless, we acknowledge that plant-derived factors may also contribute and cannot be entirely ruled out. We have revised the text in the discussion of the revised manuscript to reflect this broader interpretation. As suggested, we have also added further details to highlight the important role of the E1 phase in aphid salivation.

      Transcriptome data shown in Figure 2 (and the experimental procedure of Figure 5b) appears to be based on three biological replicates. However, these replicates appear to have been harvested at the same time in the experiment, and this makes them technical replicates, not biological replicates. The inclusion of true biological replicates that include samples from time series experiments done on different days should be considered.

      Thank you for your concern regarding the biological replication in our transcriptome analysis. Our experimental design included multiple independent pools of aphids collected at each time point. Specifically, each replicate consisted of a unique group of aphids collected from different leaf positions across multiple host plants, such that no individuals were shared among replicates. As a result, these samples represent biologically independent populations rather than technical replicates, which are defined as repeated measurements of the same biological sample used to estimate technical noise. Although all samples were collected within the same experimental time course, this approach is commonly used in time-series transcriptomic studies to minimize confounding variation associated with differences in insect age, entrainment history, or environmental conditions, all of which can obscure rhythmic gene expression patterns. By maintaining tightly controlled and consistent conditions across the sampling period, we aimed to ensure that observed transcriptional differences primarily reflected diurnal regulation rather than uncontrolled day-to-day variability.

      We acknowledge that conducting time-series experiments on different days could provide additional insight into biological variability. However, our approach aimed to reduce potential confounding effects caused by day-to-day environmental fluctuations – such as minor changes in temperature or humidity - which could significantly influence gene expression in insects. By maintaining consistent conditions, we sought to ensure that observed transcriptional differences were due to diurnal rhythms rather than uncontrolled variation. Similar designs or strategies have been employed in studies examining diurnal and circadian gene expression in both insects and plants. We have revised the Methods section to clarify our replication strategy.

      The authors conducted knockdown experiments targeting aquaporin 1 and gut sucrase 1 in aphids, resulting in reduced nymph production and decreased honeydew secretion. It is concluded that these results indicate significant roles of aquaporin 1 and gut sucrase 1 in diurnal regulation. However, it is essential to consider that these genes likely play crucial roles in aphid physiology beyond diurnal rhythms. Consequently, reduced expression would naturally impair aphid performance. The dsAQP1 and dsSUC1 aphids consistently produced less honeydew, regardless of the time of day, indicating a broader impact of gene knockdown. The observed increase of the phenotype at night may not be attributable to the specific roles of these genes in diurnal regulation but rather due to heightened aphid activity during that time (as evidenced by increased honeydew secretion) that could magnify the impact of the knockdown effect, making it easier to observe. Therefore, the knockdown of aquaporin 1 and gut sucrase 1 may exert a general negative influence on aphid fitness, independently of diurnal factors.

      We agree that these genes likely play fundamental roles in aphid physiology beyond diurnal rhythms, and that reduced expression may affect overall aphid performance. However, it is important to highlight that if the observed effects are solely due to general fitness impairments, we would more likely expect a comparable reduction across time points rather than a disproportionately stronger impact at night. We agree that the increased honeydew excretion at night is likely due to heightened aphid excretory activity. However, since this excretory behavior is downstream of osmoregulatory functions, such as water cycling to the midgut and digestion, polymerization, and excretion of oligosaccharides, the increased nighttime phenotype is likely a result of an increased nighttime regulation of osmoregulation in aphids. This hypothesis is further supported by our functional analysis, where the knockdown of AQP1 and SUC1 resulted in a loss of diurnal variation in honeydew production (Fig. 2h), indicating that the observed effects are not merely a general impact on aphid fitness but are likely tied to the genes' roles in regulating diurnal physiological processes. We have revised the discussion to clarify that our findings do not exclude general physiological roles for these genes but instead suggest that their functions intersect with diurnal rhythms to influence aphid feeding and excretion patterns.

      To analyze the roles of genes in diurnal regulation, additional controls should be incorporated. This could involve the knockdown of genes with essential functions that are not influenced by diurnal rhythms, providing a baseline comparison. Furthermore, consider including genes known to be involved in diurnal regulation in other insects, as documented in the existing literature, in the experimental design.

      We agree that incorporating appropriate controls in the RNAi experiments would provide a useful baseline for comparisons, helping to distinguish between general physiological effects and specific diurnal effects. Unfortunately, given the current limited knowledge on rhythmic genes in aphids, particularly in the context of aphid-plant interactions, it is challenging to identify appropriate rhythmic and non-rhythmic controls that can be definitively linked to or unaffected by diurnal regulation within aphids. We will ensure to consider this valuable suggestion in our future experiments.

      The same arguments as for aquaporin 1 and gut sucrase 1 above may be made for knockdown of effector genes (Figure 4). It has already been shown that knockdown of C002 impacts aphid performance, and the data herein may be explained by a general lower performance of aphids rather than a specific function of these effectors in diurnal regulation. It is also expected that knockdown of the effectors has less impact on aphids feeding from artificial diets. This does not necessarily indicate the role of the effectors in diurnal regulation.

      Our response to this comment mirrors that expressed in our earlier response regarding AQP1 and SUC1.

      In the abstract and elsewhere, the authors assert priority by stating, "...the first evidence of...". However, it's important to note that priority claims are often challenging to verify across many fields. Instead of relying solely on claims of precedence, the evidence presented in the research could stand on its own merit.

      We understand that priority claims can be difficult to substantiate across various fields, and we appreciate the importance of allowing the evidence to speak for itself. Considering this, we have revised the language in the abstract.

      Conclusion:

      The study presents intriguing new findings, particularly in the realms of honeydew analysis, EPG, and transcriptome analysis. However, the interpretation of subsequent studies employing gene knockdowns needs further consideration.

      We thank the reviewer for the thoughtful and constructive feedback. We appreciate the positive assessment of our findings on honeydew analysis, EPG, and transcriptome profiling. We have carefully revised the section on gene knockdown experiments to provide clearer interpretation and additional context, and we hope the concerns raised have now been appropriately addressed.

      Reviewer #2 (Public Review):

      Summary:

      The authors conducted a time-course of whole-body transcriptional analysis of a pest aphid, Rhopalosiphum padi, and identified four major clusters of the genes that show diurnal rhythmicity in transcription. In addition, they conducted the analysis of aphid feeding behaviour and showed that aphids salivate longer from the end of the day toward the beginning of the night while their phloem feeding time does not change throughout the day. The genes upregulated at night time were enriched with the genes involved in metabolic activities, collaborating with the results showing a higher number of honeydew excretion at night. The authors identified the list of candidate salivary genes that show diurnal rhythmicity in the transcription and silenced a salivary gene C002 and the candidate salivary gene E8696. Silencing of these genes reduced aphid fecundity and survival rate on the host plant but not on the artificial diet.

      Thank you for your thoughtful review and valuable comments on our study.

      Strengths:

      The time-course transcription study and its analysis will be of interest to researchers studying diurnal rhythms in insect biology. Also, the analysis of aphid feeding behaviour at different times of day is interesting. This study provides variable resources for those who study insect biology.

      Weaknesses:

      It is not clear to me which data was used to define the putative salivary effectors for R. padi, but the candidate salivary gene list made by Thorpe et al consists of the aphid genes encoding secreted proteins that are up-regulated in the head samples compared to the body samples. Although some proteins were confirmed to be secreted into the aphid saliva, many genes in the list are not confirmed to be expressed in the aphid salivary glands, and their products are not confirmed to be secreted into the saliva and the plant. Is E8696 expressed in the aphid salivary glands and secreted into its host plant? Without the data confirming the expression of the gene in the salivary glands and its secretion into the saliva and into the host plant, we cannot call the protein a salivary protein. Furthermore, without the observation that E8696 has some effect on plant biology, we cannot call it an aphid effector. Therefore, I cannot agree with the parts of the manuscript that refer to E8686 as an aphid salivary effector.

      We have revised the text in the Methods to clarify the database used for defining putative salivary effectors. We have also added a sentence in the discussion to indicate that these are putative effectors. The putative effector E8696 was confirmed to be expressed in the salivary glands; however, its secretion into saliva and the host plant remains undetermined due to the lack of E8696-specific antibodies. Over the past year and a half, we have been creating an antibody for E8696. However, the antibody we generated is non-specific, and as a result, we are still unable to demonstrate that E8696 is secreted into host tissue and functions as an effector. While our functional analysis provided strong evidence of E8696’s impact on aphid fecundity and mortality on host plants but not on artificial diets, we agree that without further confirmation of its secretion and effect on the host plant, E8696 should be considered only a putative salivary effector. We expect to address these important questions in future research. To prevent any confusion, we have revised our manuscript to reflect that E8696 is only a putative effector.

      It is interesting to know that some candidate salivary gene expression showed a diurnal rhythm. However, without the knowledge of the functions of the salivary effectors, especially their targets, it is not possible to conclude that the rhythmical expression is important for the aphid performance. In addition, I wonder whether the increase in gene expression is directly correlated with the increase of protein secretion into the saliva and the plant.

      The primary goal of this study was to determine whether aphid genes, particularly those associated with osmoregulation and salivary effectors, exhibit diurnal patterns of expression and whether disrupting these rhythms affects aphid performance. While we agree that the precise molecular targets of these effectors in host plants remain to be identified, our functional assays provide evidence that rhythmic expression is biologically relevant for aphid physiology. Our results demonstrate that silencing rhythmic effector genes resulted in increased aphid mortality, reduced fecundity on host plants, and, more importantly, the disruption of diurnal honeydew excretion patterns, especially for C002. As honeydew excretion is a critical physiological process for aphids, the alteration of this behavior suggests that the rhythmic expression of these genes is functionally important for aphid physiology. We believe our results provide compelling evidence that rhythmic expression plays a critical role in aphid biology. We agree that rhythmic transcript abundance does not necessarily imply proportional changes in protein secretion into saliva, and direct measurements of effector protein dynamics will be an important direction for future work. However, the observed physiological and performance consequences of disrupting rhythmic gene expression support the conclusion that temporal regulation of these salivary genes is functionally important for aphid biology, even in the absence of detailed target identification.

      Finally, the authors examined aphid survival, fecundity, and feeding behaviour. Those are important for overall aphid performance, but they do not "shape" aphid colonization. Aphid colonisation is shaped by the mechanisms by which aphids find and select their host plant and start to feed on it. Therefore, I do not agree with the title of this manuscript and some parts of the discussion.

      We agree with your perspective and have revised the title and discussion to more precisely reflect the scope of our findings, focusing on aphid performance rather than colonization. The revised title now reads “Diurnal rhythmicity in metabolism and salivary effector expression shapes aphid performance on host plants”.

      I would like the authors to develop how the knowledge of the diurnal rhythm of aphid feeding can contribute to optimise pest management. I see that there are some differences in aphid metabolism and feeding behaviour between day and night, but I would like to hear how such knowledge can optimise pest management strategies.

      We have expanded the Discussion to address how knowledge of diurnal rhythms in aphid physiology and feeding behavior could inform the optimization of pest management strategies. Specifically, we discuss how time-of-day variation in aphid feeding activity and metabolism may influence the efficacy of control measures and how chronobiological insights could be integrated into future pest management frameworks.

      Recommendations for the authors:

      Reviewing Editor:

      Based on comments from two reviewers, here are the six key areas that need to be addressed to improve the manuscript.

      Clarity and Specificity:

      (1) Salivary effectors: The manuscript defines "salivary effector" loosely. The reviewer argues for stricter criteria - a protein can only be called a salivary effector if it's confirmed to be produced in the salivary glands and/or secreted into the plant with saliva and function in or around the plant.

      We have addressed this comment and clarified the definition in the revised manuscript.

      (2) Diurnal rhythm: The paper finds a daily rhythm in aphid gene expression, but doesn't explain how these genes affect the plant. The reviewer argues that without understanding the function of these genes, the significance of the rhythm is unclear.

      We have addressed this comment and clarified that the scope of our study is to elucidate diurnal rhythmicity in aphid gene expression and to evaluate the functional importance of rhythmic genes for aphid performance. We agree that understanding how these genes interact with host plants is essential for fully elucidating their molecular functions under diurnal regulation; however, this is beyond the scope of the current study and will be pursued in future research.

      (3) Knockdown experiments: The reviewer suggests the observed effects of knocking down certain genes (aquaporin, sucrase, effectors) might be due to their general importance, not necessarily their role in the day-night cycle. They recommend including control genes and genes known to be involved in circadian rhythms for a more robust comparison.

      We have addressed this comment and clarified the interpretation of these experiments in the revised manuscript.

      Technical Issues:

      (4) Honeydew production: The explanation for nighttime honeydew production needs more exploration. Plant changes at night might also play a role, and the daytime saliva delivery phase deserves more attention in the analysis (Figure 1).

      We expanded the description and interpretation of the salivation phase by incorporating additional detail in the revised manuscript.

      (5) Gene expression data: The current data (Figure 2 & Figure 5b) lacks proper biological replicates. Replicates collected at different times are essential for stronger conclusions.

      We have addressed this comment and clarified the experimental design and replication strategy in the manuscript.

      (6) Priority claims: The reviewer advises against focusing on claiming novelty ("first evidence"). The research should be impactful based on its own merit, not just being the first to find something.

      We revised the sentences to avoid making claims of priority throughout the manuscript.

      Reviewer #2 (Recommendations For The Authors):

      Figures 2 f,g, and h : according to the legend, these experiments seemed to have a low number of replicates (n=3-5). However, Figure 2h has many data points. I understood that here n means the number of experimental replications, but it may be better to show the number of aphid samples examined.

      You are correct that the n refers to the number of experimental replicates, with each replicate comprising multiple individual aphids. Because our analyses were performed on replicate-level averages across multiple days, rather than on individual aphids, we believe this notation most accurately reflects the experimental design. To improve clarity, we have revised figure legends to explicitly state that each replicate includes several individual aphids.

      Are the orthologous proteins of E8696 expressed in aphid salivary glands or detected in saliva? Such data will strengthen the claim that E8696 is a salivary protein of R.padi.

      E8696 is expressed in aphid salivary glands, but it is not confirmed to be secreted into saliva or host plants due to the lack of specific antibodies. We have revised our manuscript to reflect that E8696 is only a putative effector. We will address this question in future research.

    1. Author response:

      The following is the authors’ response to the previous reviews

      We thank the reviewers for their additional feedback. Below, we provide detailed responses to each reviewer’s major concerns. In addition, we identified an error in the previously submitted Fig. 6C and have corrected the X-axis labels accordingly.

      Public Reviews:

      Reviewer #1 (Public review):

      Motion-related signal in ACC: the new Fig. 2E looks good, but it is hard to visualize how it is just a reordering of the old Fig. 5C.

      We thank the reviewer for this feedback. Fig. 2E and the original Fig. 5C do bear resemblance. The primary difference is the temporal window and organization of the data. In the original Fig 5C, the time window was only ± 5 sec whereas Fig. 2E is ± 30 sec. The main objective we aim to highlight is that ACC shows both activation and inhibition in response to shuttle on an extremely prolonged order, up to 30 sec. Data is sorted to separate inhibition and activation to illustrate the sustained activity persists for both populations.

      All categories in the new Fig. 4D appear to respond to shuttle initiation, with less than 1s latency. For example, type 2a/2b consists of 40% of the population and their response to movement onset is apparent. Thus, it is not clear whether most neurons respond to shuttle crossing as described in the manuscript.

      We thank the reviewer for drawing attention to this discrepancy. It was not our intention to strike comparison between shuttle initiation versus shutting crossing responses across neurons, and we do not dispute that ACC responds to both events. While shuttle initiations and crossings provide a consistent temporal alignment point, they do not define the temporal focus of much of our analyses. Given that most shuttle responses terminate within ~2 sec, the extended windows analyzed (i.e. ± 5 sec; Fig. 4) largely reflect post-action ACC activity. Overall, although ACC neurons show mixed responses to initiations or crossings, the most consistent feature is prolonged modulation that persists beyond shuttle termination. We have revised the text to reflect this focus.

      Given this and the reviewer’s feedback, we further examined whether ACC activity is more strongly aligned with shuttle initiation, crossing, or termination. To determine which shuttle event (initiation, crossing, or termination) captured the most acute changes in ACC neuronal firing, we conducted an event-locked modulation analysis (Fig. S4). Our results showed that shuttle crossing was associated with the largest fraction of significantly modulated ACC neurons (Fig. S4). These findings suggest that shuttle crossing represents the most prominent event for ACC engagement during shuttle behaviors.

      Could the authors use relatively simple analysis, such as comparing spike rate before and after crossing, or before and after initiation, to quantify the response properties of each neuron? This could also help validate the classification analysis performed in Fig. 4.

      As mentioned above, we have added a new supplemental figure to directly address this question (Fig. S4).

      Reviewer #2 (Public review):

      I think the authors did a very admirable job revising the manuscript. It is much improved. However, I believe a formal analysis of action-state versus action-content neurons on A-->B versus B-->A crossing is still warranted. I appreciate the fact that this analysis may not be as reliable with smaller ensemble sizes, but with careful pseudo-ensemble and resampling approaches, such an analysis would go a long way towards increasing the strength of evidence.

      At present, we are not sure what the reviewer means as “formal analysis”. Below is our best effort in addressing this concern.

      Firstly, in our first revised manuscript, we implemented a generalized linear model-based classification of action-content and action-state neurons using direction specific regressors. Specifically, this analysis classified neurons as action-content or action-state based on coefficient contrasts (Δβ), with appropriate statistical testing and multiple comparison correction (see Methods; Fig. 7 C–E). Neurons were classified as action-content neurons if the corrected p-value for Δβ was significant and the absolute effect size exceeded a predefined threshold (|Δ β |> 0.5). Neurons were classified as action-state neurons if Δβ was not significant but both β1 and β2 were individually significant after correction. We believe our generalized linear model-based classification offers a sophisticated and formal classification of these two neurons classes.

      Subsequently, we performed an SVM decoder to distinguish A→B from B→A shuttles. Decoding accuracy depended on action-content neurons, as their removal drastically decreased decoding accuracy, whereas removal of non-action-content neurons had no effect, further strengthening the conclusion that these populations encode distinct information.

      In the updated revision, we performed an additional SVM decoding analysis while controlling for unequal neuronal population sizes between action-state and action-content neurons (Fig. S8). Specifically, we constructed pseudo-ensembles by randomly resampling neurons within each category and training SVM decoders on size-matched ensembles. Decoder performance was evaluated across repeated resamples to generate distributions of accuracy. We found that only decoders using action-content neuronal activity predicted shuttle content with high accuracy (>95%), whereas decoders trained using non-action-content neurons performed at chance levels (Fig. S8).

      Reviewer #3 (Public review):

      The only remaining comment that was not addressed pertains to anatomy and recording details. Some electrodes appear to be clearly in M2 (Fig 2A), and the tetrodes were driven each day. I would strongly suggest that this be included as a further limitation, particularly given the statement on line 178.

      We thank the reviewer for this feedback. In the previous revision, we added a supplemental figure showing tetrode locations for each mouse (Fig. S2) and described recording details in the Methods (Lines #481–488). We agree that this should also be noted as a limitation, and we have now added this to the Discussion (Lines #384–388).

    1. Author response:

      The following is the authors’ response to the original reviews.

      As the reviewers noted, the evidence we provide is the strongest on the mechanistic link between hepatic cardiolipin deficiency and electron leak from the electron transport chain. This narrative is supported by our assessment of site-specific electron leak as well as reconstitution of exogenous cardiolipin in the small unilamellar vesicles deficient with CL. On the other hand, as pointed out by the Reviewer 2, the mechanistic link between cardiolipin to MASLD/MASH is less robust. At this moment, we have not experimentally demonstrated that the MASLD/MASH induced by CLS deletion can be rescued by replacement of mitochondrial CL in vivo. Taken together, our current narrative makes an incomplete loop between CL deficiency, electron leak, and MASLD/MASH. Nevertheless, as indicated by all the reviewers, this manuscript highlights a previously undescribed role that CL potentially plays in MASH pathology, particularly with the data that human MASH coincides with reduction in liver mitochondrial CL. We focused this revision primarily on additional descriptive experiments in CLS-LKO mice that were requested by the reviewers. Even though it is not a component of the current manuscript, we have recently successfully developed mice with hepatocyte-specific CLS overexpressing mice and began performing experiments to test causality of CL deficiency to MASLD/MASH which we hope to complete in a few years. We are hopeful that the MASLD/MASH research community will still find evidence on CL contained in this manuscript plausible, and that it provides critical information to our understanding of mechanisms for MASH pathogenesis.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Brothwell and colleagues describes a central role for hepatic cardiolipin deficiency in MASH. The authors identify cardiolipin as a mediator of two long-standing problems in the field: how dysregulated lipid metabolism relates to altered mitochondrial metabolism during MASLD, and what the innate changes are in the steatotic liver that cause the increased respiration. The authors identified reduced liver cardiolipin in humans with MASH and in a variety of mouse models with MASH. When they knocked out hepatic cardiolipin synthesis, mice developed steatosis and inflammation. These mice also recapitulated the elevated hepatic oxidative metabolism and oxidative stress found in obese humans with MASLD. Some of the in vivo functional data related to glucose homeostasis and substrate metabolism could be stronger, and interpretation of the in vitro flux data needs some clarification, but in both cases, the data are not essential to the main conclusions of the manuscript. Overall, the study offers compelling evidence that cardiolipin is reduced in MASLD and that impaired cardiolipin synthesis is sufficient to recapitulate many features of MASLD.

      We thank the reviewer 1 for the positive feedback emphasizing novel and important findings in our manuscript.

      Strengths:

      The main strengths of the study are:

      (1) The identification of reduced cardiolipin levels in the liver of humans with MASLD and in a variety of mouse models of MASLD.

      (2) The finding that loss of cardiolipin synthesis recapitulates steatosis and inflammation in MASH.

      (3) The finding that loss of cardiolipin increases mitochondrial respiration, ROS production, and fat oxidation (in a separate hepatocyte cell line), again recapitulates several previous studies in obese humans with MASLD.

      (4) Evidence, though less definitive, that cardiolipin deficiency promotes electron leak by disrupting respiratory supercomplexes and preventing CoQ reduction.

      Weaknesses:

      (1) Figure 3A-D tries to make the point that liver CLS KO causes defects in substrate handling in vivo, based on glucose and pyruvate tolerance tests. The KO mice have a blunted response to a glucose tolerance test, but the pyruvate tolerance test showed very little (almost no) effect on glucose levels in either WT or LKO mice. The small blunting of the response in the LKO is impossible to interpret (if it's real), since the ability to clear glucose is also increased, and no tracers were used. It might be useful to monitor pyruvate and lactate levels during the experiment. However, this reviewer doesn't think the data is essential to prove the authors' main points.

      Thank you for pointing this out. We have now revised our manuscript to correctly reflect our findings on GTT and PTT. In our initial submission, we failed to clearly articulate that CLS deletion appeared to increase systemic glucose handling, which is the opposite of what one might expect in liver with steatosis. We agree that additional experiments would be helpful to better understand the systemic substrate handling in the CLS-LKO mice. As the reviewer indicates, we decided to focus this particular manuscript on intracellular and mitochondrial metabolism because of cardiolipin’s known localization to mitochondria, and the central role that this organelle plays in the pathogenesis of MASLD.

      (2) After presenting convincing evidence that respiration is elevated in isolated mitochondria from CLS KO liver, the authors follow up the findings by investigating whether 13C-palmitate and 13C-glucose oxidation are altered by CLS knockdown in murine Hepa1-6 cells (Figure 4).

      A few comments are worth mentioning about Figure 4:

      (a) It is not clear why the authors chose to use a hepatoma cell line rather than primary hepatocytes from LKO mice. The latter would be more convincing, since there could be important differences in metabolism between hepatoma cells and hepatocytes (e.g., preference for fatty acids vs glucose). Nevertheless, I think the approach is sufficient to test the general effect of loss of CLS on substrate metabolism.

      We appreciate the sentiment and agree that primary hepatocytes would have been a better model. We simply have not had prior expertise to culture primary hepatocytes and do not have the system working. We completely agree that it’s important to discuss the limitation of hepa1-6 cells as a hepatoma cells and now discuss this in our manuscript.

      (b) The authors use the M+2 enrichments of TCA cycle intermediates to infer rates of oxidation of [U-13C] palmitate or [U-13C] glucose. It is important to note that this kind of data reports fractional carbon sources (i.e., substrate preference) rather than rates of oxidation. For example, data from the 13C-palmitate experiment indicates that the CLS KD cells increase the fractional contribution from 13C palmitate (compared to glucose, for example) to the TCA cycle, but the actual rate of palmitate oxidation is not implicit in the data. However, it is reasonable to suggest that, in combination with the increased rates of O2 consumption observed in isolated mitochondria, this data supports increased fat oxidation.

      We agree with the reviewer that the nuances are important: that M+2 enrichments from [U-13C] palmitate or [U-13C] glucose reflects the fractional contributions of labeled substrates to the TCA cycle rather than oxidation. We have now revised the text to clarify that the data represent carbon incorporation patterns.

      (c) I have some concern that the [U-13C] glucose experiment is more complicated to interpret than the description implies. I'm not sure what happens in this cell line, but in the liver, most labeling from pyruvate (i.e., originating from glucose in this case) enters the TCA cycle via pyruvate carboxylase, with smaller amounts entering via PDH (depending on the nutritional state). Since one could expect pyruvate carboxylase to contribute M+3 labeled TCA cycle intermediates initially, and M+2 on the first turn of the cycle, it's hard to conclude what the data indicates about glucose oxidation. The authors could generalize the conclusion by framing the TCA cycle enrichment data as the contribution of glucose carbons and noting in Figure 4A that pyruvate carbons can enter the TCA cycle via PDH or pyruvate carboxylase, without attempting to assign their relative contributions. There are better ways to do it, but it's a small nuance here since the authors aren't making a critical point about the pathways.

      This expert comment is much appreciated. We have revised the text to more broadly describe glucose carbon entry into TCA cycle through PDH and PC. We also revised the schematic to reflect this notion.

      Reviewer #2 (Public review):

      In this study, the authors show that alterations in the lipid composition of the inner mitochondrial membrane, particularly changes in cardiolipin (CL) content, lead to defects in electron transport, supercomplex formation, and oxidative stress. Using liver-specific CLS knockout mice, which are characterized by dysfunctional capacity for cardiolipin synthesis, the authors highlight an underappreciated role for CL in MASH pathology. Overall, this is an interesting study highlighting the importance of functional/physiological electron transport (and in this context, electron leakage) in MASH pathophysiology. Despite that, this manuscript has several weaknesses that require attention.

      We thank the reviewer 2 for the constructive criticisms and identifying areas of weakness were additional data or explanations can improve the manuscript.

      (1) For all LKO studies, it is stated that the decrease in hepatic CL is causal for the observed phenotype. However, it is evident that many other lipids are impacted by CLS KO, including a marked increase in hepatic PG. In this respect, the authors show no evidence that the observed metabolic phenotype is indeed due to the reduction in CL and not to other accompanying changes.

      Thanks for this comment. We agree that because deletion of CLS promotes changes in mitochondrial lipids other than CL, we cannot conclusively attribute phenotypes we observed to CL and not to other lipids such as PG. In our experience, rescuing mitochondrial phospholipids by exogenous supplementation is problematic as they most certainly are not exclusively destined to the tissue of interest, nor to the organelle of interest, and often metabolized to produce other lipids, etc, making it difficult to interpret the data. We now have mice that conditionally overexpress CLS, which could be used to address this question, but the study is in its early phase and are outside the scope of the current study.

      The one experiment we performed is the ex vivo CL supplementation by SUV fusion to mitochondria, which has an ability to rescue electron leak. While they do not demonstrate the role of CL in all phenotypes found in the CLS-LKO mice, we think that bioenergetic phenotype associated with CLS deletion is therefore likely due to the reduction in CL. We now provide these additional discussions in lines.

      (2) In the results, the authors highlight that 'MASLD has been shown to alter the total cellular lipidome in liver.' Given that this study focused on CL, it would be useful to include specific studies that pointed to changes in hepatic CL content in MASLD/MASH/fibrosis.

      We now provide citations for these studies (PMID: 30042157, PMID: 34257827).

      (3) The initial human mitochondrial lipidomics studies show a reduction in mitochondrial CL and PG content. What was the content/expression of CL synthase and PGP synthase in these samples? If this cannot be assessed, is there any association of CLS or PGPS expression and MASLD/fibrosis (etc) in publicly available databases (e.g, GEP liver) that may explain the reduction in mitochondrial PG and CL content?

      Thanks for this suggestion. Quantification of mitochondrial lipidome require a good amount of tissue, and we do not have sufficient biomaterials left to quantify gene expression. Upon our survey of publicly available database (including GepLiver), we did not find that human MASLD was associated with an increase in CLS or other enzymes of CL biosynthesis compared to healthy controls.

      (4) The validation of MASH in patients (Figure 1B) is not convincing (ie., no quantification/scoring provided). NAS /fibrosis scoring (according to Kleiner) would help to define if all patients have indeed MASH, and what subset has fibrosis. Could the reduction in CL/PG content be (also) associated with fibrosis? In addition, Masson's Trichrome should be added to Figure 1B.

      The diagnosis was based on obvious bridging fibrosis and/or regenerative nodules on H&E staining (see additional zoomed-out images in Figure 1 – figure supplement 1). Due to the severity of these cases, formal NAS scoring was not applied. We do not have the Trichrome staining available but all MASH samples had fibrosis. Thus, it is possible that reduced CL/PG is related to fibrosis. We now added more descriptions on this point.

      (5) In human lipidomics, the authors suggest that reductions are observed in tetralinoleoyl CL (Figure 1C). However, Figure 1C only shows the combined FA acyl chain length + unsaturation, therefore not allowing for FA-specific ID (unless such data are available from the LC/MS analysis).

      Thanks for pointing this out. Per lipidomic nomenclature guideline we assign combined FA acyl chain length + unsaturation when MS2 is not performed. We have validated that our 72:8 peak corresponds to TLCL, but we do not perform MS2 on every lipid species for every sample. We now clarify this point in our manuscripts.

      (6) Figures 1 J/K/I. It is obvious that the background in all murine immunoblotting analysis has been altered. The authors should provide unaltered images for these immunoblots.

      We apologizes with the confusion. In Figure 1J/K/L/M, each panel actually represents two western blots (not one, similar to Figure 3H). The above represents a western blot with OXPHOS antibody cocktail (CV, CIII, CIV, CII, and CI), while the bottom represents the second western blot with citrate synthase (CS). Thus, we had not manipulated parts of the western blot to look different. To clarify, we now place an outline in each of the western blot to clearly demarcate individual blots to avoid confusion (new Figure 1J-M).

      (7) For Figure 1, it is unclear what is meant by 'we performed all mitochondrial lipidomic analyses by quantifying lipids per mg of mitochondrial proteins'. Was the murine lipidomics carried out on fractionated mitochondria or whole liver? If whole liver, then how were the data corrected, particularly given that PG is not a mitochondria-specific lipid?

      The data are all from lipidomic analyses performed in isolated mitochondria.

      (8) While total CL content seems indeed decreased across the different mouse models, this is mostly due to 1-2 CL species showing a pronounced reduction, with the remainder being unaltered. This should at least be acknowledged in the results. This is similarly the case in the LKO livers.

      Thanks for pointing this out. We now provide additional clarification in the text.

      (9) Figure 2. A secondary biochemical analysis of changes in lipid content should be provided, e.g., total triglyceride content, particularly given that the histology analysis does not show any major changes in hepatic lipid droplets/steatosis. In addition, the Masson's Trichrome staining shows almost no collagen deposition.

      We now provide a quantification of triglycerides in Figure 2J.

      (10) Figure 3. 'CLS deletion modestly reduced glucose handling' should be reworded. The LKO mice show improved glucose tolerance (despite the MASH phenotype), which is not evident from the above wording.

      We modified our text accordingly.

      (11) Looking at the mechanism behind the increase in hepatic steatosis, the authors state that lipid accumulation can occur due to increased lipogenesis, or dysfunctional VLDL secretion or beta oxidation, and subsequently assessed the relevant proteins/pathways. What about fatty acid uptake, which is also one of the four major pathways impacted in MASLD? This should be included in this assessment in Figure 3.

      Thank you for this comment. We now provide data for genes involved in fatty acid uptake, which was not reduced with CLS deletion (Figure 3E).

      (12) For Figure 5A, it is simply stated 'CLS deletion promotes liver fibrosis in standard chow-fed condition', and it is unclear what is highlighted within the selected EM images and what the arrows refer to. The authors should clarify this within the text.

      We have modified the text accordingly.

      Reviewer #3 (Public review):

      Summary:

      Mitochondrial oxphos causes lipid accumulation, leading to MASH, although the mechanism has been poorly understood. In this study, Funai and colleagues identify that reductions in cardiolipin in the mitochondria cause disruptions in the electron transport chain. Knockout of cardiolipin synthase was sufficient to drive MASH phenotypes, increase respiratory capacity, and cause electron leak at complexes II and III. It is well established that loss of cardiolipin increases ROS. Studies to date have been performed on whole tissue lysates, but to rule out which changes in mitochondrial lipids are driven by changes in mitochondrial number versus lipid synthesis/turnover, the authors uniquely purified mitochondria from human and mouse livers in MASH and NASH models for this study. This study provides critical information to the field that will inevitably help us better understand the mechanisms underlying MASH and NASH onset. The evidence provided is both convincing and compelling. With further suggested revision experiments, this study has the potential to change our understanding of MASH and NASH pathogenesis.

      We would like to thank the reviewer 3 for the highly-encouraging feedback.

      Strengths:

      The authors use a unique approach of lipidomics on purified mitochondria. They also analyze many distinct MASH models and provide a unique resource for the field of comprehensive lipidomics analysis of the different ways in which MASH can be induced. The use of human tissue elevates the impact/significance of the findings.

      Weaknesses:

      The data on the super complexes was the least compelling, and frankly, I do not think the authors needed those data to make a compelling argument! The authors should shift their focus more to the compelling electron leak data they have collected. If possible, it would also strengthen the work to include cardiolipin rescues on more of the experiments. Finally, expanding their explanations of the model systems would be very helpful for the readership.

      Thank you for this comment. We have now revised our argument to highlight the electron leak data and less emphasis on the supercomplexes.

      Reviewer #4 (Public review):

      Summary:

      Here, the authors wish to shed light on factors that contribute to the development of liver disease in what used to be called 'the metabolic syndrome'. This is a human-health problem of considerable significance, and the insights they provide, namely the implication of a defect in mitochondrial cardiolipin (CL) content to the progression from metabolic dysfunction associated steatotic liver disease to steatohepatitis, are plausible.

      We would like to thank the reviewer 4 in an encouraging feedback.

      Strengths:

      The experimental evidence proffered is derived from the observation of lower levels of (CL) in mitochondria from the liver of patients undergoing liver transplant or resection due to endstage steatohepatitis compared with mitochondria derived from livers of patients with other conditions. This correlation is buttressed by observations made in mice with liver-selective compromise in CL synthesis and which suggest a pathological environment associated with mitochondrial dysfunction and enhanced oxidative stress, features deemed to play a role in the progression from steatotic liver disease to steatohepatitis.

      The paper is well written, and the findings are well explained and superficially convincing.

      Weaknesses:

      It is unclear how much can be learned from compromising a key enzyme that produces a key mitochondrial lipid in a busy metabolic organ like the liver - isn't the discovery of a mitochondrial defect in such a context rather trivial? And how reliably can these findings be related to the human observations? Most importantly, the chain of causality implied by the title is unproven: the key question of whether or not (somehow) preventing the drop in cardiolipin content affects the course of steatohepatitis remains unanswered.

      We agree with the reviewer that the current manuscript does not directly provide evidence that reduction in CL causes MASLD in humans, which as the reviewer describes, must be tested by rescuing CL content in the context of MASLD. We have now obtained mice with conditional overexpressor and have begun the experiments, but findings from these mice are beyond the scope of the current study. We have modified our title to “Cardiolipin deficiency disrupts electron transport chain AND drives steatohepatitis” to reduce the implication for causality.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The manuscript states that loss of mitochondrial respiration is expected in MASLD. Forexample, line 187 "MASLD is known to be associated with reduced mitochondrial oxidative capacity". A more accurate statement is that "MASH" is known to be associated with reduced mitochondrial oxidative capacity and increased ROS production in humans. As you correctly cite later for an ex vivo human mitochondrial respiration study, early MALSD, especially with obesity, is associated with elevated mitochondrial respiration (40). Since those measurements are maximal respiration rates, which might not reflect actual in vivo flux, you might also make readers aware that your data is consistent with in vivo human studies that found increased hepatic oxidative flux (TCA cycle flux) in obese subjects with moderate steatosis (PMID: 22152305), which appears to wane with severe steatosis and/or inflammation (PMID: 31012869, PMID: 40272888).

      Thank you for these suggestions. We have made the suggested changes to the text.

      Reviewer #3 (Recommendations for the authors):

      (1) Throughout the manuscript, the authors refer to the inner mitochondrial membrane, although they never perform assays to distinguish the inner vs outer mitochondrial membrane. It would be better to just refer to the cardiolipin being measured as "mitochondrial."

      Thank you. We made these changes.

      (2) In figures showing changes in cardiolipins, not all of them change; only a handful of them are reduced in NASH. Could the authors add commentary in the manuscript about what is known about these different cardiolipin species, and speculate as to why certain CLs are changing while others are not?

      Thank you. Reviewer #2 had similar comments and we provided additional discussions.

      (3) In the human tissues, what do the other mitochondrial inner membrane lipids (PC, PE, PI, PS, LPC, LPE) look like in the healthy vs NASH patients (Figure 1A-D)?

      Thank you for this request. We did not include these data in the manuscript as we have a separate ongoing study (the second author is the lead author on this paper) where we are following up on hepatic mitochondrial PS and PE, which we found to be decreased in human MASH samples compared to healthy livers. This turned out to be a convoluted story so we decided not to include it in the paper.

      (4) The descriptions of the different MASLD/MASH models are a little sparse. Especially needing more detail is the model for carbon tetrachloride injection, causing NASH. The authors should explain how each of these models typically induces MASLD/MASH.

      We now provide these details.

      (5) In figures 2E and F, total body mass is unchanged in CLS-LKO mice, but liver mass is decreased; yet on the chow diet, there appears to be lipid accumulation in the liver as well; I am wondering what the authors' reasoning is for this decreased liver mass.

      It is difficult to say conclusively, but we suspect it is due to cell death evidenced by fibrosis. It’s important to note that while there is lipid accumulation in the liver, steatosis is relatively mild and the increase in liver triglyceride is quite marginal (Figure 2J).

      (6) The lipidomics analysis and comparison of livers in these different models is a wonderful dataset that needs far more depth in terms of unpacking and describing the findings. For example, all the models of MASH show similar changes in most of the lipid species analyzed. NASH appears to be quite different than MASH. This, among other trends, is certainly worth highlighting as it will be of interest to the field.

      Thanks for this comment. We agree that while CL phenotype were common to mouse and human MASH samples, there were other changes that we observed in other lipids that may be biologically significant. As described above, we have an ongoing study pursuing mitochondrial PS in the liver.

      (7) Figure 2B - It is interesting that the CLS KO only impacts certain CLs. The 72:8 CL, which is regulated by CLS, is also a CL that appears to change in the human patient samples. The information on the specific CL that is changing seems critical to the mechanism of the role of the CL in the disease. Throughout the manuscript, it is important to specify which specific CL is being referred to, instead of broadly characterizing the changes to cardiolipins, especially since most of the cardiolipins shown do not change; only a handful of them do.

      Thank you for this suggestion. We have included additional discussions on 72:8 CL in the manuscript.

      (8) One potential non-specific mechanism whereby CLS knockout can cause MASH would be if the mice change their overall food consumption. It is an important control to test if the total food intake is different in WT vs KO mice to formally rule out this possibility.

      The food intake was not different between the group (Figure 2E).

      (9) To determine the extent to which de novo cardiolipin synthesis underlies the change in MASH/fatty liver observed in the HFD, GAN, and CCl4 models in Figure 1, the authors should also put the CLS KO mice on these diets and perform liver histology, analysis of inflammation markers, and analyze immune cell infiltration. Alternatively, the authors could try to rescue the CLS KO model by supplementing cardiolipin in the diet or by injection.

      Thank you. We have an ongoing experiment to examine the effect of hepatocyte-specific CLS overexpression on protection from GAN-induced MASLD.

      (10) Figure 3F shows a decrease in UQCRC2 by RNA but no change at the protein level in Figure 3H. The authors should comment a bit more on this disparity, and the data in Figure 3F don't mean much for the main point of the study if the levels of the proteins are unchanged.

      The reviewer is correct. We initially performed RNAseq in trying to broadly capture how CLS knockout influences liver health, which implicated that transcriptional program for mitochondrial proteins were downregulated. Nevertheless, gold standard measurements of mitochondrial content (mitochondrial protein or mtDNA) did not show change in the abundance with CLS deletion.

      (11) The increase in respiration and spare respiratory capacity upon CLS KO shown in Figure 3J is extremely interesting! The explanation of the experiment and its meaning should be significantly expanded upon.

      Thank you. We included additional discussion on this point.

      (12) Figure 4 - It is interesting that the fraction of the TCA cycle metabolites labeled is increasing with the palmitate tracer and decreasing with the glucose tracer. This implies a "fuel switch," such that more of the TCA cycle carbons originate from fatty acids than glucose upon loss of CLS. The authors should make note of this point. Also, to understand if the total molar quantity of labeling in the TCA cycle from palmitate and glucose is changing, the authors should also report the relative abundance (instead of just the fraction labeled) of the labeled metabolites and unlabeled metabolites.

      Thanks for this suggestion, we have now added this discussion.

      (13) In Figure 5C-F, the authors show that CLS deletion can activate the caspase pathway, but do not see any change in cytochrome c localization. Can the authors clarify if CLS deletion is sufficient to induce apoptosis?

      CLS deletion certainly causes cell death that induces tissue fibrosis. Activation of the caspase pathway suggests that the cell death may be due to apoptosis but we did not see changes in cytochrome c localization. Our lab is currently performing additional experience to test the possibility that CLS deletion may induce ferroptosis.

      (14) Figure 6A-C- The authors discuss the I + III2 + IV supercomplex substantially and consistently decreasing in the CLS-KO mice, however, the quantifications do not look statistically significant. Can the authors confirm if these changes are or are not significant and adjust the text accordingly?

      The reviewer is correct. Abundances of I+III2+IV supercomplexes are decreased in CLS-LKO mice compared to control mice when quantifying with supercomplex antibody cocktail or with UQCRSF1 (complex III subunit) antibody, but not with complex I antibodies. The discrepancy for these results are not entirely clear but it’s likely a combination of antibody sensitivity and a tricky nature to dissolve high molecular weight protein complexes.

      (15) The most compelling data to indicate electron leakage increasing upon CLS knockout is in Figures 7A-E. I would suggest the authors decrease their emphasis on the rearrangement of the supercomplexes and focus their discussion on the very compelling results of Figure 7.

      Thanks for this suggestion. We have modified our text.

      (16) Figure 7D shows that a major site of electron leak is from site II, and these results also fit with the profound succinate-induced respiration observed in earlier experiments. It would be nice if the authors could test the ability of cardiolipin to rescue these phenotypes, similar to the assay in Figure 5I. Assessing this rescue on the CoQ redox state would also strengthen the claims.

      Thank you for this comment. We are encouraged with your suggestions. We have thought about this quite extensively during the preparation of the manuscript but we refrained from making conclusive statements regarding complex II because the magnitude of the increase in electron leak is equally elevated at complex II and III. It’s true that CLS deletion increases succinate-induced respiration, but this might also be because succinate elicits the highest increase in respiration even in wildtype mice (see values in Figure 3K and L compared to other substrates). It would be intriguing to examine the influence of CLS deletion on complex II/III electron leak as well as succinate-induced respiration in tissues where succinate is not a preferred substrate. We have attempted cardiolipin rescue in SUV but unfortunately, we could not get this assay to work for site-specific electron leak measurements.

      (17) In Figure 7G-H, it would be nice to see a ratio of oxidized to reduced CoQ, in the CLS deletion mice and in human NASH livers, if samples are available.

      Thanks for this suggestion. Data shown (Figure 7- figure supplement 1P-S).

      (18) CoQH2 can also deliver electrons to complex II (via its reversal). Complex II shows a remarkable contribution to the electron leak phenotype (Figure 7D). Also, as the complex II monomer showed much larger changes in the native gels of Figure 6 than the complexes involving complex III. A more likely model is that oxidized CoQ accumulates in the CLS knockout model because of increased CoQH2 leak via complex II.

      Perhaps. We also thought about this but we are not sure if this fits with the observation that CLS deletion increases succinate-induced respiration, which suggests increased succinate to fumarate conversion, a notion that I am not sure can be congruent with increase CoQH2 reversal to complex II. Overall, I think we lack the tools or evidence to conclusively implicate whether CLS deletion primarily acts on complex II or III. Nevertheless, we appreciate the reviewer’s enthusiasm on these topics as we perform additional experiments on the mechanism of interactions between CL and the ETC.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      Cruz-Gonz´alez and colleagues draw on DNA methylation and paired genetic data from 621 participants (n=308 controls; n=313 participants with Alzheimer’s Disease). The authors generate a panel of epigenetic biomarkers of aging with a primary focus on the Horvath multi-tissue clock. The authors find weaker correlations between predicted epigenetic age and chronological age in subgroups with higher African ancestry than within a subgroup identified as White. The authors then examine genetic variation as a potential source for between-group differences in epigenetic clock performance. The authors draw on a large collection of publicly available methylation quantitative trait loci datasets and find evidence for substantial overlap between clock CpGs located within the Horvath clock and methQTLs. Going further, the authors show that methQTLs that overlap with Horvath clock CpGs show greater allelic variation in African ancestral groups pointing to a potential explanation for poorer clock performance within this group.

      Thank you for this summary.

      Strengths:

      This is an interesting dataset and an important research question. The authors cite issues of portability regarding polygenic risk scores as a motivation to examine between-group differences in the performance of a panel of epigenetic clocks. The authors benefit from a diverse cohort of individuals with paired genetic data and focus on a clinical phenotype, Alzheimer’s disease, of clear relevance for studies evaluating age-related biomarkers.

      Thank you.

      Weaknesses:

      While the authors tackle an important question using a diverse cohort the current manuscript is lacking some detail that may diminish the potential impact of this paper. For example:

      (1) Information on chronological ages across groups should be reported to ensure there are no systematic differences in ages or age ranges between groups (see point below).

      Thank you for pointing out this omission. The distributions are now presented in Supplementary Figure 1. While there is some variation in median age, the age ranges are similar across cohorts (median 73.1 to 79.3). The small differences do not explain the differences in accuracy between the cohorts, e.g., the median age of the African Americans (76.4) is lower than the median age for the White cohort (77.7).

      (2) The authors compare correlations between chronological age and epigenetic age in sub-groups within to correlations reported by Horvath (2013). Attempting to draw comparisons between these two datasets is problematic. The current study has a much smaller N (particularly for sub-group analyses) and has a more restricted age range (60-90yrs versus 0-100 yrs). Thus, is an alternative explanation simply that any weaker correlations observed in this study are driven by sample size and a restricted age range? Reporting the chronological ages (and ranges) across subgroups in the current study would help in this regard. Similarly, given the lack of association between AD status and epigenetic age (and very small effect in the white group), it may be of interest to examine the correlation between chronological age and epigenetic age in each group including the AD participants: would the between-group differences in correlations between chronological age and epigenetic be altered by increasing the sample size?

      Our conclusions about the reduced accuracy of the clock in admixed individuals are based on the comparison within the MAGENTA cohorts, not a comparison of MAGENTA to previously published studies. We find significantly reduced accuracy in the admixed cohorts compared to the White MAGENTA cohort. Further supporting this conclusion beyond he MAGENTA cohort, we analyzed three independent whole blood methylation datasets. Two focused on African American individuals—the Grady Trauma Project (n = 422) and the GENOA study (n = 1,394)—and one focused on White Swedish individuals (n = 729). As observed in MAGENTA, the Horvath clock had significantly lower accuracy for the African American cohorts (Figure 3 than for the White Swedish cohort.

      When comparing results across studies, the reviewer is correct that lower correlations are generally seen for older cohorts. Indeed, other studies applying the Horvath clock have seen similar correlations in older cohorts to those observed in MAGENTA (Marioni et al., 2015, Horvath 2013, and Shireby et al., 2020). We now also include the chronological age distributions of the cohorts in this study, along with their mean and standard deviations (Supplementary Figure 1). This shows that the distribution of chronological ages for White individuals is similar to the cohorts where the clocks did not perform as well. Finally, as suggested, we correlated chronological and epigenetic age with the inclusion of AD cases in each cohort for the Horvath clock. The significantly lower performance of the clock on Puerto Ricans and African Americans, relative to White individuals, remains even after including all individuals in each cohort. Thus, combining cases and controls did not qualitatively change the performance relationships for the African Americans and Puerto Ricans relative to the Whites (Supplementary Figure 3).

      (3) The correlation between chronological age and epigenetic age, while helpful is not the most informative estimate of accuracy. Median absolute error (and an analysis of MAE across subgroups) would be a helpful addition.

      We used correlation because it is commonly used to evaluate the performance of epigenetic age clocks, but we agree that other error quantification metrics provide a complementary perspective. We now include MAE and MSE comparisons across sub-groups in the revision (Supplementary Table 1). We find that across all accuracy metrics, the African American and Puerto Rican cohorts perform worse than the White and Peruvian cohorts. Interestingly, the Cubans show relatively high error despite a high correlation between predicted and chronological age. However, there are only 21 non-demented Cuban controls. In addition, we evaluated the same metrics in three replicate datasets (two African American cohorts and one for White Swedish individuals) and found the same patterns of lower accuracy across metrics in African ancestry individuals, albeit with some variation in accuracy between cohorts (Supplementary Table 2). Notably, as discussed above, this is not driven by differences in chronological age distributions: when we subset to older individuals (≥ 55 years old) in order to facilitate comparisons to MAGENTA study individuals, the median age for the White Swedish individuals (70 years old) is higher than that of the GENOA (62.7 years old) and Grady (58 years old) individuals. Despite the difference in median ages, the clock performs better on White Swedish individuals across all accuracy metrics than the African ancestry cohorts with younger individuals.

      (4) More information should be provided about how DNAm data were generated. Were samples from each ancestral group randomized across plates/slides to ensure ancestry and batch are not associated? How were batch effects considered? Given the relatively small sample sizes, it would be important to consider the impact of technical variation on measures of epigenetic age used in the current study. The use of principal Component-based versions of these clocks (Higgins Chen et al., 2023; Nature Aging https://doi.org/10.1038/s43587-022-00248-2) may help address concerns such concerns.

      Thank you for pointing out the need for additional context on data generation. We have added details to the Methods. All omics data from the MAGENTA study were generated using standard protocols that ensure minimal technical artifacts and batch effects. Samples were randomized across plates and chips to ensure that ancestry, age, and sex were not confounded with each batch. We also performed a principal components analysis of the normalized methylation data used as inputs for all MAGENTA analyses. We found that the samples did not stratify by sample plate, cohort, ethnicity, or ascertainment center along the principal components (Supplementary Figure 2).

      We also thank the reviewer for their suggestion to apply the principal component clock to account for potential technical variation. As outlined in the new section “Principal component versions of the methylation clocks also have lower age prediction accuracy for genetically admixed individuals,” using the principal component version of the Horvath clock did not result in consistent improvement in age prediction accuracy or generalization across MAGENTA cohorts (Supplementary Figures 4 and 5). The lower accuracy for age prediction in individuals with substantial African ancestry was present for the PC clock in the replication cohorts, just as in the MAGENTA cohorts (Supplementary Figure 6).

      (5) Marioni et al., (2015) found a very weak cross-sectional association between DNAm Age and cognitive function (r∼0.07) in a cohort of >900 participants. Given these effect sizes, I would not interpret the absence of an effect in the current study to reflect issues of portability of epigenetic biomarkers.

      We agree that previous links between DNAm Age and AD or cognitive function have been relatively small in magnitude. For example, the PhenoAge paper (Levine et al., 2018) and a study using the Horvath clock (Levine et al., 2015) found age acceleration of less than a year in AD patients relative to non-demented individuals. Similar results have also been observed in studies with smaller sample sizes (e.g., 700 for Levine et al. 2015 and 604 for Levine et al. 2018). Given these small effect sizes, we agree that accounting for statistical power is essential for interpretation of our results. We performed power calculations based on an effect of the size observed in previous studies (0.5 year acceleration). We have 86% power in the full MAGENTA data set to detect an effect of this size. Stratifying by cohorts, we have 75% power for the African Americans, 72% for the Puerto Ricans, 72% for the Whites, 65% for the Peruvians, and 47% for the Cubans. Thus, we believe we have high enough power that the consistent lack of association outside of the White cohort in MAGENTA is likely meaningful. Based on these calculations, there is only a 1% chance that we would not observe an effect in any of the other cohorts if the effect was present across cohorts. Nonetheless, we have added caveats about power and the small sample size to our suggestion that the reduced accuracy of the clocks contributes to the lack of AD association outside of Whites.

      (6) The methQTL analyses presented are suggestive of potential genetic influence on DNAm at some Horvath CpGs. Do authors see differences in DNAm across ancestral groups at these potentially affected CpGs? This seems to be a missing piece together (e.g., estimating the likely impact of methQTL on clock CpG DNAm).

      We agree. Thank you for this suggestion. We have added Figure 6 in the main text to address this gap. In short, we analyzed additional whole blood methylation data from inidividuals with African ancestry and found that a substantial proportion of the CpGs in methylation clocks are differentially methylated in African ancestry individuals relative to European ancestry individuals. In the case of the Horvath clock, we find that 84/353 (23.8%) of the clock CpGs are differentially methylated between ancestries. In parallel, we found that 56 of these differentially methylated clock CpGs are also affected by meQTL, many of which are at different frequencies between populations. We also investigated whether the meQTL-affected clock CpGs are associated with increased clock error in the MAGENTA individuals. We found 56 clock CpGs whose methylation levels associated with increased clock error, and 42 of these have at least one meQTL. Thus, while meQTL are not the only factor to affect the portability of methylation clocks across global populations, we suggest that they are a significant contributor, especially in the case of the Horvath clock.

      Reviewer #2 (Public review):

      Summary:

      This paper seeks to characterize the portability of methylation clocks across groups. Methylation clocks are trained to predict biological aging from DNA methylation but have largely been developed in datasets of individuals with primarily European ancestries. Given that genetic variation can influence DNA methylation, the authors hypothesize that methylation clocks might have reduced accuracy in non-European ancestries.

      Strengths:

      The authors evaluate five methylation clocks in 621 individuals from the MAGENTA study. This includes approximately 280 individuals sampled in Puerto Rico, Cuba, and Peru, as well as approximately 200 self-identified African American individuals sampled in the US. To understand how methylation clock accuracy varies with proportion of non-European ancestry, the authors inferred local ancestry for the Puerto Rican, Cuban, Peruvian, and African American cohorts. Overall, this paper presents solid evidence that methylation clocks have reduced accuracy in individuals with non-European ancestries, relative to individuals with primarily European ancestries. This should be of great interest to those researchers who seek to use methylation clocks as predictors of age-related, late-onset diseases and other health outcomes.

      Thank you for this summary.

      Weaknesses:

      One clear strength of this paper is the ability to do more sophisticated analyses using the local ancestry calls for the MAGENTA study. It would be valuable to capitalize on this strength and assess portability across the genetic ancestry spectrum, as was recently advocated by Ding et al. in Nature (2023). For example, the authors could regress non-European local ancestry fraction on measures of prediction accuracy. This could paint a clearer picture of the relationship between genetic ancestry and clock accuracy, compared to looking at overall correlations within each cohort.

      Thank you for this suggestion. To model portability across genetic ancestry as a spectrum, we regressed the Horvath clock error on the proportions of African ancestry in the genomes of the MAGENTA individuals, adjusting for chronological age. The proportion of African ancestry is significantly associated with increased Horvath clock error (p = 0.039), with the clock making less accurate age predictions by 1.46 years for individuals with full African ancestry compared to no African ancestry. We have added this new analysis to the Results.

      The authors present two possible reasons that methylation clocks might have reduced accuracy in individuals with non-European ancestries: genetic variants disrupting methylation sites (i.e., ”disruptive variants”) and genetic variants influencing methylation sites (i.e., meQTLs). The authors conclude disruptive variants do not contribute to poor methylation clock portability, but the evidence in support of this conclusion is incomplete. The site frequency spectrum of disruptive variants in Figure 4 is estimated from all gnomAD individuals, and gnomAD is comprised of primarily European individuals. Thus, the observation that disruptive variants are generally rare in gnomAD does not rule them out as a source of poor clock portability in admixed individuals with non-European ancestries.

      In the revision, we now additionally report ancestry-specific allele frequencies to demonstrate the rarity of CpGclock disrupting variants (Supplementary Figure 9). The global allele frequencies were so low that even if they all occurred in individuals of non-European ancestries, they would still be extremely rare.

      It is also unclear to what extent meQTLs impact methylation clock portability. The authors find that the frequency of meQTLs is higher in African ancestry populations, but this could reflect the fact that some of the analyzed meQTLs were ascertained in African Americans. The number of meQTL-affected methylation sites also varies widely between clocks, ranging from 6 to 271; thus, meQTLs likely impact the portability of different clocks in different ways. Overall, the paper would benefit from a more quantitative assessment of the extent to which meQTLs influence clock portability.

      We agree that the meQTL likely influence the clocks in different ways and that the ascertainment of the meQTLs in different populations makes direct comparisons challenging. To more directly link meQTL to clock performance, we identified 56 Horvath clock CpG sites whose methylation levels significantly associate with increased clock error in the MAGENTA study individuals. Of these, 42 (75%) are affected by an meQTL, including nine that are affected by an African ancestry-differentiated meQTL. As such, meQTL, and specifically meQTL that were likely not present in the training data of the Horvath clock, associated with both the methylation of CpG sites and clock error. However, as the reviewer suggests, determining causality among these factors is challenging. Given our incomplete knowledge of meQTL in different ancestries, we have added caveats to our conclusions about the effect of meQTL on clock portability.

      The paper implies that methylation clocks have an inferior ability to predict AD risk in admixed populations relative to white individuals, but the difference between white AD patients and controls is not significant when correcting for multiple testing. This nuance should be made more explicit.

      We agree that the signal is not strong in the white cohort; however, it is similar in magnitude to previous studies. As outlined in response to Reviewer 1’s Point 5, we have now added power calculations that indicate reasonable power (≥72%) to detect small effect sizes (0.5 year increase) in the white, Puerto Rican and African American cohorts. We now interpret the AD association tests in the context of these power calculations and multiple testing correction.

      Finally, this paper overlooks the possibility that environmental exposures co-vary with genetic ancestry and play a role in decreasing the accuracy of methylation clocks in genetically admixed individuals. Quantifying the impact of environmental factors is almost certainly outside of the scope of this paper. However, it is worth acknowledging the role of environmental factors to provide the field with a more comprehensive overview of factors influencing methylation clock portability. It is also essential to avoid the assumption that correlations with genetic ancestry necessarily arise from genetic causes.

      We entirely agree and have now clarified the scope of our analyses and importance of environmental factors in the revision. We intersected clock CpGs with enviromental-factor-associated CpGs from multiple epigenome-wide association studies (EWAS) and found overlaps that suggest an environemtnal contribution to differences in clock CpG methylation. However, given the lack of environmental data on the MAGENTA study individuals, as well as the lack of datasets for replication, we cannnot directly compare the environmental and genetic contributions to clock accuracy. Nevertheless, the new analyses in the revision highlight the contribution of both genetic and environmental factors to lack of portability for certain methylation clocks.

      Reviewer #2 (Recommendations for the authors):

      (1) Line 64: An association between methylation patterns and genetic ancestry does not presuppose that meQTLs vary in frequency between genetic ancestries; environmental factors could also play a role. It would be nice to comment on this further in the Introduction.

      We agree that environmental factors likely play a role in the decrease in methylation clock performance in admixed populations. We have added text highlighting this in the revised Discussion. Regarding meQTL, we agree that associations between methylation patterns and genetic ancestry do not necessarily imply that meQTL will vary in frequency between genetic ancestries. However, our new analyses in the revision find African-ancestry differentiated meQTL that associate with Horvath clock CpG methylation levels and overall clock error (Figure 6E-F and Supplementary Figure 13).

      (2) Line 116 implies Puerto Ricans have “substantial amounts of African ancestry” but the median ancestry is 15% (which is not much more than the Peruvian and Cuban cohorts).

      Thank you for pointing this out. We have clarified this statement in the text. While the median proportion of African ancestry in Puerto Ricans is 15% (vs. 6% and 2% for the Peruvian and Cuban individuals in MAGENTA), there are many individuals with substantially higher African ancestry. The upper quartile is >25% and several Puerto Ricans have >50% African ancestry.

      (3) In Figure 2B, Puerto Ricans have worse accuracy than Peruvians but a higher proportion of inferred CEU ancestry, which is interesting and defies intuition - is there any hypothesis for why this might be the case?

      In light of our new meQTL analyses, we hypothesize that the African ancestry differentiated meQTL that affect Horvath clock CpGs drive the increase in clock error for these individuals, despite having more European ancestry across their genome. Given that the Peruvians (and Cubans, for that matter) hold very little African ancestry, and also very few of the African-differentiated meQTL, this could explain some of the large difference in clock errors for the cohorts.

      (4) Figure 2C would be improved with confidence intervals.

      We thank the reviewer for this suggestion and have added confidence intervals for Figure 2C.

      (5) It’s interesting that the correlation with Cubans is positive in Figure 3B (for one clock, significantly so). Is there any rationale for this?

      We noticed this as well, but have not been able to come to a definitive conclusion. It is possible that environmental factors contribute. However, the Cuban cohort is the smallest in MAGENTA (22 cases and 21 controls) and the none of the differences are statistically significant, so more investigation in a large cohort is required.

      (6) Line 231: Which population(s) is allele frequency estimated in?

      This is the global frequency reported in gnomAD, which is calculated across all populations in gnomAD v3.0. As noted above, we now also report allele frequencies by gnomAD population (Supplementary Figure 9).

      (7) Were the meQTLs pruned? How many independent variants are there per methylation site? It would be nice to see a distribution for the sites in the Horvath clock.

      We now report the distribution of meQTL across clock CpG sites. The mean number of variants is 108; the median is 36; and the maximum is 1,699. We have now included a plot of the distribution for all 271 (out of 353) Horvath clock CpG sites (Supplementary Figure 14). We did not perform any pruning in these initial results for several reasons. First, we sought to demonstrate the great potential for meQTL to influence these CpGs and to compare the distributions of these common meQTL across populations (based on gnomAD data). Second, identifying the causal variant or variants is challenging. Given that many of these meQTLs likely reflect redundant signals, for the new analyses of African-differentiated meQTL, we restrict to a single variant per clock CpG site. We focus on the variant with the greatest absolute beta, as reported by the original meQTL study from which the variant originates.

      (8) Figure 5C might benefit from a geom density rather than overlapping bar plots; the trends are hard to see.

      We appreciate the reviwer’s suggestion and have now reworked the figure and based it on just the density curves so that readers may better appreciate the differences in allele frequencies.

      (9) Several figures would be more legible with larger font sizes.

      We appreciate this recommendations and have made the font sizes for all plots larger and more legible.

      Reviewer #3 (Public review):

      This manuscript examines the accuracy of DNA methylation-based epigenetic clocks across multiple cohorts of varying genetic ancestry. The authors find that clocks were generally less accurate at predicting age in cohorts with large proportions of non-European (especially African) ancestry, compared to cohorts with high European ancestry proportions. They suggest that some of this effect might be explained by meQTLs that occur near CpG sites included in clocks, because these variants may be at higher frequencies (or at least different frequencies) in cohorts with high proportions of non-European ancestry relative to the training set. They also provide discussions of potential paths forward to alleviate bias and improve portability for future clock algorithms.

      The topic is timely due to the increasing popularity of DNA methylation-based clocks and the acknowledgment that many algorithms (e.g., polygenic risk scores) lack portability when applied to cohorts that substantially differ in ancestry or other characteristics from the training set. This has been discussed to some degree for DNA methylationbased clocks, but could of course use more discussion and empirical attention which the authors nicely provide using an impressive and diverse collection of data.

      Thank you for this summary.

      The manuscript is clear and well-written, however, some key background was missing (e.g., what we know already about the ancestry composition of clock training sets) and most importantly several analyses would benefit from being taken one step further. For example, the main argument of the paper is that ancestry impacts clock predictions, but this is determined by subsetting the data by recruitment cohort rather than analyzing ancestry as a continuous variable. Extending some of the analyses could really help the authors nail down their hypothesized sources of lack of portability, which is critical for making recommendations to the community and understanding the best paths forward.

      Thank you for this suggestion. As noted in our response to Reviewer 2’s Point 1, we have analyzed ancestry as a continuous variable and found that the proportion of African ancestry in the genomes of the MAGENTA individuals significantly associates with increased difference in chronological and predicted age, even after controlling for chronological age (1.46 years more error for 100% vs. 0% African ancestry; p = 0.039). As outlined below, we have also added details on the training of previous clocks and the important additional previous work highlighted by the Reviewer.

      Reviewer #3 (Recommendations for the authors):

      Major comments

      There is previous literature addressing who is in the training set for methylation clocks. To my knowledge, this work has been primarily led by Nancy Krieger. It would be a valuable addition to discuss her work (and any similar work by other investigations) in the introduction. In other words, what do we currently know about the degree of bias in the training sets for methylation-based clocks? The assumption of the introduction is that the training sets are overwhelmingly European ancestry (which I assume is true) but I think some quantitative information about this would be helpful for understanding the source and magnitude of the problem.

      We thank the reviewer for bringing the work of Dr. Nancy Krieger to our attention. It directly supports the rationale for this study: the sociodemographic characteristics of the individuals used to train these clocks are poorly reported, limited to outdated population descriptors (for example, the use of “Caucasians” to describe some of the individuals used to train the Horvath and the Hannum clocks) or race and ethnicity labels. Moreover, where labels are available for training individuals, they tend to underrepresent the individuals of diverse backgrounds, as in the Horvath clock. We have incorporated Dr. Krieger’s work into the Introduction, including details of how this supports the rationale and purpose of our study.

      Related to the above comment, there has been pretty extensive previous work on the effects of race and ethnicity on epigenetic clock estimates (e.g., https://genomebiology.biomedcentral.com/articles/10.1186/s13059-016-1030-0), and that seems like it could be more explicitly weaved into the introduction and discussion.

      We thank the reviewer for highlighting this relevant article. We have added discussion of it into the Introduction. Several factors make direct comparison with our results challenging. First, the grouping of individuals based on race and ethnicity without consideration of genetic ancestry complicates comparisons. Race and ethnicity commonly do not match genetic ancestry components (see Gouveia et al., 2025 https://www.cell.com/ajhg/fulltext/S00029297(25)00173-9). Second, the study reports differences in epigenetic age accelerations (intrinsic and extrinsic) in individuals from various race and ethnic groups. It does not directly evaluate the accuracy of the epigenetic age predictions in these groups. Thus, it is challenging to interpret whether the differences in acceleration are driven by biological factors or biases in the performance of the clocks themselves.

      The main analysis that felt like it was missing was asking whether the age deviations are larger for individuals with greater proportions of African ancestry. The authors have the ability to analyze ancestry as a continuous variable, but instead performed analyses in various a priori subsets of the data; the subsets do have average differences in ancestry, but also there is heterogeneity within groups. Given that the authors calculated admixture proportions already, it seems like a missed opportunity not to use these estimates. This would also sidestep the issue of the problematic labels applied to the subsets, which mix ancestry, nationality, and race terms (note that I thought the legacy reasons why these labels are used were well-explained, but they are nevertheless problematic for biological explanations that center on ancestry/genetic information as the driver of bias).

      We appreciate the reviewer’s suggestion to investigate clock accuracy in the context of African ancestry proportions. As noted in the response to Reviewer 2’s Point 1, we modeled the clock error as a function of the fraction of African ancestry of each individual, adjusting for an individual’s chronological age. The proportion of African ancestry is significantly associated with increased Horvath clock error (p = 0.039), with the clock estimated to give less accurate age predictions by 1.46 years for individuals with 100% African ancestry compared to no African ancestry. We now report this in the Results.

      Another missed analysis opportunity occurs in lines 259-261, where the authors state “Thus, the clock with the largest decrease in performance in admixed cohorts (in terms of predicting chronological age and identifying age acceleration in AD) has the most and largest fraction of meQTLs influencing its CpGs.” This is another place where the authors make generalizations about a given cohort based on average ancestry rather than testing the claim empirically on an individual basis (e.g., by examining the number of meQTL variants a given individual is heterozygous for or has the non-European allele for).

      We thank the reviewer for this comment. This feedback motivated us to evaluate the relationship between differences in meQTL frequencies and methylation clock error. We found differences in meQTL frequency in the MAGENTA individuals, specifically many of the clock CpG affecting meQTL are most common in the African American cohort, consistent with our theory (Figure 6E,F). Nonetheless, there are 84 Horvath clock CpGs (24%) that are differentially methylated in AFR individuals, and 56 of these are affected by an meQTL, including 11 that are affected by an African ancestry-differentiated meQTL (Figure 6G). Finally, we find that 42 Horvath clock CpG sites in MAGENTA individuals with methylation levels that are significantly associated with increased clock error, and that are also affected by an meQTL (Figure 6B). However, at the individual level we do not find a clear relationship between the number of meQTL or ancestry-differentiated meQTL and methylation clock error. In light of these data, we have reframed our conclusions to state that meQTL likely contribute to clock error, while also being clear that they are not the sole cause.

      Can the authors explain or offer an investigation into why predicted age is often better in Cubans than Whites? They gave much attention to the opposite effect (of similar magnitude) in African Americans and Puerto Ricans but didn’t really discuss the surprisingly accurate prediction in Cubans.

      We did not focus on the results in the Cuban cohorts for several reasons. As discussed in response to Reviewer 2’s comment, the Cuban cohort had the smallest sample size (22 cases and 21 controls). Thus, while the correlation between methylation age and chronological age is similar to Whites, and in a few cases higher, the differences were not statistically significant. Second, looking at other error metrics, like mean absolute error, the clocks are comparatively less accurate in Cubans than on the White cohort (Supplementary Table 2). Finally, the clocks consistently find that Cubans with AD have lower predicted age than controls, though this is only significant for the ZhangEN clock. However, given these inconsisencies and the very small sample size, we caution against over-interpretation of these results. We clarify this in the manuscript and suggest that more work is needed on larger Cuban cohorts before any clear conclusions can be made.

      I was not a conceptual fan of the ensemble clock. The clocks are trained on very different things (e.g., chronological age versus clinical biomarkers) and are designed to capture different aspects of biology. Without more validation and motivation, I don’t think it makes sense to average values that are not designed to measure the same thing.

      We agree that combining the first and second-generation clocks for the task of age prediction is not sensible. However, for AD risk stratification, combining values from multiple clocks that capture different aspects of biology and aging could be beneficial. As mentioned in the main text, we took inspiration from approaches in polygenic risk scores, as well as the broader machine learning field, where ensembling often makes for better predictors. Nonetheless, consistent with the Reviewer’s intuition, we do not see improvement here.

      Minor comments

      (1) Typo in line 91.

      Thank you for bringing this to our attention. Fixed.

      (2) Lines 111-115, sample sizes would be helpful.

      We have added the sample sizes of the non-demented controls that were used to calculate these correlations in each cohort.

      (3) Line 137-138, the correlation stats would be helpful here. This is a common issue throughout the paper, more in-text statistics would help readers to evaluate the authors’ claims. For example, lines 249-251 as well. The authors refer the reader to Figure 5C, which itself has no statistics, this has two plots so it’s unclear which the authors are putting forward as the primary evidence.

      We have added more statistical details in the text and figures to address this comment. In this instance, we have removed the referenced figure.

      (4) Lines 258 and 261, I believe the authors report the same result in both these lines.

      Thank you for pointing out this lack of clarity. These lines report different, but related, results about the frequency of clock-affecting meQTL in different ancestral contexts. The first reports the frequency of clock CpGaffecting meQTL in individuals of African ancestry across all of gnomAD. The second result gives the frequency of those meQTL in different local ancestry backgrounds in admixed individuals. This is distinction is relevant since admixed individuals’ genomes are mosaics of multiple genetic ancestries. As such, a genetic variant might be present in haplotype whose ancestry is not in line with expectations based on global ancestry (e.g., an African American individual inherits a genetic variant within a European ancestry block). This local ancestry difference could modify the effect of the variant or obscure causal variants. Given the potential for confusion and similar results considering global and local ancestry context in this case, we have focused on the first result in the Main Text.

      (5) Somewhere, it would be helpful to provide the distribution/range of ages broken by cohort. Similarly, I didn’t see the breakdown of AD versus control cases within each cohort. Both of these features will impact power within a given cohort for certain analyses.

      We have added the distribution of ages by cohort in Supplementary Figure 1. Table 1 provides a breakdown of cases versus controls for each of the cohorts in the MAGENTA study.

      (6) Figure 3 is pretty hard to read. It would also be helpful if the authors put the white cohort in Figure 3A as a ’baseline’ comparison, as they use this as the baseline comparison in the text.

      We have made these changes to the figure and used larger text overall.

      (7) The various acronyms in the labels in Figure 5 are not explained. For Figure 5C - this is over-plotted and therefore hard to see.

      We have added the full population descriptors from gnomAD to the boxplots showing allele frequencies (Figure 6E). In addition, what used to be Figure 5C has been simplified and moved to Supplementary Figure 12.

      (8) The authors correct for cell type heterogeneity, which is known to vary across populations and can impact clock estimates. However, as far as I can tell, the cell type proportion estimates are coming from the DNA methylation data. The deconvolution algorithms for cell type proportions also have the same problem as the clocks of being trained on a very specific subset of human genetic and environmental diversity. Do the authors have any empirically derived estimates of cell type heterogeneity to sanity-check these deconvolution estimates? At the very least, it would be helpful to acknowledge this limitation.

      We thank the reviewer for commenting on this. There are no empirically derived estimates of cell type counts for the samples in the MAGENTA study. This is an inherent limitation of our study, and we have included text to make note of this.

      (9) There are very different sample sizes for each group, did the authors consider that their null results for the AD analyses in different cohorts are just a lack of power? This could be evaluated with power analyses or by comparing against sample sizes from similar studies in the literature.

      We agree that this is an important analysis and have added it to the manuscript. Given these small effect sizes, accounting for statistical power is essential for interpretation of our results. We performed power calculations based on an effect of the size observed in previous studies (0.5 year acceleration). Considering the full study, we have 86% power to detect an effect of this size. Stratifying by cohorts, we have 75% power for the African Americans, 72% for the Puerto Ricans, 72% for the Whites, 65% for the Peruvians, and 47% for the Cubans. Thus, we have high enough power that the consistent lack of association observed outside of the White cohort in MAGENTA is likely meaningful. Based on these calculations, there is only a 1% chance that we would not observe an effect in any of the other cohorts if the effect was present across cohorts. Nonetheless, we have added caveats about power and the small sample size to our suggestion that the reduced accuracy of the clocks contributes to the lack of association outside of Whites.

      (10) There has been a fair amount of discussion recently that single CpG-based clocks are much more variable than clocks that combine information across CpG sites, either using PC-based or window-based approaches. For example, the PC clock R package from the Levine Lab (https://github.com/MorganLevineLab/PC-Clocks) is very easily implemented and generally gives much less variable age estimations than site-level clocks. It would be nice to consider integrating or discussing these later-generation clocks as ways to improve clock performance in diverse human groups.

      We thank the reviewer for their suggestion to apply the principal component clock to account for potential technical variation. As outlined in the new section “Principal component versions of the methylation clocks also have lower age prediction accuracy for genetically admixed individuals,” using the principal component version of the Horvath clock did not result in consistent improvement in age prediction accuracy or generalization across MAGENTA cohorts (Supplementary Figures 4 and 5). The lower accuracy for age prediction in individuals with substantial African ancestry were present for the PC clock in the replication cohorts, just as in the MAGENTA cohorts (Supplementary Figure 6)

    1. Author response:

      The following is the authors’ response to the original reviews.

      Thank you very much for handling our revised manuscript and for the careful and constructive comments from the reviewers. We are grateful for the detailed feedback, which has helped us improve both the experimental presentation and the framing of the study. In response to the comments, we have substantially revised the manuscript, updated the figures and supplementary figures, and clarified several points in the text. We have also added new experimental analyses, which were essential to strengthen the manuscript.

      We would like to highlight the major changes in the revised version:

      Added the late phenotype analysis of the ror2 mutant, including loss of nasal and maxillary barbels and altered adult jaw morphology by microCT, strengthening the disease-model relevance.

      Added new data on a further target locus (wls) showing 46 bp attP insertion by PEn and comparison with HDR-mediated knock-in at the same site.

      Expanded the analysis of insertion performance at adgrf3b and clarified comparison with previously reported PE2 data.

      Added the analysis of HDR-mediated knock-in and prime editing substitution to generate ror2 W722X allele.

      Added comparative off-target analysis for PE2, PEn and HDR at three predicted off-target sites for the ror2 target.

      Resolved the cloning/NGS inconsistency for ror2 by increasing clone analysis

      We have also moderated several statements in the manuscript, for example, that editing efficiency is locus- and edit-dependent, and that broader comparison of germline transmission efficiencies between prime editing systems will require future work.

      A few reviewer suggestions would have required substantial additional experimental work that is technically demanding and beyond the immediate scope of the present methods-focused resubmission, for example, a direct side-by-side germline comparison of PE2 and PEn across several loci, or systematic cost benchmarking against HDR across multiple edit classes. Rather than overstate these points, we have acknowledged these limitations directly in the revised manuscript and narrowed our claims accordingly.

      Public Reviews:

      Reviewer #1 (Public review):

      From the work presented, it is unclear how prime editing could be used to transiently model human pathogenic variants, given the low frequency of precision edits in somatic tissue, or to isolate stable germline alleles of variants that are potentially dominant negative or gain-of-function in nature. Without a direct comparison with CRISPR/Cas9 nuclease HDR-based methods that use oligonucleotide templates to introduce edits, the advantage of prime editing is unclear. A cost comparison between prime editing and HDR methods would also be of interest, particularly for integration of longer DNA sequences

      We thank the reviewer for this important comment. In response, we added a direct comparison between PEn-mediated editing and HDR-mediated knock-in at the ror2 locus and the wls locus using insertion of a 46 bp attP sequence. This new dataset shows that PEn can achieve programmed insertion at a higher efficiency in ror2 and comparable efficiency in wls to HDR at the same target site, thereby providing a more direct benchmark within zebrafish embryos. We also revised the Discussion to better position prime editing as a practical donor DNA-free approach rather than as a universally superior method. We agree that a formal cost comparison would be informative; however, such an analysis would depend strongly on locus, edit size, optimisation burden, and local reagent production pipelines, and we believe this is beyond the scope of the present manuscript. Instead, we now discuss these practical considerations more cautiously in the revised Discussion.

      (1) In Figure 3, the data indicate a significant increase in precise edits of the 3 bp TGA using PE2 RNP (11.5%) vs. PE2 mRNA (1.3%). At the adgrf3b locus, only PEn mRNA was tested for introducing the 3 bp and 12 bp insertions. The previous study testing PE2 for 3 and 12 bp insertions was mentioned, but the frequency was not listed, and the study wasn't cited (lines 204 - 207). A comparison of germline transmission rates using PE2 vs. PEn would support the conclusion that PEn allows precise integration of longer templates and recovery of germline integration alleles.

      We appreciate this point. We revised the adgrf3b section to include the relevant reference and explicitly state the previously reported PE2 frequencies, allowing clearer comparison with our PEn data. We added our own experimental data to compare PE2 and PEn with mRNA or RNP form in adgrf3b locus (Figure 3i and j). We also refined the wording of our conclusions so that we do not imply a direct germline comparison between PE2 and PEn where such data are not available. In the revised manuscript, we now state that our germline transmission results apply to PEn-mediated insertions in the loci tested here. A full side-by-side germline comparison between PE2 and PEn across multiple loci would indeed be valuable, but this would require substantial additional animal work and time and is beyond the scope of the present resubmission.

      (2) Figure 4 shows the results of introducing a TGA stop codon that is predicted to result in nonsense-mediated decay. Testing the ability to also isolate different substitution mutations in the germline would be useful information for identifying the most effective approach for generating human disease variant models.

      We agree that this would be useful. In the present study, we focused experimentally on establishing stable lines for the insertion-based edits, while the substitution experiments were used to compare PE2 and PEn performance in somatic editing at the crbn locus. We also tested the generation of ror2 W722X allele by prime editing substitution (Supplementary Figure 3). We have therefore revised the manuscript to clarify the scope of the disease-modelling claim and now state more explicitly that our data support the generation of disease-relevant alleles in cases where short, programmed substitutions or insertions are sufficient.

      A comparison with the prime editing variant knock-in frequencies reported in the recent publication by Vanhooydonck et al., 2025, Lab Animal should be included in the Discussion.

      We have added this study to the revised manuscript and now discuss our findings in relation to the frequencies reported by Vanhooydonck et al. (2025).

      Reviewer #2 (Public review):

      The comparative analysis between PE2 and PEn systems suffers from limited evidentiary support. The comparison relies on single loci for substitutions (crbn) and insertions (ror2), raising concerns about generalizability. Additional validation across multiple loci is necessary to support broad conclusions about PE2/PEn performance

      We appreciate this concern. To strengthen the manuscript, we added new experimental data at an additional target locus, wls, where we tested insertion of a 46 bp attP sequence and compared PEn with HDR-mediated knock-in. We also included the adgrf3b insertion data more prominently. At the same time, we revised the wording throughout the manuscript so that our conclusions are more carefully limited to the loci tested here.

      Reviewer #3 (Public review):

      (1) The logic for introducing two nucleotide changes (at +3 and +10) to change a single amino acid (I378) should be explicitly explained in the main body of the manuscript. It is indeed self-explanatory when looking at Supplementary Figure 1. One way of doing it could be to include Supplementary Figure 1a in Figure 1.

      We thank the reviewer for pointing this out. We have now explained this directly in the main text. Specifically, we state that one nucleotide change introduces the desired missense mutation, whereas the second was included to reduce potential pegRNA misfolding caused by complementarity between the spacer and the PBS/RT template region.

      (2) It is not clear why a 3-nucleotide insertion was used to generate W722X. The human W720X is a single-nucleotide polymorphism, and it should be possible to make a corresponding zebrafish mutant by introducing two nucleotide changes.…

      We agree that this point and have now explained in the main text that the 3 bp stop-codon insertion was chosen as a proof-of-principle strategy for generating a precisely truncated protein through programmed insertion, a type of edit that can be broadly applied to target loci. We also tested the generation of ror2 W722X allele by prime editing substitution (Supplementary Figure 3). We also clarify that prime editing substitution was tested separately here.

      (3) Lines 137-138: T7 Endonuclease assay used in Figure 2d detects all polymorphisms, both precise changes and indels. Thus, if this assay were performed on embryos shown in Figure 1c-d, the overall percentage of modified alleles would be similarly higher for PEn over PE2 (add up precise prime edits and indels). The conclusion in the last sentence of the paragraph is, therefore, incorrect, I believe.

      We agreed with this point and revised the sentence accordingly. The text now states that no obvious cleavage was observed with the PE2/pegRNA condition, suggesting fewer editing events compared with PEn, rather than implying greater precision from the T7E1 result alone.

      (4) Use of terminology. "Germline transmission" is typically used to refer to the fraction of F0s transmitting desired changes (or transgenes) to their progeny, while "germline mosaicism" refers to the fraction of F1s with the desired change in the progeny of a given F0. "Germline transmission" in line 217 should be replaced with "germline mosaicism".

      We have replaced the terminology accordingly in the revised manuscript.

      (5) Lines 253-255: The fraction of injected embryos that had mosaic nuclear expression of GFP, indicative of NLS insertion, should be clarified. It should also be clarified whether embryos positive for nuclear GFP were preselected for amplicon sequencing and germline transmission analyses. This is extremely important for extrapolation to scenarios like epitope tagging, where preselection is not possible.

      We agree and have clarified this in the revised manuscript. We now state the fraction of injected embryos showing mosaic nuclear GFP expression, and we explicitly note that embryos were not preselected prior to sequencing or founder analysis. We further explain that preselection was not practical because the transgene is multicopy and individual fibres showed variable ratios of nuclear to cytoplasmic GFP, which made reliable scoring difficult.

      (6) Statistical analyses. It would be helpful to clarify why different statistical tests are sometimes used to assess seemingly very similar datasets (Figures 1c, 1d, 2b, 2c, 2f).

      We have clarified this in the Materials and Methods section and now state that the choice of statistical test depended on the normality and variance structure of the experimental data.

      (7) Discussion. Since authors suggest that PEn might be especially beneficial for insertion of additional sequences, it is important to stress locus-to-locus variability of success. While the precise +3 insertion was indeed tremendously efficient at both tested loci (ror2 and adgrf3b), +12 addition into adgrf3b was over 10 times less efficient. In contrast, +30 into smyhc:GFP using the shorter pegRNA was highly efficient again. Longer pegRNA did not work nearly as well. As dangerous as it is to extrapolate from small datasets, perhaps these observations indicate that optimization of RT template and PBS may be needed for each new locus in order to significantly outperform oligonucleotide-mediated HDR? If so, would the cost of ordering several pegRNAs and the effort needed to compare them factor in when deciding which method to use?

      We fully agree and have substantially revised the discussion to reflect this point. We now emphasise more clearly that editing efficiency is locus- and edit-dependent and likely influenced not only by insertion length but also by spacer sequence and pegRNA complexity. We cite the relevant literature on prime editing determinants and discuss that locus-specific optimisation may be required. We also softened our concluding claims so that the manuscript presents PEn as a practical donor DNA-free approach rather than as a universally high-efficiency solution.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Because this is a genome editing methods paper, including frequency or percentages of somatic and germline editing in the abstract, in comparison to previously published studies, it would be useful information for the intended audience

      We agree and revised the abstract to include concrete editing frequencies. We now indicate the strongest insertion efficiencies observed. We also retained the statement that edited alleles were transmitted to the next generation.

      Reviewer #2 (Recommendations for the authors):

      (2) Please include additional loci for substitutions and insertions to strengthen conclusions about PE2/PEn efficiencies.

      In response, we added further substitution data at the ror2 (Suppl. Data 3) and insertion at the wls locus (Suppl. Data 6) and strengthened the presentation of the adgrf3b insertion data: first, by adding new locus data where feasible; and second, by narrowing the wording of our conclusions so that they are explicitly limited to the loci tested here.

      (3) Please provide direct comparisons between zebrafish ror2 W722X phenotypes and human Robinow syndrome symptoms to support disease modeling claims.

      We addressed this by adding analysis of the late ror2 phenotype. In the revised manuscript, zygotic and maternal-zygotic mutants are reported to lack nasal and maxillary barbels, and one-year-old mutants show altered jaw morphology with a less protrusive lower jaw (Figure 4).

      (4) The substitution of two nucleotides (+3 G→C and +10 A→G) to target residue I378 of crbn is not justified. It is unclear why two substitutions were required to model thalidomide sensitivity or validate editing efficiency. Please explain why dual nucleotide substitutions were necessary in the crbn experiments and whether single substitutions would suffice.

      We now explain in the main text that the second substitution was introduced to reduce potential inhibitory intramolecular interactions within the pegRNA, while the primary substitution generated the intended amino-acid change. This clarification is now stated explicitly in the Results.

      (5) The reported 10.3% precise editing efficiency for PEn/pegRNA at ror2 conflicts with Supplementary Figure 2, where none of the 20 clones from PEn/pegRNA showed precise edits, while one clone from PEn/springRNA did. Please address the inconsistency between NGS and cloning results at ror2, possibly by increasing sample size or reanalyzing sequencing data.

      We addressed this directly by repeating and expanding the clone analysis. The revised Supplementary Figure 2 now includes the updated clone dataset, and the result is in much better agreement with the NGS-based frequency estimates.

      (6) Figure 3d highlights edits from PEn/springRNA but omits PEn/pegRNA results, despite the latter being described as superior. This creates ambiguity about the relative performance of pegRNA vs. springRNA. Please include PEn/pegRNA results in Figure 3d to fairly represent pegRNA performance.

      We agree. We therefore revised Figure 3e so that it now includes alignment data for PE2/pegRNA, PEn/pegRNA and PEn/springRNA, allowing more direct visual comparison of the editing outcomes.

      (7) The study does not specify the version of PEn used, or introduce some background of PE2 and springRNA. Comparisons to prior PE work in zebrafish, base editing, or HDR efficiencies are absent, obscuring the novelty of this approach. Please specify the PEn variant used, describe springRNA/PE2 structures, and compare results to prior zebrafish PE studies, BE, and HDR efficiencies for similar edits, contextualizing where PE2/PEn offers unique advantages.

      We thank the editors for this helpful suggestion. We have clarified the PEn and PE2 systems in the manuscript, specified the nuclease-based PEn used, and improved the background text introducing these editing strategies. We added the data to directly compare prime editing and HDR in the ror2 locus (Figure 3). We also expanded the Discussion to place the current findings in the context of prior zebrafish prime editing, HDR-based knock-in and base-editing work. We did not test all alternative systems experimentally in the current study, but we now discuss their relevance and clearly define the specific contribution of the present work.

      (8) The manuscript does not explore advanced PE variants (e.g., PE3, PEmax), codon optimization, or scaffold modifications to improve efficiency. Please discuss whether codon optimization, PE3/PEmax systems, or pegRNA modifications were tested or could improve outcomes.

      We agree that this should be discussed and we added recent work on zebrafish prime editing optimisation, codon optimisation, pegRNA engineering and related advances to the discussion, and explain that these are promising avenues for improving efficiency in future studies.

      (9) No data compares the off-target effects of PE2 and PEn, a critical consideration for evaluating specificity and safety. Please perform comparative off-target analyses for PE2 and PEn to assess specificity.

      In response, we performed comparative off-target analysis for the ror2 target and analysed three predicted off-target sites. These data are now included in Supplementary Figure 3 and show no significant increase in non-specific editing for the prime editing conditions tested.