10,000 Matching Annotations
  1. Jun 2026
    1. eLife Assessment

      This paper demonstrates that a genetic code expansion to tag two amyotrophic lateral sclerosis (ALS) proteins associated with stress granules is useful in an experimental context. The data are solid and demonstrate the feasibility of using ANAP-fluorescence for live cell imaging.

    2. Reviewer #1 (Public review):

      Summary:

      The authors utilize genetic code expansion to tag TDP-43 and G3BP1, and evaluate this protein tagging system (ANAP) compared to antibodies and evaluate protein trafficking and stress granule formation in response to stress with sodium arsenite treatment. They find similar staining to antibodies in HeLa cells, mouse embryonic stem cells and primary mouse cortical neurons. By incorporating the intrinsically fluorescent noncanonical amino acid Anap at carefully selected sites, the authors enable live-cell and neuronal visualization of protein localization, stress-induced redistribution, and dynamic behavior without the structural and functional compromises often associated with large fluorescent protein tags. The work provides technical framework that will be useful for live imaging of tagged proteins.

      Strengths:

      A key strength is the demonstration of the specificity of the Anap fluorescence signal through appropriate controls and the agreement between Anap labeling and antibody-based detection across multiple cell types, including primary neurons. The ability to visualize stress-induced redistribution of both G3BP1 and TDP 43 in living cells highlights the practical value of this approach.

      The functional validation of TDP 43-Anap is compelling. The rescue of both cell viability and RNA splicing defects in TDP 43 knockout models provides evidence that Anap incorporation preserves core protein functions. This is important, as functional disruption is a central concern for any alternative tagging strategy applied to aggregation-prone or RNA-binding proteins.

      Weaknesses:

      While some inherent limitations of genetic code expansion remain (e.g., variable amber suppression efficiency and the inability to directly assess endogenous protein behavior), these are acknowledged and discussed appropriately. Importantly, these limitations do not undermine the central contributions of the study.

    3. Reviewer #2 (Public review):

      In this manuscript, Chen and colleagues describe a novel means of labeling two RNA binding proteins, G3BP1 and TDP-43, using genetic code expansion. Overexpressed constructs that incorporate the intrinsically-fluorescent non-canonical amino acid Anap redistribute to cytoplasmic granules upon application of external stressors such as sodium arsenite. Similar labeling and redistribution of overexpressed G3BP1 and TDP-43 was observed in cultures of mouse primary neurons.

      Genetic code expansion and non-canonical amino acid labeling have many advantages over traditional fusion proteins for tracking protein redistribution in living cells. The authors show that they are able to label exogenous G3BP1 and TDP-43 with the non-canonical amino acid Anap, and follow labeled proteins in living cells with and without stress.

      I suspect that this method could be incredibly valuable to many investigators studying the dynamics and interactions of proteins that are difficult to label or detect by conventional methods.

      Comment on revised version:

      The revised manuscript is significantly improved, with added controls and experiments to confirm expression and Anap labeling of G3BP1 and TDP-43.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      Amyotrophic lateral sclerosis (ALS) affects nerve cells in the brain and spinal cord. The authors' approach to use genetic code expansion to tag two ALS proteins associated with stress granules has value and should be useful in the ALS field. Parts of the work are well done, but there are concerns that the evidence is incomplete overall, and additional controls would strengthen the study.

      We thank the editors and reviewers for their thoughtful assessment and for highlighting the potential value of applying genetic code expansion (GCE) to study ALSassociated proteins involved in stress granule biology. Our goal in this work was to establish and validate a minimally perturbative labeling strategy using the noncanonical amino acid Anap to monitor the localization and stress-dependent behavior of TDP-43 and G3BP1.

      We agree that additional controls can further strengthen the conclusions. In the revised manuscript, we have clarified the experimental design and added essential controls to better support the reliability of the Anap labeling approach (Supplementary Fig. 1).

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors utilize genetic code expansion to tag TDP-43 and G3BP1, and evaluate this protein tagging system (ANAP) compared to antibodies, and evaluate protein trafficking and stress granule formation in response to stress with sodium arsenite treatment. They find similar staining to antibodies in HeLa cells, mouse embryonic stem cells, and primary mouse cortical neurons. This is a useful study that demonstrates the utility of ANAP tagging to evaluate ALS proteins.

      We sincerely thank the reviewer for the positive assessment of our work and for recognizing the utility of the Anap-based GCE system for studying ALS-associated proteins.

      Strengths:

      Rescue of cell survival by ANAP-tagged TDP-43 is compelling

      We appreciate the reviewer’s highlighting of this point. Demonstrating that TDP43-Anap can rescue cell survival was an important validation in our study, as it indicates that incorporation of the noncanonical amino acid does not substantially disrupt the biological function of TDP-43. Additionally, we also tested the RNA splicing function recovery potency of TDP-43-Anap. As shown in Fig. 1K and 1L, a recovery of expression of PFKP, a protein undergoing cryptic exon when TDP-43 lost its function [1], was observed when expressing TDP-43-Anap in TDP-43 knockout Hela cells.

      Weaknesses:

      While the ANAP-tagged proteins had similar distributions to antibody staining, there were some discrepancies that may be more explained by the technique than by novel findings, as the authors suggested. The inclusion of additional controls to evaluate this would be helpful.

      This is a helpful suggestion. To ensure that the fluorescence signal observed in our experiments was specifically derived from site-specific Anap incorporation rather than background fluorescence, we performed three control conditions. Specifically, we tested: (1) cells cultured with Anap supplement, (2) cells expressing the Anap incorporation system with the addition of Anap, and (3) cells expressing both the TAG-mutated protein plasmid and the Anap incorporation system but without the addition of Anap. These control experiments were performed for both TDP-43 and G3BP1, and no observable fluorescence signal was detected under any of these conditions (Supplementary Fig. 1). We have clarified this control experiment in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Chen and colleagues describe a novel means of labeling two RNAbinding proteins, G3BP1 and TDP-43, using genetic code expansion. Overexpressed constructs that incorporate the intrinsically fluorescent non-canonical amino acid Anap redistribute to cytoplasmic granules upon application of external stressors such as sodium arsenite. Similar labeling and redistribution of overexpressed G3BP1 and TDP43 were observed in cultures of mouse primary neurons.

      We are grateful for the reviewer’s accurate summary of our study and recognition of the value of GCE strategy for labeling the RNA-binding proteins G3BP1 and TDP-43.

      Strengths:

      Genetic code expansion and non-canonical amino acid labeling have quite a few advantages over traditional fusion proteins for tracking protein redistribution in living cells. The authors show that they are able to label exogenous G3BP1 and TDP-43 with the non-canonical amino acid Anap and follow labeled proteins in living cells with and without stress.

      We acknowledge the reviewer’s comment on the advantages of GCE-based noncanonical amino acid labeling for studying protein dynamics in living cells.

      Weaknesses:

      The authors do not convincingly leverage the advantages of genetic code expansion in the current study. There is no specific question posed by the authors that can be or is answered using this approach, and several of the experiments lack critical controls. This is also not the first example of TDP-43 labeling by genetic code expansion (see PMID: 38290242). As a result, the study as a whole adds little to our understanding of protein trafficking and behavior under stress.

      We thank the reviewer for raising these important points. Although as reviewer mentioned, genetic code expansion has previously been applied to TDP-43 [2], it mainly employed the photocaged lysine incorporation system to optogenetic control of TDP-43 translocation, and the protein was still labeled by mRubby. Our paper has totally different goal, to establish and validate a minimally perturbative labeling strategy using the intrinsically fluorescent noncanonical amino acid Anap to monitor the localization and stress-dependent behavior of both TDP-43 and G3BP1. And our work extends this approach in several important ways.

      First, we demonstrate that Anap incorporation enables visualization of stress-dependent redistribution of both TDP-43 and G3BP1, two key proteins involved in stress granule biology. Importantly, we validate this approach across multiple cellular systems, including HeLa cells, mouse embryonic stem cells, and primary mouse cortical neurons, which broadens the applicability of this labeling strategy.

      Second, we provide functional validation of the Anap-tagged protein, showing that TDP43-Anap rescues both cell survival and RNA splicing activity in TDP-43 knockout cells, including restoration of PFKP expression, a known cryptic exon target of TDP-43. These results support that Anap incorporation does not substantially disrupt protein function.

      We performed additional control experiments to ensure the specificity of the labeling system. Specifically, we tested three control conditions: (1) cells cultured with Anap supplement, (2) cells expressing the Anap incorporation system with the addition of Anap, and (3) cells expressing both the TAG-mutated protein plasmid and the Anap incorporation system but without the addition of Anap. These control experiments were performed for both TDP-43 and G3BP1, and no observable fluorescence signal was detected under any of these conditions (Supplementary Fig. 1).

      We agree that the manuscript would benefit from clearer articulation of the advantages of genetic code expansion in this context. Accordingly, we have revised the manuscript to more explicitly emphasize how Anap labeling provides a minimally perturbative alternative to large fluorescent protein fusions, which can alter the phase behavior and localization of stress granule proteins.

      “Conventional fluorescent protein tags have enabled visualization of TDP-43 and G3BP1 in living cells; however, these approaches can perturb the native biophysical properties of the proteins being studied. For example, GFP or other fluorescently tagged TDP-43 usually requires additional modifications, such as deletion of the nuclear localization signal (NLS) [3, 4], to induce cytoplasmic inclusion formation. Such manipulations introduce non-physiological conditions that may alter the native trafficking and aggregation behavior of TDP-43. As for G3BP1, tags like GFP may also cause unexpected effects on the phase separation or other dynamics of the protein. In contrast, Anap based GCE strategy allows the minimally perturbative labeling and visualization of protein localization and stress-induced redistribution while preserving native protein architecture and function of both proteins. Importantly, the approach provides a generalizable genetically encoded platform for quantitatively examining the behavior of ALS-associated proteins in living cells. By enabling faithful monitoring of protein trafficking and stressgranule dynamics without extensive protein engineering, Anap-based GCE can offer a powerful strategy for probing molecular-scale mechanisms underlying ALS-linked proteinopathies”.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 1A

      The authors report that the nuclear staining of G3BP1 by ANAP labeling shows the presence of nuclear pools of G3BP1 that aren't detected with antibody staining. However, unspecific nuclear staining by aminoacylated tRNAs bound to synthetases has been described. It would be important to have a control to evaluate for this possibility.

      This is an important point. We agree that the nuclear ANAP signal should be carefully controlled to exclude the possibility of nonspecific staining arising from the Anap incorporation machinery itself, such as aminoacylated tRNAs and/or synthetases.

      To address this concern, in methods and material part, we note that after DPBS washes to remove excess Anap, cells were incubated in fresh medium for 2 hours to allow sufficient time for the decay of unstable aminoacylated tRNAs, which are generally cleared within minutes to tens of munites [5].

      Also, we performed three control conditions for both TDP-43 and G3BP1: (1) cells cultured with Anap supplement, (2) cells expressing the Anap incorporation system with the addition of Anap, and (3) cells expressing both the TAG-mutated protein plasmid and the Anap incorporation system but without the addition of Anap. Under all three conditions, we observed no detectable fluorescence signal (Supplementary Fig. 1).

      In addition, as shown in Fig. 1I, the nuclear signal of G3BP1-Anap partially colocalizes with the nuclear signal of TIA-1 in several condensate-like structures. This observation further supports that the nuclear Anap signal reflects protein-associated localization rather than nonspecific fluorescence, as it overlaps with a known RNA-binding protein that can form nuclear condensates under certain conditions.

      (2) Figure 1A, 1B

      Anap labeling appears to stain fewer cytoplasmic structures compared to antibody staining for both G3BP1 and TDP-43 after sodium arsenite treatment. Quantification would be useful to address whether this is the case. If so, might this be due to unincorporated/truncated proteins competing with Anap-labeled proteins?

      We appreciate the reviewer’s helpful suggestion. To address this point, we performed quantitative colocalization analysis using Fiji/ImageJ, calculating the Pearson correlation coefficient (R) for regions of interest between the Anap signal and antibody staining. These analyses indicate a strong overall agreement between the two detection methods under stress conditions, supporting that Anap labeling reliably reports the localization of both G3BP1 and TDP-43 (see Fig1. A, B).

      Regarding the possibility that truncated or unincorporated proteins could influence the observed signal, we note that fluorescence from Anap depends on successful amber suppression and incorporation of Anap at the engineered TAG site. Proteins that fail to incorporate Anap, such as truncated products generated by premature termination, would not produce fluorescence, and therefore would not contribute to the Anap signal. Thus, the Anap fluorescence selectively reports the population of successfully labeled full-length proteins, whereas antibody staining detects both labeled and unlabeled protein pools. This difference may partially explain why antibody staining appears to label a larger number of cytoplasmic structures.

      (3) Figure 1F

      FRAP of G3BP1-GFP in stress granules is slower than in previous publications. The underlying reasons for this should also be addressed.

      We thank the reviewer for this important observation. Differences in FRAP recovery kinetics of G3BP1 in stress granules may arise from several experimental variables that are known to influence stress granule dynamics. These include differences in cell type, expression levels of G3BP1-GFP, and imaging or photobleaching parameters. In our experiments, FRAP measurements were performed under specific conditions optimized for our experimental system, which may lead to recovery kinetics that differ from those reported in previous studies.

      (4) Figure 1H

      A full-size Western blot would be useful to evaluate for amount of truncated protein for G3BP1 and TDP-43. Could truncated proteins be competing with and altering ANAPtagged G3BP1 and TDP-43 localization in response to stress? This should be addressed.

      We acknowledge this important point. Full-size Western blotting can provide information on the overall presence of truncated species in the transfected population; however, it represents a bulk measurement and does not capture cell-to-cell variability in amber suppression efficiency at the single-cell level. We therefore cannot exclude the possibility that truncated products are present at varying levels in individual cells and may contribute, directly or indirectly, to differences between antibody staining and Anap fluorescence.

      Importantly, we observe that cells with successful Anap incorporation consistently exhibit strong antibody staining for TDP-43 or G3BP1, indicating that full-length protein is the predominant species in these cells. Because Anap fluorescence depends on successful amber suppression, it selectively reports the full-length protein population, whereas truncated products are not detected in the imaging assay. The concordance between Anap fluorescence and antibody staining therefore argues against a major contribution of truncated species to the observed localization patterns (Supplementary Fig. 1).

      Accordingly, we interpret the Anap signal as reflecting the localization of successfully labeled full-length protein, while acknowledging that heterogeneity in suppression efficiency is an important limitation of the current approach.

      (5) Figure 3

      This is a well-designed diagram.

      We are grateful for the reviewer’s positive feedback on the diagram and are pleased that the schematic effectively illustrates the experimental design and the principles of the genetic code expansion strategy used in this study.

      Reviewer #2 (Recommendations for the authors):

      The authors present a one-sided viewpoint concerning the connection between stress granules and disease (lines 45-46). A more balanced discussion is recommended, including data arguing against a role for abnormal stress granules in neurodegeneration.

      This is an important suggestion. We agree that the relationship between stress granules and neurodegeneration remains an active area of investigation and that evidence both supporting and questioning a causal role of stress granules in disease has been reported. In the revised manuscript, we have modified the Introduction to provide a more balanced discussion of this topic.

      “Altered stress-granule dynamics have been associated with ALS/FTD [6, 7]; however, whether stress granules directly drive neurodegeneration remains debated, as several studies suggest that stress granules primarily function as protective stress responses [8].”

      (1) A central rationale for the study is missing. The authors state only that G3BP1 and TDP-43 'undergo dynamic stress-dependent redistribution, making them ideal candidates for minimally invasive, site-specific fluorescent labeling.' Is there a controversy or question that can be resolved using these approaches?

      We thank the reviewer for raising this important point. The central motivation of this study is that the dynamic behavior and phase separation properties of stressgranule proteins are highly sensitive to protein modifications and tagging strategies.

      “Conventional fluorescent protein tags have enabled visualization of TDP-43 and G3BP1 in living cells; however, these approaches can perturb the native biophysical properties of the proteins being studied. For example, GFP or other fluorescently tagged TDP-43 usually requires additional modifications, such as deletion of the nuclear localization signal (NLS) [3, 4], to induce cytoplasmic inclusion formation. Such manipulations introduce non-physiological conditions that may alter the native trafficking and aggregation behavior of TDP-43. As for G3BP1, tags like GFP may also cause unexpected effects on the phase separation or other dynamics of the protein.”

      (2) Related to this, there is little context for how or why genetic code expansion is utilized for these studies

      We agree that the rationale for using genetic code expansion should be more clearly explained. In this study, genetic code expansion was employed to enable sitespecific incorporation of the small fluorescent noncanonical amino acid Anap, allowing minimally perturbative labeling of proteins of interest.

      “Anap based GCE strategy allows the minimally perturbative labeling and visualization of protein localization and stress-induced redistribution while preserving native protein architecture and function of both proteins. Importantly, the approach provides a generalizable genetically encoded platform for quantitatively examining the behavior of ALS-associated proteins in living cells. By enabling faithful monitoring of protein trafficking and stress-granule dynamics without extensive protein engineering, Anapbased GCE can offer a powerful strategy for probing molecular-scale mechanisms underlying ALS-linked proteinopathies.”

      (3) The justification for the criteria for selecting the site for incorporation of non-canonical amino acids in G3BP1 or TDP-43 is missing.

      We acknowledge this important comment and agree that the rationale for selecting the incorporation sites should be stated more clearly.

      “For TDP-43, the incorporation site was selected to avoid the major functional domains involved in RNA binding, nuclear localization, and aggregation-related behavior, thereby reducing the likelihood that Anap incorporation would perturb its native trafficking or function. For G3BP1, the selected site was chosen to minimize interference with domains important for stress granule assembly, RNA binding, and protein-protein interactions. More generally, we aimed to place the ncAA at positions likely to be solventaccessible and tolerant of substitution, while avoiding highly conserved or functionally essential residues.”

      (4) Studies in Figures 1 and 2 lack essential controls, including background signal from Anap in non-transfected cells, or those transfected with plasmids lacking the tRNA or tRS.

      This is an important point, also raised by Reviewer 1. To evaluate potential background fluorescence arising from Anap or the labeling system, we performed several control experiments. Specifically, we examined three conditions: (1) cells cultured with Anap supplement, (2) cells expressing the Anap incorporation system with the addition of Anap, and (3) cells expressing both the TAG-mutated protein plasmid and the Anap incorporation system but without the addition of Anap. Under all three conditions, we observed no detectable fluorescence signal (Supplementary Fig. 1).

      (5) Another marker of stress granules should be used for confirming the identity of G3BP1-Anap (+) or TDP-43-Anap (+) structures, including TIA1, TAF15, or polyA RNA.

      We appreciate this helpful suggestion. To further confirm the identity of the stress granule structures observed in our experiments, we performed colocalization analysis with TIA-1, a well-established marker of stress granules. The results have been included in revised manuscript.

      “Additionally, we examined the colocalization of G3BP1-Anap with TIA-1, another established stress granule marker. Under stress conditions, G3BP1-Anap largely colocalized with TIA-1 within stress granules. Interestingly, under basal conditions, the nuclear signal of G3BP1-Anap, which was not detected by antibody staining, appeared to partially colocalize with TIA-1 in several condensate-like structures. (Fig. 1I).”

      (6) There is no information on the number of granules bleached or the number of cells selected for FRAP studies. There is no information on the shaded areas in Figure 1F or 1G, and no information on statistical comparisons between regressions in Figure 1F.

      We thank the reviewer for pointing out these omissions. We have revised the figure legends to clarify these details.

      “One granule from each of three independent cells was selected and photobleached for FRAP analysis.”

      “Here, error bars with filled area are used for better data presentation. FRAP recovery curves were compared using two-way ANOVA.”

      (7) Protein dynamics measured by FRAP are highly dependent on the concentration and/or expression level of each protein. Because of this, the authors need to control for expression level in all FRAP studies.

      We agree that protein concentration and expression level can influence FRAP recovery kinetics. Since Anap incorporation is based on amber suppression, and the suppression rate in each cell varies, so it is difficult to control the expression of Anap labeled proteins, however, to minimize this potential effect, we performed FRAP measurements on cells exhibiting comparable fluorescence intensities, which served as a proxy for similar expression levels of the labeled proteins. In addition, FRAP analyses were conducted on individual granules within cells expressing moderate levels of the protein, avoiding cells with unusually high fluorescence intensity that might reflect overexpression.

      Furthermore, fluorescence recovery was normalized to the pre-bleach intensity of the selected granules, which reduces variability arising from differences in overall expression levels between cells.

      (8) There is no point of reference for TDP-43-Anap FRAP results in Figure 1G. Additional studies using variants harboring a mutated NLS (mNLS) can be used in place of TDP43-YFP.

      This is a helpful suggestion. In response, we have performed additional FRAP experiments using TDP-43<sup>ΔNLS</sup>, a commonly used construct that promotes cytoplasmic localization and facilitates analysis of TDP-43 granules. The results from TDP-43<sup>ΔNLS</sup> have now been included as a reference for the FRAP measurements of TDP-43-Anap in the revised manuscript (Fig. 1D, 1G).

      “We then used YFP-tagged nuclear localization signal (NLS)-deleted TDP-43 (TDP43<sup>ΔNLS</sup>-YFP) as a reference and performed FRAP analysis to compare the mobility of TDP-43-Anap and TDP-43<sup>ΔNLS</sup>-YFP. Fluorescence recovery of TDP-43-Anap reached ~45% within 20 s after photobleaching, consistent with liquid-like dynamics. In contrast, TDP-43<sup>ΔNLS</sup>-YFP showed only ~22% recovery, suggesting more solid-like dynamics (Fig. 1D, 1G). These results are consistent with previous reports describing relatively immobile aggregates formed by TDP-43<sup>ΔNLS4</sup>and illustrate the advantage of Anap-based labeling, which preserves native protein properties and enables real-time assessment of protein dynamics without introducing disruptive mutations.”

      (9) There is no point of reference for comparing FRAP results from G3BP1-GFP to G3BP1-Anap. What is the 'gold standard'? Without this, it is difficult to conclude that "... Anap labeling better preserved the native mobility and biophysical properties of G3BP1 than the conventional GFP tag."

      We acknowledge this important point and agree that there is currently no definitive gold standard for measuring the native mobility of endogenous G3BP1 within stress granules in living cells. Our intention was not to claim that the Anap-labeled protein definitively represents the native state, but rather to compare the relative effects of different labeling strategies.

      Thus, we rewrite the sentence as “These results suggest that G3BP1-Anap displays higher mobility compared with G3BP1-GFP, indicating that Anap labeling may provide a less perturbative approach for monitoring G3BP1 dynamics.”

      (10) The WB in Figure 1H is overexposed, making it difficult to compare expression levels between WT and V100Anap-transfected cells. In addition, there is no similar assay for confirming G3BP1-Anap expression.

      Thank you for pointing this out. In the revised manuscript, we have replaced the image with a properly exposed Western blot to allow clearer comparison of protein expression levels.

      In addition, we have now included a corresponding western blot analysis to confirm the expression of G3BP1-Anap in G3BP knockout U2OS cell (Fig. 1H). These results verify that the Anap-labeled proteins are expressed at detectable levels and support the interpretation of the imaging and FRAP experiments.

      (11) Although survival studies in Figures 1I and J are promising, a more convincing demonstration of functional replacement of TDP-43 would involve an assessment of cryptic exon splicing, comparing WT to TDP-43 KO, V100Stop- and V100Anaptransfected cells.

      This is a valuable suggestion.

      “We also evaluated TDP-43-dependent RNA splicing activity by examining the expression of PFKP, a well-established target that undergoes cryptic exon inclusion upon loss of TDP-43 function17. As shown in Figures 1K and 1L, expression of TDP-43Anap in TDP-43 knockout HeLa cells restored PFKP expression, indicating that the Anap-labeled protein retains functional RNA splicing activity. These results demonstrate that TDP-43-Anap is capable of functionally compensating for endogenous TDP-43, supporting that the incorporation of Anap does not substantially disrupt the protein’s biological function.”

      (12) Tuj1 staining in Figure 2 is inconsistent and often fails to confirm neuronal identity.

      We thank the reviewer for this important comment. We acknowledge that Tuj1 staining in Figure 2 is variable and, in some cases, does not clearly delineate neuronal identity. Notably, the reduced Tuj1 signal is primarily observed in neurons that express Anap-labeled proteins under sodium arsenite treatment, which likely reflects the combined effects of transfection-associated stress and oxidative stress on neuronal morphology and cytoskeletal integrity.

      In addition, transfection efficiency in primary neurons is inherently low and variable, and cells that successfully express the constructs may represent a more stress-sensitive subpopulation, further contributing to variability in staining quality. Despite optimization efforts, these technical constraints limit the consistency of Tuj1 labeling under these experimental conditions.

      (13) Close-up images and correlation scatter plots in Figures 1 and 2 do not add very much information.

      We thank the reviewer for this comment. To address the reviewer’s concern, we have revised the figure legends to better clarify the purpose of these panels and how they support the quantitative analysis presented in the manuscript.

      For scatter plot, “Colocalization threshold analysis was performed in Fiji/ImageJ to calculate the Pearson correlation coefficient (R) for each region of interest (A, B, I, J). The X- and Y-axes represent the fluorescence intensity values of the red and green channels, respectively. When signals are colocalized, pixels with high intensity in one channel correspond to high intensity in the other, forming a diagonal distribution. In contrast, non-colocalized signals cluster along the axes. A higher R value indicates a greater degree of colocalization. Scale bar, 3 μm.”

      Same information was added to figure legend of figure 2.

      For the scheme, please see line 412-413 in the revised manuscript.

      Reference:

      (1) Rothstein, J.D. et al. Sporadic ALS induced pluripotent stem cell derived neurons reveal hallmarks of TDP-43 loss of function. Nature Communications 16, 7092 (2025).

      (2) Shadish, J.A. & Lee, J.C. Genetically encoded lysine photocage for spatiotemporal control of TDP-43 nuclear import. Biophys Chem 307, 107191 (2024).

      (3) Gasset-Rosa, F. et al. Cytoplasmic TDP-43 De-mixing Independent of Stress Granules Drives Inhibition of Nuclear Import, Loss of Nuclear TDP-43, and Cell Death. Neuron 102, 339–357.e337 (2019).

      (4) Yan, X. et al. Intra-condensate demixing of TDP-43 inside stress granules generates pathological aggregates. Cell 188, 4123–4140.e4118 (2025).

      (5) Walker, S.E. & Fredrick, K. Preparation and evaluation of acylated tRNAs. Methods 44, 81–86 (2008).

      (6) Kassouf, T. et al. Targeting the NEDP1 enzyme to ameliorate ALS phenotypes through stress granule disassembly. Science Advances 9, eabq7585 (2023).

      (7) Van Nerom, M. et al. C9orf72-linked arginine-rich dipeptide repeats aggravate pathological phase separation of G3BP1. Proceedings of the National Academy of Sciences 121, e2402847121 (2024).

      (8) Wolozin, B. & Ivanov, P. Stress granules and neurodegeneration. Nat Rev Neurosci 20, 649–666 (2019).

    1. eLife Assessment

      This important study investigates how the hippocampus distinguishes between reactive escape and anticipatory withdrawal during approach-avoidance conflict in rats performing a naturalistic decision-making task. Solid evidence supports the main finding that hippocampal neuronal representations differ during different types of defensive behaviors, although the evidence for some of the claims in the paper could be strengthened. The study will be of interest to researchers studying memory, navigation, and decision-making in the presence of competing rewards and threats.

    2. Reviewer #1 (Public review):

      Summary:

      This study by Damphousse, Calvin, and Redish investigates how the hippocampus represents competing future outcomes during approach-avoidance conflict. Using an ethologically relevant robotic predator foraging paradigm, the authors aimed to dissociate hippocampal activity associated with reactive defensive responses (escape) from that linked to anticipatory withdrawal decisions. The central finding is that dorsal hippocampal representations differentiate these two modes of defensive behavior within a single naturalistic assay. Specifically, the authors show that attack-triggered retreats and mid-track aborts differ in movement dynamics and hippocampal spatial decoding despite sharing a common behavioral endpoint, that hippocampal representations during pauses predict subsequent behavioral outcomes, and that these representational biases emerge before overt behavioral divergence. The main importance of the study lies in moving beyond viewing the hippocampus as merely encoding spatial location or threat salience, instead suggesting that hippocampal ensemble activity dynamically tracks and differentially weights threat-related, reward-related, and safety-oriented future states to bias behavior before overt action occurs.

      Strengths:

      The study has several notable strengths. First, the behavioral decomposition into retreats, mid-track aborts, and mid-track continues is rigorous and provides a highly interpretable analytical framework. Second, replication across two independent cohorts - despite differences in arena configuration, robot design, and extinction procedures - meaningfully strengthens confidence in the robustness of the findings. Third, the unified reanalysis pipeline across cohorts reflects strong analytical discipline, and the Bayesian decoding framework is well-suited to addressing the central representational questions. Fourth, the ethological relevance of the robotic predator paradigm is a major advantage, allowing the authors to examine a richer repertoire of defensive and decision-related behaviors than is possible in conventional fear-conditioning assays. Overall, the experiments are well designed, the data are clearly presented, and the findings make a valuable contribution to understanding how the hippocampus supports decision-making under threat.

      Weaknesses:

      The study is technically strong, but a few modest revisions would further enhance it.

      (1) First, the abstract mentions extinction and reinstatement effects, but neural analyses focus primarily on the attack phase. It would be helpful to clarify or adjust the abstract accordingly.

      (2) Second, some interpretive language ("guide," "bias") leans toward causal phrasing. Given the correlational data, using "predict" or "correlate with" would be more precise.

      (3) Third, given the relationship between running speed and hippocampal theta, considering speed-related contributions to decoding differences would be useful.

      (4) Fourth, reporting turnaround positions for mid-track abort and continue trials (Figure 7) would provide helpful context.

      (5) Fifth, a figure comparing stimulated vs. non-stimulated sessions in cohort 2 would support the claim that closed-loop stimulation had no measurable effect.

      (6) Finally, reporting effect sizes for key decoding comparisons would add clarity.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript extends previous work from Calvin et al. and examines hippocampal representations during approach-avoidance conflict in a robotic predator foraging task. The paradigm itself is very interesting and addresses an important but relatively understudied question in the navigation and foraging literature: how the brain balances risk versus reward during goal-directed behavior. While hippocampal representations of positively valenced goals and future intentions have been extensively studied, much less is known about how these representations evolve during risk-reward tradeoffs involving threat.

      The authors use a relatively simple and interpretable decoding approach together with thoughtful behavioral comparisons to ask whether future behavioral outcomes can be read out from hippocampal activity before behavior diverges. The most compelling comparison is between mid-track aborts (MTAs) and mid-track continues (MTCs), where the animals initially exhibit very similar pause behavior but ultimately either abort or continue the trajectory. The authors show that decoded location during these pauses differs prior to the overt manifestation of the behavioral decision, suggesting that hippocampal representations may reflect evolving internal evaluation processes during approach-avoidance conflict.

      Strengths:

      A major strength of the work is the behavioral paradigm itself. This type of risk-reward conflict task is relatively uncommon in the hippocampal navigation literature and provides a rich framework for examining defensive decision-making during naturalistic foraging behavior.

      The decoding analyses are also relatively simple and easy to interpret. Rather than relying on highly complex modeling approaches, the authors use straightforward comparisons of decoded spatial representations across behavioral conditions, making the results accessible and conceptually clear.

      Another strength is the use of behavioral controls to isolate comparisons between related behaviors. In particular, the comparison between MTAs and MTCs is compelling because the animals exhibit similar pause states before the behavioral outcomes diverge. This provides a useful framework for asking whether hippocampal activity reflects future behavioral outcome before the decision is overtly expressed.

      Overall, the study asks an interesting question using a novel paradigm and provides evidence that hippocampal representations during approach-avoidance conflict may reflect future behavioral trajectory.

      Weaknesses:

      The main weakness is that many of the reported effects are relatively subtle and are not sufficiently controlled for differences in speed, trajectory structure, and other behavioral variables across conditions. While the subtraction plots (green versus purple decoding differences) appear visually striking, the actual effect sizes are fairly small, making it difficult to assess how robust or behaviorally meaningful these differences are.

      Relatedly, many of the most interesting questions in this task concern how behavior unfolds dynamically within a trial, yet much of the analysis averages across events and trajectories. As a result, potentially important aspects of the behavior may be obscured.

      In particular, the manuscript would benefit from richer characterization of the animals' actual movement trajectories and spatial strategies. Because the analyses rely heavily on linearized position, it is difficult to determine whether animals behave differently in two-dimensional space across conditions. For example, during continued approaches, do animals preferentially hug the wall opposite the robot? Do different behavioral conditions show distinct lateral occupancy or trajectory structure? These types of analyses would make the behavioral interpretation substantially more compelling.

      More generally, while the results are suggestive and interesting, the relatively small decoding differences and substantial behavioral confounds make it difficult to conclude that the observed effects reflect distinct internal evaluative or threat-related states.

    4. Reviewer #3 (Public review):

      Summary:

      The study reanalyzes data from a previously published cohort together with an additional cohort to investigate hippocampal activity during approach-avoidance conflict. Unlike many prior studies that isolate reward- or threat-based learning, this task requires animals to evaluate reward and threat concurrently. The central finding is that hippocampal representations differ between hesitant behaviors that lead to approach versus avoidance outcomes, with representations of the attack zone more likely during pauses preceding abort decisions. This is an important extension of prior work on hippocampal activity and deliberation, suggesting that the hippocampal content may help shape the eventual outcome.

      Strengths:

      All behavioral findings are replicated independently across cohorts, making the behavioral results highly convincing. The design is robust, and the task is especially valuable for studying approach-avoidance conflict. The behavioral paradigm is complex and rare, and neuronal recordings in such a paradigm are of great value.

      The major strength of the study is the comparison of neural activity during hesitant behavior leading to different outcomes, namely, pauses followed by the animal aborting the approach (mid-track aborts), and pauses followed by the animal committing to the approach (mid-track continues). Hippocampal activity differed between the two pauses: the attack zone was more likely to be represented during mid-track aborts. The same effect was observed on the journey before the pause: even before the animal hesitates, hippocampal activity before a pause that led to a mid-track abort was more likely to represent the attack zone than hippocampal activity before pauses that led to continued approach. This analysis suggests that hippocampal content before and during deliberative behavior is predictive of the animal's decision.

      Weaknesses:

      The interpretation of the retreat-related decoding results is less clear. The study compares two sets of retreating behavior: on the one hand, retreat after being attacked, and on the other hand, retreat after hesitation in the absence of an attack (a mid-track abort). Hippocampal activity represents the attack zone more after the animal is attacked. However, these two retreating behaviors originate from different spatial locations: retreats always start past the "attack threshold", while mid-track aborts always start before this threshold. Given that hippocampal decoding is strongly location-dependent, this difference in position makes the neural decoding results difficult to interpret. The increased representation may be due to differences in physical location, rather than the distinct processing of immediate threat and an anticipatory return state.

    1. eLife Assessment

      The authors describe the important finding that the Legionella-containing vacuole is surrounded by a dense ubiquitin "cloud" that is highly resistant to detergent extraction. The study provides compelling evidence that this structure is generated through a combination of canonical ubiquitination mediated by the SidC Legionella ligase and phosphoribosyl-ubiquitination mediated by the SidE family of Legionella ligases. These findings provide insight into how Legionella remodels the host vacuolar environment through complex ubiquitin modifications.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the reviewers' comments adequately and revised the manuscript accordingly.]

      Summary:

      In the submitted manuscript, Steinbach et al describe the formation of a detergent-resistant "cloud" around the Legionella-containing vacuole (LCV) that functions as a protective barrier. The authors show that formation of the "cloud" barrier is contingent upon the phosphoribosyl-ubiquitination activity of the SidE/SdeABC effector family, and is temporally regulated, with the assembly and subsequent disassembly of the "cloud" coinciding with replication and vacuolar expansion. The authors postulate a model of "cloud" barrier formation that relies upon a wave of initial ubiquitination by the SidC effector family, after which the SidE/SdeABC family expands the ubiquitination and forms cross-links that render the ubiquitin cloud resistant to harsh detergents. Additionally, Steinbach et al. also demonstrate that Rab5 is recruited to the LCV and remains associated for a considerable period.

      Strengths:

      This manuscript is very well written, with clear justification provided for experiments that make it very easy to follow along with the experimental logic. The figures have clearly been designed with much thought and are easy to interpret. Steinbach et al have also done a commendable job of addressing the previous reviewers' comments, even though some may suggest that some of these comments could be viewed as slightly unreasonable. This work would be of interest to both the Legionella and ubiquitin fields. Legionella researchers would potentially be interested to explore the proposed barrier model as the function for the ubiquitin "cloud," whereas ubiquitin researchers may be interested in exploring the mechanisms underlying SidE's crosslinking ability.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript "Canonical and phosphoribosyl ubiquitination coordinate to stabilize a proteinaceous structure surrounding the Legionella-containing vacuole" by Steinbach et al. is well written and presents strong evidence that satisfactorily supports the main hypothesis and research objectives. The authors have clearly demonstrated the presence of cloud-like, detergent-resistant GTPase Rab5 surrounding the LCV, and formation of the structure is dependent on the SidE family of effectors. The study provides insights into the relevant (associated with described phenotype) ubiquitination pathways. The findings advance our understanding of Legionella pneumophila vacuole remodeling during intracellular infection and open directions for future research to establish broader implications of this structure on Legionella pathogenesis.

      Strengths:

      The manuscript convincingly demonstrates the presence of a cloud-like, detergent-resistant GTPase Rab5 surrounding the LCV through elegant microscopy. The experimental evidence about the dependence of the observed phenotype on the SidE family of effectors is compelling and presented with strong scientific rigor. The introduction is well-written, and the discussion is thorough and satisfactory. The article is thought-provoking and shows preliminary evidence for ubiquitin-mediated protection and spatial organization of the LCV.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript by Mukherjee and colleagues extended earlier studies on the coordination of the SidC and SidE effector families on the generation of a unique ubiquitin layer on the surface of the vacuoles containing the bacterial pathogen Legionella pneumophila (LCV).

      Strengths:

      The main strength of the manuscript is the identification of the small GTPase Rab5 as a major "carrier" of these differently modified ubiquitin and ubiquitin chains, which was nicely quantified.

      Weaknesses:

      The results are mostly descriptive, based on mechanistic studies from earlier works.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the submitted manuscript, Steinbach et al describe the formation of a detergent-resistant "cloud" around the Legionella-containing vacuole (LCV) that functions as a protective barrier. The authors show that formation of the "cloud" barrier is contingent upon the phosphoribosyl-ubiquitination activity of the SidE/SdeABC effector family, and is temporally regulated, with the assembly and subsequent disassembly of the "cloud" coinciding with replication and vacuolar expansion. The authors postulate a model of "cloud" barrier formation that relies upon a wave of initial ubiquitination by the SidC effector family, after which the SidE/SdeABC family expands the ubiquitination and forms cross-links that render the ubiquitin cloud resistant to harsh detergents. Additionally, Steinbach et al. also demonstrate that Rab5 is recruited to the LCV and remains associated for a considerable period.

      Strengths:

      This manuscript is very well written, with clear justification provided for experiments that make it very easy to follow along with the experimental logic. The figures have clearly been designed with much thought and are easy to interpret. Steinbach et al have also done a commendable job of addressing the previous reviewers' comments, even though some may suggest that some of these comments could be viewed as slightly unreasonable. This work would be of interest to both the Legionella and ubiquitin fields. Legionella researchers would potentially be interested to explore the proposed barrier model as the function for the ubiquitin "cloud," whereas ubiquitin researchers may be interested in exploring the mechanisms underlying SidE's crosslinking ability.

      Weaknesses:

      While the work is important and describes the physical nature of the ubiquitin cloud on the Legionella vacuole, it is somewhat descriptive in nature and does not dig deeply into what purpose this cloud serves. This is a complicated topic that will certainly stimulate additional research in this area.

      We thank Reviewer #1 for positive assessment of our work. We acknowledge that our study leaves many mechanistic questions open and, as suggested by Reviewer #1, we hope that our data is thought-provoking for researchers studying Legionella, ubiquitin signaling, or both. We are greatly looking forward to the results of future experimentation on the role of the “cloud” surrounding the bacterial vacuole.

      Reviewer #2 (Public review):

      Summary:

      The manuscript "Canonical and phosphoribosyl ubiquitination coordinate to stabilize a proteinaceous structure surrounding the Legionella-containing vacuole" by Steinbach et al. is well written and presents strong evidence that satisfactorily supports the main hypothesis and research objectives. The authors have clearly demonstrated the presence of cloud-like, detergent-resistant GTPase Rab5 surrounding the LCV, and formation of the structure is dependent on the SidE family of effectors. The study provides insights into the relevant (associated with described phenotype) ubiquitination pathways. The findings advance our understanding of Legionella pneumophila vacuole remodeling during intracellular infection and open directions for future research to establish broader implications of this structure on Legionella pathogenesis.

      Strengths:

      The manuscript convincingly demonstrates the presence of a cloud-like, detergent-resistant GTPase Rab5 surrounding the LCV through elegant microscopy. The experimental evidence about the dependence of the observed phenotype on the SidE family of effectors is compelling and presented with strong scientific rigor. The introduction is well-written, and the discussion is thorough and satisfactory. The article is thought-provoking and shows preliminary evidence for ubiquitin-mediated protection and spatial organization of the LCV.

      Weaknesses:

      The manuscript is well-organized and detailed, and it is hard to find weaknesses under the set goals of the research. A few weaknesses are that the molecular determinants or the regulatory mechanisms that drive selective versus non-selective incorporation of host proteins into this structure are unclear, and, as the authors mentioned, further work is required to establish the precise biophysical basis of the detergent resistance and expansive morphology of the ubiquitinated GTPase "cloud". Currently, the function or purpose of the structure is completely speculative. The effects or importance of the structure on bacterial replication is also not established in the current study. Figure 2D, right panel, Western blot results, the authors suggested the signal present in all four lanes between 37 and 25 kDa is 'nonspecific', which is probably a 'too intense' signal to be called so. Mass spec analysis would be interesting in order to identify sources of such intense signals. With these few limitations, the research presented in this manuscript is experimentally rigorous and opens avenues for future research.

      We thank Reviewer #2 for their positive assessment and constructive criticism of our study. We agree that the degree of selectivity of incorporation of proteins into the “cloud” is of great interest, as are the molecular details of the cloud structure, and we expect that future experimentation in this area will provide insight into these key questions.

      Reviewer #2 rightly points out that our study did not address the role of the LCV associated “cloud” in supporting bacterial replication. We note that previous studies have reported growth defects for knockout strains lacking SidC/SdcA (PMID 24483784) and the SidE family (PMID 27049943). However, given the multiple roles that these effector families appear to play during infection, we cannot ascribe these defects in bacterial growth solely to the absence of the LCV associated “cloud”.

      As for the band present in the four lanes in Fig 2D, we suggest that this band is non-specific (most likely detection of the light chain of the antibody used for immunoprecipitation) because we do not observe this band in the input lanes, and we also see this band in the IP samples in Fig 2C (uninfected samples), including the vector control in which no PR-ubiquitination is observed. In Fig 2C, the non-specific bands in the IP samples appear lower intensity because the HA signal is relatively intense in comparison to the infection experiment in 2D, as overexpression of SidE family effectors results in far more PR-ubiquitination than in infection.

      Reviewer #3 (Public review):

      Summary:

      This manuscript by Mukherjee and colleagues extended earlier studies on the coordination of the SidC and SidE effector families on the generation of a unique ubiquitin layer on the surface of the vacuoles containing the bacterial pathogen Legionella pneumophila (LCV).

      Strengths:

      The main strength of the manuscript is the identification of the small GTPase Rab5 as a major "carrier" of these differently modified ubiquitin and ubiquitin chains, which was nicely quantified.

      Weaknesses:

      (1) The results are mostly descriptive, based on mechanistic studies from earlier works.

      (2) The majority of the work was dedicated to the characterization of the unique ubiquitin layer on the LCV. One important question was ignored: what is the role of Rab5 in this process? Is the GTPase activity of Rab5 required for its ubiquitination by SidC and SidE? The authors should create a Rab5 KO cell line, complement the line with different mutants of Rab5, and examine their ubiquitination and association with the LCV.

      (3) The finding that Rab5 is associated with the LCV supports the notion that the LCV has characteristics of endo- or/late endosomes. The positioning of the LCV in the endocytic pathway should be discussed in the context of earlier studies (e.g.,PMID: 38739652; PMID: 11067875; PMID: 11067875).

      We thank Reviewer #3 for their constructive criticism of our work. While we appreciate this reviewer’s interest in Rab5, our data is not consistent with Rab5 being a primary “carrier” of ubiquitin species; many more LCVs are ubiquitin-positive than Rab5-positive during early infection, and in our live imaging experiments we observe many ubiquitin-positive, Rab5-negative LCVs. We used Rab5 as a model substrate in this study because it allowed us to compare modification at the LCV membrane between the WT and avirulent dotA strain. Our data is more consistent with a model in which Rab5 is one of many small GTPases, and likely other host proteins, caught in a crosslinked mesh around the LCV. However, we agree that discussing the interaction of the LCV with the endolysosomal system is relevant; while this is discussed at length in our previous publication (PMID 38117589), we have expanded the discussion in this study to include new publications and contextualize our latest findings.

      We agree with Reviewer #3 that assessing the role of nucleotide binding state in Rab5 ubiquitination is of interest. While creating a Rab5 KO cell line was not feasible given time and technical constraints, we conducted overexpression experiments with nucleotide binding mutants that exhibit dominant phenotypes (Q79L and S34N) and find that these mutants are still recruited to the LCV and ubiquitinated during infection (see new figure S1).

      Recommendations for the authors:

      Reviewing Editor Comments:

      There are suggestions from the reviewers to further address the role of Rab5 in LCV-associated ubiquitination, including whether its GTPase activity is required for modification by SidC and SidE, and the mechanism underlying the dissolution of the ubiquitin cloud during vacuolar expansion.

      Reviewer #1 (Recommendations for the authors):

      To improve upon the manuscript and its impact, the authors could consider the following:

      Major concern:

      The temporal regulation of the ubiquitin cloud is fascinating. The authors nicely demonstrate that SidE- and SidC-type ligases cooperate to form the cloud, but how is it dissolved during vacuolar expansion? They demonstrate that ectopic expression of DopA can do this, but do DopA and DopB regulate this process natively?

      We thank Reviewer #1 for their suggestions. While we agree that this line of experimentation is absolutely of interest, it is not feasible for our lab to carry out these experiments on a reasonable timeline to include in the current work.

      Minor concern:

      A syntax error on line 267 of the manuscript should be addressed.

      This has been corrected.

    1. eLife Assessment

      This study provides important findings on the expression of glutamate receptor (GluR) subunits across developmental stages and muscle types in Drosophila. It shows that adult muscle differs in GluR composition from larval body wall muscles, which have been the focus of most past studies. The study, while convincing, could be strengthened by acknowledging that it relies on heterogeneous methods and the absence of positive signals to infer receptor loss, which limits confidence in some of its claims. The findings illuminate how Drosophila excites muscles in diverse tissue types at different life stages, and are of interest to researchers across neuroscience.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript by Sustar et al. takes a methodical approach to document the types of glutamate receptor subunits that reside in Drosophila muscles, examining developmental stages spanning from larvae to adults. Prior work thoroughly documented the subunits operating in Drosophila larval body wall muscles. Most subsequent research focused on the glutamate receptor heterotetramers found in the body wall, composed of GluRIIA/C/D/E or GluRIIB/C/D/E subunits, along with auxiliary subunits like isoforms of Neto.

      For the current work, the authors report that the larval muscle glutamate receptor composition is not universal for all Drosophila muscles. They examine the following muscle systems: larval body wall, adult abdomen, adult leg coxa, and adult indirect flight. They also briefly examine adult muscle structures associated with the proboscis, neck, and haltere. The authors find that the receptor subunits in the adult abdomen (mostly) match those in the larval body wall. This makes sense given that the adult abdominal muscles are derived from the larval body wall. Yet not much else matches the larval body wall. For example, all (or most) of the GluRII-type subunits are missing from the adult indirect flight muscles. Leg muscles have GluRII-type subunits, but they do not have all of them expressed prominently, and they are missing GluRIIB. Additionally, leg muscles express a glutamate-gated chloride channel, which could be a source of inhibitory glutamatergic transmission. Interestingly, when it comes to non-abdominal adult muscles, one general theme seems to be an active promoter (GAL4 driver) for the kainate-type glutamate receptor called Clumsy. The authors propose that Clumsy could be key to understanding how functional GluR complexes are assembled in adult insects.

      Strengths:

      (1) Documenting the types of glutamate receptors that operate in diverse insect muscle systems is important because it uncovers fundamental information.

      (2) Much of the prior research focus has been on how the body wall muscle tetramers assemble and operate. It is a strength to demonstrate the other receptor solutions used by adult NMJs.

      (3) The work uses GAL4 drivers and immunohistochemistry (when possible) in combination to draw conclusions.

      (4) The muscle anatomical analyses are of high quality. This allows the research group to reach refined conclusions.

      (5) The confocal-level images of synaptic active zones and their apposed glutamate receptor clusters are of high quality.

      Weaknesses:

      (1) There is a strawman argument that is used repeatedly to highlight the significance of the work. The argument implies that the field broadly assumes (or "tacitly" assumes) that the larval body wall glutamate receptor composition extrapolates to all muscles of the fly, including the adult. This reviewer cannot find evidence that this assumption or argument has been explicitly promulgated by others. More likely, others have not examined these muscles directly, and thus, they have not speculated one way or the other.

      (2) Related - to the extent that there has been any tacit assumption about GluRIIC/D/E-anchored receptors being ubiquitous among adult muscles, tacit doubt was raised by Rivilin et al., 2004 (cited by the authors but not as a source of doubt) and by RNAseq datasets like FlyAtlas from 2022 (replicated in Figures s11 and s12). To be clear, the current analysis is better than a bulk transcript analysis from adult tissues. But rather than "overturning" a field or being paradigm-shifting, the current data seem confirmatory of FlyAtlas - and confirmatory of Rivlin et al., 2004, which explicitly concluded that larval and adult NMJs were different .

      (3) One can draw expression-level conclusions from these data. But genetic tests (e.g., would clumsy losses of function impair leg muscles?) could help the authors and the field draw stronger conclusions about the roles of some of these glutamate receptor gene products. The current dataset falls short of definitively establishing the function of alternate glutamate receptor modules.

      (4) The confocal synaptic images are of high quality. They are good enough that one could analyze how well Brp directly apposes a specific glutamate receptor subunit for all the associated imaging data underlying Figures S1-S8. No such analysis is done, but understanding what components seem to directly oppose the site of release could lead to better conclusions.

      Overall Assessment and Discussion:

      The data in this study are of high quality, and the results support the main conclusion: adult muscle glutamate receptor clusters do not recapitulate the "canonical" larval body wall clusters. This is important, and the data stand on their own. That is the most important part. This reviewer does have suggestions on how to put the current work in proper context; the current draft appears to overstate the novelty of the findings. Additionally, some sentences need editing for accuracy. None of those concerns impeach the excellent foundational data.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript presents a broad survey of glutamate receptor composition at the neuromuscular junction in Drosophila across developmental stages and muscle types. The topic is clearly important, and the central observation-that adult muscles differ substantially from the canonical larval NMJ-is interesting and potentially impactful. The dataset is extensive and will likely be of value to the community. However, in my view, there are significant limitations in how the data are generated and interpreted, which at present reduce the strength of the conclusions.

      Strengths:

      The study addresses a relevant and timely question and provides a large and systematic dataset. The finding that adult muscles diverge from larval NMJ organization is compelling and challenges a widely held assumption in the field. The breadth of approaches, including genetic reporters, immunohistochemistry, endogenous tagging, and transcriptomic data, is, in principle, a strong aspect of the work and allows for a broad overview of receptor expression across tissues and developmental stages. Even in its current form, the manuscript provides useful descriptive information that will be of interest to the community.

      Weaknesses:

      A major concern is the reliance on a heterogeneous combination of detection methods (GAL4 reporters, antibody staining, endogenous tagging, and RNA), which are treated largely as equivalent lines of evidence. These approaches differ substantially in what they measure and in their sensitivity and specificity. While convergence across methods can in principle be convincing, here this convergence is often inferred from the shared absence of signal. This is problematic because all methods used are susceptible to false negatives for different reasons. As a result, the repeated conclusion that specific GluR subunits are "absent" from adult muscles, including those previously considered essential, is not fully justified by the data presented.

      This issue is not only theoretical. The manuscript itself seemingly contains examples where methods disagree, demonstrating that detection is incomplete and method-dependent. These discrepancies could be better integrated into the interpretation. Instead, negative results across methods are often taken as strong evidence for absence, which overstates the certainty of the findings.

      In addition, antibody validation appears to rely largely on prior work in larval tissue. Given the structural and biochemical differences in adult muscles, it is not clear that staining performance is equivalent, particularly in cases where the signal is weak or undetected. This further complicates the interpretation of negative results.

      More generally, the manuscript moves in several places from descriptive observations to functional or mechanistic implications that are not directly supported. The suggestion that adult muscles operate with fundamentally different receptor assemblies is intriguing, but remains speculative without functional validation. At a minimum, the distinction between observation and interpretation should be made more explicit.

      I thus think that the current conclusions need to be more carefully constrained. Ideally, the study would be strengthened by at least one functional experiment, such as electrophysiological recordings from adult NMJs or perturbation of candidate receptors like GluClα or Clumsy. This would help to anchor the expression data in synaptic function.

      In summary, this is an interesting and potentially important study, but the current manuscript somewhat overinterprets heterogeneous and partly indirect evidence. It will already be useful in its present form, but could be more convincing if the authors more rigorously account for methodological limitations and moderate their claims accordingly.

    4. Reviewer #3 (Public review):

      The Sustar et al. manuscript catalogs glutamate receptor composition across distinct Drosophila NMJs: larval and adult abdominal NMJs, as well as NMJs on adult leg and flight muscles. This work is important and probably overdue. The larval NMJ is the exemplar NMJ in this system, and the identity of "essential" and "alternative" subunits at this stage is assumed by many to hold across developmental stages and NMJ types. Here, the authors show that there is surprising diversification among NMJ types and that the notion of essential/alternative subunits only holds true at larval NMJs.

      The study will generate interest in the Clumsy GluR subunit, which has not been well-characterized at all, but is widely expressed at adult NMJs. They also find striking extrasynaptic expression of glutamate-gated chloride channel GluRClalpha in adult leg and flight muscles, raising questions about its role. The study is interesting, logical, and well-written. The figures are clear, and the discussion was particularly thoughtful. I have a couple of comments that the authors could consider.

      (1) They cite Rivlin et al., (2004) in the Introduction as the sole previous study to investigate the molecular composition of adult NMJs, but do not mention this work again. In the Discussion, it would be helpful to compare/contrast their finding with those of the earlier work.

      (2) Were these analyses done in adults of consistent ages? It seems possible that the GluR subunit composition could be different in very young adults or in aged flies. The age of the animals should be mentioned in the Methods.

      (3) The broad expression of GluCl:V5 in adult leg and flight muscles is surprisingly robust and appears to light up the edges of all muscle fibers. Would the authors comment on the controls that were done to ensure that this staining is real and specific to animals carrying that V5 endogenous tag?

      (4) The snRNAseq data in Figure S12 differ a bit from the IHC/GAL4 data summarized in the table in Figure 2. In particular, the data suggests that Ukar and Grik are widely expressed in adult muscles. Is there a reason not to include an "snRNA seq" column in Figure 2 alongside the data from GAL4 lines and IHC? To my mind, it is about as reliable as GAL4 lines that often capture only a subset of the full expression pattern. In this case, the snRNAseq data suggest that Ukar/Grik are likely at adult flight muscle NMJs, which might be important since NMJ was negative for everything except Neto-beta by IHC.

    1. eLife Assessment

      This work introduces a new paradigm for modeling decision-making under time pressure: rather than having to infer the evidence accumulated by the subjects, experimenters can directly measure it on a trial-by-trial basis. This is an important advance, as it has the potential to address questions that are off limits to the standard paradigm. The methodology and analyses are convincing, especially the ones that manipulate the reward structure. Additional analyses - in particular, a deeper comparison to an ideal observer model - would strengthen and broaden the conclusions.

    2. Joint Public Review:

      Summary:

      Kalburge et al. investigate a task in which human subjects make a decision based on the accumulation of noisy evidence. Tasks like this have been studied for decades, but always with the same essential ingredient: noisy moment-by-moment evidence has to be integrated internally by the subjects, and so is not observed by the experimenter.

      In this study, the authors depart from this scenario and make the evidence visible. Specifically, subjects see a pigeon moving stochastically on a screen, and they have to determine whether the net motion is to the right or to the left. This provides the experimenter direct access - on a trial-by-trial basis - to the bounds the subjects use to make their decision.

      The authors apply this paradigm across a range of tasks, each one differing in how the signal-to-noise ratio (SNR; defined to be the ratio of the drift rate of the pigeons to the standard deviation of the noise) changes over time and across trials. The tasks range from the standard case of constant SNR to the non-standard case where the SNR changes abruptly in the middle of the task.

      The authors determined, on a trial-by-trial basis, the bounds used by the subjects. Setting the bounds optimally when the SNR changes over time or across trials is a non-trivial problem; not surprisingly, then, the subjects were suboptimal. However, they weren't very suboptimal; instead, their behavior was "satisficing" (in the words of the authors), meaning their bounds were reasonably close to the optimal ones. Since the loss is relatively flat near the maximum, and finding the optimal bounds is hard, this is a sensible strategy.

      Strengths:

      The main strength of this work is the introduction of a new paradigm that supports a trial-by-trial measure of the decision bound. This allows direct measurement of the bound at decision time within individual trials. This, in turn, allows experimenters to determine whether the decision bound differs across decision time or fluctuates for the same decision time across trials. This is harder, although not impossible, to do with tasks in which decision bounds have to be estimated across multiple trials, especially when the SNR is changing.

      The authors use this paradigm to show that the decision bounds are mostly constant when the SNR is constant within and across trials. This has been shown indirectly before by fitting models with different parametric boundary shapes, but not directly by measuring the boundary separately for different decision times (but see Kira, Yang, and Shadlen, 2015). They also demonstrate that variability in these bound estimates arises from measurement noise rather than trial-by-trial variability in bound heights, something that could not have been done with previous paradigms.

      They furthermore replicate findings that subjects adjust their bounds, including weak collapse, to changing reward contingencies and SNRs, further validating their paradigm. And finally, the work demonstrates an apparent within-trial bound change if the SNR changes (predictably) mid-trial, as predicted by their previous work (Barendregt et al., 2022). This is -- to our knowledge -- the first confirmation of this prediction.

      Weaknesses:

      There are two non-technical weaknesses.

      First, comparison to optimal behavior was mainly qualitative; a quantitative comparison would greatly strengthen the work.

      Second (although not exactly a weakness), the work does not leverage the full potential of trial-by-trial estimates of the decision bound, which is a missed opportunity. To our understanding, the only finding that relied on trial-by-trial access to the bound was that the variability in the bound estimate was a major source of measurement noise. Their finding that the bound changes to reward contingencies and SNR, on the other hand, did not require such a trial-by-trial estimate. However, with this task (and not standard paradigms), the authors could determine how the bounds change during learning, which would give insight into the learning rules that participants use to adjust their bounds.

      There are also a few technical issues.

      (1) The authors argue that they don't observe a collapsing bound when the SNR varied across blocks (Figure 5). However, they only seem to perform this analysis on the difference in boundaries between trials with different SNRs (Figsures 5B, D). Observing a zero difference implies that the boundary shape is the same across SNRs, but does not rule out a collapse.

      (2) The evidence for a within-trial boundary change for conditions with a within-trial SNR change could be stronger. The data shown in Figures 6C, D is very noisy, and there are no error bars. For individual participants, is the estimated change in bound larger than the variability in bound estimates before and after the SNR changepoint? Are there potentially other measures that could be used to make the point of a clear change in boundary within individual trials more convincing?

      (3) The work assumes that bound height estimates are biased due to the bounded accumulation nature of the decision process, and it corrects for these biases with a simulation-based correction (Methods and Figure 7). To our understanding, this correction assumes that the decision time is the first time that this boundary is crossed. However, the authors do not demonstrate that this is the strategy that participants use; they need to explicitly rule out the possibility that there are significant pigeon excursions across the boundary before the decision time.

      (4) The authors did not consider other stopping rules, such as a decision based on the last few trials. Showing that a stopping rule based purely on the bound fits the data better than other possible rules would strengthen the manuscript.

    1. eLife Assessment

      This important study examines the benefits of spatial cognition in a wild population of mountain chickadees. Using robust genetic analyses and experimental design, the authors show with compelling evidence that females seeking out extra-pair copulations prefer males with strong spatial cognition, and that these males have a reproductive advantage over other males. This work is of broad interest to evolutionary and behavioural biologists.

    2. Reviewer #1 (Public review):

      This manuscript presents compelling evidence from a wild chickadee population linking heritable spatial cognition to extra-pair paternity success, supporting sexual selection via good genes in a food-caching species. The integration of RFID cognition tests with ddRAD paternity assignment is methodologically strong and timely for behavioral ecology, though causal mechanisms and confounds warrant clarification.

      Overall, a major revision of the manuscript is recommended, addressing the points below.

      (1) Confirmation of manipulation and treatment effects. The central claim hinges on spatial cognition driving EP siring, but direct evidence that cognition predicts observed copulations (vs. post-copulatory mechanisms) is absent. While territories do not cluster by performance (Figure S4), quantify male aggression/movement data during fertile periods to rule out intrusion-based EPP. The authors should provide metrics like nearest-neighbor distances for EP sires or playback responses linking cognition to dominance, as in prior chickadee work. Without this, causal female preference remains correlational.

      (2) Female cognition-EPY link inconsistency. Poor female cognition predicts more EPY (first-20-trials: offspring-level χ²=6.21, P=0.013; nests: χ²=6.79, P=0.009), but not for full-task (P>0.5). The authors should discuss why (e.g., learning speed vs. memory stability) and add exploratory correlations (female errors vs. EPY proportion). They should soften claims in the Discussion section of "female-driven" without consistent support and should frame this as a hypothesis.

      (3) Cognitive task sensitivity and validity. Mean errors aggregate learning curves effectively, but single feeder-assignment (non-preferred) confounds neophobia/motivation with spatial ability. The authors should report trial-by-trial improvements (Figure S7 subset) or criterion-to-learn metrics. Justify excluding high-error birds (<3 mean); sensitivity analysis needed to check bias toward high performers.

      (4) Paternity assignment robustness. ddRAD-CERVUS with bimodal LODs (Figure S8) is solid, but unassigned EPY (social-genotyped but no sire) implies missing sires (~?% of EPY?). Include all alive males as candidates yearly? Test power simulations for LOD thresholds. 2019 exclusion justified, but multi-year SNP alignment could boost resolution.

      (5) Mechanistic speculation vs. data. Discussion invokes hippocampus genes (GWAS priors) and good genes, but no offspring cognition/survival data. Label as hypotheses; suggest tracking EPY recruitment. No brood size costs for EP sires is key, but monitor long-term nest investment (e.g., feeding rates).

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the authors ask whether spatial cognition is under sexual selection in mountain chickadees. To do so, the authors examined a large dataset that includes a) spatial cognition data for both males and females (obtained via use of a clever RFID-based feeder system) and b) social and extra-pair paternity nesting data. As predicted, males with higher spatial cognition sired more extra-pair offspring, and extra-pair sires had, on average, higher spatial cognition scores than the males they cuckolded. Interestingly, females with lower spatial cognition scores were more likely to seek extra-pair copulations, potentially to compensate for their own low spatial cognition. Surprisingly, there was no difference in spatial cognition scores between males that sired their own offspring and those that lost paternity at the nest. Also surprising was the fact that there were no differences in patterns of extra-pair paternity and spatial cognition between high- and low-elevation sites. The latter is particularly surprising in that spatial cognition should be under stronger selection at the high elevation site. Overall, this is a fascinating study that demonstrates that spatial cognition - a trait under natural selection as it directly impacts winter foraging and survival behaviour -is also under sexual selection.

      Strengths:

      The authors have a robust dataset (n = 732 offspring sampled over 3 years), high-quality spatial cognition data collected with a procedure that has been well-honed over the years, and couple the data with solid statistical procedures that address many potential covariates and potentially confounding factors. In addition, the authors are careful in the discussion to elaborate on the many potential alternative explanations from the results and questions that are likely to arise in the minds of readers (e.g., how are females assessing male spatial ability?)

      Weaknesses:

      Overall, no major weaknesses were identified in this study. As always, there are editorial issues that I would encourage the authors to consider, including presentation of data/results and clarification on some statistical issues. Overall, however, this is an excellent study that will make an important contribution to our understanding of the evolution of cognition and targets of sexual selection.

    4. Reviewer #3 (Public review):

      Summary:

      The authors presented evidence that spatial cognition in this population is under sexual selection, with extra-pair males, primarily chosen by the females, having better spatial cognition than males they cuckolded and males with better spatial cognition having more extra-pair young.

      Strengths:

      This cognitive ecology study was conducted on a well-known long-term study population of free-ranging mountain chickadees. This strong base, alongside a thorough study design and extensive statistical analyses, enabled the authors to address research questions that few other labs can address, making this a potentially powerful study of broad general interest.

      Weaknesses:

      Throughout the manuscript, there is a focus on the "mean number of location errors per trial over the first 20 trials". Performance changes across trials, so why weren't learning vs peak performance analyzed separately? Similarly, authors also describe results in the context of the entire task, but sometimes in the context of the first 20 trials - why is one prioritised over the other, and why is the emphasis not always consistent? Are the results across the two generally the same? A more thorough explanation addressing all these points is necessary.

      Lines 429-432: Why was a categorical (i.e., chi-square test) and not a numerical comparison implemented? A numerical statistical test would capture more of the variation (i.e., the number of years separating the social and EPY males).

    1. eLife Assessment

      This important study provides evidence that plateau pikas, at moderate densities, can facilitate yak nutrition by suppressing a poisonous plant, offering a helpful perspective on reciprocal interactions between small mammal ecosystem engineers and large herbivores. The evidence is solid, supported by a manipulative field experiment and appropriate measurements of intermediary ecological processes, although some claims about density dependence, competition, and stress-gradient mechanisms are not fully supported by the experimental design. The work will be of interest to ecologists, conservation biologists, and rangeland managers, particularly those studying grassland herbivore interactions and livestock management on the Qinghai-Tibetan Plateau.

    2. Reviewer #1 (Public review):

      Summary:

      This is important and significant work because it helps describe the complexity of interactions between system components where two herbivores interact with vegetation. Whereas other studies have shown that the larger ungulate (yaks, Bos grunniens, in this case) can facilitate the abundance and population growth of the smaller (the semi-fossorial lagomorph, Ochotona curzoniae, plateau pika hereafter), this study flips the tables and shows that, at least under some conditions, moderate densities of the plateau facilitate the nutritional condition of yaks.

      The study was not designed to investigate the reasons that pikas clip Stellera chamaejasme. That said, based on other studies and general knowledge of the ecology of these pikas, it is likely that they clip (although do not eat) this plant because its relatively large size hinders predator detection. This species of pika does better where vegetation height is low than where it is higher.

      Strengths:

      Notably, the strong inference the authors can claim for their results is supported by the careful experimental design. A weaker paper would have simply noted correlations between pika burrow density and yak feeding efficiency without experimental removal. This paper, to its credit, not only used experimental removals but also documented the various intermediary results that support the ultimate conclusions. The statistical approaches used appear to be appropriate. (Readers are encouraged to read the full Materials and Methods, which are available in the Supplementary Materials section.)

      Weaknesses:

      Although the study was well designed and executed, and its conclusions appear strongly supported, readers interested in the management implications of the Qinghai-Tibetan Plateau should be mindful of its limitations. First, the study site, at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera chamaejasme becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. Thus, it would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace S. chamaejasme as the problematic plant for pastoralists. Second, the authors make no mention of wild ungulates, so it is unclear what, if any, role they may have played in this system. At least one study in Qinghai Province, albeit at a slightly higher elevation, showed that not only pikas, but also Tibetan gazelles (Procapra picticaudata), which were commonly observed on grazed pastures, grazed more frequently on some dicots avoided by domestic sheep than did the livestock themselves (Harris et al. 2015). It would also be instructive to learn if similar facilitation as observed here applied to the other principal livestock species in the area, domestic sheep (which are often herded together with smaller numbers of domestic goats). Finally, as suggested by this study, the interactions between all components of the system are complex and interactive. If pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition?

      Citation:

      Harris RB, Wang, WY, Badinqiuying , Smith AT, Bedunah DJ (2015) Herbivory and Competition of Tibetan Steppe Vegetation in Winter Pasture: Effects of Livestock Exclosure and Plateau Pika Reduction. PLoS ONE 10(7): e0132897. doi:10.1371/journal.pone.0132897

    3. Reviewer #2 (Public review):

      Summary:

      This study uses a combination of field sampling and manipulative experiments to test for facilitative impacts of pikas on yaks via suppression of a poisonous forb. The authors found that, when Stellera forbs were present, yak weight increases over the growing season were greater in the presence of pikas compared to in their absence. This occurred because, although pikas do not consume Stellera, they clip it and use it in nest/burrow construction, thereby decreasing its relative abundance in the plant community. Thus, overall, the study contributes to our understanding of how herbivores of different size classes indirectly affect each other via the use of shared resources.

      Strengths:

      It is well known that large herbivores on grasslands impact smaller animals, but the reciprocal interaction is rarely tested. Thus, this study asks a valuable question, and the experiment is well-designed to test it. The authors also do a good job of demonstrating the potential conservation impacts of their research.

      Weaknesses:

      What the authors tested is really cool, but their claims go far beyond what they can say based on their experimental design. For example, the authors claim to show that pika impacts on yaks display density-dependent transitions from competition to facilitation. However, their experiment only looked at the presence (at moderate densities) and absence of pikas, and they only tested for facilitation, not competition.

      The paper would also benefit from changes to the framing in the introduction and discussion. For example, the authors pitch the work as a test of the stress-gradient hypothesis. However, there is no abiotic stress gradient in the study, which is an essential component of the SGH. They also pitch the work in terms of density dependence, but there is no significant variation in population densities beyond the presence-absence binary. The paper would be stronger if they focused their framing around the literature on facilitative interactions across mammals of different size classes, especially indirect facilitation via use of shared resources, which is what this paper is really about.

      Finally, the paper has significant weaknesses in the experimental and statistical methodology. Most importantly, there are inconsistencies in what is visualized in the figures compared to the model results. For example, the results section in several places notes a lack of significant interaction terms in the model but shows interactions in the p-values on the figures. The authors also plot smoothed lines rather than their model results and then draw interpretations from those lines that cannot be tested in the models that they used. There are also missing details that are important for model interpretation, including the distributions used and the sample sizes. Another major concern with experimental design is in the forage nutrient analyses. The authors picked plants along a grazing trail, then measured nutrient content without standardizing based on plant species, so any differences across treatments could be because of what they happened to grab rather than overall forage quality.

    4. Author response:

      eLife Assessment

      This important study provides evidence that plateau pikas, at moderate densities, can facilitate yak nutrition by suppressing a poisonous plant, offering a helpful perspective on reciprocal interactions between small mammal ecosystem engineers and large herbivores. The evidence is solid, supported by a manipulative field experiment and appropriate measurements of intermediary ecological processes, although some claims about density dependence, competition, and stress-gradient mechanisms are not fully supported by the experimental design. The work will be of interest to ecologists, conservation biologists, and rangeland managers, particularly those studying grassland herbivore interactions and livestock management on the Qinghai-Tibetan Plateau.

      Thank you very much for these positive assessments of our work, below we provided the point-by-point responses to the comments from the 2 peer reviewers, and we hope these revisions are satisfied.

      Reviewer #1 (Public review):

      Summary:

      This is important and significant work because it helps describe the complexity of interactions between system components where two herbivores interact with vegetation. Whereas other studies have shown that the larger ungulate (yaks, Bos grunniens, in this case) can facilitate the abundance and population growth of the smaller (the semi-fossorial lagomorph, Ochotona curzoniae, plateau pika hereafter), this study flips the tables and shows that, at least under some conditions, moderate densities of the plateau facilitate the nutritional condition of yaks.

      The study was not designed to investigate the reasons that pikas clip Stellera chamaejasme. That said, based on other studies and general knowledge of the ecology of these pikas, it is likely that they clip (although do not eat) this plant because its relatively large size hinders predator detection. This species of pika does better where vegetation height is low than where it is higher.

      Strengths:

      Notably, the strong inference the authors can claim for their results is supported by the careful experimental design. A weaker paper would have simply noted correlations between pika burrow density and yak feeding efficiency without experimental removal. This paper, to its credit, not only used experimental removals but also documented the various intermediary results that support the ultimate conclusions. The statistical approaches used appear to be appropriate. (Readers are encouraged to read the full Materials and Methods, which are available in the Supplementary Materials section.)

      We appreciate these positive comments on our work.

      Weaknesses:

      Although the study was well designed and executed, and its conclusions appear strongly supported, readers interested in the management implications of the Qinghai-Tibetan Plateau should be mindful of its limitations. First, the study site, at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera chamaejasme becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. Thus, it would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace S. chamaejasme as the problematic plant for pastoralists.

      Agree! We will acknowledge this limitation in the Discussion, by adding the paragraph below (see the Third point):

      “Despite of these, several questions remain deserve further investigation. First, our study examined pika–yak interactions only during the summer period, when food resources are most abundant. Whether such facilitative effects weaken or even shift toward competition under more stressful conditions—for example, when forage becomes limited during autumn or winter—remains to be tested. Second, if pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition (Yang et al., 2026)? Third, our study site located at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. It would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace Stellera as the problematic plants for pastoralists (Lu et al., 2012; Li and Zhao, 2025). Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      Second, the authors make no mention of wild ungulates, so it is unclear what, if any, role they may have played in this system. At least one study in Qinghai Province, albeit at a slightly higher elevation, showed that not only pikas, but also Tibetan gazelles (Procapra picticaudata), which were commonly observed on grazed pastures, grazed more frequently on some dicots avoided by domestic sheep than did the livestock themselves (Harris et al. 2015). Citation:

      Harris RB, Wang, WY, Badinqiuying , Smith AT, Bedunah DJ (2015) Herbivory and Competition of Tibetan Steppe Vegetation in Winter Pasture: Effects of Livestock Exclosure and Plateau Pika Reduction. PLoS ONE 10(7): e0132897.

      doi:10.1371/journal.pone.0132897

      Agree! We will add more details about the study site, particularly regarding wild ungulates, in the Methods section. Specifically, we will include the following sentence: “Wild ungulates, such as Tibetan gazelles (Procapra picticaudata) (Harris et al., 2015), and other small mammals such as rabbits and zokors, occur rarely in the area.” This key reference will also be cited in this section.

      It would also be instructive to learn if similar facilitation as observed here applied to the other principal livestock species in the area, domestic sheep (which are often herded together with smaller numbers of domestic goats).

      Agree! The same as mentioned above. We will acknowledge this limitation in the Discussion, by adding the paragraph below (see the Final point):

      “Despite of these, several questions remain deserve further investigation. First, our study examined pika–yak interactions only during the summer period, when food resources are most abundant. Whether such facilitative effects weaken or even shift toward competition under more stressful conditions—for example, when forage becomes limited during autumn or winter—remains to be tested. Second, if pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition (Yang et al., 2026)? Third, our study site located at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. It would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace Stellera as the problematic plants for pastoralists (Lu et al., 2012; Li and Zhao, 2025). Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      Finally, as suggested by this study, the interactions between all components of the system are complex and interactive. If pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition?

      Agree! The same as mentioned above. We will acknowledge this limitation in the Discussion, by adding the paragraph below (see the Second point):

      “Despite of these, several questions remain deserve further investigation. First, our study examined pika–yak interactions only during the summer period, when food resources are most abundant. Whether such facilitative effects weaken or even shift toward competition under more stressful conditions—for example, when forage becomes limited during autumn or winter—remains to be tested. Second, if pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition (Yang et al., 2026)? Third, our study site located at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. It would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace Stellera as the problematic plants for pastoralists (Lu et al., 2012; Li and Zhao, 2025). Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      Reviewer #2 (Public review):

      Summary:

      This study uses a combination of field sampling and manipulative experiments to test for facilitative impacts of pikas on yaks via suppression of a poisonous forb. The authors found that, when Stellera forbs were present, yak weight increases over the growing season were greater in the presence of pikas compared to in their absence. This occurred because, although pikas do not consume Stellera, they clip it and use it in nest/burrow construction, thereby decreasing its relative abundance in the plant community. Thus, overall, the study contributes to our understanding of how herbivores of different size classes indirectly affect each other via the use of shared resources.

      Strengths:

      It is well known that large herbivores on grasslands impact smaller animals, but the reciprocal interaction is rarely tested. Thus, this study asks a valuable question, and the experiment is well-designed to test it. The authors also do a good job of demonstrating the potential conservation impacts of their research.

      We appreciate these positive comments on our work.

      Weaknesses:

      What the authors tested is really cool, but their claims go far beyond what they can say based on their experimental design. For example, the authors claim to show that pika impacts on yaks display density-dependent transitions from competition to facilitation. However, their experiment only looked at the presence (at moderate densities) and absence of pikas, and they only tested for facilitation, not competition. The paper would also benefit from changes to the framing in the introduction and discussion. For example, the authors pitch the work as a test of the stress-gradient hypothesis. However, there is no abiotic stress gradient in the study, which is an essential component of the SGH. They also pitch the work in terms of density dependence, but there is no significant variation in population densities beyond the presence-absence binary. The paper would be stronger if they focused their framing around the literature on facilitative interactions across mammals of different size classes, especially indirect facilitation via use of shared resources, which is what this paper is really about.

      We agree that our work had explored only the facilitative effects of pikas on yaks, rather than the density-dependent balance between competition and facilitation, and the Stress Gradient Hypothesis (SGH).

      We plan to make the major revisions below to address this important concern.

      (1) We will revise the title as “Moderate density of small mammalian herbivores facilitates livestock growth in grasslands ”.

      (2) We will delete all the statements about density-dependent transition of facilitation and competition and the SGH in the Abstract, Introduction, Discussion, and the References sections.

      Finally, the paper has significant weaknesses in the experimental and statistical methodology. Most importantly, there are inconsistencies in what is visualized in the figures compared to the model results. For example, the results section in several places notes a lack of significant interaction terms in the model but shows interactions in the p-values on the figures.

      In the Results section, there are only two locations where we discussed non-significant interactions: Line 148–149 “Pikas and Stellera had no interactive effects on abundance of sedges, forbs, and neutral detergent fiber (NDF) of total forage for yaks (Fig. 3F,I, fig. S1, table S3,5).” and Line 161–162 “Pikas and Stellera had no interactive effects on yaks’ foraging efficiency on forbs (fig. S2, table S7).”.

      We have cross-checked both the manuscript as submitted and the website, and in every instance we are consistent in not reporting interactions as non-significant when the model output shows significance.

      We will confirm these details in the revised version as “Pikas and Stellera had no interactive effects on abundance of sedges, forbs, and neutral detergent fiber (NDF) of total forage for yaks (Fig. 3F,I, Fig. S1, Table S5, S8). ”; and “Pikas and Stellera had no interactive effects on yaks’ foraging efficiency on forbs (Fig. S2, Table S10).” in the Results section.

      The authors also plot smoothed lines rather than their model results and then draw interpretations from those lines that cannot be tested in the models that they used.

      Agree! There are only two figures in which we used generalized additive models (GAMs) to plot smoothed lines: Figure 2C and Figure 3C.

      For Figure 2C, the supplementary table for the GAMM associated with the smoothed line was not originally included, but we will add it as Table S4 in the revised version. For Figure 3C, we explicitly fit a GAMM corresponding to the plotted line, and the model results will be reported in the Table S7 in the revised version.

      There are also missing details that are important for model interpretation, including the distributions used and the sample sizes.

      Agree! We will provide the Table S13 to summarize all statistical models used in the study, including the distributions used and the sample sizes in the Supplementary Materials. We will also add a sentence of “A summary of all statistical models used in the study is available in table S13.” in the Statistical analyses section to indicate this information.

      Another major concern with experimental design is in the forage nutrient analyses. The authors picked plants along a grazing trail, then measured nutrient content without standardizing based on plant species, so any differences across treatments could be because of what they happened to grab rather than overall forage quality.

      We will revise this section to provide more details on how forage samples were collected and their quality were analyzed. Specifically, five forage samples were collected per grazing plot, focusing on the two dominant plant species —one sedge and one grass—that were most frequently grazed by yaks. To ensure comparability across plots and treatments, we mixed the two species at equal dry mass (5 g). We will revise this section as below.

      “To assess forage quality, five forage samples were collected from each grazing plot to quantify their nutritive values. To obtain samples that reflect the forage actually consumed by yaks, we tracked the animals along their grazing paths and collected the plant tissues of the two most frequently consumed species: the dominant sedge Kobresia humilis and the dominant grass Elymus nutans (Fig. 2B; Pan et al., 2019). The collected tissues of each species were dried in a forced-air oven at 60 °C for 48 h, then ground through a 1-mm mesh. Subsequently, 5 g of each dried and ground species were combined in a 1:1 dry mass ratio, and the resulting mixture was stored in plastic bags for subsequent analyses.”

    1. eLife Assessment

      The authors combine a modeling approach, using a digital twin, with electrophysiological evidence in two species to assess the role of inhibition in shaping selectivity in the visual cortex. The results provide a fundamental advance beyond the classic view of sensory coding by proving compelling evidence that many neurons in visual areas exhibit dual-feature selectivity. Overall, the work compellingly showcases how in silico experiments can generate concrete hypotheses about neuronal coding that are difficult to discover experimentally.

    2. Reviewer #1 (Public review):

      The multi-species approach of testing the model in macaque and mouse is excellent, as it improves the chances that the observed findings are a general property of mammalian visual cortex. It would be useful to delineate however any notable differences between these species, which are to be expected given their lifestyle.

      The overall performance of the model appears to be excellent in V1, with over 80% performance, but falls substantially in V4. It would be important to consider the implications of this finding; for example, in the context of studying temporal lobe structures that are central to recognizing objects. Would one expect that model performance decreases further here, and what measures could be taken to avoid this? Or is this type of model better restricted to V1 or even LGN?

      While the manuscript delineates novel axes of inhibitory interactions, it remains unclear what exactly these axes are and how they arise. What are the steps that need to be taken to make progress along these lines?

      Comments on revised version.

      The authors have adequately addressed the points I raised in my review during the revision.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This manuscript used deep learning to highlight the role of inhibition in shaping selectivity in primary and higher visual cortex. The findings hint at hitherto unknown axes of structured inhibition operating in cortical networks with a potentially key role in object recognition.

      The multi-species approach of testing the model in macaque and mouse is excellent, as it improves the chances that the observed findings are a general property of mammalian visual cortex. However, it would be useful to delineate any notable differences between these species, which are to be expected given their lifestyle.

      The overall performance of the model appears to be excellent in V1, with over 80% performance, but it falls substantially in V4. It would be important to consider the implications of this finding; for example, in the context of studying temporal lobe structures that are central to recognizing objects. Would one expect that model performance decreases further here, and what measures could be taken to avoid this? Or is this type of model better restricted to V1 or even LGN?

      While the manuscript delineates novel axes of inhibitory interactions, it remains unclear what exactly these axes are and how they arise. What are the steps that need to be taken to make progress along these lines?

      Reviewer #2 (Public review):

      The classic view of sensory coding states that (excitatory) neurons are active to some preferred stimuli and otherwise silent. In contrast, inhibitory neurons are considered broadly tuned. Due to the gigantic potential image space, it is hard to comprehensively map the tuning of individual neurons. In this tour de force study, Franke et al. combine electrophysiological recordings in macaque (V1, V4) and mouse (V1, LM, LI) visual cortex with large-scale screens based on digital twin models, as well as beautiful systems identification (most/least activating stimuli). Based on these digital twins, they discover dual-feature selectivity (which they validate both in macaques and mice). Dual-feature selectivity involves a bidirectional modulation of firing rates around an elevated baseline. Neurons are excited by specific preferred features and systematically suppressed by distinct, non-preferred features. This tuning was identified by excellently combining advances in AI & high-throughput ephys.

      The study is comprehensive and convincing. Overall, this work showcases how in silico experiments can generate concrete hypotheses about neuronal coding that are difficult to discover experimentally, but that can be experimentally validated! I think this work is of substantial interest to the neuroscience community. I'm sure it will motivate many future experimental and computational studies. In particular, it will be of great interest to understand when and how the brain leverages dual-feature selectivity. The discussion of the article is already an interesting starting point for these considerations.

      Strengths:

      (1) Using computational models to predict neuronal responses allowed them to go through millions of images, which may not be possible in vivo.

      (2) The cross-species and cross-area consistency of the results is another major strength. Pointing out that the results may be a fundamental strategy of mammalian cortical processing.

      (3) They show that the feature causing peak excitation in one neuron often drives suppression in another. This may be an efficient coding scheme where the population covers the visual manifold. I'd like to understand better why the authors believe that this shows that there are low-dimensional subspaces based on preferred and non-preferred stimulus features (vs. many more, but some axes are stronger).

      We thank the reviewers for their constructive and helpful feedback on our manuscript. We are delighted that they found the study to be “comprehensive and convincing” and a “tour de force” in its combination of electrophysiological recordings with large-scale digital twin screening. We appreciate that the reviewers highlighted the strengths of our multi-species approach and the “cross-species and cross-area consistency” of the results, noting that the work showcases how in silico experiments can generate concrete, experimentally validatable hypotheses. Overall, we agree with the assessment of the reviewers. We have performed the following changes to the text to clarify and strengthen the manuscript, without introducing new analyses or altering the conclusions. 

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Page 3: The authors state that RFs were mapped using sparse noise, with the goal to ensure that the RFs align with the visual stimulus, but no data appear to be shown regarding this alignment. It would be important to provide a full analysis of the sparse noise-mapped RFs for both V1 and V4. Also, is it correct that the V4 data analyzed here came from a single animal? This could potentially be problematic and would need to be addressed, for example, by performing analyses also in V1 for participant animals separately. Please elaborate.

      We have added a sentence to the Results section clarifying the sparse noise RF mapping procedure, noting that probe insertions were targeted orthogonal to the cortical surface so that neurons sampled along the probe depth share overlapping receptive fields, allowing a single stimulus configuration to adequately drive the entire recorded population. We have also corrected the text to clarify that V4 data were collected from 2 animals (not 3 as previously stated in an earlier draft), consistent with the Methods section.

      (2) Page 4: Only half the neurons in V4 are "high confidence" in terms of test image performance, which seems a little low and probably significantly lower than the corresponding value for V1 of 84%. It is unclear how to interpret this confidence, but it seems to suggest that half of the V4 neurons are not well captured by the model. If true, this fraction appears large enough to cast doubt on the validity of the V4 results. Please elaborate.

      We have expanded the text to explicitly discuss the lower proportion of high-confidence in-silico neurons in V4 relative to V1. We attribute this to the greater complexity of V4 tuning compared to V1, as well as missing contextual information such as image surrounds and sequential image context—factors that likely limit model performance in higher visual areas. We note that our restriction of analyses to high-confidence neurons provides resilience against these limitations, and that the goal was not to maximize predictive performance per se but to identify response patterns—dual-feature selectivity—that are robust across neurons, areas, and species.

      (3) Page 5: It seems that identical L2 norms are valid for discounting contrast variations, particularly if the neural responses are linear, since the L2 norm is computed on the entire RF. It might be judicious to attenuate the claim that contrast variation has no effect.

      We have softened the claim that contrast variation has no effect. The revised text now states that L2 normalization controls for root-mean-squared contrast but does not fully equate effective contrast in nonlinear cells, whose responses depend on the spatial structure of the stimulus beyond its total energy. We note that residual contrast dependent effects, particularly in the suppressive regime, cannot be entirely excluded.

      (4) Page 6: The authors acknowledge that, at least for simple cells, a phase shift in the grating and concomitant ON-OFF overlap is an inhibitory axis, which is correct. It does not really become clear what other axes were found, and whether any of these represent a novel discovery about V1.

      We have clarified the description of inhibitory axes in V1, noting that while phase-shifted stimuli represent a well-established suppressive axis for simple cells reflecting linear On-Off subfield structure, and complex cells exhibit no coherent suppressive pattern due to phase pooling, neither model class accounts for the multidimensional suppressive structure we observe. We have made explicit that our unbiased approach reveals suppressive structure spanning simultaneous changes across orientation, spatial frequency, phase, and texture, exceeding what any single known suppressive mechanism predicts.

      (5) Page 7: Dreamsim is based on human similarity judgements, whereas the data is from macaques. Is there any evidence suggesting that macaque similarity judgements might be similar to those of humans?

      We have added a paragraph to the Discussion acknowledging that DreamSim was trained on human perceptual similarity judgments while our neuronal data are from macaques. We note that this cross-species application is supported by the deep homology between primate ventral visual streams, and that natural-image similarity judgments have been found to be highly consistent across macaques and humans. Importantly, we clarify that we deploy DreamSim not as a model of macaque perception but as an image feature embedding to test whether stimuli that cluster in perceptual space evoke similar neuronal responses—a use that is robust to the precise calibration of the metric. We also note that we are developing custom macaque-specific embeddings for future work.

      (6) Page 7: How many images were in the test set?

      We have added the number of test images to the relevant text (n=75 for V1, n=150 for V4) and to the Figure 1 caption.

      (7) Page 8: As mentioned above, performing the analysis on V1 data of individual subjects and demonstrating similar digital twins might be an additional way to confirm the models' accuracy.

      We have added text noting that for V4, 1digital twin models were fit independently per neuron without sharing information across animals, and that extreme image sets identified by the model elicited correspondingly extreme responses in neurons from the other animal, confirming that identified selectivity patterns are not idiosyncratic to individual subjects.

      (8) Page 11: The mouse data is presented very briefly only, and the authors seem to imply that there is a high degree of coding similarity between this rodent species and macaques and, by extension, humans. Were there any notable differences between the mouse and macaque data?

      We have added text explicitly noting that while macaque and mouse visual cortex differ substantially in their functional organization and the complexity of neuronal selectivity, the broader principle—that non-sparse neurons are jointly defined by distinct excitatory and suppressive feature sets—generalizes across mammalian visual systems. We clarify that this does not imply that mouse and macaque visual cortex share similar functional organization or equivalent complexity of neuronal selectivity; rather, within the representational regime of each area, neurons are organized such that excitatory and suppressive feature sets are jointly structured and distinct.

      (9) Page 13: One main finding of the study is that inhibition appears to operate along additional dimensions that had not been previously recognized, but what is the nature of these dimensions, how do they arise and relate to known inhibitory effects in V1 such as centre-surround effects? The fact that suppression is tuned in response to natural images or other complex objects is not a new finding, and there is plenty of published work along these lines; the authors may want to cite Tamura et al 10.1152/jn.01267.2003. I am not sure introducing the term "dual feature selectivity" is really a major conceptual advance.

      We have added a citation to Tamura et al. (2004) in the Discussion, alongside other prior work documenting suppression by non-optimal stimuli. We have also expanded the Discussion to more carefully position our findings relative to existing work on feature-selective suppression, noting that while prior work has established that inhibition can be structured and feature-selective, our results suggest a broader organizing principle: within each visual area, there exists a set of feature combinations from which individual neurons draw both their excitatory and suppressive preferences.

      (10) Page 14: The authors enumerate a number of technical limitations, which is to be commended. It would be useful for them to comment on the particular advantages of the digital twin model, compared to a more traditional analysis of the responses to the thousands of natural images that were experimentally obtained. It seems likely that the main finding, i.e. tuned inhibition, is also evident directly in this population (?). While the digital twin is to some degree validated by the test images, its responses to the much larger set of images studied are not validated, and one must trust that the ResNet50 indeed captures V4 selectivity. It would be useful to discuss some of these points, and highlight a potential way that digital twins (maybe as a shared model between laboratories) can learn from a large number of animals and datasets, and maybe even be used to generate novel visual stimuli suitable to test emergent hypotheses.

      We have added a paragraph to the Discussion explicitly contrasting the advantages of digital twin models with direct analysis of experimentally recorded responses, noting that digital twins enable screening of more than one million images per neuron in silico, gradient-based synthesis of stimuli precisely optimized to drive or suppress individual neurons, and cross-model verification of identified selectivity patterns—a test that has no analog when working with fixed experimental image sets.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) Call out Figure 1/b in the main text. 

      We have added a callout to Figure 1b in the main text

      (2) Can you make a supplementary figure illustrating more examples with skewness around the middle (e.g. 1.5, 2, 2.5)? Namely, you state that 2 is a good threshold for deciding if it is non-sparse, but you only present clear-cut cases in Figure 2 (with <0.75 and >3.5). I am wondering if 2 is a good threshold?

      We have revised the text to clarify that the skewness threshold of 2.0 is adopted purely for analytical convenience to focus subsequent analyses on neurons with sufficiently graded response distributions, and that the key findings are not dependent on the exact threshold chosen. We explicitly note that the underlying distribution of sparsity is continuous, consistent with recent findings (Gondur et al., 2025).

      (3) The reference "A tale of two tails: Preferred and anti-preferred natural stimuli in visual cortex." Has no authors. I know it's anonymous, but maybe put that for now? I also congratulate including a paper that is anonymously under review at ICLR 2026. I don't find Unk, 2025 in the list of references. Perhaps related?

      We have updated the reference “A tale of two tails” to include the authors (Gondur et al., 2025) and ensured it appears consistently in the reference list. We have also resolved the missing “Unk, 2025” citation, which now correctly refers to this same work.

      (4) Why do you use a different model for the analysis in Figure 8?

      We have added text to the Methods and Results clarifying why a distinct architecture was used for the V4 evaluator model in Figure 8. Specifically, the V4 generator model uses a fixed, pretrained ResNet50 backbone whose weights are deterministic; any re-trained model sharing this backbone would not constitute a genuinely independent evaluation. By contrast, for V1, the ConvNeXt core is fine-tuned from different random initializations, producing architecturally equivalent but computationally independent models. A truly independent V4 evaluator therefore required a fundamentally different architecture.

    1. eLife Assessment

      This fundamental work substantially advances our understanding of tissue deformation and growth patterns during the earliest stages of mammalian heart development. One of the strengths of the work is the compelling quantitative approach to analyzing time-lapse imaging data using an original computational pipeline, which goes beyond the current state of the art and provides new insights into heart tube formation. Overall, this rigorous study will be of broad interest to computational and developmental biologists studying tissue dynamics.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed all the comments raised in the previous round of review. The revised manuscript includes new labeling experiments revealing boundary compression at the cardiac poles consistent with the authors predicted dynamic model of heart tube formation.]

      Summary:

      The study by Raiola et al. conducted a quantitative analysis of tissue deformation during the formation of the primitive heart tube from the cardiac crescent in mouse embryos. Using the tools developed to analyze growth, anisotropy, strain, and cell fate from time-lapse imaging data of mouse embryos, the authors elucidated the compartmentalization of tissue deformation during heart tube formation and ventricular expansion. This paper describes how each region of the cardiac tissue changes to form the heart tube and ventricular chamber, contributing to our understanding of the earliest stages of cardiac development.

      In order to understand tissue deformation in cardiac formation, it is commendable that the authors effectively utilized time-lapse imaging data, a data pipeline, and in silico fate mapping. The study clarifies the compartmentalization of tissue deformation by integrating growth, anisotropy, and strain patterns in each region of the heart.

    3. Reviewer #2 (Public review):

      The authors address an important challenge in developmental biology: the quantitative description of tissue deformation during organogenesis. They have developed a new pipeline to quantify early heart tube morphogenesis in the mouse, with cellular resolution. They adopt an elegant approach by integrating multiple 3D time-lapse datasets into a dynamic atlas of cardiac morphogenesis in order to compute spatio-temporal deformation patterns. The main findings highlight a strong compartmentalization of cell behaviors, with tissue growth and anisotropy exhibiting complementary and spatially segregated patterns. Using these data, the authors developed an in-silico fate mapping tool to interrogate cell displacement within the myocardium. This virtual model provides new mechanistic insights into how the bilateral cardiac primordia converge and transform into a three-dimensional heart tube. The authors identify "belt-like" constraints at the arterial and venous poles that prevent tissue expansion and thus shape the ventricular barrel morphology.

      The computational framework is highly innovative and impressive, providing an unprecedented 3D model of tissue deformation during heart morphogenesis. It also opens avenues for testing hypotheses regarding tissue growth and the forces that cause cell motion.

      Overall, this carefully performed study provides a new model for exploring tissue deformation during organogenesis and will be of broad interest to computational and developmental biologists.

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript by Raiola and colleagues entitled "Quantitative computerized analysis demonstrates strongly compartmentalized tissue deformation patterns underlying mammalian heart tube formation" takes a highly quantitative approach to interrogating the earliest stages of cardiogenesis (12 hours, from early cardiac crescent to early heart tube) in a new and innovative way. The paper presents a new computational framework to help identify both regional and temporal patterns of tissue deformation at cellular resolution. The method is applied to live embryo imaging data (newly generated and from the group's previous pioneering work). In the initial setup, the new model was applied directly to raw time-lapse data, and the results were compared to actual cell tracks identified manually, showing close correlations of the model with the manual tracking. Next, they integrated spatial and temporal information from different embryos to generate a new model for tissue movement, driven by parameters such as tissue growth and anisotropy. Key findings from their model suggest that there are distinct compartments of tissue deformation patterns as the bilateral cardiac crescent develops into the linear heart tube, and that the ventricular chamber forms by a defined expansion pattern, as a 'hemi-barrel shape', with the arterial and venous poles (IFT and OFT) acting as the harnessing belts constraining the expansion of the chamber further. Lastly, the model is tested for its ability to predict future residence of cardiac crescent cells in the heart tube, which it seems to be able to do successfully based on fate tracking validation experiments.

      The manuscript provides an exceptionally careful analysis of a critical stage during heart development - that of the earliest stages of morphogenesis, when the heart forms its first tube and chamber structures. While numerous studies have interrogated this stage of heart development, few studies have performed time-lapse imaging, and, to my knowledge, no other report has performed such in in-depth quantitative analysis and modeling of this complex process. The computational model applied to normal heart development of the myocardium (labelled by Nkx2-5) has revealed multiple new and interesting concepts, such as the distinct compartments of tissue deformation patterns and the growth trajectories of the emerging ventricle. The fact that the model operates at cellular resolution and over a nearly continuous time period of approximately 12 hours allows for unprecedented depth of the analysis in a largely unbiased manner. Going forward, one can imagine such models revealing additional information on these processes, performing analyses of subpopulations that form the heart, and maybe most importantly, applying the model to various perturbation models (genetic or otherwise). The manuscript is very well written, and the data display is accessible and transparent.

      No major weaknesses are noted with the study. It would have been very exciting to see the model applied to any kind of perturbation, for example, a left-right defect model, or a model with compromised cardiac progenitor populations. However, the amount of live imaging required for such analyses renders this out of scope for the current study.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      The study by Raiola et al. conducted a quantitative analysis of tissue deformation during the formation of the primitive heart tube from the cardiac crescent in mouse embryos. Using the tools developed to analyze growth, anisotropy, strain, and cell fate from timelapse imaging data of mouse embryos, the authors elucidated the compartmentalization of tissue deformation during heart tube formation and ventricular expansion. This paper describes how each region of the cardiac tissue changes to form the heart tube and ventricular chamber, contributing to our understanding of the earliest stages of cardiac development.

      Strengths:

      In order to understand tissue deformation in cardiac formation, it is commendable that the authors effectively utilized time-lapse imaging data, a data pipeline, and in silico fate mapping.

      The study clarifies the compartmentalization of tissue deformation by integrating growth, anisotropy, and strain patterns in each region of the heart.

      Weaknesses:

      The significance of the compartmentalization of tissue deformation for the heart tube formation remains unclear.

      While it is obvious that the patterns of deformation should be relevant to model the cardiac crescent into the primitive cardiac tube, we do not provide direct evidence that changing these patterns affects heart tube formation. In this sense, the Reviewer is correct and this is a limitation of the study.

      Reviewer #1 (Recommendations for the authors):

      (1) It is interesting that growth rate and anisotropy are anticorrelated. However, the functional significance of this anticorrelation in heart formation remains unclear. It may be worthwhile to analyse the importance of the relationship between the two by adding inhibitors to cultured embryos or using mutant mouse models.

      We appreciate this thoughtful suggestion and agree that such experimental approaches, involving inhibitors or mutant mouse models, could provide powerful validation of the proposed relationship. However, generating the appropriate lines and performing the necessary quantifications would represent a substantial effort that extends beyond the scope of the current study. Our focus here is to establish the correlation and its potential implications, leaving these more in-depth mechanistic investigations for future work.

      (2) The authors claim to have analysed tissue deformation at the cellular level. Although cell labelling of specific regions using Tat-Cre and DiI injection and tracking of their fate have been performed, this still gives the impression of tissue-level analysis. An analysis "at the cellular level" would be expected to describe morphology, proliferation, polarity, etc., at the single-cell to multi-cell level.

      We thank the reviewer for the comment. Our analysis does not involve single-cell characterization (e.g., morphology, proliferation, polarity) but focuses on quantifying tissue motion. The motion extracted from the images achieves cellular-level precision, as demonstrated by testing the registration algorithm and validating it against cellular tracking experiments. The accuracy of the method is therefore at the cellular scale. The goal of our study is to describe tissue dynamics during heart development, not to perform detailed cellular analyses. The novelty of our approach is that it enables tissuescale quantification in developmental mouse heart imaging, where cell density and image resolution make automated single-cell tracking unfeasible. By using fluorescence labelling as markers, we obtain cellular-level accuracy in tissue motion quantification.

      (3) It is stated that cardiomyocytes, cardiac mesodermal cells, and SHF cells were labeled with Nkx2.5-GFP/Nkx2.5-Cre, Mesp1-Cre, and Islet1-Cre, respectively; however, the results of the labeling using these mice are not presented, and the reason for using different mouse strains is not apparent. Information on these mouse strains is missing in the Materials & Methods section. In particular, attention must be paid to mice of the same name but different strains. Islet1-Cre mice are not SHF-specific and exhibit activity in part of the left ventricular progenitors. Sparse labeling induced by low-dose tamoxifen administration is also unclear regarding the timing and concentration of tamoxifen administration. The authors should provide data on labeling efficiency and region, and also discuss the usefulness of analyses using different mouse strains.

      We thank the reviewer for raising these important points. In this study, the use of different mouse strains was not driven by a biological comparison or lineage-specific analysis, but by the availability of high-quality cardiac imaging datasets generated in the laboratory. The primary goal of this work is methodological: to demonstrate that developmental cardiac imaging data can be reused within an engineering framework to quantify tissue deformation. For this purpose, we do not track individual cells but instead use fluorescence labelling as a versatile strategy to follow tissue motion without requiring a strain- or lineage-specific labelling.

      We acknowledge that Islet1-Cre mice are not SHF-specific and exhibit activity in part of the left ventricular progenitors. However, this limitation does not affect our analysis, as the specificity of the labelled cells is not used in the image processing or deformation quantification pipeline.

      Regarding tamoxifen, we clarified the dosage and administration in the revised Experimental workflow section. Importantly, tamoxifen treatment does not influence the proposed image analysis framework, since labelled cells are employed solely as fiducial references to provide ground-truth validation of the tissue motion estimated from image registration.

      We made these points clearer in the Results section and in Materials and Methods.

      (4) It is noteworthy that the authors have utilized many new analytical methods that they have developed. In the analysis presented in this paper, it is understandable that the methods described in another paper by the authors (Raiola M et al., 2025) are utilized; however, it is important to note that this causes some overlap. It is necessary to clearly distinguish and describe whether the novelty of the methods is based on those developed in this paper or those described in the paper by Raiola M et al. (2025).

      We thank the reviewer for this important observation. We agree that it is essential to clearly separate the methodological developments reported in Raiola M et al. (2025) from the present work. As described in Raiola M et al. (2025), the methods have already been fully developed and validated. In this paper, our focus is to apply these approaches to cardiac development in order to generate and describe the new biological insights. In the revised version, we made this distinction more explicit in the Results and Discussion, highlighting the methodological continuity with our previous work and the biological contribution of the present study.

      Minor points

      (1) In Figure 4, the labels and legends for A, A', B, and B' are reversed. Similar colours are used for C through F, making it difficult to distinguish between them.

      We thank the reviewer for noticing this error. We have corrected it.

      (2) In Figure 5, the labels start with B.

      We thank the reviewer for noticing this error. We have corrected the labels in Figure 5.

      Reviewer #2 (Public review):

      The authors address an important challenge in developmental biology: the quantitative description of tissue deformation during organogenesis. They have developed a new pipeline to quantify early heart tube morphogenesis in the mouse, with cellular resolution. They adopt an elegant approach by integrating multiple 3D time-lapse datasets into a dynamic atlas of cardiac morphogenesis in order to compute spatio-temporal deformation patterns. The main findings highlight a strong compartmentalization of cell behaviors, with tissue growth and anisotropy exhibiting complementary and spatially segregated patterns. Using these data, the authors developed an in-silico fate mapping tool to interrogate cell displacement within the myocardium. This virtual model provides new mechanistic insights into how the bilateral cardiac primordia converge and transform into a three-dimensional heart tube. The authors identify "belt-like" constraints at the arterial and venous poles that prevent tissue expansion and thus shape the ventricular barrel morphology.

      The computational framework is highly innovative and impressive, providing an unprecedented 3D model of tissue deformation during heart morphogenesis. It also opens avenues for testing hypotheses regarding tissue growth and the forces that cause cell motion. However, the proposed model of ventricular chamber formation with the two constraining belts remains hypothetical, lacking biological validation and requiring strengthening or modulation.

      Overall, this carefully performed study provides a new model for exploring tissue deformation during organogenesis and will be of broad interest to computational and developmental biologists.

      We agree with the Reviewer on the limitations of the proposed model due to limited experimental validation. In the revised version of the manuscript we provide further experimental evidence that strengthens the biological validation of the proposed barrel model with two transversal “belts” generating the barrel shape of the primitive ventricle.

      Reviewer #2 (Recommendations for the authors):

      (1) The study proposes a new model of heart morphogenesis by identifying two regions of tissue contraction at the arterial and venous poles. Although the fate map tool has been validated using two ex vivo approaches (DyeI microinjection and TAT-Cre genetic labelling), the conclusions regarding the two belts still need to be demonstrated using in vivo/ex vivo experiments and quantification of cell movements.

      We thank the reviewer for this important suggestion. We agree that experimental validation of the two contraction belts is essential to strengthen the conclusions of the study. In the revised manuscript, we have addressed this point by adding new experimental data directly supporting the existence and dynamics of both D1 and D2 contraction boundaries.

      Specifically, we performed microinjection experiments in which four anchor points, two along D1 and two along D2, were labeled in living embryos and tracked after 10–14 hours (Figure 5C–E, Table S2). For D2, the Euclidean distance between the two anchor points was computed from multiphoton microscopy images acquired at t0 and tfinal (voxel size: 0.57 × 0.57 × 2.5–6.0 µm). In all three embryos analyzed, the D2 anchor points converged over time, with the segment retaining on average 0.27 ± 0.14 of its initial length (range: 0.13–0.40), confirming the lateral compression predicted by the model. For D1, the in-plane geodesic distance between anchor points was measured at t0 and after 10–14 hours. Given the difficulty of imaging the arterial pole at high resolution by whole-mount microscopy, cryosections were used for these measurements (pixel size: 0.65 × 0.65 × 0.042 µm). The D1 segment similarly underwent contraction, retaining on average 0.50 ± 0.22 of its initial length (range: 0.23–0.77). Together, these results provide direct experimental evidence that both boundaries undergo compression during heart tube formation, consistent with the contraction dynamics predicted by the virtual model and supporting the existence of the two belts described in the study.

      We acknowledge that the quantitative values show variability across embryos, which reflects two main sources of uncertainty: (i) the exact position of microinjection along D1 and D2 could not be perfectly standardized; (ii) embryos were not staged at exactly stage 2 at t0 nor did they all reach exactly stage 8 at tend, introducing stage-dependent variability. The primary goal of this experiment was therefore not to precisely quantify compression rates, but to demonstrate that tissue contraction along both boundaries occurs in vivo, consistent with the barrel model predictions. The fact that contraction was observed in all six embryos analyzed, despite the inherent variability of the experimental setup, supports the robustness of this conclusion. These points have been discussed in the revised manuscript.

      (2) The region labelled as OFT appears to correspond instead to the right ventricle primordium, as demonstrated previously by cell labelling of the anterior heart field (Zaffran et al., 2004, PMID: 15217909). The nomenclature should be corrected in the figures and the text. Alternatively, the term "arterial pole" may be useful.

      We thank the reviewer for this observation. We aligned our nomenclature with the literature, correcting the labelling in figures and text.

      (3) The integration of 12 different time-lapses into the model is very impressive. However, while the early stages (2 to 5) are very well covered, the number of replicates for the later stages is much lower. Figure S4 highlights variability between some of the samples, but this is not commented on in the results or the discussion. How does this impact the averaging of tissue deformation patterns and the subsequent model predictions? We thank the reviewer for this comment. We acknowledge that the number of specimens is lower and more variable at later stages. This limitation primarily arises from technical constraints associated with long time-lapse imaging. Because embryo positioning could not be actively tracked during growth, manual repositioning was required, and since embryo development proceeded overnight, maintaining perfect alignment throughout the acquisition was challenging. As a result, several embryos gradually drifted out of the imaging volume and had to be excluded due to incomplete coverage. In addition, at later stages the onset of uncoordinated and subsequently coordinated cardiomyocyte contractions introduces motion-related blurring, which further limits image quality at the acquisition frequency used. These technical limitations were already discussed in the context of the imaging methodology and Limitation and Future Directions section in Raiola et al. (2025).

      As shown in Figure S4, variability between embryos is present and reflects natural biological diversity. Figure S4 also indicates that this variability is highly localized, whereas the regions identified as anticorrelated growth and anisotropy zones remain consistently preserved across embryos. The variability observed in Figure S4, we note that while inter-embryo variability is present, it mainly affects the magnitude of tissue deformation rather than the spatial pattern of deformation. As shown in the additional analyses presented in Figures S5 and S6, the overall organization of deformation, both in terms of growth and anisotropy, is consistently preserved among embryos within the same stage group, within the expected range of natural intra-embryonic variability.

      Finally, regarding the in-silico fate map, our model was not constructed as a statistical average but as a descriptive framework obtained from the concatenation of selected representative embryos. Constructing a statistical model was not feasible due to the limited number of embryos at later stages and the frequent occurrence of incomplete datasets (e.g., randomly missing inflow or arterial pole regions). Under such conditions, only the left ventricular primordium and the inner curvature would have been consistently preserved, thereby limiting the analysis to a very restricted and less informative region. We emphasized these points more clearly in the revised Result section.

      (4) Since the growth rate appears to be highly regionalized, could the authors provide a molecular mechanism for one of these growth patterns?

      We thank the reviewer for this insightful suggestion. Although correlating growth patterns with specific molecular mechanisms would greatly enhance the study, such an effort necessitates extensive additional experimentation, including spatial transcriptomics and detailed molecular analyses. As this falls outside the scope of the present work, we have chosen not to incorporate molecular mechanism data in this manuscript, reserving it for future research.

      (5) Could the model be used to predict new experimental outcomes? For example, could the author simulate a perturbation and validate it through in vivo experiments using mouse mutants?

      We thank the reviewer for this interesting suggestion. At this stage, the model cannot be used to predict new experimental outcomes, as it was designed as a descriptive rather than a statistical or predictive framework. The predictive potential of the model, including the simulation of perturbations, was discussed in detail in Raiola et al. (2025), where this aspect was indicated as a direction for future work.

      We clarified this more explicitly in the revised Results and Discussion sections.

      Minor points

      (1) The readership of eLife is diverse. The methodology and figures could be further annotated, and the axes (A/P, D/V, L/R) could be labelled in all figure panels.

      We thank the reviewer for this helpful suggestion. We revised the figures to include clearer annotations and ensure that the axes (A/P, D/V, L/R) are consistently labelled across all panels.

      (2) It is sometimes difficult to follow without reference to the companion paper. For example, machine learning is mentioned in the summary but is not described in this paper.

      We thank the reviewer for this comment. We clarified in the revised manuscript that the staging system is machine learning-based, using morphometric features to align specimens over time, and indicate that full methodological details are provided in Raiola et al. (2025). This will help readers understand the approach while keeping the focus on the biological findings.

      (3) The authors state the versatility of the model in the introduction, but this is not really addressed in the manuscript; please modulate.

      We thank the reviewer for this feedback. We agree that the versatility of the model was not sufficiently demonstrated throughout the manuscript. In the revised version, we rephrased the Summary to ensure that our claims are aligned with the descriptive scope shown in the current study.

      (4) The authors describe a rightward rotation of the ventricle in stage 9, which they relate to the arterial pole rotation described by Le Garrec et al., 2017. However, this event was reported to occur at E8.5f (which would be equivalent to stage 7). Please modulate or modify.

      We thank the reviewer for this observation. Heart tube rotation is a gradual process that begins at earlier stages, including stage 7, depending on embryo developmental variability. In our study, using the Atlas-based framework described by Esteban et al. (2022), this rotation becomes clearly detectable and morphologically prominent at stage 9, as illustrated in Figure 6d of Esteban et al. At this stage, rightward rotation of the ventricle emerges as the dominant feature in terms of tissue deformation and associated growth patterns, providing a robust reference point to describe and quantify the process. Thus, the description of stage 9 does not indicate the initiation of ventricular rotation, but rather the stage at which the process is most evident and measurable. We moderated it into the revised manuscript to avoid potential ambiguity.

      (5) Some rationales are missing. Why aren't all of the initial 16 time-lapses used for the cumulative deformation pattern analysis? Please explain the impact on the virtual fate mapping of using either labelling of cell clusters or cell continuums. Explain how the Strain Agreement Index neighborhood size (6-7 cells) was chosen, and whether the results are robust at other scales.

      We thank the reviewer for raising these important points. We agree that this section requires clarification and will expand it in the revised Results and Discussion. Not all 16 time-lapses could be included in the cumulative deformation analysis, as this approach relies on concatenating individual embryos into the Atlas framework while preserving the largest possible overlap of tissue. A technical limitation of our recordings was the nonsystematic loss of cardiac tube extremities (inflow tract or arterial pole) due to embryo drift during acquisition. Consequently, several time-lapses provided incomplete tissue coverage and were excluded to avoid an inconsistent assessment of cumulative deformation. In fact, some regions of the tissue would have reflected the contribution of multiple embryos, whereas others would not. Moreover, the registration required to align anatomical regions across stages and embryos would have yielded inaccurate correspondences. For these reasons, we decided to exclude such cases. We commented on this point in more detail in the revised manuscript. For the Strain Agreement Index, the choice of a 6–7 cell neighbourhood size represented a balance between local resolution and robustness. This scale was small enough to allow the tissue to be computationally flattened, while larger neighbourhoods would have included folded regions and created artefacts during the flattening step. Conversely, smaller neighbourhoods would have produced fragmented, “salt-and-pepper” patterns lacking generalization. We commented on this point in more detail in the revised manuscript.

      (6) Figure 5: The panels are mislabelled (B-C versus A-B).

      We thank the reviewer for noticing this mistake. We have corrected the panel labels in Figure 5 to ensure consistency.

      (7) Figure 5C: The red region in stage 2 within the IFT is missing.

      We thank the reviewer for this observation. We have corrected Figure 5C accordingly.

      (8) Typo in Figure 1 legend (p.5): "Our dataset includes multiple specimens raging from E7.75 to E8.25" - should be "ranging".

      We thank the reviewer for pointing this out. We have corrected the typo in the Figure 1 legend.

      (10) Figure S3 legend should state: "Deformation analysis for stage 7, stage 8, and stage 9."

      We thank the reviewer for pointing this out. We have revised the Figure S3 legend accordingly.

      Reviewer #3 (Public review):

      Summary:

      The manuscript by Raiola and colleagues entitled "Quantitative computerized analysis demonstrates strongly compartmentalized tissue deformation patterns underlying mammalian heart tube formation" takes a highly quantitative approach to interrogating the earliest stages of cardiogenesis (12 hours, from early cardiac crescent to early heart tube) in a new and innovative way. The paper presents a new computational framework to help identify both regional and temporal patterns of tissue deformation at cellular resolution. The method is applied to live embryo imaging data (newly generated and from the group's previous pioneering work). In the initial setup, the new model was applied directly to raw time-lapse data, and the results were compared to actual cell tracks identified manually, showing close correlations of the model with the manual tracking. Next, they integrated spatial and temporal information from different embryos to generate a new model for tissue movement, driven by parameters such as tissue growth and anisotropy. Key findings from their model suggest that there are distinct compartments of tissue deformation patterns as the bilateral cardiac crescent develops into the linear heart tube, and that the ventricular chamber forms by a defined expansion pattern, as a 'hemi-barrel shape', with the aterial and venous poles (IFT and OFT) acting as the harnessing belts constraining the expansion of the chamber further. Lastly, the model is tested for its ability to predict future residence of cardiac crescent cells in the heart tube, which it seems to be able to do successfully based on fate tracking validation experiments.

      Strengths:

      The manuscript provides an exceptionally careful analysis of a critical stage during heart development - that of the earliest stages of morphogenesis, when the heart forms its first tube and chamber structures. While numerous studies have interrogated this stage of heart development, few studies have performed time-lapse imaging, and, to my knowledge, no other report has performed such in in-depth quantitative analysis and modeling of this complex process. The computational model applied to normal heart development of the myocardium (labelled by Nkx2-5) has revealed multiple new and interesting concepts, such as the distinct compartments of tissue deformation patterns and the growth trajectories of the emerging ventricle. The fact that the model operates at cellular resolution and over a nearly continuous time period of approximately 12 hours allows for unprecedented depth of the analysis in a largely unbiased manner. Going forward, one can imagine such models revealing additional information on these processes, performing analyses of subpopulations that form the heart, and maybe most importantly, applying the model to various perturbation models (genetic or otherwise). The manuscript is very well written, and the data display is accessible and transparent.

      Weaknesses:

      No major weaknesses are noted with the study. It would have been very exciting to see the model applied to any kind of perturbation, for example, a left-right defect model, or a model with compromised cardiac progenitor populations. However, the amount of live imaging required for such analyses renders this out of scope for the current study.

      We agree with the Reviewer on the relevance of applying this pipeline to mutant conditions. We are engaged on those experiments but they represent a major effort beyond the scope of this manuscript, as also indicated by the Reviewer.

      Reviewer #3 (Recommendations for the authors):

      (1) Application of the model to defective heart development:

      While including perturbation models seems out of scope for the present work, some discussion on how the model might benefit our understanding of early cardiac defects, or any currently unknown mechanisms acting at this stage of development, could be included in the discussion of the manuscript. This would help highlight the enormous power that this new model could bring to understanding these critical steps during heart development, in a quantitative and unbiased manner.

      We thank the reviewer for this insightful comment. Our approach is a deterministic, descriptive framework that integrates individual tissue motion into a common spatiotemporal Atlas, providing a quantitative description of early HT morphogenesis. The primary goal of this framework is to establish a robust baseline of normal HT development under wild-type conditions.

      This baseline is essential for studying heart defects, as deviations from normal tissue motion and deformation patterns can reveal developmental defects like altered growth or aberrant morphogenetic trajectories. Currently, the limited number of embryos per developmental stage (typically 2-4) does not allow the construction of statistically robust inferential models. Nevertheless, by mapping all embryos into a unified reference system and providing quantitative descriptors of tissue motion, our framework already enables meaningful comparisons between normal and abnormal development.

      We have clarified this point in the Discussion section.

      (2) Confusion with Raiola et al., 2025:

      The manuscript frequently references an accompanying manuscript, which is currently a preprint on bioRxiv. The relationship of these two papers is not clear from the description. Not only is the majority of the data shared between the reports, but some figures seem to overlap quite substantially. The methods state that "the computational workflow is detailed in Raiola et al 2025". Any clarification on this would be helpful.

      We thank the reviewer for raising this important point and we appreciate the opportunity to clarify the relationship between the two manuscripts. The two studies indeed rely on the same underlying dataset; however, their aims and scope are fundamentally different. Raiola et al. (2025) is a purely methodological study, whose sole objective is to describe, validate, and benchmark a computational framework for spatiotemporal alignment, motion integration, and in-silico fate mapping. That work deliberately avoids biological interpretation, as the proposed approach is designed to be general and transferable to other organs or developmental systems.

      In contrast, the present manuscript represents the biological application of this validated framework. Here, the computational model is used as a tool to extract, characterize, and interpret biologically meaningful information about early heart morphogenesis, including myocardial motion patterns, regional growth and anisotropy, and fate relationships, supported by experimental validation.

      To avoid any ambiguity, we revised the Introduction and Materials and Methods to explicitly state this distinction and clarify why the methodological details are provided in Raiola et al. (2025), while the current manuscript focuses on biological insight rather than computational development.

      (3) Additional point:

      Concerning overlap with the authors' related manuscript in Bioarchive on the computational workflow: the number of specimens analysed should be noted without referral to the second manuscript (as currently mentioned in the figure legends). Is the "b" necessary when referring to the second manuscript?

      We thank the reviewer for this suggestion. We included the number of specimens analysed directly in the revised manuscript to improve clarity for the reader. Regarding the citation format, the "b" in Raiola et al. (2025) is used to distinguish between two manuscripts from the same group published in the same year.

    1. eLife Assessment

      This study presents valuable findings on the differential effects of RNA on the phase separation, aggregation dynamics, and bioactivity of PSMα3 and LL-37. The authors provide solid evidence from complementary biophysical and cell-based experiments that RNA influences peptide assembly and associated in vitro activities. The study is of interest for understanding interactions between amyloidogenic peptides and nucleic acids.

    2. Reviewer #1 (Public review):

      [Editors' note: This version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      The manuscript by Rayan et al. aims to elucidate the role of RNA as a context-dependent modulator of liquid-liquid phase separation (LLPS), aggregation, and bioactivity of the amyloidogenic peptides PSMα3 and LL-37, motivated by their structural and functional similarities.

      Strengths:

      The authors combine extensive biophysical characterization with cell-based assays to investigate how RNA differentially regulates peptide aggregation states and associated cytotoxic and antimicrobial functions.

    3. Reviewer #2 (Public review):

      In this paper, Rayan et al. report that RNA influences cytotoxic activity of the staphylococcal secreted peptide cytolysin PSMalpha3 versus human cells and E. coli by impacting its aggregation. The authors used sophisticated methods of structural analysis and describe the associated liquid-liquid phase separation. They also compare to the influence of RNA on aggregation and activity of LL-37, which shows differences to that on PSMalpha3.

      Major comments on the previous version:

      (1) The premise, as stated in the introduction and elsewhere, that PSMalpha3 amyloids are biologically functional, is highly debatable and has never been conclusively substantiated. The property that matters most for the present study, cytotoxicity, is generally attributed to PSM monomers, not amyloids. The likely erroneous notion that PSM amyloids are the predominant cytotoxic form is derived from an earlier study by the authors that has described a specific amyloid structure of aggregated PSMalpha3. Other authors have later produced evidence that, quite unsurprisingly, indicated that aggregation into amyloids decreases, rather than increases, PSM cytotoxicity. Unfortunately, yet other groups have in the meantime published in-vitro studies on "functional amyloids" by PSMs without critically challenging the concept of PSM amyloid "functionality". Of note, the authors' own data in the present study that show strongly decreased cytotoxicity of PSMalpha3 after prolonged incubation are in agreement with monomer-associated cytotoxicity as they can be easily explained by the removal of biologically active monomers from the solution.

      In their revision and in the rebuttal, the authors have further described their concept regarding what they call "functionality" of PSMalpha3 amyloids. They now admit that monomers are the active cytolytic form, like other researchers have stressed, whereas amyloids are not. This represents a considerable difference to earlier papers in which they ascribed functionality, i.e. cytolytic capacity, to PSMalpha3 amyloids, a claim that has raised considerable controversy. Now, they use the term "functional " to describe that PSMalpha3 amyloids, while not cytolytic, can be reversed to a cytolytic monomeric state, calling them a "dynamic reservoir". There is no evidence that such a reservoir is necessary for the cytolytic activity of the monomers to be established; also, there is no evidence that in a biological system, such an amyloid reservoir exists. To continue calling PSMalpha3 amyloids "functional" based on this - considerably changed - concept of the authors appears inappropriate, given the finally admitted absence of cytolytic activity of the PSM amyloids in addition to the continuing complete lack of evidence of any biological relevance of PSM amyloid formation.

      (2) That RNA may interfere with PSM aggregation and influence activity is not very surprising, given that PSM attachment to nucleic acids - while not studied in as much detail as here - has been described. Importantly, it does not become clear whether this effect has biologically significant consequences beyond influencing, again not surprisingly, cytotoxicity in vitro. The authors do show in nice microscopic analyses that labeled PSMalpha3 attaches to nuclei when incubated with HeLa cells. However, given that the cells are killed rapidly by membrane perturbation by the applied PSM concentrations, it remains unclear and untested whether the attachment to nucleic acids in dying cells makes any contribution to PSM-induced cell death or has any other biological significance.

    4. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This study presents valuable findings on the differential effects of RNA on the phase separation, aggregation dynamics, and bioactivity of PSMα3 and LL-37. The authors provide solid evidence from complementary biophysical and cell-based experiments that RNA influences peptide assembly and associated in vitro activities. The study is of interest for understanding interactions between amyloidogenic peptides and nucleic acids, although the physiological significance and some aspects of the mechanistic interpretation would benefit from further clarification.

      We are grateful for the positive assessment. The two outstanding concerns about physiological significance and mechanistic interpretation are addressed in detail below through Reviewer #2's comments. We have made targeted revisions throughout the manuscript, and have been careful to distinguish genuine clarifications from reframing that would misrepresent what the data show.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Rayan et al. aims to elucidate the role of RNA as a contextdependent modulator of liquid-liquid phase separation (LLPS), aggregation, and bioactivity of the amyloidogenic peptides PSMα3 and LL-37, motivated by their structural and functional similarities.

      Strengths:

      The authors combine extensive biophysical characterization with cell-based assays to investigate how RNA differentially regulates peptide aggregation states and associated cytotoxic and antimicrobial functions.

      Weaknesses:

      While the study addresses an interesting and timely question with potentially broad implications for host-pathogen interactions and amyloid biology, some aspects of the experimental design and data analysis require further clarification and strengthening.

      We thank Reviewer #1 for the positive assessment. Previous revision round incorporated all major quantitative additions requested:

      Quantitative EMSA binding analysis with Kd values and Hill coefficients (Fig. S1)

      Quantitative FRAP recovery curves with mobile fractions and half-times (Figs. S4, S8, S12)

      Colocalization metrics — Pearson's correlation coefficient and Manders' overlap coefficients (Fig. S5)

      Quantification of AmyTracker630 amyloid signal intensity (Fig. S6)

      Explicit acknowledgment of limitations regarding phase diagram boundaries and csat

      Revised interpretation clarifying nucleolar localization as phenomenological, not causal

      Reviewer #2 (Public review):

      In this paper, Rayan et al. report that RNA influences cytotoxic activity of the staphylococcal secreted peptide cytolysin PSMalpha3 versus human cells and E. coli by impacting its aggregation. The authors used sophisticated methods of structural analysis and describe the associated liquidliquid phase separation. They also compare to the influence of RNA on aggregation and activity of LL-37, which shows differences to that on PSMalpha3.

      That RNA impacts PSM cytotoxicity when co-incubated in vitro becomes clear. However, I have two major problems with this study:

      The premise, as stated in the introduction and elsewhere, that PSMalpha3 amyloids are biologically functional, is highly debatable and has never been conclusively substantiated. The property that matters most for the present study, cytotoxicity, is generally attributed to PSM monomers, not amyloids. The likely erroneous notion that PSM amyloids are the predominant cytotoxic form is derived from an earlier study by the authors that has described a specific amyloid structure of aggregated PSMalpha3. Other authors have later produced evidence that, quite unsurprisingly, indicated that aggregation into amyloids decreases, rather than increases, PSM cytotoxicity. Unfortunately, yet other groups have in the meantime published in-vitro studies on "functional amyloids" by PSMs without critically challenging the concept of PSM amyloid "functionality". Of note, the authors' own data in the present study that show strongly decreased cytotoxicity of PSMalpha3 after prolonged incubation are in agreement with monomerassociated cytotoxicity as they can be easily explained by the removal of biologically active monomers from the solution.

      In their revision and in the rebuttal, the authors have further described their concept regarding what they call "functionality" of PSMalpha3 amyloids. They now admit that monomers are the active cytolytic form, like other researchers have stressed, whereas amyloids are not. This represents a considerable difference to earlier papers in which they ascribed functionality, i.e. cytolytic capacity, to PSMalpha3 amyloids, a claim that has raised considerable controversy. Now, they use the term "functional " to describe that PSMalpha3 amyloids, while not cytolytic, can be reversed to a cytolytic monomeric state, calling them a "dynamic reservoir". There is no evidence that such a reservoir is necessary for the cytolytic activity of the monomers to be established; also, there is no evidence that in a biological system, such an amyloid reservoir exists. To continue calling PSMalpha3 amyloids "functional" based on this - considerably changed - concept of the authors appears inappropriate, given the finally admitted absence of cytolytic activity of the PSM amyloids in addition to the continuing complete lack of evidence of any biological relevance of PSM amyloid formation.

      That RNA may interfere with PSM aggregation and influence activity is not very surprising, given that PSM attachment to nucleic acids - while not studied in as much detail as here - has been described. Importantly, it does not become clear whether this effect has biologically significant consequences beyond influencing, again not surprisingly, cytotoxicity in vitro. The authors do show in nice microscopic analyses that labeled PSMalpha3 attaches to nuclei when incubated with HeLa cells. However, given that the cells are killed rapidly by membrane perturbation by the applied PSM concentrations, it remains unclear and untested whether the attachment to nucleic acids in dying cells makes any contribution to PSM-induced cell death or has any other biological significance. Overall, the findings can be explained in a much more straightforward way with the common concept of cytotoxicity being due to monomeric PSMs, and the impact of nucleic acids on cytotoxicity being due to lowering of the concentration of that active form by RNA attachment. Further limiting the significance of the findings, whether this interaction has any biological significance on the physiology or infectivity of the PSM producer remains largely unexplored.

      We thank the reviewer for the detailed comments. We appreciate the opportunity to further clarify our interpretation of the relationship between PSMα3 assembly, cytotoxicity, and RNA-mediated regulation. In the revised manuscript, and building on the previous revision round, we substantially expanded and refined the Discussion and Introduction to more clearly distinguish between mature fibrils, transient assembly intermediates, and broader assembly state-dependent mechanisms. We also incorporated additional literature representing different perspectives from the field. The revised manuscript presents a model in which biological activity is governed by dynamic assembly pathways and membrane-associated intermediates whose formation, persistence, and structural organization are modulated by environmental conditions, including RNA.

      A central point raised by the reviewer is the suggestion that the RNA effects observed here can be explained simply by sequestration of active monomeric PSMα3. We respectfully disagree that this interpretation can account for the data. A monomer-depletion model makes a clear experimental prediction: conditions that promote aggregation should proportionally reduce activity by reducing the free monomer pool. However, our data show the opposite behavior. RNA promotes PSMα3 aggregation, induces liquid–liquid phase separation, and reshapes fibril morphology into distinct polymorphic assemblies, yet preserves cytotoxic and antimicrobial activity over incubation periods during which peptide alone progressively loses activity. Thus, activity does not correlate with suppression of aggregation or maintenance of soluble peptide. Instead, the data indicate that assembly trajectory and supramolecular organization are functionally relevant parameters. We state this point explicitly in the section “RNA preserves PSMα3 bioactivity,” where we added text clarifying that RNA does not prevent aggregation but redirects the assembly pathway toward structurally and functionally distinct states.

      To further clarify our interpretation, we substantially revised the section “PSMα3 cytotoxicity arises from dynamic assembly intermediates.” This section now integrates multiple independent lines of evidence supporting an assembly-state-dependent model. Together, these observations argue against a simple binary model in which either monomers alone or mature fibrils alone determine activity. Instead, they support a framework in which transient intermediates formed along the assembly pathway contribute to membrane disruption and cytotoxicity. Consistent with this interpretation, our confocal and super-resolution microscopy experiments directly show PSMα3 accumulation and aggregation at bacterial and cellular membranes (Figs. 5, 6C, S10), supporting a model in which assembly occurs in direct association with membrane interfaces rather than exclusively in bulk solution prior to membrane contact. We expanded the Discussion accordingly.

      We acknowledge the reviewer’s alternative interpretation that the nucleolar/nucleic-acid association observed in HeLa cells may reflect post-lysis binding following membrane permeabilization. We agree that this is a valid consideration at the cytotoxic concentrations used here, where membrane disruption is rapid (Figs. 5–6, Movies S1–S2). The Discussion therefore clarifies that nucleolar localization under these conditions is unlikely to represent a distinct intracellular toxic mechanism, but instead reflects the intrinsic nucleic-acid binding capacity of PSMα3 after cellular entry. We accordingly do not claim that intracellular nucleic-acid interactions contribute causally to cell death in these experiments. The potential biological relevance of PSMα3–nucleic acid interactions at sub-cytotoxic concentrations, where membrane disruption does not dominate, remains an important question for future investigation.

      We additionally revised the manuscript to clarify the significance of the EGCG comparison. We agree with the reviewer that the EGCG data alone do not demonstrate “amyloid-mediated cytotoxicity,” and we do not make that claim. Rather, the comparison between EGCG and RNA provides evidence that different assembly trajectories produce different functional outcomes. EGCG redirects PSMα3 into amorphous, non-fibrillar assemblies that lose activity, whereas RNA promotes aggregation while preserving activity and generating distinct supramolecular morphologies. If activity depended solely on monomer concentration, both conditions would be expected to reduce activity similarly through sequestration. Instead, the divergent outcomes support the conclusion that assembly architecture and assembly pathway are functionally important.

      In response to the reviewer’s concern that the manuscript overstates the concept of “functional amyloid,” we explicitly distinguish between mature fibrils and dynamic assembly processes, and we avoid wording that could be interpreted as implying that mature fibrils themselves are the active cytotoxic entities. At the same time, we note that the broader concept of functional amyloid-like assembly pathways is widely used in biology to describe assemblies whose formation regulates storage, localization, stabilization, or timing of bioactive states, including hormone-storage amyloids, RNA-binding protein assemblies, and bacterial curli systems. Within this framework, our interpretation is that PSMα3 assembly dynamics modulate the availability and lifetime of bioactive species rather than that mature fibrils themselves are directly toxic.

      Importantly, we also broadened the manuscript substantially by incorporating independent studies from multiple unrelated systems supporting the principle that supramolecular organization influences biological function. These additions include: studies showing that structured fibrillar assemblies of LL-37 are required for specific antibacterial activities; work demonstrating that the nanoscale organization of β-defensin–nucleic acid complexes governs immunostimulatory potency; studies correlating α-helical solid-state conformations with cytotoxicity across fibril-forming antimicrobial peptides; salt-induced PSMα3 polymorphism studies showing distinct toxicities for amorphous versus fibrillar assemblies; and real-time AFM work demonstrating that membrane-associated protofibrillar intermediates are more disruptive than mature fibrils. We also added discussion of recent cryo-EM structures showing that RNA acts as a structural cofactor shaping tau fibril polymorphism at atomic resolution, as well as two-dimensional infrared spectroscopy studies demonstrating coexistence of cross-α and cross-β PSMα3 polymorphs. Together, these orthogonal observations from multiple systems support the broader principle that assembly architecture is a major determinant of biological behavior.

      We also addressed the reviewer’s concern regarding biological relevance. We agree that direct in vivo validation remains an important future direction and state this explicitly in the revised Discussion. However, we respectfully submit that establishing the mechanistic principle that RNA regulates PSMα3 assembly state and functional output is itself a meaningful contribution independent of immediate in vivo confirmation. To better contextualize potential physiological relevance, we expanded the “Biological and therapeutic implications” section to discuss biologically plausible extracellular environments in which PSMα3 may encounter nucleic acids, including biofilms enriched in extracellular RNA, extracellular vesicles, damaged host tissues, inflammatory milieus, and host-derived extracellular RNA released as DAMPs.

      Overall, the revised manuscript reflects a substantially expanded discussion of PSMα3 assemblystate-dependent activity, the role of RNA in modulating assembly trajectories, and the broader conceptual implications for membrane-active peptide assemblies.

      Further remarks:

      (1) Circumstantial evidence based on the "amyloid inhibitor", EGCG: The results with EGCG, which has been shown to have a moderate amyloid-reducing effect on PSMalpha 1 and PSMalpha4, should not be taken as evidence for amyloid-based cytotoxicity. While increased concentrations of EGCG reduced the cytotoxic effect of PSMalpha3, it is not convincingly shown that this is due to a lower concentration of amyloid vs. monomeric PSM.

      We agree that the EGCG data alone should not be interpreted as evidence that mature amyloid fibrils are the directly cytotoxic species. Our interpretation is more limited and focuses on the effect of assembly redirection. Specifically, EGCG redirects PSMα3 into amorphous, non-fibrillar assemblies that lose activity, whereas RNA promotes aggregation while preserving activity and producing structurally distinct assemblies. The key conclusion is therefore that functional outcome depends on the nature and trajectory of assembly rather than on aggregation versus non-aggregation alone. We clarified this distinction in the revised Discussion section addressing RNA- versus EGCG-mediated modulation of PSMα3 assembly.

      (2) It is appreciated that the authors refrain from presenting the unsubstantiated concept of "functional" PSM amyloids in the discussion. However, wording in that direction must also be removed from other parts of the manuscript (e.g. "bioactive fibrillar polymorphs". "The formation of cross-alpha amyloids has been correlated with toxic activity", etc.), generally refraining from uncritically implying that amyloid formation underlies PSM biological activity, and rather discussing that the much more likely explanation of the findings is a lowering of cytolytically active, monomeric PSM concentration.

      In the Introduction, the phrasing 'may enable dynamic switching' has been used to soften the mechanistic claim regarding cross-α assemblies. The phrase 'bioactive fibrillar polymorphs' was revised in the previous round. At the same time, statements such as “cross-α amyloid formation has been correlated with toxic activity” are retained because they describe experimental observations reported in multiple studies (including Tayeb-Fligelman et al., 2017, 2020; Malishev 2018), without implying direct causality (correlation is not causation). We now explicitly frame these observations within a broader discussion of transient assembly intermediates and assembly-state-dependent toxicity.

      (3) Discussion: "PSM alpha3 interaction with nucleic acids within human cells ...supports a comparable mechanism...". Delete. Unsubstantiated.

      This sentence was removed in the previous revision round and remains absent from the current manuscript.

      (4) The authors should cite papers that have argued against their hypothesis and not only their own manuscripts.

      We appreciate this suggestion and agree that alternative interpretations should be represented explicitly. In the revised manuscript, we added and discussed studies including Zheng et al. (2018) and Yao et al. (2019) (already cited in both earlier versions), which support models in which advanced amyloid formation reduces cytotoxicity and active species are prefibrillar. These studies are now discussed substantively in both the Introduction and Discussion alongside our own work and that of others.

      More broadly, we revised the manuscript to present the current understanding of PSMα3 toxicity as an actively debated question in the field rather than as a settled model. At the same time, we note that citing our prior studies remains necessary where the present work directly builds upon previously reported structural, biophysical, and mechanistic observations.

      If the reviewer has additional specific references in mind, we welcome them and will incorporate them.

    1. eLife Assessment

      This fundamental study provides compelling evidence for the functional segregation of the sensorimotor cortex into precisely delineated areas, and highlights a rapid transition in functional properties at the boundaries between these areas. This result further confirms and extends recent work on the diversity of neural response specificities across cortical areas in the context of complex behavioral tasks. This work will be of interest to neuroscientists studying sensory-motor functions.

    2. Reviewer #1 (Public review):

      Summary:

      Here the authors address the organization of reach-related activity in layer 2/3 across a broad swath of anterodorsal neocortex that included large subregions of M1, M2, and S1. In mice performing a novel variant water-reaching task, the authors measured activity using two-photon fluorescence imaging of a GECI expressed in excitatory projection neurons. The authors found a substantial diversity of response patterns using a number of metrics they developed for characterizing the PETHs of neurons across reach conditions (target locations). By mapping single-neuron properties across cortex, the authors found substantial spatial variation, only some of which aligned with traditional boundaries between cortical regions. Using Gaussian mixture models, the authors found evidence of distinct response types in each region, with several types prominent in multiple cortical regions. Aggregating across regions, four primary subpopulations were apparent, each distinct in their average response properties. Strikingly, each subpopulation was observed in multiple regions, but subpopulation members from different regions exhibited largely similar response properties.

      Strengths:

      The work addresses a fundamental question in the field that has not previously been addressed at cellular resolution across such a broad cortical extent. I see this as truly foundational work that will support future investigation of how the rodent brain drives and controls reaching.

      The quantification is thoughtful and rigorous. It is great that the authors provide explanation for and intuition behind their response metrics, rather than burying everything in the Methods.

      The Discussion and general contextualization of the Results is thorough, thoughtful, and strong. It is great that the authors avoid the common over-interpretation of classical observations regarding cortical organization that are endemic in the field.

      All things considered, this is the best paper regarding spatial structure in the motor system I have ever read. The breadth of cellular resolution activity measurement, the rigor of the quantification, and the clear and open-minded interrogation of the data collectively have produced a very special piece of work.

      Weaknesses:

      There are two important issues left unaddressed that the authors plan to address in their future work. The first is the relation between observed neural activity patterns and movement kinematics, and in particular how much the activity variation across target locations may relate to the kinematic differences across these different conditions, as opposed to true higher-order movement features like reach direction. The second issue is how to interpret the results in relation to existing ideas about behavioral organization in motor/premotor cortex.

      Comments on revised version:

      The authors have done an excellent job addressing my previous concerns. I have no additional concerns with the manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      The functional parcellation of cortical areas is a critical question in neuroscience. This is particularly true in frontal areas in mice. While sensory areas are relatively well characterized by their tuning to sensory stimuli, the situation is much less clear for motor areas. This has become even more ambiguous since recent studies using large-scale neuronal recordings consistently report mixed sensory and motor-related activity throughout the brain and motor mapping studies have shown that movements evoked by cortical stimulation are by no means limited to motor areas alone. Here, the authors use a correlation approach combining large-scale functional imaging at cellular-resolution with movement-tracking in mice executing a reaching task. Across multiple recording sessions in the same animals, the authors have imaged a large portion of the sensorimotor cortex at cellular resolution in mice performing a reaching task, recording the activity of nearly 40,000 neurons. By aligning the calcium signal of each neuron to three task events-the Go cue triggering the reach, the onset of paw lift, and the contact between the paw and the target-for different target positions, the authors identified different response patterns distributed differently across cortical areas. They defined a set of features that describe the neurons' response pattern, representing the temporal dynamics and tuning properties for the different target positions. These features were used to construct cortical maps, and the authors show that, interestingly, gradient maps obtained from the first derivative of the feature maps reveal sharp discontinuities at the boundaries between anatomically defined cortical areas. Using dimensionality reduction of the neuronal response features, the authors found that, despite clear differences in their average response properties, individual neurons from the same cortical areas do not form distinct clusters in the reduced-dimensional space. In fact, most areas contain heterogeneous neuronal populations, and most neuronal populations are present in multiple areas, albeit in different proportions. Interestingly, the authors identified four neuronal subpopulations based on the distance between the components of the Gaussian mixture model used to model the distribution of neurons within each area. One of these subpopulations is almost exclusively represented in the anterior M2 cortex, while another is broadly distributed across the different areas.

      Strengths:

      This article is based on an impressive dataset of nearly 40,000 neurons covering a large portion of the sensorimotor cortex and on innovative analytical approaches. This study is likely the first to clearly demonstrate boundaries between cortical areas defined based on the responses of individual neurons. This innovative approach to functional mapping of cortical areas potentially opens up new perspectives for higher-resolution mapping of frontal cortical areas, using a broader repertoire of sensory and motor evoked responses.

      Weaknesses:

      One limitation of this study - inherent in most cell imaging studies - is that it only takes into account the activity of neurons in superficial cortical layers. One might think that taking into account neuronal activity across the different layers would allow for an even finer functional cortical segmentation.

      Comments on revised version:

      The authors have answered all my questions and this new version has largely improved in clarity.

    4. Author response:

      The following is the authors’ response to the original reviews.

      In preparation for release of the analysis code used in the paper, we made many analyses more parallel to one another in their exact preprocessing. This resulted in very slight changes to many panels, but these changes are nearly invisible and conclusions did not change. In one case, though, we realized that the way we were presenting data was potentially misleading (the timing plot in Figure 3A). The original plot was of the distribution of pixel values from the spatially smoothed map instead of distributions over individual neurons. We have now swapped it out for better interpretability and changed the accompanying text accordingly.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Here, the authors address the organization of reach-related activity in layer 2/3 across a broad swath of anterodorsal neocortex that included large subregions of M1, M2, and S1. In mice performing a novel variant water-reaching task, the authors measured activity using two-photon fluorescence imaging of a GECI expressed in excitatory projection neurons. The authors found a substantial diversity of response patterns using a number of metrics they developed for characterizing the PETHs of neurons across reach conditions (target locations). By mapping single-neuron properties across the cortex, the authors found substantial spatial variation, only some of which aligned with traditional boundaries between cortical regions. Using Gaussian mixture models, the authors found evidence of distinct response types in each region, with several types prominent in multiple cortical regions. Aggregating across regions, four primary subpopulations were apparent, each distinct in its average response properties. Strikingly, each subpopulation was observed in multiple regions, but subpopulation members from different regions exhibited largely similar response properties.

      Strengths:

      The work addresses a fundamental question in the field that has not previously been addressed at cellular resolution across such a broad cortical extent. I see this as truly foundational work that will support future investigation of how the rodent brain drives and controls reaching.

      The quantification is thoughtful and rigorous. It is great that the authors provide an explanation for and intuition behind their response metrics, rather than burying everything in the Methods.

      The Discussion and general contextualization of the results are thorough, thoughtful, and strong. It is great that the authors avoid the common over-interpretation of classical observations regarding cortical organization that are endemic in the field.

      All things considered, this is the best paper regarding spatial structure in the motor system I have ever read. The breadth of cellular resolution activity measurement, the rigor of the quantification, and the clear and open-minded interrogation of the data collectively have produced a very special piece of work.

      Thank you! We really, really appreciate this!

      Weaknesses:

      The behavioral task is very impressive and an important contribution to the field in its own right. However, given that it appears substantially different from the one used in the previous paper, the characterization of the behavior provided in the Results is too brief. More illustration of the behavior would be helpful. For example, it is rather deep into the paper when the authors reveal that the mice can whisk to help localize the target location. That should be expressed at the outset when the behavior is first described. Other suggestions for elaborating the behavior description are included below.

      Thank you. Although the task will be treated in greater detail in the next paper (where we more closely relate neural activity to the kinematics), we have added more exposition of the task here. In particular, we now include a figure with a characterization of the trial-to-trial variability across reaches to the same target versus across reaches to different targets (Figure 2-figure supplement 1B). This supports the idea that the mice aimed their reaches. We have also expanded that text.

      Regarding whisking, we have now revised that text to make clear that we do not know how the mice localize the spout. The original work by Galinanes and Huber argued that they find the spout by sniffing the water; they may do the same here, or may find it via whisking. It is also possible that the whisking they do is simply because the spout moves in and they are excited, or startled, or do it by reflex. We simply have no evidence one way or another. We have therefore revised the text to make it clearer that whisking-related activation could have occurred for a variety of reasons.

      Statistical support for key claims is lacking. For example, "The five areas of interest varied in the fraction of neurons that were modulated: M2 had 14%, M1 had 23%, S1-fl had 30%, S1-hl had 25%, and S1-tr had 27%" - I cannot locate the statistical tests showing that these values are actually different. Another example is Figure 7, where a key observation is that distributions of PETH features are distinct across regions. It is clear that at least some distributions are not overlapping, but a clearer statistical basis for this key claim should be provided.

      Good idea. For the proportions, we have now added first a Chi-square test for homogeneity to show that there is variation in the proportions, then shown the results of pairwise two-proportion Z tests (Bonferroni-corrected for multiple comparisons) as a binary matrix in Figure 3-figure supplement 1B. For the area distributions in the t-SNE space (Figure 7), we have added a 2-dimensional Kolmogorov-Smirnov test, again corrected for multiple comparisons, with p-values quoted in the text.

      I understand that the authors are planning a follow-up study that addresses the relation between activity patterns and kinematics. One question about interpreting the results here though, is how much the activity variation across target locations may relate to the kinematic differences across these different conditions, as opposed to true higher-order movement features like reach direction.

      We agree this is a very important question. However, having done many of the analyses to examine the question for the next paper in the series, we do not know of a shortcut to the right answer. This question requires thorough treatment, and so we leave it to be covered in subsequent work. Instead, after our speculation about how responses suggest function, we are now explicit that these hypotheses needs testing:

      “In each of these cases, determining the relationships of the observed activity patterns to function will require specific attempts to link the activity to kinematics, target location, sensory feedback, and more; these relationships will be addressed in future work.”

      Reviewer #2 (Public review):

      Summary:

      The functional parcellation of cortical areas is a critical question in neuroscience. This is particularly true in frontal areas in mice. While sensory areas are relatively well characterized by their tuning to sensory stimuli, the situation is much less clear for motor areas. This has become even more ambiguous since recent studies using large-scale neuronal recordings consistently report mixed sensory and motor-related activity throughout the brain, and motor mapping studies have shown that movements evoked by cortical stimulation are by no means limited to motor areas alone. Here, the authors use a correlation approach combining large-scale functional imaging at cellular resolution with movement-tracking in mice executing a reaching task. Across multiple recording sessions in the same animals, the authors have imaged a large portion of the sensorimotor cortex at cellular resolution in mice performing a reaching task, recording the activity of nearly 40,000 neurons. By aligning the calcium signal of each neuron to three task events-the Go cue triggering the reach, the onset of paw lift, and the contact between the paw and the target-for different target positions, the authors identified different response patterns distributed differently across cortical areas. They defined a set of features that describe the neurons' response pattern, representing the temporal dynamics and tuning properties for the different target positions. These features were used to construct cortical maps, and the authors show that, interestingly, gradient maps obtained from the first derivative of the feature maps reveal sharp discontinuities at the boundaries between anatomically defined cortical areas. Using dimensionality reduction of the neuronal response features, the authors found that, despite clear differences in their average response properties, individual neurons from the same cortical areas do not form distinct clusters in the reduced-dimensional space. In fact, most areas contain heterogeneous neuronal populations, and most neuronal populations are present in multiple areas, albeit in different proportions. Interestingly, the authors identified four neuronal subpopulations based on the distance between the components of the Gaussian mixture model used to model the distribution of neurons within each area. One of these subpopulations is almost exclusively represented in the anterior M2 cortex, while another is broadly distributed across the different areas.

      Strengths:

      This article is based on an impressive dataset of nearly 40,000 neurons covering a large portion of the sensorimotor cortex and on innovative analytical approaches. This study is likely the first to clearly demonstrate boundaries between cortical areas defined based on the responses of individual neurons. This innovative approach to functional mapping of cortical areas potentially opens up new perspectives for higher-resolution mapping of frontal cortical areas, using a broader repertoire of sensory and motor evoked responses.

      Thank you!

      Weaknesses:

      The second part of the article, which presents multimodal responses in the cortical areas, seems to be a perhaps overly complicated way of showing what has already been demonstrated in numerous recent publications, but these new analyses expand upon these previous observations by revealing an interesting functional organization of the sensorimotor cortex, highlighting interesting similarities and differences between certain areas.

      We understand the concern: a number of recent papers have also noted different neuron response characteristics distributed throughout the motor system. We compare and contrast in greater detail following the more specific comments on this below, but we briefly summarize here. The way previous work handled the data – for example, starting with PCA – mixes what neurons are tuned for and when they are tuned for it with what we refer to as the “response format”: properties like tuning sharpness, response duration, etc. We focused primarily on this response format, and designed our features to be mostly independent of tuning preferences or peak response timing. We therefore pick up on different properties of neurons’ responses than those prior works. In addition, no previous work we know of examined these properties across large swathes of cortex at single-cell resolution in the context of forelimb control. Together, these aspects of our work allowed us to produce high-resolution mapping of response properties in a way we have not seen in any prior work.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      In addition to addressing the weaknesses stated above, I suggest the authors also consider the following.

      The one big question left unresolved here is whether we should be thinking about these four subpopulations as distinct types with a biological basis and importance, or just reflections of activity pattern heterogeneity. The authors say that "we did not observe tight clusters in feature space separated by gaps," but their discussion here is light and a bit unclear, and their engagement with the issue of types versus heterogeneity, in my view, could be improved. We do not need "gaps" where the density goes to zero in parameter space, but we do need reproducible troughs between peaks. The authors should clarify if there are substantial and reproducible troughs in the parameter space between their four subpopulations.

      This is a great idea, and we have added three analyses and additional text to address it. We break this concern down into two more specific questions, based on the next comment by this reviewer.

      (1) Are the clusters well separated / do they have troughs between them? (Note that even with troughs, clustering might not be stable if the clustering algorithm is poorly matched to the shapes of the clusters.)

      (2) Is the clustering stable? (It can be stable even without troughs, if, for example, the distribution has a long tail and a GMM needs one Gaussian for the body of the distribution and a second for the tail.)

      First, to directly address the presence or absence of troughs between clusters, we have added Figure 9-figure supplement 2A and 2B. For each pair of subpopulations, we trained a logistic regression classifier to separate the 5D feature vectors of the neurons in one subpopulation from the feature vectors of the neurons in the other subpopulation, then projected the feature vectors onto this axis. Note that because the subpopulations are defined by GMMs, which have nonlinear boundaries, the (linear) logistic classifier does not typically produce perfect classification. Nevertheless, this analysis provides a window onto how well separated each cluster is from each other cluster in feature space. In 5 of the 6 pairwise comparisons, it is obvious that the distributions are different and have at least some dip in the distribution density at the boundary. The one pair of clusters without a trough between them were the forelimb somatomotor and hindlimb somatomotor subpopulations. This was surprising to us, given that their likelihood maps are so strongly distinct, but this presumably reflects trying to capture a nonlinear classifier boundary with a linear one (see below). Overall, this analysis argues that the clusters do have fuzzy edges that blend into one another, but reflect concentration of mass near the centers of the clusters we identified.

      Second, to address the same question with a different nonlinear method, we have added a version of the t-SNE plot from Figure 7 that is instead colored and contoured by subpopulation identity instead of area (Figure 9-figure supplement 2B). Agreement with the GMMs is not a given here either, because t-SNE is a fundamentally different and independent nonlinear transform from that performed by the GMM classification. Nevertheless, the subpopulations were again nicely separated – though not with troughs, possibly thanks to the inherent difficulty of interpreting point density with t-SNE. Interestingly, here the hindlimb somatomotor subpopulation was the best separated from the other subpopulations, supporting the idea that the lack of separation we observed above with the logistic projections was indeed due to a nonlinear boundary. This analysis again argues that neurons are more likely to have features that lie near the center of a cluster, but that the edges of the clusters run into one another. Additionally, this analysis makes clear that treating the hindlimb somatomotor subpopulation as a second cluster can be supported by other analyses, even if not by the logistic regression projection.

      Third, to address the question of cluster stability, we have performed random splits of our data, GMM clustered the two halves independently, applied the GMM from one half to the other, and asked how similar the clusterings are using the Adjusted Rand Index. This produced a value of 0.856, which for this sensitive measure argues that the clustering is rather stable (at least for the three clusters that can be found with all data together, which does not include the smaller-in-size Anterior subpopulation). Note that we did not perform this analysis on the more complicated version where we fit a GMM to each area separately then cluster those; in our main analysis, the hierarchical clustering agreed with what we found by eye, but determining the number of clusters for hierarchical clustering is in general very unstable and so we did not have an objective way to determine the “true” number of clusters.

      In addition to these new analyses, we note that three analyses we had already included bore strongly on this issue. Regarding separability of the clusters, the fact that our likelihood maps (Figure 9C-F) were quite distinct for different subpopulations argues that we picked up on ‘real’ differences. Second, Figure 9B found that when clustering non-overlapping data – different cells from different areas – we obtained clusters that were nearly identical in their feature distributions. Third, Figure 10E used the clusterings from different areas’ data to create likelihood maps, and found that they were extremely similar. These analyses together argue strongly that we are finding ‘types’ in a meaningful sense; given that we know the areas do have different distributions of properties, if there weren’t types then clustering would yield different clusters for different areas. Given the importance of the question, however, we are grateful that the reviewer encouraged us to find additional ways to make this point!

      The original t-SNE plot is beautiful and quasi-fractalic, but it does not show clear signs of four cell types. The single-neuron activity profiles are clearly heterogeneous in very interesting ways, but heterogeneous does not imply a strong or reproducible multimodality that would indicate meaningful cell types. Clustering algorithms will always spit out an answer. If you just have elements uniformly distributed across a parameter space, plus some noise, when you ask for X clusters, you will get X clusters that have different centroids. When you ask an algorithm to cluster without defining the number of clusters, noise can lead the algorithm to produce a particular number of clusters that again will have distinct centroids. The salient question, though, is whether in the present case there is a parameter space in which the clusters are substantially and or reproducibly distinct. Distinct here would mean that peaks in the density across some parameter space are separated by troughs - again, we don't need true gaps. The more substantial the differences between clusters are (again, not the differences between centroids but the prominence of the density troughs between them), the more biologically meaningful the clustering is likely to be. Reproducibility here could be addressed with resampling methods (e.g., how often do two separate halves of the cells produce the same clusters?).

      Please see the reply above, which includes our addressing of this concern.

      The Introduction is generally good, but it could further develop existing ideas about how function is distributed across cell types and regions. We would like to be able to imagine different answers to the question of how activity patterns are organized that might have divergent implications for how the circuit works. I understand we have very little to go on in terms of data, but I think it would be helpful for readers to be given more of a sense of what *could* be important.

      Good idea. We have added such a paragraph to the Introduction:

      “To frame possible outcomes, consider that single neuron responses can vary along many dimensions. Cells could differ according to which movements or time periods they are recruited for (tuning), what movement parameters their activities reflect (encoding), or how their responses are structured across different movements (e.g., nonlinear encoding structure). Further, differences in these response properties across cells could be distributed over the cortical sheet in a variety of ways. Cells could form distinct “categories” or clusters that are spatially well-aligned to the boundaries of anatomically defined regions. Or, categories of neurons might span area boundaries in spatial footprints that do not relate obviously to area boundaries, and that either abut or overlap. At a fine-grained scale, cells with similar responses could be physically located near one another as in primate and feline visual cortex, or similarly-responsive neurons might be salt-and-pepper intermingled as seen in rodent visual cortex or in primate motor cortices during reaching behaviors.”

      It should be clarified in the Results how the cue relates to the target location. Most would assume a different cue for each location, but this does not appear to be the case. The authors should clarify whether there was some amount of searching for the precise target location after the reach, or else how the block structure or other sensory information allowed mice to learn where exactly the target would be. In the absence of target-specific cues, some sense of how the mice achieved target-specific reach trajectories should be offered.

      Related to this, in Figure 1, it would be good to see some individual trajectories, as they all overlap near the target in the current plot. Clearly, the reaches were targeted, but it is unclear how targeted. Some of the adjustments at the end may reflect searching or palpation to resolve the precise spout location. It is very much ok if the mice were not reaching with micron precision each time to each of 15 different targets, but it would be good to provide the reader a better sense of what the mice were doing.

      These are important points. First, to clarify, the Cue is just a Go cue, and was the same for all targets. It is now described in the Results as “non-target-specific”. For additional explanation about supplemental analyses to assess “aiming”, see replies to Reviewer #1 Public Review comments above. Finally, regarding how the mice locate the target: we just don’t know. As discussed above, Galiñanes and Huber found evidence for the mice using stereo sniffing, but whisking, listening to the motors, or some other strategy are also conceivable. We simply don’t have data to weigh in on this. We now make this limitation clear where we describe the task.

      In Figure 1A, CFA does not look well aligned with Tennant et al. (2011). CFA should only extend to +1 AP. The overlap of CFA and RFO seems strange. RFO also does not totally align with the injection coordinates used in An et al, biorxiv 2022.

      Thank you for your attention to these points. Our designation of the name CFA to the red dashed outline in Figure 1A was consistent with an earlier version of our previous work (Grier et al 2026) wherein we referred to the anatomical outline “MOp-ul” from Munoz-Casteneda et al 2021 as CFA. We have since revised that nomenclature to now refer to the outline as M1-fl, or the forelimb representation of primary motor cortex.

      Our placement of RFO was obtained by aligning the Allen CCF from Figure 1K of An et al 2022 to our version of the Allen CCF and outlining the hotspot of RFO with a circle. We have slightly adjusted the location of RFO posterior and medial to more closely align with the injection coordinates reported in the methods of An et al. 2022 of “1.5-1.88 mm anterior from Bregma, 2.25-2.63 mm lateral from the midline.” Because (as far as we understand) the injection coordinates and the map are not perfectly in register, we show a compromise between the two.

      We stress that the Figure 1A map is meant to be descriptive in its illustration of the variety of organizational zones that have been identified across mouse sensorimotor cortex.

      Discrepancies in the alignment procedure, animal strain, and mapping modality all introduce heterogeneity across mapping attempts that we do not aim to reconcile or resolve here.

      Related to this, aspects of the results do seem consistent with the distinction between RFA and CFA, but this is not acknowledged or discussed. For example, the barriers in Figure 6H that lie along the M1/M2 border - these seem consistent with the gap between RFA and CFA. The same could be said for the dim trough along the M1/M2 boundary that appears to separate RFA and CFA in Figure 3B. A slightly more rostral and lateral location of CFA compared to Tennant's definition or the regions backlabeled from cervical spinal injections (see Wang, Maunze et al. J Nsci, 2018) could be expected if flattening the brain under the coverslip for imaging effectively stretches the ML axis, and Bregma (notoriously hard to define reliably at this spatial scale) was defined a bit more caudally here than in other studies. Related to this, it would be better for the field if people described their method for defining Bregma in the Methods. I suggest the authors do this here.

      We appreciate the suggestion and have acknowledged the suggested correspondence in the discussion. Given the difference in our approach from those that originally characterized RFA (through ICMS and deep layer projection tracing) we have avoided making overly strong conclusions about this correspondence in our data. See the quoted text below.

      “The spatial distribution of modulated cells in Figure 3 suggests a distinction between the caudal forelimb area (CFA, involving M1 and S1-fl) and the rostral forelimb area (RFA) in M2, while the feature gradient boundaries suggest a distinction between M1 and M2 more generally. The absence of a clearly delineated RFA was surprising, given its distinct projection patterns (Carmona et al. 2024; Hira et al. 2013b; Wang et al 2018) and functional differences from CFA (Kristl et al. 2025; Morandell and Huber 2017; Saiki-Ishikawa et al. 2025), but our results might suggest that the activity in layer 2/3 of RFA does not differ markedly from other nearby subregions of M2.”

      Regarding bregma, we did not use it for atlas alignment here. Alignment was accomplished through a combination of paw vibration mapping and the location of the central sinus. Bregma’s location was only relevant for our injection of tdTomato labeling, and that labeling was used here only to stabilize the image plane. We include an estimate of it on the map solely in an attempt to be helpful, but we cannot claim we have the most reliable method for defining it.

      The authors focus on activity aligned to cue timing. This is sensible, but it could be meaningful to know how this choice affects the definition of organization. If response clustering is largely different across time, it would seem important. I understand that addressing this question may be beyond the scope of this paper. I just wanted to raise the issue with the authors for their future consideration.

      We agree that this is important to address directly. There are two aspects to this comment: (1) does it matter if activity from approximately the same time period is aligned to the paw lift or contact instead of the cue? (2) What changes if we use data from a different period of time?

      Regarding the first question (alignment), if we switch to aligning our data based on lift or contact, we have more statistically modulated neurons (see Figure 3C), but everything else is qualitatively similar with one exception: the GMM optimization doesn’t separate out the Anterior subpopulation from the Forelimb Motor subpopulation. The Anterior subpopulation only has a relatively small number of members, and they mostly exhibit the strongest peaks in their PETHs when Cue-aligned, so this makes sense. We now show the modulation maps for all of the locking events (Figure 3-figure supplement 1).

      The issue of the time window is a little more complicated. There are many choices we made in this work, of course, not least of which are the task we used and the features we chose based on hand-inspection of thousands of PETHs. As we noted in the Discussion, different tasks or different features would likely distinguish more subpopulations from one another. We think of the time window as a feature choice, albeit an implicit one. We chose not to include later time points because this begins to strongly include reward signals, which are known to be large (Levy et al 2020) and can dominate other aspects of the responses. The largest differences we noted when trying time windows that extended later are that mouth-related areas are separated out in the subpopulation analyses, perhaps because of later licking/consummatory responses, but we have not explored fully enough to speak confidently on this point without much more work and another 10 figures. To keep the scope of the paper manageable, we now call out this choice explicitly (see text below). We thank the reviewer for raising these important points.

      “Crafting additional PETH features, or using end-to-end neural network approaches to discover other features, might enable the discovery of additional structure (Minderer et al. 2019; Wang et al. 2023b). For example, our PETH features were chosen to be invariant to the onset time of activity, but these onset times were markedly later in lateral M1 than in adjacent M2 or S1-fl. Including onset times, using a wider window of time that includes more of the reward/licking period, aligning data to other behavioral events, or adding other PETH features would presumably result in finer subdivisions of sensorimotor cortex.”

      The map in Figure 4 is very cool, and the spatial structure is quite striking. In terms of the actual values of the onset times in each region, I am a little concerned with a dependence on the level of reach-related activity modulation, especially relative to the level of background activity (potentially related to posture). Less reach-related activity and more background activity, which we might expect for trunk and hindlimb regions, could seemingly skew the onset times earlier. We could be getting the right answer, or an answer that makes intuitive sense, for the wrong reason. Can this potential confound be excluded with some sort of control analysis?

      The previous text wasn’t clear. We have now clarified what we meant, very much in line with the reviewer’s thoughts. In addition, note that our change to what is displayed in the histogram (now neurons, previously pixel values) makes clearer that there is a multi-peaked distribution of onset times and it is mostly the prevalence of each peak in each area that varies. The text now reads:

      “These distributions over neurons revealed clear differences in the overall profile of activation: early onsets were more prevalent in S1 trunk and hindlimb regions, perhaps due to activity related to the animal stabilizing itself even if the neurons became more active later; then M2, and finally S1-fl and M1. Nevertheless, each area contained neurons activated at any given time in the trial.”

      The "Peak time variation" metric could potentially vary with activity level, with lower, noisier activity levels making cells appear less persistent. Perhaps a control analysis, based on SNR or some reasonable assumptions of the linkage between calcium signals and spiking, could be performed to measure the extent to which this could be creating differences between regions.

      Good idea. We have now performed this analysis, and the reviewer was correct: the correlation between peak time variation and a simple metric of SNR (assessed as range of PETH / max s.e.m.) was substantial: ⍴=-0.53. We now report this correlation and describe in the Results that this metric is driven by both true peak time variation and trial-by-trial variation. Thank you for this!

      “Peak time variation. To quantify whether a neuron’s firing peaked at the same time for every target or varied by target, we found the peak firing rate of the response to each target, then computed the standard deviation of these peak times across targets. This value is therefore higher if the peak time varied and nearly zero if the timing was consistent. Notably, this measure correlated substantially with overall signal-to-noise ratio of a neuron’s PETH (Spearman’s ⍴=-0.53; Methods), and thus partly measures trial-to-trial variability, not just true peak timing variability. This metric was quite low in M1, indicating highly consistent timing of the activity peak (and reliable responses), and was highest in the posteromedial part of M2 (presumably corresponding to the hindlimb representation) and the posterior tip of S1-hl (Figure 5B).”

      One could argue that the likelihood calculations illustrated in Figure 8 are biased higher for neurons within each region since they were used for defining the likelihood for that region. I think these likelihood calculations should be done for separate neurons other than the ones used to compute the mixture model for each region.

      We agree with the point about bias: the by-area GMM in Figure 8 is biased toward cells within the area, though the effect is probably quite mild given the large numbers of neurons and modest number of parameters. However, this model was intended to make the point that even if you give an area an unfair advantage, you still can’t cleanly isolate it. This was intended to help motivate the following analysis of subpopulations, and we have now made this logic clearer. Doing it this way has the advantage that the GMM components are identical between Figures 8 and 9, while if we held out the test neurons it would not be possible to make them the same without some complicated version of bagging on the GMM components. The reviewer is right that we should make this bias explicit, though, and we have now done so:

      “This mapping approach is explicitly biased toward finding feature differences between areas, allowing for a direct test of the hypothesis that response profile distributions are area-specific.”

      To me, the last Results section (Spatial overlaps between subpopulations indicate intermingled members) does two things: it shows you get the same results when you map each cell to a subpopulation independently of its area, and it shows that defining the subpopulations with cells from each area gives you essentially the same results, arguing against spatial variation of properties within subpopulations. I worry that these two points are getting merged together or not made clearly enough here, especially the first one. In general, the logic of this section does not seem well conveyed.

      Thank you for the feedback. In particular, your first point is made by Figure 9-figure supplement 4 when we fit an area-agnostic GMM to all modulated cells in the five main areas. However, your second point is one of the two main goals of the last Results section, along with the demonstration of the spatial distributions of cells after hard-clustering them by subpopulations. We have tried to clarify these main points further through substantial edits of the results section for Figure 10.

      One set of ideas that is highly relevant and should be raised concerns an ethological organization of the motor cortex. Since the observations of Graziano, there has been a steady stream of results describing ethological organization in rodents as well. This literature is briefly reviewed in Kristl et al., Nature Communications, 2025. For example, because of the potential for a differential involvement of grasping movements across different target locations, some of the variation in neuronal tuning described in the present manuscript may stem from a region preferentially involved in grasping.

      We agree that the Graziano literature, and the substantial literature in rodent that was inspired by Graziano’s work, is highly relevant to understanding the organization of motor areas. Kristl 2025 handles these issues very thoughtfully. The challenge here is that there are many possible different reconciliations of the stimulation results with ours, and some seriously unresolved challenges in doing so. To name a few:

      Our subpopulations and high-gradient boundaries both give quite different pictures than microstimulation does in rodent motor and sensory cortices. In particular, microstim produces more subregions that evoke different movements than we identify, and the borders don’t generally line up. This implies that the mapping between the two approaches is probably complicated.

      There is a completely alternative possibility to explaining the Graziano-like results: microstimulation is thought to preferentially hit axons, and some of these projections reach the medullary motor regions. Given that the medullary motor regions have known topography in the movements they evoke (Yang et al 2023) – but may or may not be driving the movements during flexible behavior – the two approaches may not be reconcilable. Or, it may require a much deeper understanding of medulla as driving the primary movement and cortex acting as a residual controller. This is an exciting set of ideas, but as yet very underdeveloped in our understanding.

      We don’t know if the subpopulation structure exists at all in L5, or in the PT cells, and if it does whether it differs. This is crucial given the frequent targeting of deep layers by ICMS stimulation protocols.

      As we caution in the Discussion, it is possible that our subpopulation findings are at least partly specific to the task we used.

      Although it is beyond the scope of this paper and will be addressed thoroughly in separate work, we have spent significant time with encoding models for joint angles and high-level target encoding in these same data. Given those results, we are fairly confident that the reviewer’s reasonable guess, of tuning variation due to intersections between body parts, does not seem to be the main driver of the subpopulation structure we find.

      After careful thought and discussion amongst the authors, we did not think that including this discussion in the paper was likely to improve interpretability of the present results for most readers. We very much agree with the point, though, and when we can narrow down the possible explanations in the future (likely in our next paper on this topic, which will address encoding) we plan to address it. We thank the reviewer for encouraging us to think through this.

      Minor:

      (1) Page 3: "densely shared" - perhaps "broadly shared"? Dense implies most/all the neurons get the same signals, which may not be true.

      Changed to “widely”.

      (2) Page 4: "data-driven approaches" - could be more specific - isn't everything we do data-driven?

      Changed to “bottom-up”.

      (3) Page 4: "spanned areas" - perhaps "spanned multiple cortical areas", since everything spans an area.

      Changed to “spanned multiple areas” (we mention cortex just a few words earlier).

      (4) Page 5: "intervals were generally fast" - awkward, "short" perhaps.

      Agreed, changed.

      (5) Page 5: "which asks whether the activity for a neuron changes over time consistently in relation to any target" - Rephrase to disambiguate between consistent temporal variation in firing for all targets and variation across targets in the firing patterns. In other words, are we talking about cells that are just modulated during reaching, or cells whose firing patterns differ across targets?

      Changed ending to “to any given target”. The ZETA measure really does simply ask whether there is a change in firing rate over time that is consistent across trials, for each target independently. A neuron that exhibits an identical bump for all targets would register as modulated. We chose this measure in part because of the number of temporally-modulated but untuned cells. This wasn’t very clear as we had written the text, so we now note this explicitly in the Methods. Thank you for pointing out that this wasn’t clear.

      “For all analyses, only neurons modulated by the relevant locking event were included. Note that this measure looks for modulation over time to any target; it is indifferent to whether the neuron exhibits tuning across targets.”

      (6) Figure 1: It seems like some of the abbreviations used in 1A have not been defined yet in the paper.

      Yes. It’s a long list, and we wanted to put the citations for the description of each area together with the definition of the acronym. Moreover, we wanted all this info together with the description of how we aligned these area descriptions from others’ work with one another on the Allen atlas. This was impractical in the caption, and would be a long digression for what is intended as a simple point in the Results, which is why we refer to the Methods here.

      (7) Page 8: "Given that these areas have known spatial organization within them and structure was apparent by eye in the spatial scatterplot of modulated neurons (Fig. 3A)," - it is not clear what spatial structure we are supposed to see in 3A.

      Good point. We have changed the parenthetical to: “(for example, the less modulated band along the M1/M2 border in Fig. 3A)”

      (8) Page 8-10: The region-wide onset analysis breaks up the flow from PETHs to the metrics used to quantify them. I suggest moving this section (Onset of neural activity varied with somatotopy and subregion) to later in the manuscript.

      We appreciate the reviewer’s input on organization. We went back and forth many times in how to organize the many results in this paper. The reviewer is right that this analysis breaks the flow, but the reason we included it where we did was threefold. First, it uses an easily-understood metric to introduce the reader to how we made maps from single-neuron features. Second, it easily introduces the power of making such maps. Finally, it makes clear that if we are not careful with how we handle time in the feature design, timing will dominate.

      All these things said, this has helped inspire us to add a result in which we re-examine timing broken down by subpopulation (Figure 9-figure supplement 2C). It shows that subpopulations timing distributions appear more distinct than distributions for areas, but there is still substantial heterogeneity in timing that is explained by location in cortex and not subpopulation membership alone.

      (9) Page 12, Target tuning linearity: This metric should be clarified in the Results. It is not clear how the 2D of targets is turned into 1D. Also, the plot in the figure has correlation on the y-axis, and it is not clear how each target location gets its own correlation value. The phrase "optimized anchor target" is unclear.

      Agreed this needed to be clearer. The text in the Results now reads: “To quantify how linearly a neuron’s activity related to target location in physical space, we correlated the 15D vector of mean activity of the neuron for each target with the 15D vector of the targets’ ordinal distances from the neuron’s preferred target (Methods).” In agreement with your suggestion, we have dropped use of the phrase “anchor target” in favor of “preferred target”, which should be clearer. We have also revised the Methods text accordingly to clarify.

      To directly answer your question, we turn the targets from 3D positions into 1D by computing the ordinal distance of each target from a preferred target. (Note that the preferred target is actually the one that maximizes the resulting correlation; this is detailed in the Methods). There therefore aren’t 15 correlations; we’re correlating two 15D vectors, where each has one element per target and the “ordinal distance” vector has a zero for the preferred target. Hopefully the new description makes this clearer.

      The figure schematic was unclear, thank you for catching that. We have updated the Y axis to read “mean activity” and the X axis now reads “dist. to pref. target.”

      (10) Page 12, paragraph beginning "We also compared our metric maps simply using the top 20 PCs." - This paragraph is unclear, since both sentences refer to using the metrics. I would guess the authors mean that the metric maps were compared with and without PCA and basis rotation, but this is not clearly stated.

      Thank you, this was unclear as written. We have changed it to:

      “We also compared our metric maps with maps generated from the top 20 PCs of the PETHs (Methods), rotated using VARIMAX to identify a sparser basis (Musall et al. 2019).”

      (11) Page 18: "These results make clear that the working hypothesis - of areas with well-separated feature distributions - is incorrect." This is the clearest statement of the impact of the results. The authors could consider including this in the Abstract or Introduction.

      Thank you for pointing this out. We agree, and have added a similar phrase to the Abstract.

      (12) Figure 9: It would be great to also just see the average PETHs for each of the four clusters to get a better sense of how their time series differ.

      Good idea. The feature computations are a many-to-one mapping, so it’s not possible to literally generate a PETH from the mean of the cluster, but we have added PETHs from well-modulated neurons that are near the means of their subpopulations (Figure 9-figure supplement 1).

      (13) Figure 9B: Colorbar has no label.

      Fixed, thanks.

      (14) Figure 9C: Need a colorbar - need to see the difference in density for locations.

      The color map is the same Figure 8B, which is now noted in the caption for Figure 9C. The scaling of likelihoods is almost totally uninformative; they’re not well-behaved like probability distributions, so you’ll note that even on Figure 8B the labels are simply “max likelihood” and “min likelihood”. The important pieces of information here are that these are log likelihoods (noted in the Figure 8 caption), and the visualization of the color map itself (from the color bar). Given these considerations, we have elected to keep the maps themselves a little larger by not trying to squeeze in a minimally-informative colorbar to all of the plots, but thank you for noting that the reference to 8B was needed.

      (15) Page 22: "additional spatial structure could be present" - The nature of the additional spatial structure here is a bit opaque. The authors could clarify what additional structure may be present.

      Good idea. This paragraph now reads:

      “The overlaps in the subpopulation likelihood maps above imply that members of different subpopulations are spatially intermingled, but it is less clear whether each subpopulation has homogeneous response profiles across space. In particular, the use of likelihoods mixes two properties: the fraction of neurons in a given neighborhood that are members of each subpopulation, and the heterogeneity of response profiles amongst members of that subpopulation. These properties could vary systematically with respect to one another, and the spatial structure shown by the likelihood map does not disentangle them.”

      (16) Figure 10E, legend: "GMM component" - I think this should be "GMM subpopulation" to avoid confusion with the previous use of "component" above, referring to the components of the GMM models for each region.

      Thank you – good catch. Changed to “Likelihood map”.

      (17) Page 24: "Note that this consistency also validates the use of clustering to combine components and identify the subpopulations in the first place." - I don't totally get this, and how this result validates the method of combining components, as opposed to just clustering all the cells from all regions at once. Perhaps the implied opposing strategy is not clear here.

      We have changed this sentence to:

      “Note that this consistency mirrors the low Bhattcharyya distances between corresponding GMM components in Figure 9B, and further validates the use of clustering to combine components from different areas.”

      Regarding the reviewer’s larger point, we have three thoughts. First, we do also show the result of fitting the GMM to all cells together (Figure 9-figure supplement 4).The result is similar, but the Anterior subpopulation is lost because its membership is low and so the ICL criterion can’t justify a fourth cluster. Second, because we imaged more neurons in some areas than others, fitting the GMMs to each area separately put their representations on a more equal footing. Finally, doing the analysis this way allowed us to most directly compare our two hypotheses, as illustrated in Figures 8A and 9A.

      (18) Page 25: "in the zones where different subpopulations overlapped" - I would omit this, since "intermingled" seems to mean exactly this.

      We included this phrase to prevent quickly-skimming readers from incorrectly concluding that the subpopulations overlapped entirely and were therefore intermingled everywhere. The reviewer is right that it’s unnecessary for a careful reader, but we aimed to prevent misinterpretation by readers that might skip to the Discussion for a results summary.

      (19) Page 25: "content of the activity, but also its format" - the difference between content and format is not entirely clear. Metaphor not quite metaphoring here. Agreed. We have added examples to clarify.

      “This makes clear that there are potentially important differences not just in the content of the activity (e.g., encoding target vs. movement commands (Grier et al. 2026)), but also its format (e.g., linear encoding vs. nonlinear, persistent vs. brief responses).”

      (20) Page 30, bottom: In the description of the behavior, more details should be provided, especially since the paradigm is new. For example, it says the block size was reduced - what was the ultimate block size?

      Targets were cued randomly in the behavior performed during neural recordings. Blocked trials were used during training and were phased out incrementally as performance improved. This and various other details have been added. Please let us know if there are other specific details you would like to see in the final version.

      (21) Page 39, citation of An, Mulcahey et al.: There is a biorxiv version with a different author list that could be cited.

      This was an error with our citation manager, and has been corrected. Thanks for catching it.

      Reviewer #2 (Recommendations for the authors):

      Overall, this is a remarkable study with well-designed in-depth analyses, and I only have some minor suggestions that could help improve the clarity of the paper.

      Thank you!

      General:

      It is not immediately clear to me why the GMM approach used in this study is more interesting than a clustering approach based on single-neuron response patterns (See Esmaeli et al., Neuron 2021 or Oryshchuk et al., Cell Report 2024). But my impression is that it led to the same observation that most clusters are widely distributed across cortical areas, with different proportions, but a few clusters are quite specific to a few areas. A noticeable difference perhaps is the number of clusters - or response profile - that seems particularly low (only 4) in the current study. Could the authors clarify and comment on that, maybe?

      The reviewer brings up an interesting point: at heart, these works ask related questions, albeit about different effectors, tasks, recording modalities, and types of information encoded. Those differences probably mean that results cannot be directly compared, but we can certainly discuss the methodological tradeoffs. The two papers mentioned take a more traditional first step, using PCA on the vectorized PETHs to reduce dimensionality, then layer on a spectral approach to improve clusterability. These are good methods; we use something similar as our alternate method, applying VARIMAX to the PCs instead of spectral methods to preserve linearity of transforms. For the kinds of responses both they and we have, PCA will tend to most strongly pick up two aspects of the responses: tuning and timing. This is because vectorized PETHs will have large values in the rows corresponding to the target/condition and time points where the high activity is, and the alignment of these profiles with those of the other neurons will capture a large fraction of the variance. For data like either theirs or ours, this would tend to cluster apart left-tuned cells from right-tuned, and (more importantly here for revealing spatial structure) early-response cells from later response cells. That intuition is consistent with what those papers report, and examining our VARIMAX’ed PC plots closely (which have sharpened in the latest version thanks to improved normalization), we can see that they break apart sub-regions largely based on timing. In our feature approach, we intentionally chose our features to be largely invariant to both tuning preferences and timing. Instead, we chose our features to pick up on what we call the single cell “response format”: response duration; peak time variation (but not absolute timing); and tuning sharpness, persistence, and linearity. These different methods pick up on different aspects of responses.

      To double-check that the PCA-then-spectral approach reveals similar structure to our use of VARIMAX on the PCs, we tried applying the suggested method to our data. We applied spectral clustering to the N x 20 PETH PC feature matrix, then fit an area-agnostic GMM to the spectral features. We plot the likelihood map for the components of a GMM with 10 modes. The GMM components did not display clear spatial structure beyond that observed in the VARIMAX’ed PCs (Figure 5-figure supplement 1) and were less interpretable than those identified by area-agnostic clustering of our response features (Author response image 1). As noted, the number of subpopulations identified by the clustering of our hand-engineered features is lower than what would be obtained from clustering the PCs of the PETHs. This is likely the result of the substantial heterogeneity in activity onset and preferred target that is preserved by PCA. Because our central approach is largely agnostic to these two sources of variation, the number of identified clusters reflects the dominant patterns of variation beyond these two sources.

      Author response image 1.

      GMM fit to spectrally transformed PETH PCs, agnostic to anatomical areas. One GMM was fit to the spectrally-embedded PC feature vectors of cells from all 5 main areas. Each component of a 10 component model is shown.

      Also, I think it would greatly help the reader to return to PETHs at some point, if possible, to show the response profiles of each identified neuronal subgroup (page 20). To what extent are they similar or different across the cortical areas (for the same neuronal subgroup)?

      This is a good idea. We have added a figure to address this question and the related question by R1 (Figure 9-figure supplement 1). In short, given the wide variety of PETHs we observed, there is of course still substantial variation within subpopulation, and some mild but systematic differences in the distribution of what we observe across areas. We now discuss the conclusions from this plot in the Results:

      “As a qualitative depiction of the response profiles identified with each subpopulation, we plotted the two highest-likelihood cells for each area/subpopulation combination (Figure 9-figure supplement 1). These examples reveal stereotypy in the subpopulation responses across areas, but also show variation across areas, especially for the two somatomotor subpopulations.”

      Specific:

      (1) Figure 2B and M&M: the 3D spatial organization of the target locations is not immediately clear. What is the spacing between target locations? What is the 'final azimuthal spacing'?

      Added, thanks. The pairwise horizontal distances between targets were between 1.72 and 6 mm apart and the vertical spacing within a column was 1 mm. “Final azimuthal spacing” just referred to the targets being closer together during training and our gradually spacing them apart to their final locations. We have also added some relevant details about the training.

      (2) Figure 2C: It would help to have a scale bar (mm).

      Added, thanks.

      (3) Figure 2C: It would be easier to appreciate the variability of the trajectories across trials to plot an overlay of trajectories to one target only (could be a Supplementary Figure).

      The reviewer has a good point: the variability and accuracy of aiming was hard to ascertain from the plot. We experimented with a few options for making this clearer most effectively. We have now added Figure 2-figure supplement 1 that shows in the third subpanel of panel A the finger centroid trajectories for one of the 15 targets highlighted for the mouse shown in Figure 2C, mouse 3. The centroid trajectories for all other mice are shown as well to illustrate similarities and differences across animals as well as the overall variability. As noted elsewhere we have also included an analysis of the variability of the centroid trajectories, showing that reaches to a given target were more similar than reaches to different targets. We think this provides a fuller picture of the behavior and intend to provide still more detail in future work. Thank you for suggesting additional detail here!

      (4) Figure 4: It would be nice to also show the amplitude-normalized grand-average PETHs for the different areas.

      This is an interesting suggestion. After careful consideration, we think that this analysis is not as effective for depicting overall timing and modulation profiles as the current ones, given the strong amount of target selectivity and response time heterogeneity (now better visible in the revised Figure 4A). When computing the grand mean of all cells within each area, the dominant features distinguishing areas are onset time and response duration. The differences across areas in these two features are better supported by the analyses of Figures 4 and 5 due to the large amount of heterogeneity in responses within each area. We thank the reviewer for encouraging this exploration; more complicated spin-offs will likely inform additional timing analysis in the next paper on these data.

      (5) Figure 7C: figure legend - although it is quite self-explanatory, please explicitly indicate which pattern corresponds to the 'Three contour levels (98%, 95%, 90%)'.

      We have now added this as a legend on the figure panel itself (here and on similar plots). Thanks for pointing this out.

      (6) Figure 8: Is there also an interesting asymmetry between sensory are motor areas, with neurons in sensory areas being more likely associated with motor areas (B and C), whereas neurons in motor regions are less likely to arise from the distribution of sensory areas (dark blue color in frontal regions in D, E, and F)?

      This is an interesting observation, but we understand it to be an artifact of colormap scaling. As mentioned above, likelihoods are not well-behaved like probability distributions are: for example, they are not bounded at 1, and their sums over a dataset can have any positive value. The only things that can be interpreted are their relative values. This makes their scaling functionally arbitrary – you’ll notice we used “min likelihood” and “max likelihood” instead of numbers, which would be nearly meaningless – and therefore presents a problem for scaling the colormaps. We don’t know of a principled way around this problem. To deal with it, we simply put the ends of our colormap at the extreme pixel values. It so happens that both the M1 and M2 maps had a handful of neurons in a less-sampled spot at the bottom of M2 that were very low-likelihood, which results in what you noticed. We debated removing those neurons for this purpose, but we had no basis on which to do that kind of manipulation, so we left it as the most honest representation of the data we could produce.

      To clarify this, we now mention in the caption “The ends of the colormap were set to the maximum and minimum likelihood values for each map.”

      (7) Figure 9B: there are two-time 'S1-hl: 1' indicated at the two bottom rows of the distance matrix. I suppose one of them should be 'S1-tr: 1' instead?

      Fixed, thanks for catching it.

      (8) Page 20: 'This hinted at a second hypothesis: that some of the 'modes' (groups of neurons) discovered separately in each area might correspond.' ???

      We had meant “mode” as in “multimodal”, but it was very unclear. We have rewritten the sentence:

      “This hinted at a second hypothesis: that a peak in the multimodal distribution from one area might correspond to a peak in the multimodal distribution of a different area.”

      (9) Figure 9S2: Please indicate for which area each map is computed.

      The caption was not clear enough about what we were doing here: we fit the GMM on all neurons together, ignoring which area they came from. We have now clarified it in the caption:

      “One GMM was fit to the feature vectors of cells from all 5 main areas. Each map plots the likelihood for all cells to each of the three components of this area-agnostic GMM.”

      (10) M&M, Subjects and surgical procedures: 'ambient temperature of 71.5 {degree sign}F', please use international units.

      Done.

    1. eLife Assessment

      This study makes a valuable contribution to the understanding of meta-learning and its neural mechanisms by distinguishing two timescales of learning rate adaptation: rapid, within-block reductions and slower, location-specific, meta-learned adjustments. Behavioural data and computational modelling provide convincing evidence that individuals adjust learning rates both rapidly in response to uncertainty and more gradually through meta-learning of environmental statistics. Neuroimaging results indicate that meta-learned learning rates are represented in orbitofrontal cortex, and that prediction errors are encoded across a distributed network including the ventral striatum, where they are modulated by expectations about error magnitude. The manuscript is timely and clearly written and opens the door to future work on how these signals contribute to adaptive behaviour.

    2. Reviewer #1 (Public review):

      Summary:

      Simoens and colleagues use a continuous estimation task to disentangle learning rate adjustments on shorter and longer timescales. They show that participants rapidly decrease learning rates within a block of trials in a given "location", but that they also adjust learning rates for the very first trial based on information accrued gradually about the statistics of each location, which can be viewed as a form of metalearning. The authors show that the metalearned learning rates are represented in patterns of neural activity in the orbitofrontal cortex, and that prediction errors are represented in a constellation of brain regions including ventral striatum, where they are modulated by expectations about error magnitude to some degree. The work opens the door to future work focusing on how exactly these signals contribute to adaptive behavior.

      Strengths:

      The authors build on an interesting task design allowing them to distinguish moment-to-moment adjustments in learning rate from slower adjustments in learning rate corresponding to slowly gained knowledge about the statistics of specific "locations". Behavior and computational modeling clearly demonstrate that individuals adjust to environmental statistics in a sort of metalearning. fMRI data reveal representations of interest including those related to adjusted learning rates and their impact on the degree of prediction error encoding in the striatum.

      Weaknesses:

      It was nice to see that the authors could distinguish differences between the OFC signals that they observed and those in the visual regions based on changes through the session. However, the linkage between these brain activations and a functional role in generating behavior remains somewhat unclear, opening the door for alternative interpretations.

      Comments on revised version.

      I appreciate the authors responses and they have largely addressed my concerns. I understand the concerns about power with regard to the individual differences/behavioral analyses included in the rebuttal. However, my personal view, which is perhaps a matter of taste, is that the paper would benefit from a description of these results - along with a clear description of why the authors are hesitant to draw a strong interpretation from the negative result.

    3. Reviewer #2 (Public review):

      Summary:

      Across two experiments, this work presents a novel spatial predictive inference paradigm that facilitates the investigation of meta-learning across multiple environments with distinct statistics, as well as more local learning from sequences of observations within an environment. The authors present behavioral data indicating that people can indeed learn to distinguish between noise levels and calibrate their learning rates accordingly across environments, even on initial trials when revisiting an environment. They complement their behavioral results with computational modeling, further bolstering claims of both local and global adaptation. Additional fMRI results support the role of OFC in this meta-learning process, with central OFC activity reflecting similarity between environments. This similarity emerges over time with task experience. Holistically, this paradigm and these data add to our understanding of how humans dynamically adapt their behavior on different timescales.

      Strengths:

      The novel paradigm represents a clever and creative expansion of spatial predictive inference tasks. The cover story was well chosen to facilitate an intuitive understanding of both the differences between environments, and the estimation of the mean within environments.

      Additionally, the authors present complementary results from two experiments, which strengthens the behavioral findings. This is especially effective as the initial experiment's results were a bit noisy, and the modifications within the second experiment increased both power and the specificity/accuracy of participant predictions. Taken together, the behavioral results provide convincing evidence that participants did distinguish environments based on their underlying statistics and adapted their initial behavior accordingly.

      Beyond this, the combination of behavioral results, computational modeling, and neuroimaging enhances the impact of the work. It paints a fuller picture of whether and how humans meta-learn the global statistics of environments, and this is an important direction for the field of adaptive learning.

      Weaknesses:

      Throughout much of the paper, the authors refer to the distinctions between environments primarily as differences in "initial learning rates" or "environment-specific learning rates." The optimal initial learning rate did indeed differ across environments -- the result of differences in underlying task statistics. These differences in task statistics result in distinct optimal initial learning rates and also vary with aspects of spatial position (e.g. vertical position in the example figure). The authors convincingly show that OFC activity increasingly reflects these variables throughout task experience. Given that these variables vary together, future work will be needed to distinguish whether particular variables drive these dynamics, or whether together they combine to evoke the representational differences.

      The current work is also quite suggestive of meaningful individual differences in both local and global adaptive learning, in line with other prior work on predictive inference. This is perhaps underexplored in this data set, but certainly leaves the topic ripe for follow up going forward.

      Finally, more information on all clusters that survived multiple comparisons correction would be useful, even in the absence of a priori hypotheses. For instance, there is commentary in the discussion section on the ACC, but this is not mentioned in the results, and it is unclear whether there were other undescribed clusters that survived correction.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It was nice to see that the authors could distinguish differences between the OFC signals that they observed and those in the visual regions based on changes through the session. However, the linkage between these brain activations and a functional role in generating behavior was left unexplored. Without further exploration, it is hard to tell exactly what role the signals might be playing, if any, in the behavior of interest.

      To link the behavioral with the fMRI data, we now correlated fMRI decoding accuracy with behavioral performance. We studied behavioral performance in two ways: the difference in high versus low noise environment learning rates, and mean accuracy (i.e., absolute prediction error). We correlated both measures with the decodability of the environment in the central OFC. Each correlation was calculated either in the full experiment, or only the second half. However, none of these correlations were significant (all p > .1). Given the difficulty of interpreting this result, and our lack of statistical power for doing individual difference analyses, we decided not to report these analyses in the final paper.

      Reviewer 2 (public review):

      (1) The authors make the distinction between meta-learned "global" learning rates and within environment learning rate adaptation in response to "local" fluctuations/observations. Though the experimental paradigm is novel, there are certainly links to prior work - for instance, though change point structures don't entail revisiting unique environments, they do require meta-learning from environmental statistics that is distinct from transient local adaptation to prediction errors. This tendency to increase one's learning rate after large prediction errors is appropriate in change point environments, though, as is true in this study, the amount of increase should be dependent on. This represents a similar kind of slower-timescale learning or reuse of more "global" parameters, and can be seen to different extents in prior work. It might benefit readers if the authors were to link the current work to previous research more explicitly to draw clearer connections between the approaches and findings.

      We thank the reviewer for their very helpful literature suggestions and now contextualize and discuss our findings in light of relevant literature.

      (2) Throughout much of the paper, the authors refer to the distinctions between environments primarily as differences in "initial learning rates" or "environment-specific learning rates." This is particularly prominent when discussing fMRI results. Though the optimal initial learning rate did differ across environments, this was the result of differences in underlying task statistics. It will be important to clarify this throughout the text, because of the confounds between task statistics and initial learning rate (and to some extent, the position on the screen), it is not possible to separate the impact of these specific variables. This is also relevant to understanding the justification for using methods like RSA to test whether brain regions represent task states similarly. If the main hypothesis is that neural activity reflects the (initial) learning rate itself, then a univariate analysis approach would seem more natural.

      We agree that task statistics are not the same as differences in learning rates. However, we do not consider this as a confound: The point of the differences in task statistics is exactly to generate differences in learning rates. With our paradigm, we deliberately tried to dissociate variations in learning rate that were induced by learned environmental differences versus local task statistics. We tried to make this dissociation more clear, especially when discussing the fMRI results.

      (3) For the neuroimaging results in particular, the specificity of some of the results (e.g. ventral striatum showing an effect of prediction error only in the low noise condition in the second half of task experience, only on the first trial) is a bit surprising. Additional justification of or context for these results would be useful to help readers gauge how expected or surprising these findings are.

      We agree some of these findings were unexpected. We now also highlight that while we expected the ventral striatum to be involved in prediction error processing, we had no strong a priori expectations regarding these further modulations by time and environment. We also tried to contextualize these interactions more.

      (4) There are some methodological details that are unclear (e.g., how were the positions of the crabs selected relative to the location they emerged from? Looking at Figure 1C, it looks like the crabs spread out unevenly, and that the single position they emerge from is not necessarily at the center of the crab locations.) Additional detail and clarity would help address some unanswered questions (more details below).

      We clarified the experimental procedure at several places, and now added a video that helps illustrate the trial timeline better.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) With regards to the primary weakness mentioned above, it would be nice to have some link between the brain signals of interest and upcoming behavior. For example, can you read something out of OFC that enables you to better predict what the participant will do next? Or even better, do so beyond any behavioral variability that is explained by the computational model?

      To link the behavioral with the fMRI data, we now correlated fMRI decoding accuracy with behavioral performance. We studied behavioral performance in two ways: the difference in high versus low noise environment learning rates, and mean accuracy (i.e., absolute prediction error). We correlated both measures with the decodability of the environment in the central OFC. Each correlation was calculated either in the full experiment, or only the second half. However, none of these correlations were significant (all p > .1; see plots in Author response image 1). Given the difficulty of interpreting this result, and our lack of statistical power for doing individual difference analyses, we decided not to report this analysis in the paper.

      Author response image 1.

      (2) A number of the learning analyses are based on splitting the session into halves. As a first pass, this seems like a reasonable thing to do, but I certainly wonder what the dynamics of the meta-learning actually look like, and it seems like the data collected would be sufficient to gain some insight into those dynamics through some sort of sliding window analysis.

      We thank the reviewer for this interesting suggestion, which was also raised by Reviewer 2. We now calculated the learning rate in a sliding window of 20 trials (i.e., trial x to x + 19), and provide revised figures for each experiment separately (Fig. 2E and Fig. 4E, respectively).

      (3) The model selection procedures described make sense, but it would still be useful if the authors justified them by showing that they work in synthetic data (ie, generate a confusion matrix). I may be confused about what delta-SE is, but I'm confused about why two models with very different fits have the same value (211) for that metric.

      We report model recovery on synthetic data, which yielded model recovery rates of 100%, and added these to our Methods section. To clarify the Reviewer’s second point, ∆SE is the standard error of the difference between a model’s LOOIC and the top ranked model’s LOOIC. There is no one-to-one mapping between the ∆SE and a model’s LOOIC.

      (4) Was the central OFC anatomical ROI overlapping with the cluster surviving in the whole brain analysis? I didn't see this mentioned in the text, and it certainly would be important for interpreting the two results together.

      The central OFC indeed overlapped with the cluster surviving whole brain analysis, which we report on page 17-18.

      (5) The authors found regions that reflected learning rate at the "island presentation" phase of the task - it could be distinguishing this analysis and its meaning from other work that has focused on representations of learning rate at the time of feedback.

      We agree that this is an important distinction worth emphasizing. Therefore, we added the following lines to our discussion paragraph: 

      “Importantly, previous studies examined neural correlates of learning rates during outcome evaluation, where learning rates may be adjusted online as a function of locally experienced prediction errors (e.g., (Behrens et al., 2007; Browning et al., 2015; Nassar et al., 2012). In contrast, our RSA analysis targeted neural activity at island presentation, before any outcome information was available. At this moment, learning rates cannot be updated based on current feedback and instead reflected the retrieval of a previously learned, environment-specific learning-rate settings. This difference reflects our hypothesis that the OFC represents the latent states in a cognitive map of the task (Knudsen & Wallis, 2022; Moneta et al., 2024; Schuck et al., 2018; Wilson et al., 2014), which are expected to activate as soon as the agents can infer which task state it is in. Several studies have identified such “partially observable” task states in the medial OFC (Bradfield et al., 2015; Schuck et al., 2016; Tan et al., 2025; Wimmer & Büchel, 2019), in line with the region identified here (but see e.g., (Ongur & Price, 2000), for important anatomical distinctions between medial and lateral OFC and (Tan et al., 2025) for an example of related functions in lateral OFC). Our finding extends this notion by suggesting a link between OFC and meta learning, wherein meta-learned information becomes encapsulated in task states (Hattori et al., 2023; Moneta et al., 2024).”

      (6) "Specifically, it showed a more negative response to larger (location) prediction errors, which is consistent with its documented role in showing a more positive response to more positive reward prediction errors (Calderon et al., 2021) - keeping in mind that being closer to the centre of where the crabs appeared (i.e., smaller location prediction errors) is less negatively or more positively surprising (i.e. smaller negative or larger positive reward prediction errors)."

      I found this sentence very hard to parse. Do PE responses in the high noise environment get "compressed" in their representation over time (ie, it takes a larger error to get the same BOLD response)? If so, this relates to claims made in Diederen 2016... but see also Mah 2024 Cell Reports, who fails to see learning rate encoded in DA system in striatum of rodents that appear to adjust their learning rates.

      Thank you for pointing to this. We agree that this sentence was hard to parse, and so we now split it in three revised sentences. We also agree with the Reviewer’s interpretation, and would like to thank the Reviewer for their useful literature suggestions which we now added to our discussion. 

      (7) Figure 7 should use a different color scheme because many of the activations just appear black, and I can't tell whether they are positive or negative. It was also notable in Figure 7A that regions are not visible, including ACC, which is typically thought to encode prediction errors in such paradigms. It would probably be useful for the authors to include a table of all clusters exceeding multiple comparisons correction and to on differences to other work examining absolute prediction errors. ACC does appear on the second trial, which made me wonder whether there were changes in the prediction error coding from first to subsequent trials. 

      Thank you for pointing this out. We now revised our color scheme which we agree makes it much clearer now. Although the ACC is frequently implicated in prediction error–related signals (e.g., Behrens et al., 2007), models suggest that ACC responses more strongly reflect unsigned prediction errors, surprise, or the need for control and model updating (Alexander & Brown, 2019; Hayden et al., 2011; Silvetti et al., 2018). In our task, ACC activity only emerged on the second trial, when participants had formed an initial estimate and prediction errors could meaningfully signal the need to update internal models or control settings. We now added a to the Discussion highlighting this distinction and relating our findings to this prior work emphasizing prediction errors and control-related signals in ACC.

      (8) The authors suggest that fast learning would presumably occur in a neural activation space, whereas slow learning would occur through weight adjustments. This makes sense, but activity-based dynamics have been suggested to do rapid adjustments by encoding a "latent state" though (Razmi 2022 j neurosci) -- and such a latent state has been shown in OFC (Schuck etc)... but here OFC is more implicated in the slow learning. I am curious about whether authors could on this a bit in the discussion. 

      Thank you for bringing up this interesting question. We can only speculate but a crucial factor is on which level of resolution tasks states operate. On the one hand “detailed” trial-level states are needed that map a specific sensory input onto a specific latent state and its value. Such states would change quickly, possibly through activation dynamics, and are in line with how they have been operationalized in Razmi or Schuck etc. On the other hand, successful task performance also needs “higher level” states that describe entire task phases or full tasks, as in the present experiment. Due to the different speeds of learning, it appears plausible that these would be learned with synaptic changes. We expand on this in the discussion as follows: 

      “Our finding extends this notion by suggesting a link between OFC and meta learning, wherein meta-learned information becomes encapsulated in task states (Hattori et al., 2023; Moneta et al., 2024). Consistently, OFC has been shown to represent task states (Moneta et al., 2024; Stalnaker et al., 2015; Wilson et al., 2014). While earlier evidence shows that the OFC represents concrete aspects of task states, such as task-relevant stimulus features (Schuck et al., 2016), we hypothesized that the OFC also represents more abstract aspects, such as learned, environment-specific learning rates. Indeed, we showed that the central OFC gradually came to represent these environment-specific learning rates (or the environment-specific statistics that drive them). While previous work speculated that these different levels could have different neural underpinnings (Sharpe et al., 2019), our findings indicate OFC might signal states on multiple levels. This does not imply identical learning dynamics; fast-changing trial-specific states might be learned through activity dynamics, while higher-level contextual states could involve synaptic plasticity.”

      (1.9) Also, as a more minor point in the same section, the sentence about blocking synaptic plasticity in OFC sounded interesting, but should have a reference.

      Thank you for noticing, we now added the reference (Hattori et al., 2023).

      Reviewer #2 (Recommendations for the authors):

      (1) Additional links to prior literature: In terms of prior work in which there is something akin to more "global" adaptation, some examples of potentially relevant prior work include:

      McGuire, Nassar, Gold, & Kable (2014) Neuron 

      D'Acremont & Bossaerts (2016) Cerebral Cortex 

      Lee, Gold, & Kable (2020) Decision 

      Bakst & McGuire (2021) JEP: General 

      Bakst & McGuire (2023) Cognition

      We would like to thank the reviewer for pointing us to these different literature suggestions which we agree help us contextualize and discuss some of our findings better. We now refer to McGuire et al. (2014) when discussing the fMRI results, and d'Acremont & Bossaerts (2016) when discussing potential alternative strategies in the high noise environment (the Reviewer’s last point). Finally, we integrated the clearly relevant works of Bakst & McGuire (2021; 2023) and Lee et al. (2020) in our discussion of meta-learning different adaptive strategies. 

      (2) Individual differences: Though not always the focus of work on predictive inference, one common finding has been that there are pronounced individual differences in behavior (see, e.g., coefficients in Figure 2 in Nassar et al. 2019 eLife, or Figure 2 McGuire et al. 2014 Neuron, or Bakst & McGuire 2023 Cognition). There appears to be substantial variability between individuals in your data as well (i.e., Figure 2B, 4B, and the modeling figures). It would be interesting to see some direct exploration of this variability: baseline learning rate appears to differ between participants to a large extent, does their rate of adaptation (across trials within a block) also differ? Does their metalearning occur at different rates (in fact, do some participants not show evidence of appropriate meta-learning at all)? 

      Relatedly, your computational modeling approach fits the six candidate models hierarchically, and therefore the reported results show the overall best fit for the group. It might be worthwhile to determine whether individuals have different best-fitting models. This could be another way to characterize the variability between individuals. 

      In concert with this, it could be a useful complement to determine whether either the strength of the OFC neural similarity results or their time course reflects aspects of behavior. Put another way, is it the case that not only does OFC activity and behavior both come to reflect task structure, but that these changes happen to a similar extent and over a similar time course across individuals?

      We agree it would be highly interesting to investigate meaningful individual differences in both fast and slow adaptations in learning rate. However, our sample was not set up and is underpowered to conduct such analyses. In response to a similar by Reviewer 1, we did run correlational analyses between differences in learning rate, performance accuracy, and the responsiveness of the OFC. However, none of these analyses yielded a significant effect. We decided to not include these results in the paper, for reasons of statistical power, but we report them in Author response image 1.

      (3) fMRI:

      (3a) The primary finding in OFC is restricted to the central OFC. The manuscript would benefit from additional explanation regarding this specific subregion. 

      Thank you for bringing up this important distinction. In the discussion we now clarify as follows: 

      “This difference reflects our hypothesis that the OFC represents the latent states in a cognitive map of the task (Wilson et al., 2014; Schuck et al. 2018; Knudsen & Wallis, 2022; Moneta et al, 2023), which are expected to activate as soon as the agents can infer which task state it is in. Several studies have identified such “partially observable” task states in the medial OFC (Schuck et al., 2016; Bradfield et al., 2015; Wimmer et al., 2019; Tan et al., 2025), in line with the region identified here (but see e.g., Öngur & Price, 2000, for important anatomical distinctions between medial and lateral OFC and Tan et al., 2025, for an example of related functions in lateral OFC).”

      (3b) Though the main clusters visible in Figure 6 are the occipital and OFC clusters, there appear to be others. Did other clusters indeed rise to statistical significance in the whole-brain analysis? If so, is there a reason they aren't included or discussed? 

      All clusters visible in Figure 6C survived FDR correction. However, we refrained from interpreting these other clusters, because we had no prior hypotheses about them like we did for the OFC.

      (3c) Why do you posit that the ventral striatum becomes less sensitive to RPE on the second trial over time? And why is the ventral striatum only sensitive to RPE in the low noise environment generally?

      We reasoned the ventral striatum should be more responsive to more positive reward prediction errors. While we further assumed this response could be modulated by both time and environment, we would like to emphasize that we had no specific hypotheses about the direction of this modulation. We now also make this clearer in the manuscript. This being said, we believe both the pattern that its responsiveness to the second trial decreases over time, and the pattern that it was most sensitive to the low noise environment, can be considered fitting with its broader involvement in coding behaviorally relevant reward prediction errors. Namely: 

      First, we believe that as the participants learn more about the global reward structure of the task, they should obtain a better understanding of the fact that, per round, all crabs always center around a fixed mean. Therefore, the first RPE is most behaviorally relevant, and every later RPE has an exponentially decreasing relevance. As participants obtain more experience with this aspect of the task over time, the VS should show a lower responsiveness to the second RPE over time.

      Second, as participants learn more about the local differences between the three different environments, they should learn that especially in the low noise environment, RPEs are most behaviorally informative. That is, in this environment it makes most sense to have a high learning rate and thus let the RPEs substantially inform the placement of the cage on the next trial. Accordingly, participants showed that the ventral striatum was most responsive to RPEs in these environments.

      (4) Methods

      (4a) This section could generally benefit from some proofreading. 

      We now proofread the method section. 

      (4b) The main results text states that 49 participants performed Experiment 1, while the methods section reports 50 participants. Which is correct? 

      (4c) Following this, on page 8, statistical results are reported with a df = 49 (which would be appropriate only if n=50). 

      The correct sample size was actually 50, we adjusted the text and degrees of freedom where incorrect accordingly (note: only text is in track changes, but degrees of freedom were also changed accordingly). 

      (4d) Additionally, I am a bit surprised by the Experiment 1 findings that learning rates on the second trial were significantly different between low and high noise conditions, in that the effect size found using all trials was stronger than both the first half of trials (no significant effect) and the second half (significant but weaker than all trials). Are these all the same type of statistical test? Double-checking the statistics might be worthwhile. 

      It is not the effect size that is larger across the full experiment, but the t-statistic. This is possible because a t-statistic depends on both effect size and noise estimate, and the latter is smaller with more data. 

      (4e) The methods and results both state that the five crabs always emerged from one position in the sand. How were the locations of the crabs selected relative to this position? Looking at Figure 1C, it looks like the crabs spread out unevenly, and that the single position they emerge from is not necessarily at the center of the crab locations. 

      The crabs did indeed spread out evenly. However, we can see how the graphic in Figure 1C can be confusing, as two crabs are shown to be caught, which breaks the symmetry of the dispersion (because some crabs can run away after the even spreading phase, see Methods). We emphasized the even spreading more clearly in the new version of the paper. We think the flow of events will be much clearer with our newly added animation (Video 1).

      (4f) The methods section states that the crabs "spread out to cover the same proportion of the screen width as the cage (18.75%)" (page 23). The corresponding visual in Figure 1C appears to show something different. 

      This looks different because the graphic illustrates the last 500 msec, where crabs can run away (see also response to 4e, and the novel animation that was added).

      (4g) Information on the timing of the trials would be useful to include in Figure 1C or similar. 

      The reader can find this information in the Methods section. We chose not to include it in the caption to avoid information overload.

      (4h) The methods section specifies that there was a 3-7s ITI after the first and second trials of each block. How was the ITI selected for each trial? Were there ITIs between the other trials? If so, what were they? 

      The ITIs were selected from a truncated exponential distribution. This selection was not random, but rather a distribution was carefully constructed for each environment (and event of interest: boat presentation, first trial of each block, second trial of each block) separately to ensure that enough longer ITIs were selected for each environment (and event of interest). Of course, the order in which the ITIs were used across blocks, was random. The same approach was used to determine the duration of the presentation of the boat at the start of each block. There were no ITIs after later trials.

      (4i) Please provide a link to the data and analysis materials on OSF in the text. 

      We now provide a link to the data and analysis materials in our methods section.

      (4j) In the methods section, there are some references to information provided "below" (page 26: "The two approaches resulted in different posterior densities (see below) for estimate uncertainties, but in similar posterior densities (see below) for learning rates..."). Where in the paper is this referencing? 

      We indeed did not detail this further as we considered it not further relevant to our main study, and now removed the references to “below”.

      (4k) The methods section specifies using uniform priors between the lower and upper bounds of the relevant parameters. This seems likely to be 0 and 1, but should be listed explicitly. 

      Thank you for noticing. We now added this to our manuscript.

      (4l) For parameter recovery, correlations are provided to indicate effective recovery. These correlations are indeed high and suggest excellent recovery, but correlations wouldn't reveal if there was systematic over- or underestimation occurring. It might be useful to provide some visualizations of the parameters and their estimates to speak to this potential issue. 

      We now visualize the parameter recovery results in Author response image 2, which show that, indeed, there was a slight underestimation of the decay rates, but not the learning rates. Importantly, our main analyses and results all pertain to the learning rates, and we never made hypotheses or conclusions about the decay rates.

      Author response image 2.

      (4m) The methods section ends with a reference to a reward localizer (page 32). This localizer doesn't appear to be mentioned/used elsewhere. 

      Indeed. We implemented the localizer because we wanted to independently identify reward processing areas. However, this localizer did not succeed in localizing a reward area (no significant results), possibly due to the fact that (1) it was performed by the end of the experiment when participants may have been fatigued, and (2) there was no learning component in this localizer task. For these reasons, we did not use it after all.

      (5) Analysis: 

      (5a) Did you consider fitting a Bai model that only allowed for environment-specific initial learning rates (with a non-environment-specific decay rate)? Given that the data (e.g., Figure 2, Figure 4) seems to support differences in initial learning rate but not necessarily a difference in the rate of change, it might be worthwhile to see whether a model like that fits best. 

      We now fitted this extra model, which we called the semi-environment-specific Bai model. See Author response tables 1 and 2 for result in experiments 1 and 2, respectively) for the results. This new model has the best (in Experiment 2) and second-to-best (in Experiment 1) LOOIC. In a way, this is not surprising, because the model formulation is entirely based on the data. We think that we can draw the same substantive conclusions with or without this extra model, so for simplicity we did not include this new model in the paper itself.

      Author response table 1.

      Note. Models are ranked in descending order according to how well they fit the data. LOOIC refers to a model’s approximated expected log pointwise predictive density. Higher values indicate higher out-of-sample predictive fit. SE refers to the standard error of a model’s LOOIC. ∆LOOIC refers to the difference between a model’s LOOIC and the top ranked model’s LOOIC. ∆SE refers to the standard error of the difference between a model’s LOOIC and the top ranked model’s LOOIC.

      Author response table 2.

      Note. Models are ranked in descending order according to how well they fit the data. LOOIC refers to a model’s approximated expected log pointwise predictive density. Higher values indicate higher out-of-sample predictive fit. SE refers to the standard error of a model’s LOOIC. ∆LOOIC refers to the difference between a model’s LOOIC and the top ranked model’s LOOIC. ∆SE refers to the standard error of the difference between a model’s LOOIC and the top ranked model’s LOOIC.

      (5b) If part of the goal is to investigate whether there is a distinct local change in LR between conditions (dependent on prediction errors), then there might be more direct ways of doing so as a complement to the modeling approach. One potential way could be to visualize the LR or change in LR as a function of PE. 

      We agree that it’s beneficial to use a direct (model-free) approach to represent learning rate as a function of condition; that is also part of our approach. For example, see Figures 2, 4, which shows learning rate as a function of condition, but in a model-free manner. We think learning rate as a function of prediction error is less informative, because the idea is that prediction error can (in Kalman-filter terminology) be indicative of either noise variance or process variance, and participants are able to distinguish between them. This is also why we constructed the conditions in such a way that on the very first trial, prediction errors were on average the same across conditions. The fact that participants did respond appropriately to prediction errors on the very first trial (i.e., larger updates or learning rates in the low noise condition), suggested they are able to assign the prediction error to process variance (in the low noise condition) versus noise variance (in the high noise condition).

      (5c) In addition to looking at the evolution of LR across trials within a block separated by task epoch (i.e., Figure 2C-D & Figure 4C-F), the structure of the task would lend itself very nicely to visualizing the evolution of the second trial LR on its own across instances. This could provide additional insight into the meta-learning process.

      We thank the reviewer for this interesting suggestion, which was also raised by Reviewer 1. We now calculated the learning rate in a sliding window of 20 trials (i.e., trial x to x + 19), and provide revised figures for each experiment separately (Fig. 2 and 4, respectively).

      (6) The environment-specific Bai model appeared to become less good at capturing participant behavior with increased environmental noise. Why do you think this is?

      We thank the reviewer for raising this point. In this environment, individual outcomes are considerably less indicative of the latent mean, which may reduce the usefulness of the trial-by-trial, prediction-error–driven learning-rate adjustments that we see in the other environments. Under such extreme conditions of variability, people may rely less on delta-rule updating and more on alternative strategies (D'Acremont & Bossaerts, 2016; Reynders et al., 2026), such as exploratory adjustments or heuristics that are not explicitly captured by the Bai model but also outside the scope of the present paper.

    1. eLife Assessment

      This important study provides solid novel evidence for a role of ripples in the hippocampus in visual short-term memory. The work is strong in employing state-of-the-art intracranial electrophysiology in epilepsy patients with multivariate pattern classifiers in the context of an elegant experiment, but several aspects of the theoretical framing, mechanistic interpretation, and analysis strategy are incomplete.

    2. Reviewer #1 (Public review):

      Summary:

      Cai et al. investigated the role of ripples in the hippocampus and coupled between the hippocampus and the neocortex in visual short-term memory (VSTM) using a similar lures match-to-sample task. The main findings are that hippocampal, but not neocortical ripples, ramp up during the maintenance period, peaking shortly before the memory response is given. This ramping-up effect was stronger for correct compared to incorrect trials. Furthermore, the authors show that stimulus category could be better decoded during coupled hippocampo-neocortical ripples compared to uncoupled ripples. These results provide compelling novel evidence for a role of ripples in supporting human visual short-term memory.

      Strengths:

      (1) State-of-the-art intracranial EEG in 13 patients during a well-designed visual short-term memory task, with simultaneous hippocampal and neocortical recordings.

      (2) Thorough analysis pipeline with validation to detect ripple events, and distinguish them from spurious ripple activity (i.e., as induced by IEDs).

      (3) Use of multivariate classifiers to resolve the neural representation of the stimuli.

      Weaknesses:

      It is difficult to find clear weaknesses in this paper, as the analyses are thorough, the results are clear, and the writing is excellent. However, some more sanity checks on the validity of ripples could have been conducted (i.e., making sure that ripple events have multiple peaks in the unfiltered raw signal at the ripple frequency). Also, the time window for coupled ripples appears to be a bit long, which makes it questionable to what degree these ripples are coupled (i.e., the time window is ~5 times longer than the duration of a ripple event). Lastly, the ramping-up effect could have been more clearly depicted in the figures, but that's a fairly minor point.

    3. Reviewer #2 (Public review):

      Summary:

      Liu et al. record intracranial EEG from the hippocampus and lateral temporal lobe in thirteen neurosurgical patients while they perform a delayed match-to-sample visual short-term memory task. The central question is whether hippocampal sharp-wave ripples (brief high-frequency oscillations well established in the long-term memory consolidation literature) also contribute to the active maintenance of visual representations over a short delay. The authors report three main findings: hippocampal ripple rates progressively ramp up across the 7-second maintenance period, hippocampal ripples temporally co-occur with ripples in the lateral temporal lobe, and these coupled events coincide with above-chance category-level decoding of the memorized stimulus in the lateral temporal lobe. The findings are interpreted within the dynamic coding framework of working memory, which predicts discrete reactivation bursts rather than sustained firing during maintenance. The question is timely, and the use of intracranial recordings affords a level of temporal and spatial resolution unavailable to non-invasive methods.

      Strengths:

      The study addresses a genuinely important and underexplored question: whether a neural mechanism best characterized in the context of offline memory consolidation is also engaged during active online maintenance. The use of intracranial recordings in humans is well suited to this question, providing the millisecond temporal resolution and regional specificity needed to detect transient high-frequency events. The dissociation from long-term memory, tested by splitting remembered trials according to whether the item was later recalled in a cued-recall test, directly addresses what would otherwise be a significant confound, and the finding that ripple dynamics during maintenance are unrelated to subsequent long-term memory performance adds specificity to the interpretation. The coupled ripple analysis is methodologically grounded, and the finding that coupled but not isolated ripples coincide with elevated memory decoding is mechanistically informative. The multivariate decoding approach applied to lateral temporal lobe spectral power provides a meaningful index of memory reactivation that goes beyond simple univariate rate measures. The control analysis and the alternative ripple detection method provide useful robustness checks. The public availability of preprocessed data and analysis code on OSF is commendable.

      Weaknesses:

      (1) Theoretical motivation for examining ripples in visual short-term memory.

      A fundamental question that the paper does not adequately address is why hippocampal ripples, a mechanism strongly associated with offline memory consolidation during sleep, where they coordinate the transfer of hippocampal representations to cortex through temporally compressed replay, should be recruited for the online maintenance of visual information over a seconds-long delay. The Introduction acknowledges this gap but does not close it. The dynamic coding framework is used to motivate the ramping-up prediction, but this framework is agnostic about the specific neural mechanism responsible for reactivation bursts. In particular, the literature cited by the authors predicts high-frequency population activity or gamma bursts, but not specifically hippocampal ripples. The reasoning that "ripples share key properties with postulated reactivation bursts" risks being circular: it amounts to saying that ripples could be the relevant mechanism because the relevant mechanism has properties that ripples also have. A stronger theoretical motivation would require either evidence that the replay or reactivation computations that ripples support during offline states are also engaged during active short-term maintenance, or a mechanistic account of how the circuit processes underlying ripple generation are recruited differently across these two contexts.

      This concern is compounded by what the authors present as one of their main controls. The finding that ripple dynamics during maintenance are not associated with subsequent long-term memory performance is treated as a reassurance that the observed effects are specific to short-term memory. But if ripples are canonically a long-term memory consolidation mechanism, the observation that they are engaged by a short-term memory task while appearing disengaged from concurrent long-term memory encoding is itself a finding that demands explanation. Resolving this tension is important for the paper's contribution to be correctly interpreted by the field.

      (2) Ripple detection and specificity.

      Even granting that ripples could in principle contribute to short-term memory maintenance, the study does not establish that the detected events are physiological sharp-wave ripples rather than broadband high-frequency activity. The detection band (70-180 Hz) substantially overlaps with the high-gamma range, which is a well-established proxy for local neural population activity and coding, and is broader than the 80-120 Hz band used by several of the cited papers, including Vaz et al. (2019), Ngo et al. (2020), Chen et al. (2021), Staresina et al. (2023), and Kunz et al. (2024). Without demonstrating that detected events have the hallmark features of physiological sharp-wave ripples, a clear narrowband spectral peak, and characteristic waveform morphology, it is difficult to conclude that the observed effects reflect a ripple-specific mechanism rather than a more general high-frequency population activity phenomenon. The reported mean rate of 0.29 Hz is somewhat higher than rates reported in some recent work, such as Chen et al. (2021, ref 74) and Kunz et al. (2024, ref 15). It is worth noting that van Schalkwijk and Helfrich (2026, Nature Communications) demonstrated that a large proportion of awake ripple detections in the human medial temporal lobe reflect false positives arising from aperiodic 1/f noise, with task-related modulations of this noise floor producing spurious detections. The authors present an 80-120 Hz control analysis as a robustness check, but this inverts the appropriate logic: if 80-120 Hz is the more validated band, as the cited literature suggests, it should serve as the primary analysis rather than a supplementary one.

      (3) Internal inconsistency with the dynamic coding framework.

      The authors invoke the dynamic coding framework, which predicts that reactivation bursts should ramp up toward the end of the retention interval in the region where memory representations are actively maintained. The hippocampal ramping-up result is presented as confirming this prediction. However, the lateral temporal lobe, the region where above-chance category decoding is found and memory reactivation is attributed, shows no corresponding ramp-up. The authors acknowledge this asymmetry but do not offer a mechanistically satisfying explanation, and the suggestion that the effect might exist in unsampled subregions cannot be evaluated with the current data. This leaves the framework's core prediction unconfirmed in the region that is claimed to maintain the representations.

      (4) Coupled ripples, directionality of hippocampal-lateral temporal coupling, and the ramping-up paradox.

      The conclusion that coupled hippocampal-lateral temporal ripples coordinate memory reactivation creates a logical tension that the paper does not resolve. If hippocampal ripples drive lateral temporal reactivation only when co-occurring with lateral temporal ripples, and hippocampal ripples ramp up in a memory-predictive fashion, then the absence of lateral temporal ripple ramping up implies that the hippocampal ramp-up is not primarily expressed through the coupled ripple mechanism, undermining the coherence of the two main findings. The coupled ripple analysis further quantifies only temporal co-occurrence and provides no evidence about the direction of influence. Without demonstrating that hippocampal ripples systematically precede lateral temporal ripples (i.e., the expected signature of hippocampus-to-cortex information flow), the central claim that hippocampal ripples drive lateral temporal reactivation remains an interpretive assumption. Directly testing whether lateral temporal ripples specifically coupled to hippocampal ripples show a ramping temporal profile during maintenance (even if overall lateral temporal ripple rates do not) is necessary to establish whether the lateral temporal lobe engages in hippocampally-gated reactivation bursts in the manner the framework predicts. Additionally, reporting the distribution of peak lags between hippocampal and lateral temporal ripple peaks, and testing whether hippocampal ripples systematically precede lateral temporal ripples, is similarly necessary to support the directional interpretation.

      (5) Trial-level analysis clarity.

      The paper reports that ripples occurred in 54%, 79%, and 27% of trials during encoding, maintenance, and retrieval, respectively, but does not state whether subsequent analyses were conducted on trials thresholded by ripple occurrence. Given that occurrence rates vary substantially across stages and conditions, this inclusion criterion has implications for interpreting rate differences and should be stated explicitly.

      (6) Statistical model specification.

      The methods describe the ramping-up analysis using both a "logistic" link function and a "Poisson link function" in different places, with the dependent variable described inconsistently as ripple occurrence and ripple count. These are not equivalent, and the distinction matters for interpreting the reported coefficients. Additionally, the regional dissociation in Figure 3 appears to be assessed by fitting separate models to each region and comparing results informally. This does not constitute a direct test of whether slopes differ between regions and risks the well-known error of inferring a difference based on one p-value being significant while another is not. A direct region × time interaction test would more cleanly support the claimed dissociation.

    4. Reviewer #3 (Public review):

      Summary:

      Liu, He, et al. present results suggesting hippocampal ripples support short-term working memory. The basic finding that hippocampal ripples increase during a 7s working memory maintenance period is intriguing and previously not shown as far as I know, but a lack of control analyses within the task, across brain regions, or as compared to alternative oscillatory signals makes the overall evidence weak. The author needs to more thoroughly evidence this signal via several analyses (suggested below) to strengthen their finding. The paper moves on to a hippocampal-cortical ripple coupling analysis that needs further methodological details and corrected statistics to make a meaningful contribution. As is, the ripple coupling results don't seem to necessarily relate to the hippocampal ripples found in the maintenance period, making the manuscript somewhat incoherent and of low impact in its current form.

      Major issues:

      (1) The framing sets up "visual short term memory" (VSTM) and "long term memory" (LTM) as two different things. A long line of research with humans possessing MTL/hippocampus damage shows the hippocampal memory system contributes to working memory only when the task is difficult enough to warrant its recruitment (see Hannula et al. 2006 J. of Neuroscience, Pertzov et al. 2013 Brain, or particularly Jeneson et al. 2012 Learning & Memory and J. of Neuroscience). This theory therefore, suggests that the hippocampus contributes to working memory via LTM mechanisms, as opposed to it possessing two different roles (VSTM and LTM). While the authors might disagree with this framing, at a minimum, they should describe this line of work. As is, it's difficult to know how their task fits into this literature since it's a cross between a pattern separation probe (identify repeats from lures), working memory (7 s delays), and subsequent cued associate recognition. Addressing why they used this combination of task features would help frame its place in the literature.

      (2) The basic idea of looking for hippocampal ripples as a marker for working memory maintenance is new, with no prior literature (that I know of in rodents or in the handful of human intracranial ripple papers) to build on. That said, I suspect hippocampal ripples act as a proxy for hippocampal activation, providing a possible explanation for the hippocampal ripple increase shown during the Maintenance period. The effect they show is well supported by the mixed effects modeling (MEM), making it a potentially meaningful finding, but considering the novelty, it's rather important that control analyses rule out alternative possibilities. I suggest two important ones and a third related to the lack of parametric manipulations in the next paragraph. First, the authors frame the paper by suggesting hippocampal ripples share features with beta/gamma burst theories of working memory maintenance. In that case, the obvious question is why use a ripple detector instead of measuring gamma (or beta) activity as in this previous work? Some work has suggested hippocampal ripples act differently than high-frequency activity (see Sakon et al. 2024 J. of Neuroscience), so an analysis contrasting ripples and gamma seems rather important. Second, and relatedly, the authors only compare the hippocampus and lateral temporal cortex (LTC), likely because these tend to be sites with strong coverage in epilepsy cases. That's ok, but typically there is also reasonable coverage in other MTL areas like entorhinal cortex and amygdala, which would serve as important controls to show what they're measuring likely relates to sharp-wave ripples (a hippocampal phenomenon) and not something more generic like gamma or HFA (as shown in Sakon et al. 2024, Howard et al. 2003 Cerebral Cortex, Axmacher et al. 2007 reference 26, Meltzer et al. 2008 Cerebral Cortex, etc.).

      (3) Related to the last point, since there are no parametric manipulations (e.g., different delay durations, different set sizes, varying lure difficulties) there's no way to assess increased hippocampal ripples with stronger loads, which would be important for determining the hippocampal dependence of their task in the first place. Do the authors have any justification for this task as an assessment of hippocampal working memory? I could imagine using a top vs. bottom tercile of lure discrimination difficulty (as assessed across all participants or control non-patients) to compare hippocampal activity. But only after the first trial, each pair is used since only then would the patient have awareness of the difficulty of the upcoming comparison. Or maybe something could be done by comparing VSTM performance by splitting patients based on how they performed at the LTM test.

      (4) Also related to the VSTM vs. LTM framing, the authors use an "LTM" cued category recognition task--presumably done at the end of the repeat/lure recognition task--as a way to argue that the hippocampal ripple effects they see relate to VSTM and not LTM. The LTM task is disappointingly underdescribed, where even in the methods (lines 588-592) I cannot figure out when this task was probed, how many trials were done in comparison to the VSTM task, etc. Considering they use the LTM task to support their VSTM interpretation, it's rather crucial to understand precisely what they did. As is, the comparison they do present relies on a statistical error, where they compare p-values (n.b. https://www.nature.com/articles/nn.2886) instead of performing a direct interaction test (lines 177-180). Specifically, if they want to say their signal relates more to VSTM subsequent memory rather than LTM subsequent memory, they need to run a model of the form: ripple_rates ~ remembered + test_type + remembered*test_type (where test_type is either their VSTM or LTM task).

      (5) As noted, the increase in hippocampal ripples during maintenance seems substantial, and the MEM confirms a significant increase over time. That said, the presentation of the data is atypical, with an example raster from one channel followed by average time courses of ALL participants below it. Why not show full raster plots for all participants? Ripples are so sparse that all the data in the task can be visualized in a single raster easily. A swarm plot indicating inter-patient variability in the maintenance signal also seems crucial. As is, there is no way to assess how much of the signal depends on a small subset of channels or patients.

      (6) To compare ripple rates across task phases, they average over the bounds of each phase (lines 657-660) and input these into their MEMs. This approach makes sense for quantifying what we see in the ripple plots (Figure 2), except for Encoding, where they average over the entire 3 s window, even though there is clear tuning only from ~0-1 s. Using the tuned region and not the entire window is standard and would be more appropriate for the comparisons to maintenance, retrieval, etc (e.g., line 147-148 doesn't check out when looking at the figure), otherwise you are averaging over a seeming ripple inhibition from 1-2 s. They perform a cluster-based permutation test as is, so that a window or something a bit wider would be appropriate.

      (7) The authors pivot to a hippocampal-cortical ripple coupling analysis to build the argument that the hippocampal ripples shown in Figure 2 support memory maintenance in the cortex. They use a window of -500 to 500 ms from hippocampal ripples to assess coupling. This is quite wide, since it doesn't seem plausible that a cortical ripple 500 ms from a hippocampal ripple means they synchronize. They cite two papers to justify the analysis, both of which use {plus minus}500 ms windows, but for spindle-ripple coupling, not ripple-ripple, so are miscited. Later in the paper, they switch to {plus minus}50 ms for another coupling analysis, raising the question of why they used {plus minus}500 ms in the previous analysis to begin with. If they want to claim cortical ripples are tuned by hippocampal ripples all the way up to 500 ms away, they should show the rasters (as in Figure 4a) and timecourse ripple rates, but going beyond {plus minus}500 ms to show that ripples in the {plus minus}50-500 ms range are above, say 500-1000 ms to justify their window selection. I will point out that there IS previous work that used {plus minus}500 ms to measure cortical-cortical ripple coupling (Dickey et al 2022 PNAS, which should be cited regardless, as I believe the first hippocampal-cortical ripple paper showing memory effects), although the figures in that paper suggest anything beyond {plus minus}250 ms returns to baseline (see Figure 2A-B).

      (8) Lines 239 to 243 comparing p-values instead of an interaction test.

      (9) I don't understand what "Further analysis based on the identified cluster" means (line 271). I see in Figure 5c that their broadband classifier identified a window of optimal decoding, but did they use only activity in this cluster to train the subsequent classifier (Figure 5d)? If so, this is not described in the methods. And if it is done that way, I don't think the logic makes sense. As mentioned in comment 6, the ripples during encoding tune to 0-1s after image presentation. So it doesn't make sense to use a 1.85-2.25 s window for ripple-locked decoding-they should just be using the 0-1 s window (or whatever their cluster-based permutation test shows in Figure 2b). Otherwise, it would appear they are studying two different phenomena.

      (10) As is, the results in Figure 5d need to be redone. First, the results described on lines 271-275 once again suffer from comparing p-values. They need to run an interaction model if they want to claim Maintenance shows stronger ripple-locked decoding than Encoding (it almost certainly will not, since Encoding appears to show some evidence of decoding (p=0.118)). Second, even if they do change the framing to say Encoding and Maintenance show significant decoding, is it meaningful if Retrieval fails to? If you cannot decode the same information at the time of retrieval as is theoretically being held in working memory during the delay, the coupled ripple reactivation story wouldn't appear to make sense. They do show significant Retrieval decoding in Figure 5a-b, but since I don't really understand how they settled on the "identified cluster" in Figure 5c, I'm not sure what to make of the difference between these decoders.

      (11) Finally, as mentioned in the summary, the analyses in Figures 2-3 seem disjointed from those in Figures 4-5. Part of this has to do with the switch to a broadband classifier, then a switch back to coupled ripples, and then, as I already mentioned, decoding results with time windows that don't align with the hippocampal ripple effects they showed earlier. Further, since the main point of Figures 2-3 is to establish a ramp in hippocampal ripples across maintenance, shouldn't they be trying to show how the decoding changes over the course of the Maintenance period? It would also help the interpretation of Figure 5 to see how the coupled ripples change over time in Figure 4 (as they showed them in Figure 2).

      Minor issues:

      (1) Instead of citing a software package like Emmeans, the statistical test being performed should be explained.

      (2) Decoding % accuracy in the heatmaps in Figure 5 and supplementary would be more intuitive, particularly since Figure 5b uses accuracy anyway.

      (3) Figure 2b is misleading with an unnecessary change in the y-axis for retrieval.

      (4) In Figure 2d, a significant cluster is mentioned, but not drawn onto the figure as in Figure 2b.

    1. eLife Assessment

      This valuable study combines sub-millimeter 7T fMRI, EEG, representational similarity analysis, and deep neural network modeling to investigate layer-specific spatiotemporal dynamics underlying human object processing in early visual cortex and lateral occipital cortex; the authors report temporally distinct signatures in superficial layers of LOC that are interpreted as reflecting sequential feedforward and feedback processing during visual recognition. The multimodal methodological approach and empirical dataset are substantial and will be of broad interest to researchers in visual neuroscience, layer-fMRI methodology, and computational vision. However, the evidence supporting the central interpretation of interareal feedback remains incomplete, as the observed dynamics could also be explained by alternative mechanisms such as within-area recurrent processing, and there are additional concerns regarding several methodological and modeling choices underlying claims about increasing representational complexity at later time points. Overall, the study provides solid evidence for layer- and time-specific neural dynamics during object processing, while the interpretation of these signals as feedback-related remains provisional.

    2. Reviewer #1 (Public review):

      Summary:

      This study combines representational similarity analysis (RSA) with 7T layer-specific fMRI and EEG to examine how neural representations in specific cortical layers of EVC and LOC correspond to the temporal dynamics of visual processing. The authors interpret these correspondences as reflecting feedforward and feedback processes, based on their relative timing and their similarity to representations in different layers of a deep neural network (DNN).

      Strengths:

      The combination of RSA with laminar fMRI is a promising approach for dissociating the functional roles and dynamics of different cortical layers within the same functional region, and it holds considerable potential for elucidating computational mechanisms both within and between levels of the visual hierarchy. However, several issues should be addressed before the authors' conclusions can be fully supported.

      Weaknesses:

      (1) The authors report that the representation in the LOC superficial layer resembles EEG-derived neural representations at ~400 ms post-stimulus, and that this similarity is best explained by representations in the higher layers of the DNN. From these two observations, they conclude that activity in the LOC superficial layer is driven by feedback signals. However, neither line of evidence directly dissociates feedforward from feedback contributions.

      Specifically, late-stage representations in LOC could instead reflect the outcome of local recurrent computation, given that the superficial layer also serves as an output layer of the local cortical circuit. Moreover, the correlation with the DNN peaks at higher layers rather than being dominated by them, and feature tuning in higher DNN layers does not necessarily map onto higher-order cortical regions such as PFC.

      While a feedback contribution to the LOC superficial layer is consistent with theoretical predictions and known cortical anatomy, the current evidence is indirect. I would recommend that the authors either tone down this conclusion or, at a minimum, explicitly clarify the strength and limitations of the evidence in the Discussion.

      (2) I could not find information regarding the fMRI slice orientation or whether temporal regions beyond LOC were covered. The reported FOV (192 × 192 mm) seems quite large if only EVC and LOC were targeted. Did the authors acquire data from other object-selective regions in the temporal cortex, and if so, did they analyze these?

      It would strengthen the feedback interpretation considerably if the RDM of the LOC superficial layer could be shown to resemble RDMs from more anterior temporal regions, which would be consistent with feedback originating from higher-order object-processing areas.

      (3) Related to the previous point, LOC is a relatively large region, and based on the figures, it appears that the LOC ROI may contain two subregions. It would be helpful for the authors to show the location and extent of the LOC ROI in example participants.

      If the ROI does indeed span two subregions, do these subregions share the same laminar profile and temporal dynamics?

      (4) The authors report no feedback-related information in EVC, which contrasts with a number of prior fMRI studies that have demonstrated object-related feedback signals in EVC. One plausible explanation for this discrepancy is task relevance: in the present study, participants performed only a fixation color-change task, whereas in previous work they were required to attend to object features or identity (e.g., Morgan et al., 2019, J Neurosci; Kok et al., 2016, Curr Biol; Mohsenzadeh et al., 2018, eLife; Hou et al., 2026, eLife). Task demands on object processing may substantially modulate the strength of feedback signals to EVC, and this possibility warrants discussion.

      (5) A substantial body of work has used specialized paradigms to dissociate feedforward and feedback signals in EVC (e.g., Williams et al., 2008, Nat Neurosci; Fan et al., 2016, PNAS; Hou et al., 2026, eLife). These studies are directly relevant to the current work but are not cited.

      (6) Multidimensional scaling (MDS) visualizations of the RDMs (as in, e.g., Mohsenzadeh et al., 2018) are not included in the manuscript. These visualizations are important for interpreting the representational format across different layers of LOC and EVC, and I would encourage the authors to include them.

    3. Reviewer #2 (Public review):

      Summary:

      Carricarte and colleagues set out to identify and functionally characterize feedforward (FF) and feedback (FB) information flow during object perception in humans, a question that has been difficult to address non-invasively because FF and FB signals overlap rapidly in time and across regions. The authors capitalize on the canonical cortical microcircuit-FF terminations primarily in middle layers, FB terminations primarily in superficial and deep layers, to spatially separate these signals using sub-millimeter (0.9 mm isotropic) GE-BOLD fMRI at 7T in early visual cortex (EVC) and lateral occipital complex (LOC). They combine these layer-resolved fMRI patterns with millisecond-resolution EEG (from a previously published dataset using the same 24 images) via representational similarity analysis-based EEG-fMRI fusion, and use a Vision Transformer (DeiT) trained on ImageNet to characterize the feature complexity of the resulting spatiotemporal signatures.

      The authors first review their approach at the macroscale, replicating the expected EVC-then-LOC temporal hierarchy and the EVC-low/LOC-high feature complexity gradient. They then apply the same framework at the mesoscale of cortical layers, reporting: (a) early middle-layer signals in both EVC (~100 ms) and LOC (~160 ms) consistent with FF processing, (b) a later superficial-layer signal in LOC (~400 ms) interpreted as FB; (c) a layer-uniform feature-complexity profile in EVC (peaking at low-mid DNN layers across all depths); and (d) a feature-complexity dissociation in LOC, where middle-layer signals correspond to mid-to-high DNN layers and superficial-layer signals to high DNN layers. They argue that this complexity shift, combined with the timing difference, indicates interareal FB into LOC.

      Strengths:

      (1) The combination of layer-fMRI at 7T, EEG, and DNN-based representational analysis is well motivated through RSA. Each modality compensates for a known limitation of the others (fMRI: poor temporal resolution; EEG: poor spatial resolution; DNN: surrogate for representational format), and the RSA framework provides a principled common currency. Relatedly, the two-step macroscale-then-mesoscale design, in which the macroscale fusion replicates established findings before the same approach is applied at the layer level, is a sound and welcome scientific strategy that strengthens confidence in the combined-modality inferences.

      (2) The authors include multiple complementary controls: partialing out lower layers to mitigate vascular draining, voxel-count matching across layers, an alternative DNN (AlexNet), an alternative time-window definition based on between-layer differences, and time-resolved commonality analyses. The convergence across these analyses is reassuring.

      (3) Methodological transparency: The authors are forthright about partial-volume effects, foveal-confluence aggregation, and the indirect nature of the temporal estimates derived from EEG-fMRI fusion.

      Weaknesses:

      The central interpretive claim-that the late (~400 ms), superficial-layer LOC signal indexes interareal feedback that increases representational complexity-is intriguing, but in my view it is not yet fully supported by the evidence presented based on the following context.

      (1) Eye movements as a possible confound for late signals. Stimuli were presented for 1 second, and fixation was enforced only behaviorally via a color-change task on a central cross. No eye-tracking is reported for either the fMRI or EEG datasets. While this approach is not uncommon, the absence of gaze monitoring introduces ambiguity when the goal is to decouple feedforward and feedback contributions at fine temporal resolution in EEG recordings. Under these conditions, multiple image-driven saccades within a trial are plausible, and saccade patterns are likely to be systematically image-specific, given the small (n = 24) and heterogeneous naturalistic stimulus set. Critically, the temporal window over which RDM correlations are interpreted as feedback coincides with the period during which observers typically make 2-4 fixations (average fixation durations of ~250-330 ms; Rayner, 1998; Henderson, 2003), meaning the late EEG-fMRI fusion peaks fall in a window where image-locked saccadic activity and successive foveation-driven feedforward responses would be expected to accumulate. Late peaks could therefore reflect cumulative feedforward responses across successive foveations rather than top-down feedback. The manuscript would be strengthened by providing eye-tracking data (if available), control analyses leveraging post-hoc indicators, or a discussion citing prior evidence that EEG/fMRI response profiles in this paradigm are robust to such eye movements.

      (2) Decoding accuracy along the visual hierarchy raises questions about whether LOC is adequately engaged. Pairwise decoding accuracy is substantially higher in EVC than in LOC (Figure 1D), and the noise ceiling for LOC RDMs is markedly lower than for EVC across all layers (Supplementary Figure 4D-F). This pattern inverts the canonical hierarchical gradient of progressively stronger object decoding along the ventral visual stream, as well as the analogous gradient observed in DNN late layers that underlies the commonality analyses. As written, it is unclear how the manuscript reconciles this with its emphasis on LOC's role in higher-order, feedback-modulated representations with greater tolerance or increased complexity--unless decoding accuracies should be understood as image-level discrimination rather than at the level of object-category discrimination. A parsimonious alternative is that the 24-image set is too small or too coarse to reveal category-level representations in LOC robustly, such that LOC RDMs may be driven by lower-level or background/contextual variance and noise. This concern has direct bearing on the mesoscale commonality analyses supporting the "feedback transmits high-complexity features" conclusion. I would encourage the authors to (a) report split-half reliability of LOC RDMs alongside the commonality analyses, and either (b) acknowledge that the feature-complexity inferences are conditional on LOC RDMs faithfully capturing object structure rather than residual contextual/low-level variance, or (c) discuss how replication with a richer stimulus set might bear on the feedback-content interpretation.

      (3) The interareal feedback interpretation could be more robustly defended against intra-areal alternatives. In EVC, the authors carefully consider non-feedback explanations for layer-specific dynamics, including lateral connections modulating gain and superficial GE-BOLD bias, and conclude these are sufficient. The same skepticism is not extended to LOC, where the corresponding superficial-layer signal is interpreted as interareal feedback, with speculative sourcing to DLPFC. Slow (unmyelinated) horizontal/lateral propagation in superficial cortical layers (e.g., Davis et al., 2024) can, in principle, produce delayed superficial-layer signals on the timescale observed here without any interareal contribution. This asymmetry is compounded by the treatment of the absence of sustained EVC activity following the middle-layer peak, which is dismissed as a "limitation of the spatial and temporal sensitivity of our measurements" (lines 388-390). If feedback to EVC truly cannot be resolved with this method, the corresponding feedback claim in LOC-imaged with the same protocol warrants comparable caution. The manuscript would benefit from either presenting positive evidence that distinguishes interareal feedback from intra-areal recurrence (e.g., frequency-band signatures, source-resolved EEG, or coupling with frontal regions), or qualifying the conclusion to "delayed superficial-layer activity consistent with either interareal feedback or intra-areal recurrence."

      (4) The predictive coding framing is invoked but not well-grounded. The Discussion (lines 349-357) includes a theoretical implication of predictive coding. Predictive coding makes content-specific claims-feedback carries predictions, feedforward carries error signals relative to those predictions, and dissociating these requires manipulations of expectation, congruence, or predictability, none of which are present in the current design. The observed layer-wise timing differences do not bear evidence for rejecting non-predictive accounts. I would suggest either removing this framing or explicitly noting that the present data neither support nor refute predictive coding.

    4. Reviewer #3 (Public review):

      Summary.

      Carricarte and colleagues use 0.9mm 7T fMRI in EVC and LOC, fused with previously collected EEG using the same stimulus set, in order to dissect feedforward and feedback contributions to human object processing through their layer-specific termination patterns. They report a feedforward signal in middle layers of EVC (~100ms) and LOC (~160ms), and a later signal in superficial LOC (~400ms) that they interpret as interareal feedback. Using commonality analysis with a Vision Transformer, they argue that this late signal carries higher-complexity features than the earlier signal, and conclude that feedback actively increases representational complexity in LOC.

      Strengths.

      The empirical work is methodologically ambitious. Sub-millimeter 7T coverage of both EVC and LOC, combined with layer-resolved EEG-fMRI fusion, represents a substantial technical achievement. The authors first reproduce established macroscale EEG-fMRI fusion patterns at 7T before extending the approach to the layer level. The figures throughout are beautifully designed and convey complex analyses with clarity. The empirical core of the paper - that LOC contains layer-distinct dynamics at distinct times, with the late signal carrying representational structure that differs in some way from the early signal - is supported by the data, though with caveats imposed by the LOC noise ceiling.

      Weaknesses.

      The authors' interpretation of these data (interareal feedback that reflects feature-complexity, related to the functional role of these signals) is not adequately supported and requires either reframing or substantial additional evidence.

      Feedback vs. recurrence. The late superficial-LOC signal is interpreted as interareal feedback, but the data are equally consistent with within-area recurrence, lateral connections, or sustained feedforward dynamics. A reader expecting evidence of higher-area signals returning to early-time middle layers - a signature of interareal feedback - finds none in either region.

      "Functional role" overclaim. The paper repeatedly claims to characterize the "functional role" of feedforward and feedback, but contains no behavioral linkage, no perturbation, and no analysis relating signals to perceptual outcomes; the fMRI task is explicitly orthogonal to object processing. What is demonstrated is spatiotemporal dynamics and representational format - both valuable, neither equivalent to functional role.

      DNN analysis. The DNN analyses use several non-standard modeling choices that introduce more uncertainty than clarity. In the main analyses, the authors only use four sampling points from a single model (DeiT-small): transformer blocks 1, 7, and 12, plus the classification head. Then, the authors make their headline claims about complexity by comparing block 12 and the classification head; within the model, this is a distinction between an embedding layer and a supervised category readout, not a feature-complexity gradient. As such, the author's interpretation conflates semantic layers with representational "complexity." A more convincing use of this modeling strategy would be to demonstrate these effects in multiple models that might disentangle these factors-e.g., supervised (ResNet/ViT), self-supervised (DINOv2), and vision-language (CLIP) models-then to visualize these brain-model relationships across all layers. Alternatively, there are many suitable model-free analyses that could demonstrate the unique representational information within LOC without introducing any model-related concerns.

      Reliability of LOC layer-resolved RDMs. The lower-bound noise ceiling for LOC mesoscale RDMs is approximately 0.05 across layers, with deep-LOC reliability essentially at zero. The central layer-resolved dissociation rests on RDMs that individual subjects barely reproduce; consequently, the deep LOC layer is dropped from the commonality analysis (Figure 4C shows only middle/superficial layers, while Figure 4B shows all three for EVC) because the data cannot support it. This is not damning, but it is consequential, and not sufficiently addressed in the manuscript.

    1. eLife Assessment

      This foundational and valuable study expands our understanding of circadian clock work in non-model taxa in wider environmental niches, using solid methods for protein and RNA detection to describe the expression pattern of PDH, cry2, and per in the central nervous system of Euphausia superba. While the anatomical annotation is extensive, support for the identification of the clock network is incomplete.

    2. Reviewer #1 (Public review):

      Summary:

      Hüppe and colleagues characterized the network of neurons in the central nervous system of Antarctic krill that contained pigment-dispersing hormone (PDH), an important output factor in the circadian clock of insects. These neurons in the brain are putative clock neurons since a subset also expressed the clock genes period and cryptochrome 2. As one of the ocean's major contributors to biomass, krill is an ecologically important marine species that experiences challenging daily and seasonal environmental fluctuations in its high-latitude habitat. A comprehensive study of krill's internal clock may help to understand the extent of its resilience to the rapidly changing climate.

      The authors used antibody staining against PDH across the whole central nervous system and additional in situ hybridization for cry2 and per mRNA, with a focus on the supraesophageal ganglion. There, they identified the major neuropils in the eye stalks and central brain of Antarctic krill. The resulting staining pattern aligns with the identified circadian clock network in insects and PDH-expressing networks in other crustaceans, making these neurons highly likely candidates for krill clock neurons.

      Strengths:

      (1) This study provides the first clues about the circadian clock architecture in a non-model organism in chronobiology, Antarctic krill, with a clear 3D reconstruction of the putative clock network.

      (2) The authors effectively place their results within the extensive body of literature on arthropod circadian clock networks to argue that the neurons they describe are likely the circadian clock in krill.

      Weaknesses:

      (1) The data presented here are not sufficient to support the claim that the described network is the circadian clock because functional evidence is missing.

      (2) Additionally, the study falls short of identifying any elements of the positive limb of the canonical circadian clock transcriptional-translational feedback loop, e.g., clk or cyc, in the PDH-expressing neurons.

      (3) No sample sizes are reported, making it difficult for readers to assess the generalizability of the presented data.

    3. Reviewer #2 (Public review):

      Summary:

      This study advances our understanding of the neuronal basis of the circadian clock in pancrustaceans. It extends our knowledge on the pigment-dispersing hormone system and provides links to information on the expression of core clock components, cryptochrome 2, and period. The data are sound and well-documented.

      Comments:

      The neuronal components of the arthropod circadian clock system have been analysed extensively in insects. Much less information on this system is available on malacostraca crustacea crustaceans. However, considering that malacostracan crustaceans and insects go back to a common pancrustacean ancestor and considering that we know that the brain architecture in these two groups shares many commonalities (see, e. g., extensive reviews by N. J. Strausfeld), we have to expect that crustaceans and insects share many of the characteristics of the circadian system. This is the case, e. g., for the network of pigment-dispersing hormone-positive neurons. The authors cite these studies, although late in the paper (discussion, line 339ff), and I suggest to move this info into the introduction: "339 ff: The arborization pattern of the PDH-network has been described in various malacostracan crustaceans, including Carcinus maenas (Alexander et al., 2020; Mangerich & Keller, 1988; Mangerich et al., 1987), Cancer productus (Hsu et al., 2008), Orconectes limosus (de Kleijn et al., 1993; Mangerich & Keller, 1988; Mangerich et al., 1987), Homarus americanus (Harzsch etal., 2009), Cherax destructor, Procambarus clarkii (Sullivan et al., 2009), and Procambarus virginalis (Luna et al., 2010)."

      The strength of this paper is that it extends our knowledge on the PDH system and brings together neuroanatomical information on PDH-positive neurons with information on the expression of core clock components, cryptochrome 2, and period. That way, it advances our understanding of the neuronal basis of the circadian clock in pancrustaceans. The data are sound and well documented, and the authors are to be applauded for the superb dissection presented in Figure 1.

      Below, please find some essential suggestions on how to further improve the paper.

      (1) Framing of the study:

      I know that krill is a key element of the Southern Ocean's food webs, but my sense is that discussing the current findings in a context of resilience of this species to global ocean change means largely overselling this study:

      - Lines 47, 48: "and the resilience of this key species in a rapidly changing Southern Ocean."

      - Lines 70 ff: "Hence, understanding the mechanisms of adaptation, including biological clocks, is crucial for predicting how species, populations, and whole ecosystems will respond to climate change."

      - 154 ff: "The Southern Ocean environment experiences rapid change (Abram et al., 2025; Meredith et al., 2019; Thomalla et al., 2023). To assess krill's resilience to environmental changes, understanding the mechanisms that govern daily and seasonal timing in krill is essential."

      - 325 ff: "The rhythmic adaptation of krill to its high-latitude environment is key to its success in the Southern Ocean, which in turn represents a cornerstone for the well-being of the whole krill centred ecosystem. To predict krill's resilience to rapid environmental changes, it is essential to understand the mechanisms that govern daily and seasonal timing in krill."

      - 597 ff: "A detailed mechanistic understanding of the flexibility of clock-based processes is therefore essential to predict krill resilience in a changing Southern Ocean."

      My understanding is that duration of day length is one of the most predictable environmental drivers, and - despite the seasonal changes of day length - nevertheless a very stable one compared to fluctuations of environmental drivers such as temperature or salinity (see, e.g. this recent review on environmental driver fluctuations on nervous system functioning in crustaceans: Stein W, Harzsch S (2021) The Neurobiology of Ocean Change - insights from decapod crustaceans. Zoology: 125887. https://www.sciencedirect.com/science/article/pii/S094420062030146X).

      I do not see how global ocean change may significantly change day length, and what this study has to do with understanding this species' resilience against ocean change. I suggest that you explain in more detail why the light day length will change in the future or strongly tone this aspect. Statements such as Line 76 ff: "Due to their disproportionate importance for ecosystem function, understanding the resilience of ecological key species is essential in assessing the fate of ecosystems in the future." are completely out of focus here and, again, trying to oversell the current study.

      (2) Uncited essential studies of crustacean neuroanatomy, missing connection to contemporary crustacean neurobiology:

      - Line 157: "despite the ecological importance of E. superba, only very little is known about its neurobiology".

      - Line 329: "However, so far, little was known about the neurobiology of krill in general."

      I agree that this species' brain is understudied, but this makes it even more important to cite the little information that IS available. Please consider this essential reading for any crustacean neurobiologist: "Sandeman, D.C., Scholtz, G., Sandeman, R.E., 1993. Brain evolution in decapod crustacea. J. Exp. Zool. 265, 112-133." to find information on the basic brain anatomy in E. superba.

      The manuscript in many places seems to reinvent the wheel and raises the impression that our knowledge of crustacean brain morphology is close to zero. The authors in places seem to operate in a vacuum, and I find it disturbing that in a study on the crustacean brain, very few references are provided to studies on crustacean brain anatomy, such as the following essential book chapter: "Schmidt, M., 2016. Malacostraca. In: Schmidt-Rhaesa, A., Harzsch, S., Purschke, G. (Eds.), Structure & Evolution of Invertebrate Nervous Systems. Oxford University Press, Oxford, pp. 529-582. https://www.researchgate.net/publication/315366157"

      In terms of brain anatomy, I would like to know if the authors have a hypothesis on whether and how their target species' brain structure may be similar or different to the brains of other "shrimps" as described, e. g., in the following studies. If so, please elaborate in the introduction:

      Krieger J, Hörnig MK, Sandeman RE, Sandeman DC, Harzsch S (2020), Masters of communication: The brain of the banded cleaner shrimp Stenopus hispidus (Olivier, 1811) with an emphasis on sensory processing areas. Journal of Comparative Neurology 528(9): 1561-1587.

      Meth R, Wittfoth C, Harzsch S (2017) Brain architecture of the Pacific White Shrimp Penaeus vannamei Boone, 1931 (Malacostraca, Dendrobranchiata): correspondence of brain structure and sensory input? Cell and Tissue Research 369(2): 255-271.

      (3) Lacking rigor and command of crustacean brain nomenclature

      I suggest that for their brain nomenclature, the authors should rigorously stick to that laid out by Sandeman et al. 1992 (not yet cited in the ms): Sandeman, D.C., Sandeman, R.E., Derby, C.D., Schmidt, M., 1992. Morphology of the brain of crayfish, crabs, and spiny lobsters: a common nomenclature for homologous structures. Biol. Bull. 183, 304-326.

      More specifically, in lines 41, 163, 199, 204, 207, and throughout the paper, the authors use the terms "Optic lobes" or "optic lobe neuropils". To the best of my knowledge, "optic lobe" is not a term used in crustacean neuroanatomy at all (as opposed to insects). Lamina, medulla, and lobula are collectively referred to as "visual neuropils" (see Krieger, J., Hörnig, M. K., Sandeman, R. E., Sandeman, D. C., & Harzsch, S. (2020). Masters of communication: The brain of the banded cleaner shrimp Stenopus hispidus (Olivier, 1811) with an emphasis on sensory processing areas. Journal of Comparative Neurology, 528(9), 1561-1587. https://doi.org/10.1002/CNE.24831). The medulla terminalis and mushroom bodies are referred to as "lateral protocerebrum". All afore-mentioned neuropils are summarized as "eyestalk neuropils" (compare nomenclature in Schmidt 2016 as referenced above).

      Line 170, 172, 175 ff, and Figure 1. "abdomen", "abdominal ganglia": Contra the book chapter by Siegel 2016 "Introducing Antarctic Krill Euphausia superba Dana, 1850", his Fig. 1.2, the "tail" of crustaceans in most books on crustacean anatomy is not called "abdomen" but instead "pleon"; hence the name "pleopods" for the appendages of the pleon (instead of "abdomipods"). What is more, I suggest using the terms "pleon ganglia" instead of "abdominal ganglia", following the terminology suggested in "Harzsch S, Sandeman D, Chaigneau J (2012) Morphology and development of the central nervous system. In: Forest J and von Vaupel Klein JC (Eds.). Treatise on Zoology - Anatomy, Taxonomy, Biology. The Crustacea Vol. 3. Brill, Leiden pp. 9-236."

      Line 174: "thoracic ganglia". In Figure 1, there is a labelling mistake as these ganglia are named "thoracaic ganglia".

      Line 176, and throughout the paper: "supraesophageal ganglion". Following the standard nomenclature for crustaceans (see, e. g., Schmidt, M., 2016. Malacostraca. In: Schmidt-Rhaesa, A., Harzsch, S., Purschke, G. (Eds.), Structure & Evolution of Invertebrate Nervous Systems. Oxford University Press, Oxford, pp. 529-582. https://www.researchgate.net/publication/315366157", this structure (as in insects) is typically called a "brain". For terminology, also consult the following nomenclature paper: "Richter, S., Loesel, R., Purschke, G., Schmidt-Rhaesa, A., Scholtz, G., Stach, T., Vogt, L., Wanninger, A., Brenneis, G., Döring, C., Faller, S., Fritsch, M., Grobe, P., Heuer, C. M., Kaul, S., Møller, O. S., Müller, C. H. G., Rieger, V., Rothe, B. H., Stegner, M., Harzsch, S. (2010). Invertebrate neurophylogeny: Suggested terms and definitions for a neuroanatomical glossary. Frontiers in Zoology, 7. https://doi.org/10.1186/1742-9994-7-29".

      Line 212, and throughout the paper - hemielliposoid body: please refer to Harzsch Krieger 2011 and the numerous references to studies by Strausfeld cited therein in crustaceans. Strausfeld has provided compelling evidence that the crustacean hemiellipsoid body is equivalent to the insect mushroom body, so this term should be replaced. Harzsch, S., & Krieger, J. (2021). Genealogical relationships of mushroom bodies, hemiellipsoid bodies, and their afferent pathways in the brains of Pancrustacea: Recent progress and open questions. Arthropod Structure & Development, 65, 101100. HYPERLINK "https://doi.org/10.1016/J.ASD.2021.101100" https://doi.org/10.1016/J.ASD.2021.101100.

      Legend, figure 2, and others, and throughout the paper: "The olfactory neuropiles comprise the lateral antennal neuropile (LAN, ochre), the olfactory lobes (OL, yellow), and the antennal neuropile (AnN, green)." This is a strange terminological mix that you should urgently revise according to the standard terminology by Sandeman et al. 1992 (as referenced above). The LAN is the lateral antenna 1 neuropil. The AnN is the antenna 2 neuropil. The AnN is NOT deutocerebral but tritocerebral.

    4. Author response:

      Reviewer #1 (Public review):

      Summary:

      Hüppe and colleagues characterized the network of neurons in the central nervous system of Antarctic krill that contained pigment-dispersing hormone (PDH), an important output factor in the circadian clock of insects. These neurons in the brain are putative clock neurons since a subset also expressed the clock genes period and cryptochrome 2. As one of the ocean's major contributors to biomass, krill is an ecologically important marine species that experiences challenging daily and seasonal environmental fluctuations in its high-latitude habitat. A comprehensive study of krill's internal clock may help to understand the extent of its resilience to the rapidly changing climate.

      The authors used antibody staining against PDH across the whole central nervous system and additional in situ hybridization for cry2 and per mRNA, with a focus on the supraesophageal ganglion. There, they identified the major neuropils in the eye stalks and central brain of Antarctic krill. The resulting staining pattern aligns with the identified circadian clock network in insects and PDH-expressing networks in other crustaceans, making these neurons highly likely candidates for krill clock neurons.

      Strengths:

      (1) This study provides the first clues about the circadian clock architecture in a non-model organism in chronobiology, Antarctic krill, with a clear 3D reconstruction of the putative clock network.

      (2) The authors effectively place their results within the extensive body of literature on arthropod circadian clock networks to argue that the neurons they describe are likely the circadian clock in krill.

      Weaknesses:  

      (1) The data presented here are not sufficient to support the claim that the described network is the circadian clock because functional evidence is missing.

      (2) Additionally, the study falls short of identifying any elements of the positive limb of the canonical circadian clock transcriptional-translational feedback loop, e.g., clk or cyc, in the PDH-expressing neurons.

      (3) No sample sizes are reported, making it difficult for readers to assess the generalizability of the presented data.

      We thank the reviewer for recognizing the contribution of this study to advancing our understanding of clock systems in non-traditional model organisms. We acknowledge that definitive functional evidence would require the generation of null mutants of core clock components, which is currently not feasible in this species. In a revised version, we will adjust our claims to more precisely reflect the evidence presented and include sample sizes to allow the reader to better assess the representativeness of the results.

      Reviewer #2 (Public review):

      Summary:

      This study advances our understanding of the neuronal basis of the circadian clock in pancrustaceans. It extends our knowledge on the pigment-dispersing hormone system and provides links to information on the expression of core clock components, cryptochrome 2, and period. The data are sound and well-documented.

      Comments:

      The neuronal components of the arthropod circadian clock system have been analysed extensively in insects. Much less information on this system is available on malacostraca crustacea crustaceans. However, considering that malacostracan crustaceans and insects go back to a common pancrustacean ancestor and considering that we know that the brain architecture in these two groups shares many commonalities (see, e. g., extensive reviews by N. J. Strausfeld), we have to expect that crustaceans and insects share many of the characteristics of the circadian system. This is the case, e. g., for the network of pigment-dispersing hormone-positive neurons. The authors cite these studies, although late in the paper (discussion, line 339ff), and I suggest to move this info into the introduction: "339 ff: The arborization pattern of the PDH-network has been described in various malacostracan crustaceans, including Carcinus maenas (Alexander et al., 2020; Mangerich & Keller, 1988; Mangerich et al., 1987), Cancer productus (Hsu et al., 2008), Orconectes limosus (de Kleijn et al., 1993; Mangerich & Keller, 1988; Mangerich et al., 1987), Homarus americanus (Harzsch etal., 2009), Cherax destructor, Procambarus clarkii (Sullivan et al., 2009), and Procambarus virginalis (Luna et al., 2010)."

      The strength of this paper is that it extends our knowledge on the PDH system and brings together neuroanatomical information on PDH-positive neurons with information on the expression of core clock components, cryptochrome 2, and period. That way, it advances our understanding of the neuronal basis of the circadian clock in pancrustaceans. The data are sound and well documented, and the authors are to be applauded for the superb dissection presented in Figure 1.

      Below, please find some essential suggestions on how to further improve the paper.

      (1) Framing of the study:

      I know that krill is a key element of the Southern Ocean's food webs, but my sense is that discussing the current findings in a context of resilience of this species to global ocean change means largely overselling this study:

      Lines 47, 48: "and the resilience of this key species in a rapidly changing Southern Ocean."

      Lines 70 ff: "Hence, understanding the mechanisms of adaptation, including biological clocks, is crucial for predicting how species, populations, and whole ecosystems will respond to climate change."

      154 ff: "The Southern Ocean environment experiences rapid change (Abram et al., 2025; Meredith et al., 2019; Thomalla et al., 2023). To assess krill's resilience to environmental changes, understanding the mechanisms that govern daily and seasonal timing in krill is essential."

      325 ff: "The rhythmic adaptation of krill to its high-latitude environment is key to its success in the Southern Ocean, which in turn represents a cornerstone for the well-being of the whole krill centred ecosystem. To predict krill's resilience to rapid environmental changes, it is essential to understand the mechanisms that govern daily and seasonal timing in krill."

      597 ff: "A detailed mechanistic understanding of the flexibility of clock-based processes is therefore essential to predict krill resilience in a changing Southern Ocean."

      My understanding is that duration of day length is one of the most predictable environmental drivers, and - despite the seasonal changes of day length - nevertheless a very stable one compared to fluctuations of environmental drivers such as temperature or salinity (see, e.g. this recent review on environmental driver fluctuations on nervous system functioning in crustaceans: Stein W, Harzsch S (2021) The Neurobiology of Ocean Change - insights from decapod crustaceans. Zoology: 125887. https://www.sciencedirect.com/science/article/pii/S094420062030146X).

      I do not see how global ocean change may significantly change day length, and what this study has to do with understanding this species' resilience against ocean change. I suggest that you explain in more detail why the light day length will change in the future or strongly tone this aspect. Statements such as Line 76 ff: "Due to their disproportionate importance for ecosystem function, understanding the resilience of ecological key species is essential in assessing the fate of ecosystems in the future." are completely out of focus here and, again, trying to oversell the current study.

      (2) Uncited essential studies of crustacean neuroanatomy, missing connection to contemporary crustacean neurobiology:

      Line 157: "despite the ecological importance of E. superba, only very little is known about its neurobiology".

      Line 329: "However, so far, little was known about the neurobiology of krill in general."

      I agree that this species' brain is understudied, but this makes it even more important to cite the little information that IS available. Please consider this essential reading for any crustacean neurobiologist: "Sandeman, D.C., Scholtz, G., Sandeman, R.E., 1993. Brain evolution in decapod crustacea. J. Exp. Zool. 265, 112-133." to find information on the basic brain anatomy in E. superba.

      The manuscript in many places seems to reinvent the wheel and raises the impression that our knowledge of crustacean brain morphology is close to zero. The authors in places seem to operate in a vacuum, and I find it disturbing that in a study on the crustacean brain, very few references are provided to studies on crustacean brain anatomy, such as the following essential book chapter: "Schmidt, M., 2016. Malacostraca. In: Schmidt-Rhaesa, A., Harzsch, S., Purschke, G. (Eds.), Structure & Evolution of Invertebrate Nervous Systems. Oxford University Press, Oxford, pp. 529-582. https://www.researchgate.net/publication/315366157"

      In terms of brain anatomy, I would like to know if the authors have a hypothesis on whether and how their target species' brain structure may be similar or different to the brains of other "shrimps" as described, e. g., in the following studies. If so, please elaborate in the introduction:

      Krieger J, Hörnig MK, Sandeman RE, Sandeman DC, Harzsch S (2020), Masters of communication: The brain of the banded cleaner shrimp Stenopus hispidus (Olivier, 1811) with an emphasis on sensory processing areas. Journal of Comparative Neurology 528(9): 1561-1587.

      Meth R, Wittfoth C, Harzsch S (2017) Brain architecture of the Pacific White Shrimp Penaeus vannamei Boone, 1931 (Malacostraca, Dendrobranchiata): correspondence of brain structure and sensory input? Cell and Tissue Research 369(2): 255-271.

      (3) Lacking rigor and command of crustacean brain nomenclature

      I suggest that for their brain nomenclature, the authors should rigorously stick to that laid out by Sandeman et al. 1992 (not yet cited in the ms): Sandeman, D.C., Sandeman, R.E., Derby, C.D., Schmidt, M., 1992. Morphology of the brain of crayfish, crabs, and spiny lobsters: a common nomenclature for homologous structures. Biol. Bull. 183, 304-326.

      More specifically, in lines 41, 163, 199, 204, 207, and throughout the paper, the authors use the terms "Optic lobes" or "optic lobe neuropils". To the best of my knowledge, "optic lobe" is not a term used in crustacean neuroanatomy at all (as opposed to insects). Lamina, medulla, and lobula are collectively referred to as "visual neuropils" (see Krieger, J., Hörnig, M. K., Sandeman, R. E., Sandeman, D. C., & Harzsch, S. (2020). Masters of communication: The brain of the banded cleaner shrimp Stenopus hispidus (Olivier, 1811) with an emphasis on sensory processing areas. Journal of Comparative Neurology, 528(9), 1561-1587. https://doi.org/10.1002/CNE.24831). The medulla terminalis and mushroom bodies are referred to as "lateral protocerebrum". All afore-mentioned neuropils are summarized as "eyestalk neuropils" (compare nomenclature in Schmidt 2016 as referenced above).

      Line 170, 172, 175 ff, and Figure 1. "abdomen", "abdominal ganglia": Contra the book chapter by Siegel 2016 "Introducing Antarctic Krill Euphausia superba Dana, 1850", his Fig. 1.2, the "tail" of crustaceans in most books on crustacean anatomy is not called "abdomen" but instead "pleon"; hence the name "pleopods" for the appendages of the pleon (instead of "abdomipods"). What is more, I suggest using the terms "pleon ganglia" instead of "abdominal ganglia", following the terminology suggested in "Harzsch S, Sandeman D, Chaigneau J (2012) Morphology and development of the central nervous system. In: Forest J and von Vaupel Klein JC (Eds.). Treatise on Zoology - Anatomy, Taxonomy, Biology. The Crustacea Vol. 3. Brill, Leiden pp. 9-236."

      Line 174: "thoracic ganglia". In Figure 1, there is a labelling mistake as these ganglia are named "thoracaic ganglia".

      Line 176, and throughout the paper: "supraesophageal ganglion". Following the standard nomenclature for crustaceans (see, e. g., Schmidt, M., 2016. Malacostraca. In: Schmidt-Rhaesa, A., Harzsch, S., Purschke, G. (Eds.), Structure & Evolution of Invertebrate Nervous Systems. Oxford University Press, Oxford, pp. 529-582. https://www.researchgate.net/publication/315366157", this structure (as in insects) is typically called a "brain". For terminology, also consult the following nomenclature paper: "Richter, S., Loesel, R., Purschke, G., Schmidt-Rhaesa, A., Scholtz, G., Stach, T., Vogt, L., Wanninger, A., Brenneis, G., Döring, C., Faller, S., Fritsch, M., Grobe, P., Heuer, C. M., Kaul, S., Møller, O. S., Müller, C. H. G., Rieger, V., Rothe, B. H., Stegner, M., Harzsch, S. (2010). Invertebrate neurophylogeny: Suggested terms and definitions for a neuroanatomical glossary. Frontiers in Zoology, 7. https://doi.org/10.1186/1742-9994-7-29".

      Line 212, and throughout the paper - hemielliposoid body: please refer to Harzsch Krieger 2011 and the numerous references to studies by Strausfeld cited therein in crustaceans. Strausfeld has provided compelling evidence that the crustacean hemiellipsoid body is equivalent to the insect mushroom body, so this term should be replaced. Harzsch, S., & Krieger, J. (2021). Genealogical relationships of mushroom bodies, hemiellipsoid bodies, and their afferent pathways in the brains of Pancrustacea: Recent progress and open questions. Arthropod Structure & Development, 65, 101100. HYPERLINK "https://doi.org/10.1016/J.ASD.2021.101100" https://doi.org/10.1016/J.ASD.2021.101100.

      Legend, figure 2, and others, and throughout the paper: "The olfactory neuropiles comprise the lateral antennal neuropile (LAN, ochre), the olfactory lobes (OL, yellow), and the antennal neuropile (AnN, green)." This is a strange terminological mix that you should urgently revise according to the standard terminology by Sandeman et al. 1992 (as referenced above). The LAN is the lateral antenna 1 neuropil. The AnN is the antenna 2 neuropil. The AnN is NOT deutocerebral but tritocerebral.  

      We thank the reviewer for acknowledging this paper's contribution to our understanding of the neuronal basis of the circadian clock in Pancrustaceans, as well as for the positive evaluation of the data documentation and presentation.

      We would like to clarify that we are aware of the existing body of literature on crustacean neuroanatomy and did not intend to present our data as a first in this field. This study intersects multiple communities (e.g., chronobiology, crustacean neurobiology, krill ecology), and the current focus of the manuscript arose from an attempt to make the paper as accessible to these communities as possible. We acknowledge, however, that the current version falls short in its engagement with the existing literature on crustacean brain anatomy. We therefore thank the reviewer for the input on crustacean neuroanatomy and its nomenclature, which will help us improve the manuscript in these respects. In a revised version, we plan to adjust the framing of the study to more precisely reflect the data presented. This will include better situating the present findings within the existing literature on crustacean neuroanatomy and its specific nomenclature, while toning down the emphasis on ecological importance and implications.

      Reviewer #3 (Public review):

      Summary:  

      A solid and very descriptive study of gene expression of three factors in krill, PDH, per, and cry2 that are important for circadian rhythms in insects. The results reveal optic areas in which PDH colocalises with each or per and cry2, and central brain areas where it does not. The authors speculate on the functional implications of their results for biological rhythms.  

      Comments:

      This manuscript describes a detailed anatomical study of the brain of krill in a circadian gene expression context. The results are well described, and the work is well done considering the obvious technical/practical difficulties of working with this species. Having stated that, the authors in their Methods write that the animals, after being caught, were placed in constant darkness. Is there any idea at all of when in ZT these brains were processed? Are the representations of gene expression taken at random around the clock? Perhaps the authors might make this explicit somewhere in the ms as it is an important point.

      The manuscript focuses mostly on PDH and its overlap or not with per or cry2. I found Figures 5 and 6 particularly confusing. The panels show PDH colocalising (or not-filled or unfilled arrows) with cry2 or with per. What they do not show (to me) is that per and cry2 colocalise. Now, of course, they probably do, but Figure 5 does not show this - or am I misinterpreting it? In Figure 6 again, I cannot see any panels with per and cry2 overlaid. Seems different sections were used for each probe? Is that what 'Areas with high per/ cry2-expression are marked by white arrowheads' means? I see that lines 493 and 494 confirm my suspicions that per/cry were not shown to be colocalised. Perhaps the authors could make this clearer up front than halfway through the Discussion, and clarify this in their legends, which are a little misleading in this respect?

      We thank the reviewer for his positive evaluation of our work, acknowledging the difficulty when working with this organism, and for the constructive comments. In a revised version of the manuscript, we will clarify the sampling time in the Methods. We will also state upfront — and in the figure legends — that per and cry2 were assessed on separate sections and their direct co-localization was therefore not demonstrated. However, as both components were independently shown to co-localize with PDH, their spatial overlap is nevertheless suggested by the shared co-localization with PDH. We will make this reasoning explicit earlier in the manuscript to avoid any misleading implications.

    1. eLife Assessment

      This important study reveals distinct representations of task-related information in the dendrites and somata of cortical neurons during sensorimotor learning and behavioral adaptation. The evidence is compelling, combining simultaneous imaging of dendritic and somatic activity during behavior to demonstrate compartment-specific encoding of sensory cues, motor actions, and corrective signals. The work will be of broad interest to neuroscientists studying dendritic computation, motor learning, and the cellular mechanisms underlying adaptive behavior.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Scheib et al. identify distinct calcium dynamics in the somata and tuft dendrites of layer 5 pyramidal cells in mice performing a licking task. Animals are trained to lick water ports on the left or right following an acoustic cue, and can adjust their targeting when the ports are displaced. For tongue premotor cortical neurons projecting to the ventromedial thalamus, calcium transients in tuft dendrites are tightly locked to the direction-instructive cue, while somatic calcium signals are more broadly dispersed and more frequently synchronized with tongue motion and port contact. Finally, when the targets are shifted, tufts exhibit a sparse but large corrective signal on an improperly-targeted first lick, and the changes in population activity in the tufts and somata differ after adaptation to the new port locations.

      Strengths:

      In my opinion, this is a very strong manuscript which reports several novel and significant observations, contains high-quality data and (for the most part) reasonable analyses, and is clear and well-written. Most prior studies of cortical sensorimotor processing have measured the output of neurons using extracellular recording - an approach which obscures potentially important signaling differences between neuronal compartments. This study leverages cutting-edge imaging techniques in mice to document large, time-dependent differences between calcium signals at cortical somata and tuft dendrites. This phenomenon could have major implications at the cellular level for synaptic plasticity, and at the systems and behavioral levels for motor adaptation. As described below, I have only one major technical concern (which should be addressable with additional analysis), along with several relatively minor suggestions for improving the manuscript.

      Weaknesses:

      At a conceptual level, the authors may wish to elaborate a bit on what sensorimotor computation they think the circuit is implementing, and how their results help explain this implementation. Several possibilities are raised: tuft activation could "prime" the pyramidal cells in advance of movement initiation (line 319ff), or could track errors to engage plasticity (line 351ff) and solve the credit assignment problem (line 362ff). It might be helpful to make one of these proposals more concrete with a computational model, but this is not strictly necessary.

      My only major technical concern relates to the analyses in Figures 4F-H, 5G-I, and 6H-K (c.f. equations 2-5). Typically, one identifies population-level factors by projecting neural activity onto fixed dimensions of interest; this makes it possible to see how activity evolves over time along interpretable coordinates. Here, however, the coding directions are redefined at each time point, so the "choice" activity at time t is actually a different signal from the "choice" activity at t+1. This procedure is a bit like comparing the activity of one neuron at one time point with the activity of a different neuron at a later time point. It also makes the physiological interpretation more complicated: if the dimensions are fixed, one can see how a downstream neuron could "read out" the signal by computing a weighted sum of the activity of upstream neurons, but it is harder to see how this could happen if the weights are always rotating.

      A few comments on the behavioral task and results. After the port shift, the error rate is quite high, and doesn't diminish much between the early and late epochs (approximately 42% and 38% error rate, respectively; Figure 1I). That is, mice do not seem to fully master the task. Clearly, animals do alter their aim, but even this does not seem to change much between early and late periods (Figure 1J). I recommend that the authors show the behavioral data at a finer level of granularity (e.g., by plotting the change in exit trajectory on all individual trials across sessions, with a loess fit) to allow an assessment of the adaptation rate and when adaptation saturates. It would also be more conventional to refer to the behavioral changes as "motor adaptation," instead of "skill learning." (The latter would be appropriate if the port offset were randomized across trials, and animals received two separate cues for direction and offset, but I suspect this task would be too difficult for mice to learn.)

      This is perhaps a semantic point, but it might not be entirely accurate to refer to the activity evoked by the directional cue as "sensory." Typically, a "sensory" response should encode some feature of a stimulus - in this case, the frequency of a tone. Here, it seems likely that the cue-aligned activity reflects the instructed lick direction, rather than the auditory information per se. (Presumably, these premotor neurons do not have well-behaved auditory tuning curves.) By comparison, in macaques performing center-out reach tasks, activity in dorsal premotor cortex rapidly ramps up following a visual cue instructing the direction of an upcoming reach, but one usually wouldn't refer to this activity as "visual" or "sensory" (though this is sometimes done). I suggest the authors either use "Instruction" or similar (e.g., in Figure 4F), or clarify in the text whether they think the activity is a genuine auditory response or something else.

    3. Reviewer #2 (Public review):

      Summary:

      The authors set out to compare functional encoding in the tuft dendrites and somata of a specific cortical cell type during motor planning and learning.

      Strengths:

      The investigation of a specific projection type (L5 ET) is a strength that aids reproducibility and interpretation. The elegant approach to increasing the depth of field of dendritic imaging is another strength. The data analyses are largely clear in their methods, scope, and interpretation. The writing is extremely clear and appropriately referenced, with an excellent Introduction, in particular.

      Weaknesses:

      It is not obvious whether the selected labeling strategy avoids labeling Layer 6 CT neurons, which would contaminate dendritic recordings. The images provided suggest enrichment in L5, but a discussion of this important potential caveat is warranted, especially since within-cell comparisons of apical dendrites to somata were not performed.

      The application of DeepInterpolation to dendritic data appears to be novel, and little detail or vetting is provided. The reader is left guessing: Was the model retrained or fine-tuned on dendritic data? How does the denoising affect the resulting segmentation and activity traces? Is denoising necessary for this workflow?

      The activity patterns of the recorded cells appear to lack the characteristic ramping during the delay epoch previously reported in both calcium imaging and electrophysiology studies. Given that a major contribution to the significance of the work is to constrain models of ALM function, a discussion of how the data aligns with previous measurements in the same circuit would improve the work.

      It would be very informative to compare differences in signals between dendrites and somata of the same cells. Consistently tracing dendrites to their respective somata would assuage worries of potential contamination from dendrites of deeper cells and enable more direct comparisons of signal transformations between dendrites and somata. It would be good to understand the relationship between dendritic calcium signals and backpropagating action potentials in this task. The authors detect less frequent calcium events in tufts versus somata; is this due to selective backpropagation of action potentials? The dynamics of this process were recently investigated by Adam Cohen's group in vivo and in vitro, and measurements in the present settings could be compared to such work.

      The Coding Direction analyses presented in this work, while consistent with previous literature on population codes in ALM, are at odds with the nature of the measurements here. The changes in representation that occur between the dendrites and soma of an individual cell are probably best thought of in terms of the dynamics of signals themselves within individual neurons, rather than in the information encoded across a population.

      This work is largely observational, describing signals that might reflect computational transformations and/or instruct plasticity, but those possibilities have not yet been deeply investigated. The manuscript does a good job of laying out these as future directions.

    4. Reviewer #3 (Public review):

      Summary:

      This article by Scheib et al. investigates how layer 5 extratelencephalic (ET) neurons in the frontal cortex encode sensorimotor information during motor learning, focusing on differences between their apical tuft dendrites and somas. The authors alternated recordings among these ET neuronal compartments in the mouse anterior lateral motor cortex (ALM) during a cued directional licking task with a target port shift. They found that while tuft dendrites predominantly encode sensory cues, with a subset selectively active during corrective actions, somatic activity was more strongly associated with action timing. Additionally, learning induced divergent plasticity: tuft dendrites increased their selectivity but decreased response gain, maintaining stable net selectivity, whereas somas showed increased net selectivity early in learning. Together, these findings reveal distinct sensorimotor representations and learning-related plasticity in dendritic and somatic compartments, providing insight into how compartment-specific activity in the frontal cortex may contribute to motor skill acquisition.

      Strengths:

      The authors developed an innovative imaging approach and a comprehensive data analysis pipeline to address a knowledge gap in the literature. By alternating imaging of dendritic tufts and somas in the same animals, they compare compartment-specific activity during motor learning and identify distinct encoding of task variables and learning-related plasticity across these compartments. Interestingly, a subset of dendritic tufts shows activity associated with corrective actions. The findings are discussed in the context of current theories of dendritic computation, credit assignment, and motor learning, providing a useful foundation for future mechanistic studies.

      Weaknesses:

      No major weaknesses were identified.

    1. eLife Assessment

      This valuable study examines whether reduced cooperation is driven by betrayal aversion beyond nonsocial loss aversion, using matched social and nonsocial risky decision-making tasks combined with computational modeling and EEG. The authors provide solid empirical evidence that social risk is processed differently from matched nonsocial risk, offering a meaningful contribution to the study of cooperation and decision-making under uncertainty. However, further justification of the computational modeling approach would strengthen some of the conclusions. This work will be of interest to researchers studying social decision-making, cooperation, trust, and the neural and computational mechanisms underlying risk and betrayal aversion.

    2. Reviewer #1 (Public review):

      Summary:

      The non-social task was a classic risky decision-making task with a binary choice between an option with a sure gain and a risky option with a probabilistic gain or loss. In the social task, the sure option was an individual gain (as in the non-social option) and the probabilities in the risky option, which were shown to participants, were framed as probabilities of other previous participants (i.e., "partners") to cooperate or not; a probabilistic gain (when the partner cooperated) also led to a gain of the partner, while a probabilistic loss meant that the partner would receive the amount lost by the participant. This loss was framed as "betrayal." The authors show differences in how probabilities and amounts (of gains/losses) affected choices, RTs, and ERPs (P3 and LPP).

      Strengths:

      Since participants faced decisions with the same individual payoffs in a non-social and a social condition, this setup made it possible to use identical standard analyses for choices, RTs, and ERPS as well as (almost) identical economic models for the two conditions.

      Weaknesses:

      (1) The task does not include many components that are usually considered central for cooperation or "betrayal" and this is not discussed appropriately. At the same time, the "emotional aspects" of the operationalized "betrayal" are not directly assessed.

      a) The standard economic game for cooperation is the prisoner's dilemma, in which participants make independent choices at the same time without getting any explicit information on the cooperation probability of their partner before they make their decisions. Furthermore, most of the time the interactions are repeated. Actually, the trust game as one other frequently used economic game, also includes a back and forth of transfers between the partners. So, here, I am not so convinced by the operationalization of a low cooperation probability, which is shown before the decision, as "betrayal." The authors should motivate and explain their rationale more clearly in reference to such other tasks.

      b) The setup of the task, especially the fake interaction with the fake partners, should be made clearer in the main text (before reporting the results). I would argue for including the task picture in the main text.

      c) In general, I am in favour of taking participants' choice behaviour as the main outcome measure. But given the strong implications of "emotional costs" made by the authors, I would have expected some ratings of "betrayal" on a trial-by-trial basis. I would at least include this as a shortcoming.

      d) Also, given the framing of the study, I would have expected some exploratory analyses regarding individual differences with respect to, e.g., social value orientation, etc. I would at least include this as an outlook.

      (2) The standard statistical analyses could be improved.

      a) It is good that the authors have rather long sections using standard regression analyses. But they are a bit lengthy, and the modelling should be more prominent.

      b) In a couple of places, the authors say something like "this is significant, but that is not." Here, it has been made very clear that the interaction term needs to be looked at. As far as I can see, this has not always been done.

      c) For this binary choice, the difference in expected value (EV) between the sure and the risky options is one crucial comparison. But the authors never take that into account. This difference does not depend on the amount, which the authors dub "principal." That is, the sure option simply has an EV of x, i.e., the amount. The risky option has the EV = p2x + (1-p)0.5x, with p being the probability of gain/cooperation. That is, the two options have the same EV at p=1/3, independent of x. This should be made clear.

      d) Relatedly, RTs should depend on the differences in EV (and not so much on p or on x per se). This can be seen by the more or less quadratic relationship between p and RTs (Fig 1A), with a peak around a p of 1/3.

      e) RTs are often log-transformed. It should be briefly mentioned why this was not done here.

      (3) The modelling evidence is relatively weak. This is my main point.

      a) (Cumulative) prospect theory should be introduced.

      b) The models seem overly complicated with many free parameters. I would have expected some simpler versions and more comparisons between models that differ in just one parameter.

      - e.g., it is really nice that the authors used a probability weighting function. BTW: Please describe this more clearly in the introduction and in the results. But for this limited range of probabilities, this might be too much.

      - e.g., why directly assume two different exponents in the utility function for gains and losses, and in addition a loss aversion parameter lambda? Only lambda would be a better starting point here.

      c) The differences in AIC (Figure 2A) seem rather minuscule, and the distribution of winning models is not very peaked. I am not convinced that Model 3 is the winning model.

      d) Crucially, and related to the previous points, judging from Fig 2C, the "betrayal" parameter kappa seems to be zero for about half of the participants. The authors should look into this.

      - Would a model just like model 3 but without kappa (i.e., kappa set to zero) perform better? Is this just model 2?

      - How is kappa set in the non-social condition?

      - This massive skew, to say the least, is never discussed.

      - A correlation is definitely not warranted.

      (4) The ERP results seem to me rather superficial. But I am not an EEG expert.

      a) The authors do not seem to look at the outcome phase, which could be interesting for differences in reward/loss processing in the two task versions.

      b) Again, differences in EV seem to be more important from a conceptual point than probabilities or amounts; see my comment 2d.

      c) Also, the authors report ERPs for the two task types separately but do not seem to run proper comparisons between them, see my comment 2b.

      (5) Preregistration: It should be made very clear early on that this study was not preregistered.

      (6) Quality checks: The authors should check if some participants are outliers in terms of the number of missed trials, always choosing the same option, etc. It is notoriously difficult to find good post hoc reasons for excluding participants (one reason why replications and preregistrations are important). In any case, the data quality should be checked and described a bit more.

    3. Reviewer #2 (Public review):

      Summary:

      This paper investigates risk and cooperation decisions by integrating computational modeling with event-related potential (ERP) measures. Participants completed two tasks involving financial risk and cooperation under possible betrayal. The comparison between social and non-social decision-making is interesting and potentially valuable. However, the conceptual framing, theoretical grounding, and modeling rationale require substantial clarification.

      Strengths:

      (1) The paper introduces comparable tasks to probe social vs. non-social decision making.

      (2) The authors use a model to identify a psychological distinction and test its validity using neural data.

      Weaknesses:

      (1) Conceptual framing and theoretical clarity

      The primary theoretical contribution of the paper is currently unclear. Specifically, it is not clear what key difference the authors hypothesize between risk and cooperation conditions. This distinction should be grounded in prior literature.

      The manuscript states: "Indeed, mutual cooperation maximizes social welfare, whereas betrayal benefits the trustee but comes at the trustor's expense in the Trust Game (Joyce et al., 1995)." However, the authors do not discuss the substantial literature on the Trust Game, which is used here but not explicitly acknowledged.

      • The original Trust Game framework and behavior in one-shot settings (e.g., Berg et al., 1995).

      • The persistence of cooperation even when defection is economically optimal (e.g., Berg et al., 1995; Fehr & Fischbacher, 2003).

      • The influence of trustworthiness of the partner on cooperation decisions has been previously studied (Ma et al., 2022).

      • Differences between social and non-social decision-making contexts have also been reported with matched tasks (Liu et al., 2024).

      (2) Distinction between constructs (risk, loss aversion, betrayal aversion)

      The introduction introduces multiple related constructs-risk aversion, loss aversion, and betrayal aversion-but does not clearly differentiate them. A theoretically grounded distinction is needed.

      In particular:

      • The manuscript introduces multiple related constructs, or maybe the terms are used interchangeably? The distinction between risk aversion, loss aversion, defection aversion, and betrayal aversion should be clearly defined.

      • Betrayal aversion versus loss aversion is introduced but not clearly differentiated. Importantly, it should be clarified that this distinction is not experimentally manipulated but instead inferred through computational modeling. This point is currently not made explicit, which leads to confusion in the introduction

      • The computational model should be introduced clearly in the introduction. Without explaining how these constructs are operationalized in the model, the framework is difficult to follow.<br /> The statement "In the risk task, losses were solely impersonal" is also unclear. It seems the authors may mean "personal or non-social" rather than "impersonal" as rewards are always personally relevant.

      (3) Hypotheses and preregistration

      The manuscript would benefit from more theoretical rationale for hypotheses. For example:

      • What is the basis for hypothesizing that financial loss aversion and betrayal aversion independently affect cooperation choices?

      • Why should these constructs be separable and modeled independently?

      • Additionally, the absence of preregistration is a limitation that should be acknowledged even more.

      • Given the flexibility of the modeling approach and number of parameters, this is particularly important.

      • For instance, the rationale for focusing on decision times is also not clearly explained and should be better motivated.

      (4) Computational modeling

      There are several concerns regarding the modeling approach:

      • The choice of model comparison metric should be justified. Why is AIC used rather than BIC, which penalizes model complexity more strongly? This is particularly relevant given the inclusion of additional parameters to capture processes not directly measured by the task.

      • Full model recovery analyses are missing. A full model recovery is necessary to demonstrate that competing models produce distinguishable behavioral patterns. This needs to be shown in order to justify the specificity of the winning model

      • How correlated are the parameters across participants, particularly loss and betrayal parameters?

      • More broadly, it is unclear how well loss aversion and betrayal aversion can be differentiated based on behavior alone. If these constructs are separable, they should predict distinct aspects of behavior.

      (5) ERP analyses

      The ERP results (e.g., P300 and LPP) seem to suggest that betrayal aversion is relevant in both time periods and similarly.

      • Do neural signals differentially reflect betrayal aversion versus loss aversion earlier and later on?

      • Are there significant interaction effects between betrayal and loss aversion for each ERP component?

    4. Reviewer #3 (Public review):

      Summary:

      In this study, the authors aim to address two questions. First, do people avoid cooperation primarily because of betrayal aversion beyond loss aversion? Second, can the effects of betrayal aversion and loss aversion be dissociated at the behavioral and neural levels? To address these questions, the authors compared individuals' choices of taking risks in a nonsocial risk task with those in a social cooperation task, with the two tasks matched in success probability and principal amount. They fitted computational models that include betrayal-aversion and loss-aversion terms and related the model parameters to ERP measures. Based on these analyses, the authors concluded that betrayal aversion has a stronger effect on cooperation than loss aversion and that betrayal is encoded earlier than loss in the brain. This is an important research question, and the attempt to combine computational modeling with ERP analysis is valuable. However, the current data analyses may not be able to support all the conclusions the authors made. For instance, the claims concerning the dissociation between betrayal aversion and loss aversion are not yet sufficiently supported by the evidence.

      Strengths:

      (1) The research question is theoretically important. Distinguishing betrayal aversion from loss aversion is important for research on trust, cooperation, and risky decision-making.

      (2) The approach of integrating behavioral measures, self-report ratings, computational modeling, and ERP data is valuable and gives the study significance.

      (3) The behavioral findings are broadly consistent. Participants reported stronger emotional responses in the cooperation task and were less willing to accept risk in the cooperation condition. These findings are generally in line with previous work on betrayal aversion and provide a reasonable manipulation check for the contrast between social and nonsocial risk.

      Weaknesses:

      (1) The manuscript states that the two tasks are matched in probability and principal amount, but the cooperation task additionally introduces partner outcomes, betrayal, and prosocial components. The Methods section states that, in the cooperation task, if both players cooperate, the principal is doubled and then split equally; if the partner betrays, half of the participant's principal is transferred to the partner. The model also includes an expected-other-reward term, namely, V_other=ω[p⋅2X+(1-p)⋅1.5X]. This raises an interpretive concern: if the two tasks differ not only in whether the source of uncertainty is social, but also in partner outcome, intentionality, and potential inequity structure, then the fitted "betrayal aversion" parameter may in fact reflect multiple motives rather than betrayal aversion alone. In the current experimental design, the "betrayal aversion" parameter may not be uniquely interpretable as a pure betrayal-specific construct, and the current evidence is insufficient to support such a specific interpretation.

      (2) Participants were informed that the cooperation probabilities were derived from previous real participants, whereas in fact these probabilities were randomly generated. In addition, six participants explicitly expressed doubts about the authenticity of the social interaction, yet the authors retained these participants with only the brief statement that this "did not affect the results." For such a critical manipulation, this explanation is too brief. I recommend that the authors report robustness analyses excluding skeptical participants. Since six participants reportedly doubted the authenticity of the social interaction, and some participants also performed poorly on the catch trials, it would be important to show whether the main behavioral, modeling, and ERP findings remain after excluding these participants. This is especially important because the manuscript's central interpretation depends on the assumption that the cooperation task was genuinely experienced as social.

      (3) The descriptions of the sample size are inconsistent across sections. The Participants section states that, after excluding one participant for misunderstanding the instructions, the final sample consisted of 49 participants; however, the behavioral results section later states that only 42 participants were included in the final analyses due to recording problems. This discrepancy is important because readers need to know clearly which sample was used for the behavioral analyses, which for the model fitting, and which for the ERP analyses; whether these analyses were conducted on the same participants; and whether the exclusion criteria were consistent across analyses. The manuscript needs a more transparent description of sample size and exclusion criteria.

      (4) The authors need to do more thorough analyses to validate their models. In addition to AIC and parameter recovery, I would encourage the authors to include other model comparison metrics where possible, such as BIC and exceedance probability, as well as model-recovery analyses. The authors should also do model-based simulation analyses to show that the winning model can capture the contextual effects observed in real data.

      (5) The authors should explain the rationales for the choice of ERP time windows and component selection in more detail. The current ERP analyses are time-locked to principal onset, and P3/LPP are extracted from fixed time windows. The authors should explain why this is the most appropriate time-locking point for examining betrayal- and loss-related computations, and why alternative time-locking points, such as probability-cue onset or other key task events, were not used. More importantly, the time windows of P3 and LPP are defined arbitrarily in the current analyses. The authors need to apply a more principled approach to define ERP components. It looks like the P3 and LPP are from the same ERP component in Figure 3.

      (6) The manuscript has several internal inconsistencies in terminology, figure references, and result descriptions. These issues weaken the clarity of the arguments and reduce the readability of the manuscript.

      (7) The authors partially achieved their aims. The study does provide evidence that social risk and nonsocial risk are not treated equivalently, and it also offers a computational framework that is informative for the field. This is an important topic, and the overall approach is promising.

    5. Author response:

      We agree that the manuscript would benefit from a more clearly articulated conceptual framing, stronger model validation, more explicit statistical and ERP comparisons, and improved transparency regarding task design, sample inclusion, and preregistration. In the revised manuscript, we plan to address these points through substantial revision of the Introduction and Discussion, along with additional robustness and validation analyses, and more cautious interpretation of the main findings.

      Reviewer #1 raised important points about the framing of the cooperation task, the interpretation of betrayal, the standard statistical analyses, the modelling, and the ERP analyses. In response, we plan to clarify that the present task captures betrayal-related social risk or anticipated partner defection, rather than betrayal in its full interpersonal and emotional sense, and to better motivate this operationalization with reference to the betrayal-aversion and trust-game literature. We will moderate our claims regarding “emotional costs,” incorporate a more explicit task overview and accompanying schematic into the main text, and frame individual differences as a key avenue for future research. In addition, we will streamline the standard behavioral analyses, make the expected-value structure of the task explicit, add EV-based analyses of choice and reaction time, strengthen the ERP analyses, clarify that the study was not preregistered, and provide a complete report of data-quality checks. For the modelling section, a central revision will be to simplify the model structure and refit the models using a Bayesian hierarchical approach.

      Reviewer #2 emphasized the need for stronger theoretical framing and more specific distinctions between related constructs. In the revised manuscript, we will substantially revise the Introduction to better situate the present task in relation to the Trust Game literature and prior work comparing social and non-social decision-making under matched payoff structures. We will also define risk aversion, loss aversion, anticipated partner defection, and betrayal-related aversion more explicitly, and clarify that the distinction between betrayal-related aversion and loss aversion is inferred through computational modelling rather than directly manipulated as separate experimental factors. We also plan to introduce the computational model earlier in the manuscript, clarify how the key constructs are operationalized, replace unclear wording such as “impersonal losses,” strengthen the rationale for our hypotheses, and acknowledge the lack of preregistration more clearly.

      Reviewer #3 highlighted the need to align our conclusions more closely with the current evidence. In the revised manuscript, we will moderate the interpretation of the betrayal-related parameter, acknowledging that the cooperation task differs from the non-social risk task not only in social versus non-social uncertainty, but also in partner outcome, intentionality, and potential inequity structure. We therefore plan to avoid treating this parameter as a pure betrayal-specific construct and to describe it more cautiously as capturing betrayal-related social risk or aversion to anticipated partner defection. We also plan to report robustness analyses excluding participants who expressed doubts about the social interaction, as well as participants with poor catch-trial performance or otherwise low-quality data, and to clarify the sample sizes and exclusion criteria used for behavioral, modelling, and ERP analyses. Finally, we will strengthen model validation and ERP reporting, including broader validation analyses and more cautious interpretation if the evidence for temporal dissociation between betrayal-related aversion and loss aversion proves weaker than currently stated.

      Across these revisions, we also intend to simplify the model structure and use Bayesian hierarchical fitting to strengthen model validation, while avoiding overly strong claims if the additional analyses provide only modest support for a single preferred model.

    1. eLife Assessment

      This important study demonstrates that extrachromosomal circular DNA and chromatin-associated proteins are components of stress granules. The data from a range of cellular and microscopy approaches are convincing, but the main conclusions would be further strengthened by demonstrating functional relevance and by extending the analysis to additional cell types. This paper will be of broad interest to cell biologists and those studying stress granule formation.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Demeshkina and Ferré-D'Amaré showed that extrachromosomal circular DNA (eccDNA) and chromatin-associated proteins are present in stress granules, based on proteomic and sequencing analyses. Using HCR-FISH combined with imaging, the authors showed the colocalization of eccDNA with stress granule proteins. Furthermore, they found that CRISPR machinery targeting the eccDNA component of stress granules disrupts stress granule assembly, and that this effect is largely independent of Cas9 endonuclease activity. Notably, expression of cytoplasmic chromatin factors restores stress granule formation in the presence of CRISPR machinery in yeasts. This also rescues the growth defect caused by hypoxic stress, which correlates with impaired stress granule formation. Together, this manuscript provides insight into the presence of eccDNA in cytoplasmic membraneless organelles, specifically stress granules, and suggests a functional role for eccDNA within these structures under stress conditions.

      Strengths:

      The authors used a panel of ribonucleases to demonstrate that stress granule cores isolated from yeast and HEK293 cells are resistant to plasmid-safe DNase, an enzyme that does not degrade circular double-stranded DNA. To further support the presence of extrachromosomal circular DNA (eccDNA) in stress granules, they performed Circle-Seq on stress granule cores. The gel electrophoresis and sequencing experiments complement each other well, providing consistent evidence for eccDNA within these granules. Overall, this study provides insight into potential cytoplasmic roles for eccDNA, an area that remains largely unexplored.

      Weaknesses:

      (1) Figure 1F suggests that stress granule cores are susceptible to DNase I but not to plasmid-safe DNase (psDNase). However, its smearing pattern in the psDNase condition appears similar to that in the DNase I treatment shown in Figure 1E, although psDNase produces more discrete bands. The authors should comment on these differences between Figures 1E and 1F, or consider revising Figure 1F to improve consistency with Figures 1E and 1D.

      (2) The authors should clearly define "colocalization". Does it refer to complete spatial overlap between two signals (i.e., VCP and T30), or partial overlap (i.e., AHNAK DNA and G3BP)? Figure 3 and the associated text are descriptive. Quantitative analysis would strengthen the conclusions. For example, the authors could analyze the fraction of molecules localized to stress granules or provide Pearson's correlation coefficient or similar measurements.

      (3) The authors used a CRISPR-based approach to target the Ty1 LTR retrotransposon, an abundant stress granule eccDNA, and they observed a loss of stress granule formation. However, this phenotype may be specific to Ty1 eccDNA rather than representative of all eccDNA species present in granules. In particular, the title "Cytoplasmic circular DNA is a key constituent of stress granules" implies a broader role. To support this claim, the authors should consider approaches that more globally deplete eccDNA rather than targeting a single eccDNA.

      (4) The authors should provide additional experimental evidence to support the claim that eccDNA is packaged in a chromatin-like state. The rescue of stress granule formation by ectopic expression of modified chromatin-associated proteins (CHD1NES and GCN5NES) following CRISPR treatment does not necessarily demonstrate that eccDNA is packaged like chromatin under basal conditions.

    3. Reviewer #2 (Public review):

      Summary:

      The authors report the presence of extrachromosomal circular DNAs (eccDNAs) within the core of stress granules purified from both yeast and mammalian cells.

      Strengths:

      This study is important for understanding the molecular mechanisms underlying stress granules containing eccDNAs and is likely to have a major impact on future research. A major strength of the study is the extensive experimental validation performed in yeast cells. In particular, cytoplasmic CRISPR-mediated targeting of eccDNAs suppresses stress granule formation and impairs recovery from hypoxic stress in yeast cells.

      Weaknesses:

      The conclusions would be further strengthened by validating the functional findings in an additional model system, such as mammalian cells.

      Comments:

      (1) Section: "Stress granule cores contain eccDNA"

      a) The presence of eccDNAs would be more convincingly demonstrated using an orthogonal validation approach, such as DNA FISH targeting MYC and Centromere 8 (CEN8) on metaphase spreads from HEK293T cells (as performed in PMID: 34819668).

      b) The study would also benefit from assessing the presence of eccDNAs in the extracellular medium. For example, DNA could be extracted from conditioned media and analyzed by PCR using primers spanning eccDNA breakpoint junctions (as performed in PMID: 40074906; PMID: 36123406).

      (2) Section: "eccDNA-CRISPR abrogates stress granules"

      These findings should be further validated under additional stress conditions, such as drug-induced stress (like methotrexate) or nutrient deprivation in the cell medium.<br /> In addition, the same set of experiments should be performed in HEK293T cells to support the broader relevance of the observations.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Demeshkina and Ferré-D'Amaré showed that extrachromosomal circular DNA (eccDNA) and chromatin-associated proteins are present in stress granules, based on proteomic and sequencing analyses. Using HCR-FISH combined with imaging, the authors showed the colocalization of eccDNA with stress granule proteins. Furthermore, they found that CRISPR machinery targeting the eccDNA component of stress granules disrupts stress granule assembly, and that this effect is largely independent of Cas9 endonuclease activity. Notably, expression of cytoplasmic chromatin factors restores stress granule formation in the presence of CRISPR machinery in yeasts. This also rescues the growth defect caused by hypoxic stress, which correlates with impaired stress granule formation. Together, this manuscript provides insight into the presence of eccDNA in cytoplasmic membraneless organelles, specifically stress granules, and suggests a functional role for eccDNA within these structures under stress conditions.

      Strengths:

      The authors used a panel of ribonucleases to demonstrate that stress granule cores isolated from yeast and HEK293 cells are resistant to plasmid-safe DNase, an enzyme that does not degrade circular double-stranded DNA. To further support the presence of extrachromosomal circular DNA (eccDNA) in stress granules, they performed Circle-Seq on stress granule cores. The gel electrophoresis and sequencing experiments complement each other well, providing consistent evidence for eccDNA within these granules. Overall, this study provides insight into potential cytoplasmic roles for eccDNA, an area that remains largely unexplored.

      Weaknesses:

      (1) Figure 1F suggests that stress granule cores are susceptible to DNase I but not to plasmid-safe DNase (psDNase). However, its smearing pattern in the psDNase condition appears similar to that in the DNase I treatment shown in Figure 1E, although psDNase produces more discrete bands. The authors should comment on these differences between Figures 1E and 1F, or consider revising Figure 1F to improve consistency with Figures 1E and 1D.

      We suggest that the appropriate comparisons are between the DNase I and psDNase treatments within each figure panel, and not between panels (e.g., Figures 1E vs. 1F). The electrophoretic gels in the different panels were run for different lengths of time, and therefore the comparison between gels would be spurious. In Figure 1E, electrophoresis after DNase I treatment results in a characteristic smear, while after psDNase treatment yields discrete bands (lanes 2–3 vs. 4–5). Electrophoretic conditions for this figure were optimized to minimize diffusion and allow quantitative evaluation. The electrophoresis shown in Figure 1F, which compares yeast and mammalian stress granule core nucleic acids, was run for a longer period — as evidenced by the greater migration distance from the loading wells — yet still clearly shows the same qualitative difference between DNase I (smear, lane 3) and psDNase (discrete bands, lanes 1–2) treatments for the yeast samples. The apparent discrepancy noted by the referee therefore simply reflects the difference in electrophoretic conditions between the gels shown in the two separate figure panels.

      (2) The authors should clearly define "colocalization". Does it refer to complete spatial overlap between two signals (i.e., VCP and T30), or partial overlap (i.e., AHNAK DNA and G3BP)? Figure 3 and the associated text are descriptive. Quantitative analysis would strengthen the conclusions. For example, the authors could analyze the fraction of molecules localized to stress granules or provide Pearson's correlation coefficient or similar measurements.

      In our considered opinion, categorizing colocalization as either "partial" or "complete" implies a level of molecular precision that is physically unattainable at the resolution limits of any current light microscopy modality, and would therefore be misleading. Our approach employs super-resolution confocal laser scanning microscopy (Airyscan) with hybridization chain reaction fluorescence in situ hybridization (HCR-FISH) or with immunofluorescence. The detection method used offers higher spatial resolution and signal-to-noise ratio than single-point detector/physical pinhole confocal (or widefield epifluorescence) microscopy used in most prior stress granule studies. Despite these enhancements, the system retains inherent diffraction-imposed limits: a lateral (XY) resolution of ~130 nm and an axial (Z) resolution of ~350–400 nm, defining the minimum separable distance between two fluorescent signals. Structures smaller than these thresholds remain unresolved within a single point spread function (PSF) maximum – a volume sufficiently large to simultaneously accommodate multiple stress granule cores or tens of thousands of individual proteins (such as G3BP) and dozens of nucleic acid molecules several thousand nucleotides in length. Consequently, any detected fluorescence signal may represent the superimposition of a large and indeterminate number of individual molecules or particles. True molecular interaction analysis remains for future studies using technologies with angstrom resolution (e.g., cryo-electron tomography, cryo-EM, X-ray crystallography, smFRET, EPR, NMR, etc.). Metrics such as Pearson's correlation coefficient report solely on the degree of signal overlap at the PSF scale (hundreds of nanometers) and would not provide any insight beyond what is already conveyed by our data.

      (3) The authors used a CRISPR-based approach to target the Ty1 LTR retrotransposon, an abundant stress granule eccDNA, and they observed a loss of stress granule formation. However, this phenotype may be specific to Ty1 eccDNA rather than representative of all eccDNA species present in granules. In particular, the title "Cytoplasmic circular DNA is a key constituent of stress granules" implies a broader role. To support this claim, the authors should consider approaches that more globally deplete eccDNA rather than targeting a single eccDNA.

      We respectfully disagree with the referee that further depletion of eccDNA would alter our conclusions. A central finding of our study is that stress granules can be abrogated cytoplasmically by co-expressing a Cas9 endonuclease, active or inactivated by point mutations (D10A /H840A), and a gRNA (which is itself a fusion of the crRNA and trcrRNA, natively separate RNAs in the source bacterium). We show in Figure 4 that when the gRNA targets the Ty1 sequences, endonucleolytically active holoenzyme co-expression in the cytoplasm results in loss of the corresponding eccDNAs, as assayed by sequencing of the relevant cytoplasmic fractions. Critically, when a catalytically inactive Cas9 protein (dCas9) is co-expressed with the gRNA instead of the wild-type endonuclease, depletion of the eccDNAs containing Ty1 sequences no longer takes place (Figures 4D and 4E), but stress granule formation is still abrogated (Figure 4C).

      In our manuscript, we indicated (as "data not shown”) that co-expression with Cas9 of a gRNA "targeting" a sequence that is absent from the S. cerevisiae genome still results in abrogation of stress granule formation. These data are shown in Author response image 1. The gRNA is targeted to the sequence 5’-agaatcgatgcattt, which is absent in the genome of the yeast strain used.

      Author response image 1.

      It follows from our experiments that stress granule abrogation (1) is not a result of the catalytically active Cas9 endonuclease; (2) is not a result of the presence of a gRNA-directed but catalytically inactive Cas9 holoenzyme, but (3) is the result of the presence of a CRISPR holoenzyme (as defined above) in the cytoplasm.

      To reiterate, abrogation of stress granules occurs when a Cas9-gRNA complex is present in the cytoplasm, regardless of whether the nuclease activity exists, or the gRNA targets a sequence that is present in the genome. Importantly, the holoenzyme is required for this phenomenon: presence of the endonuclease or the gRNA alone does not abrogate stress granule formation (Figures S5).

      It is because of this unexpected observation that we next hypothesized that activities of the Cas9-gRNA complex other than sequence-specific gRNA-targeted endonucleolytic activity is driving the suppression of stress granule formation. The best documented such activity is DNA sequence sampling (1-dimensional diffusion). We think that 1-dimensional diffusion of the Cas9-gRNA holoenzyme is displacing from the cytoplasmic eccDNA interactors whose association with the DNA is required to drive stress granule assembly. The fact that the stress-granule suppressive effect of cytoplasmic Cas9-gRNA expression can itself be suppressed by two completely unrelated proteins whose only shared feature is action on chromatin (CHD1 and GCN5) strongly supports this hypothesis (Figures 4G, 4H and S6; also response to point 4, below), in addition to confirming that cytoplasmic eccDNA is packaged by histones in a conformation that CHD1 and GCN5 can both recognize.

      (4) The authors should provide additional experimental evidence to support the claim that eccDNA is packaged in a chromatin-like state. The rescue of stress granule formation by ectopic expression of modified chromatin-associated proteins (CHD1NES and GCN5NES) following CRISPR treatment does not necessarily demonstrate that eccDNA is packaged like chromatin under basal conditions.

      We would like to reiterate the temporal order in our experimental design (detailed in full in Methods and summarized in Results). Cas9<sub>NES</sub>-gRNA and CHD1<sub>NES</sub> (or GCN5<sub>NES</sub>) were expressed simultaneously (not sequentially) in the cytoplasm. This was intentional, so as to give each player ample opportunity to engage its preferred substrate under non-stress conditions, prior to the brief oxidative stress. The referee appears to believe that cytoplasmic eccDNA was pre-exposed to Cas9<sub>NES</sub>-gRNA, and then the bound endonuclease challenged with chromatin-modifying enzymes.

      Our experimental design accounts for the contrasting substrate specificities of CRISPR and chromatin-modifying enzymes. Cas9-gRNA (holoenzyme) binds to nucleosome-free DNA with sub-nanomolar dissociation constant (Kd 0.1–1 nM) but its association with chromatinized DNA is impeded 5- to 100-fold (Isaac et al., 2016; Yarrington et al., 2018; Strohkendl et al., 2021). In contrast, whereas CHD1 binding to DNA is strictly nucleosome-dependent — its chromodomains actively block engagement with protein-free DNA (Hauk et al., 2010), and its productive binding (Kd 10–200 nM) relies on obligate multivalent contacts with the histone octamer, H4 tail, and wrapped DNA (Farnung et al., 2017; Sundaramoorthy et al., 2018).

      Our observation that stress granule formation was unperturbed following oxidative stress is most parsimoniously interpreted as CHD1<sub>NES</sub> outcompeting the CRISPR machinery for cytoplasmic binding to eccDNA by virtue of the latter existing in a histone-bound state that is recognized as chromatin by CHD1 –simultaneously favoring CHD1<sub>NES</sub> engagement and impeding Cas9 access. Thus, our experiment in effect employs stress granule formation as a readout for differential binding to chromatin or chromatin-like eccDNA.

      Farnung, L., Vos, S.M., Wigge, C., and Cramer, P. (2017). Nucleosome-Chd1 structure and implications for chromatin remodelling. Nature, 550(7677), 539–542.

      Hauk, G., McKnight, J.N., Nodelman, I.M., and Bharat, T.A.M. (2010). The chromodomains of the Chd1 chromatin remodeler regulate DNA access to the ATPase motor. Mol Cell, 39(5), 711–723.

      Isaac, R.S., Jiang, F., Doudna, J.A., Lim, W.A., Narlikar, G.J., and Bhatt, D.L. (2016). Nucleosome breathing and remodeling constrain CRISPR-Cas9 function. Nature Struct Mol Biol, 23(12), 1097–1103.

      Strohkendl, I., Saifuddin, F.A., Gibson, B.A., Bhatt, D.L., Russell, R., and Bharat, T.A.M. (2021). Inhibition of CRISPR-Cas9 by bacteriophage-encoded proteins. Mol Cell, 81(8), 1665–1679.

      Sundaramoorthy, R., Hughes, A.L., Singh, V., Wiechens, N., Ryan, D.P., El-Mkami, H., Petoukhov, M., Svergun, D.I., Treutlein, B., Sproll, P., and Owen-Hughes, T. (2018). Structural reorganization of the chromatin remodeling enzyme Chd1 upon engagement with nucleosomes. eLife, 7, e35720.

      Yarrington, R.M., Verma, S., Schwartz, S., Trautman, J.K., and Carroll, D. (2018). Nucleosomes inhibit target cleavage by CRISPR-Cas9 in vivo.PNAS, 115(38), 9450–9455.

      Reviewer #2 (Public review):

      Summary:

      The authors report the presence of extrachromosomal circular DNAs (eccDNAs) within the core of stress granules purified from both yeast and mammalian cells.

      Strengths:

      This study is important for understanding the molecular mechanisms underlying stress granules containing eccDNAs and is likely to have a major impact on future research. A major strength of the study is the extensive experimental validation performed in yeast cells. In particular, cytoplasmic CRISPR-mediated targeting of eccDNAs suppresses stress granule formation and impairs recovery from hypoxic stress in yeast cells.

      Weaknesses:

      The conclusions would be further strengthened by validating the functional findings in an additional model system, such as mammalian cells.

      Comments:

      (1) Section: "Stress granule cores contain eccDNA"

      (a) The presence of eccDNAs would be more convincingly demonstrated using an orthogonal validation approach, such as DNA FISH targeting MYC and Centromere 8 (CEN8) on metaphase spreads from HEK293T cells (as performed in PMID: 34819668).

      The relationship between eccDNA dynamics and stress granule assembly across distinct cell cycle phases remains an important and poorly explored question. To our knowledge, no published data currently describe how stress response mechanisms are regulated during mitotic division, particularly in metaphase. Our identification of eccDNA as a component of stress granule cores can provide a first tractable framework to investigate this relationship. However, a systematic and in-depth characterization of this phenomenon warrants a dedicated future investigation.

      (b) The study would also benefit from assessing the presence of eccDNAs in the extracellular medium. For example, DNA could be extracted from conditioned media and analyzed by PCR using primers spanning eccDNA breakpoint junctions (as performed in PMID: 40074906; PMID: 36123406).

      We agree with the referee that eccDNA biology represents a fascinating and rapidly evolving area of research, particularly given the emerging role of eccDNA in oncogenesis. In this context, our identification of eccDNA as a core structural component of stress granules opens a novel avenue for exploring the connection between stress-dependent translational regulation and disease-associated eccDNA dynamics. While we acknowledge the importance of this direction, a rigorous investigation of this relationship requires extensive multifaceted experimentation that falls beyond the scope of the current study.

      (2) Section: "eccDNA-CRISPR abrogates stress granules"

      These findings should be further validated under additional stress conditions, such as drug-induced stress (like methotrexate) or nutrient deprivation in the cell medium. In addition, the same set of experiments should be performed in HEK293T cells to support the broader relevance of the observations.

      We agree with the referee that the composition and dynamics of stress granules arising from different stressors is an important endeavor. However, given the range of stressors documented to result in stress granule formation, those studies fall well beyond the scope of this manuscript. We will note however that the presence of eccDNA in stress granules of yeast and human cells is strong evidence for conservation of function(s). We think that exploration of the role of eccDNA in stress granule formation across the kingdoms of life (stress granules were first observed in heat-shocked tomato plants), cell cycle stages, stressors, etc. will be important research programs for the future.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figures 3D and 3I: The use of magenta and red makes it difficult to distinguish between the two labeled signals. Consider using more contrasting colors to improve visual clarity.

      We appreciate the comment regarding color choices in the figures. In our view, magenta and red are sufficiently distinguishable as nucleic acid labels, particularly when combined with the green signal representing G3BP in these panels.

      (2) Figures 3F and 3G: Do the authors have an explanation for why AHNAK or MAPT DNA (white) does not colocalize with the anti-DNA immunofluorescence signal?

      Immunofluorescence (IF) is standard for detecting protein antigens but has limitations when the target is a non-protein molecule such as DNA, owing to its compacted chromatinized state. Anti-DNA antibodies can miss a significant fraction of their targets because the DNA backbone remains largely inaccessible, a limitation that DNA-FISH overcomes by directly hybridizing probes to denatured DNA sequences with high specificity. The fixation step required for both IF and FISH imaging can introduce additional steric barriers that disproportionately restrict antibody access compared to small nucleic acid probes. Even under optimized conditions, the IF signal with anti-DNA antibodies is inherently reflective of a subset of the total cellular DNA content.

      (3) Adding a subtitle on page 12 ("The abundant histones in purified stress granule...") would improve the overall structure and readability of the manuscript.

      We think that an additional subtitle would not substantially improve the readability of what is, admittedly, a very dense manuscript that employs a diversity of experimental approaches.

      (4) It would strengthen the analysis if statistical significance were included for the different time points in Figure 5C.

      We appreciate the reviewer’s suggestion. Figure 5C shows the largest difference at 40–45 hours after stress recovery, which is statistically significant between Cas9NES-gRNA (or dCas9NES-gRNA) and Cas9NES or gRNA only (two-tailed Student’s t-test, *, p ≤ 0.05). All primary experimental data are publicly available (FigShare) so further analyses can be performed by interested future parties.

    1. eLife Assessment

      This useful study examines how the swimming of a lab strain of Escherichia coli changes under laboratory conditions designed to mimic disordered, porous environments: a setting that is biologically relevant and less understood than chemotaxis in bulk liquid. By combining experimental evolution in soft agar, inducible control of run duration, visualization of flagella, and theoretical modeling, the authors demonstrate that the optimal mean run duration for migration in semisolid agar is substantially shorter than in liquid and decreases with increasing agar concentration. The authors provide incomplete evidence regarding optimal chemotaxis in porous media as emerging through evolution from tradeoffs that involve newly discovered trapped states.

    2. Reviewer #1 (Public review):

      In this manuscript, the authors study optimal chemotactic navigation of bacteria in disordered environments. Most previous work has studied bacterial chemotaxis in free liquid, but navigation in obstructed environments is gaining more attention. Here, the authors first used the classic swim plate assay to select E. coli for chemotaxis in soft agar at two agar concentrations. In the higher concentration, they observed that the population's migration speed increased and the mean run duration decreased over selection cycles. Importantly, the growth rate did not change, so the change in migration speed was due to improved chemotaxis. Then, using a strain in which they could control the mean run duration with an inducible promoter, they measured population migration speed as a function of mean run duration, observing a peak. In liquid, theory predicts a peak when the run duration is comparable to the time scale of rotational diffusion. Here, the peak is at a much shorter run duration, and the optimal run duration decreased with agar concentration. A key feature in previous studies of bacterial motion in obstructed environments has been the dynamics of cell trapping and escape via tumbling. By directly visualizing the flagella in single cells, the authors found that the majority of trap events in semisolid agar did not end with a tumble. This is important because it means that the peak in the migration speed has a different origin from the peak typically seen in the diffusion coefficient, which is due to a balance between longer runs and less time spent trapped. Instead, using a minimal theoretical model, the authors argue that the peak in the migration speed is due to a balance between longer runs, which improve chemotaxis, and having those runs terminate with a tumble rather than a trap event, because runs that end with trapping do not result in up-gradient bias. Qualitatively similar behavior is seen in simulations of a more complex model of chemotaxis.

      Overall, we find the results to be significant and the evidence to be strong. We have some comments, which the authors need to address to improve/clarify their work:

      (1) The authors' model predicts that, because cells spontaneously escape traps without tumbling, the diffusion coefficient should depend monotonically on mean run length even though the chemotaxis coefficient is non-monotonic. It would strengthen the paper if the authors could show this to be true in experiments. Part of the reason for this comment is that the flagella labeling experiments were done in agar that was rapidly cooled in a freezer and then thawed, whereas the migration experiments were performed in agar cooled at room temperature. Our (anecdotal) understanding is that the cooling rate dramatically affects the properties of the agar mesh. Verifying that diffusivity is monotonic in mean run length would therefore show that cells' spontaneous escape from traps is not an artifact of the cooling protocol.

      (2) Two agar densities were used in their study (0.2%, 0.3%). As shown in Figure 1, while cells in the 0.3% agar showed significant improvements during the directed evolutionary experiments, the cells in 0.2% agar didn't. Correspondingly, the evolved average run time did not show significant changes in the 0.2% agar, but it decreased in the 0.3% agar. What is the reason for this difference? Does it mean the cells are already optimized for the 0.2% agar medium?

      (3) Related to the previous comment, the comparison between Figure 1 and Figure 2 should be made clearer. In Figure 2, a peak performance at an intermediate run time is shown, with the optimal run time decreasing with the agar density. Qualitatively, this result, i.e., the existence of the peak performance, gives the evolution experiments shown in Figure 1 a nice explanation. However, quantitatively, the run times shown in Figures 1 and 2 are quite different. For example, for the 0.3% agar case, the change of run time decreases from ~0.6sec. in cycle-1 to ~0.4sec in cycle-40. However, in Figure 2, the optimal run time is ~0.9sec., which means that the migration speed would decrease if the run time is decreased from 0.6sec to 0.4sec. We understand this may only be considered as a qualitative result. However, it does raise the question of what the molecular mechanisms are that drive the directed evolution, which the authors should address.

      (4) In Figure 3B, the distributions of speed in different media (liquid versus agar) for cells with bundled and split flagella are shown. While the distribution for the bundled flagella shows nicely the emergence of the trapped state (peak near zero speed), the distribution for the split flagella shows a significant shift of the distribution. Does this mean the agar medium also changes the tumble state significantly? In fact, we are puzzled by the observation that in bulk liquid, the run speed distribution for cells with split flagella seems to be quite similar to that of cells with bundled flagella, which might indicate problems in determining run speed.

      (5) Finally, none of the points plotted have error bars. Error bars would allow the readers to evaluate i) whether the changes in mean run speed during selection are significantly resolved and ii) whether the peaks in the migration speeds are significantly resolved.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Bai and colleagues investigates how Escherichia coli navigates and explores agar gels through chemotaxis and what parameters of bacterial swimming are tuned under selection pressure for rapid migration (i.e., reaching the edge of the agar plate quickly). Prior studies have examined related questions to a substantial degree. Examples include "Migration of Chemotactic Bacteria in Soft Agar: Role of Gel Concentration" (https://pmc.ncbi.nlm.nih.gov/articles/PMC3145277) and numerous other studies in this area (e.g., "Migration of bacteria in semi-solid agar" https://www.pnas.org/doi/10.1073/pnas.86.18.6973). From such studies has emerged the paradigm/model that reorientation (i.e., tumbling) is essential when bacteria navigate agar, which is considered a model for "complex" environments, because run-only bacteria become trapped in the agar matrix and are unable to migrate far. This new manuscript provides some evidence that this paradigm may be overly simplified or incomplete. As I understand it, the authors propose that migration is influenced to a greater extent by bias in the chemotactic run, where runs up attractant gradients are longer. The authors incorporate these data into a new model for chemotactic navigation and claim that this work establishes a general principle for how bacteria optimize active transport through complex environments.

      I will first note to the editor and authors that I am not qualified to assess the detailed mathematics of the model, and my review therefore focuses on the biology and phenotypes described. Nevertheless, in my view, this manuscript, in its current form, has several important limitations. For each point, I provide suggestions for additional experiments that could strengthen the rigor of the work and clarify the claims.

      Strengths:

      A strength of this work is the use of microscopy and automated methods to characterize an extremely large number of bacterial cells, which strengthens the authors' claims. However, substantially greater detail on these approaches is needed for the analysis to be reproducible and to allow verification that the analyses were performed correctly.

      Weaknesses:

      Major concerns

      (1) Claims are overly broad, and the experimental system is too artificial to support general conclusions about bacteria, chemotaxis, or evolution.

      E. coli MG1655 is a longstanding model organism in the chemotaxis field, and agar chemotaxis assays are also widely used. However, the authors make very broad claims about how phenotypic changes observed during selection in 0.2% or 0.3% agar relate to bacterial chemotaxis and evolution more generally. In essence, the experimental foundation on which the authors build a complex theoretical framework is limited to a domesticated laboratory strain of E. coli and a highly artificial environment consisting of agar in a Petri dish. Although E. coli is well studied, its motility and taxis behaviors are not necessarily representative of bacteria across nature. In addition, natural environments are dynamic, and bacteria rarely experience stable gradients for extended periods, such as the 24-hour time-frame used here. The authors have also only focused on responses to attractant gradients with undefined complex growth media, and not assessed if this is also true for repellent gradients. This is important to consider because E. coli also generates repellent gradients (indole) that are not considered here. E. coli also generates AI-2, sensed as an attractant, that would be an opposing force for migration. For these reasons, it is not clear that the data and theory presented here generalize to diverse bacterial species, to natural environments, or to chemotaxis broadly.

      The authors should acknowledge that further work is needed to generalise their findings by testing additional organisms, such as non-laboratory E. coli isolates, other enteric bacteria, and species with fundamentally different motility systems (e.g., Campylobacter jejuni). Further work could also expand beyond agar by examining chemotaxis in a biological matrix such as mucin, as well as testing responses to defined attractants and repellents.

      (2) No genetic component is identified, so claims about evolution are not supported.

      Evolution requires heritable genetic changes that produce phenotypes advantageous under a given selection pressure. The authors state that bacteria were selected for rapid migration and that this selection produced progressively more efficient migrators. However, no sequencing analyses of the evolved isolates were performed, no genetic changes were identified, and no mechanism underlying this phenotypic shift was described. Without identifying genetic alterations, they cannot substantiate the claim that evolution occurred. Whole-genome sequencing of the evolved isolates is necessary to determine whether specific mutations underlie the observed phenotypes.

      (3) The predictive power of the model is not tested.

      The authors develop a model with post-dictive capability, meaning the model reproduces behaviors similar to those observed in the data used to construct it. However, the manuscript does not demonstrate that the model has predictive power. Demonstrating predictive performance would substantially increase the value of the model. For example, the authors could perform an additional round of selection and predict the resulting bacterial behavior under a condition not used during model construction (such as a different agar concentration or predicting the behavior of different bacteria). Otherwise, the authors should tone down the claims.

      (4) Limited novelty and impact of the environmental difference studied.

      A central point of the manuscript is the difference between evolution in 0.2% versus 0.3% agar and how this difference relates to the proposed model. However, this represents a relatively minor change in the environment experienced by the bacteria. Developing an extensive theoretical framework and proposing that bacterial evolution is highly sensitive to these parameters based on this narrow experimental system may be premature. This would be addressed by the suggested broadening of experiments described above.

      (5) The manuscript is too brief, and some data and methods are insufficiently described, particularly related to the machine learning analysis.

      The manuscript addresses a complex topic, yet the main text, methods, and figures are very brief, which need not be the case. As a result, it is often difficult to understand exactly what was done and how the data support the authors' claims. More detailed descriptions of the experimental approaches and analyses are necessary.

      One example is the machine learning approach used for cell tracking. This method is only briefly described, and no validation data are presented that would allow readers to evaluate whether the approach performs accurately. If the method is robust, it would be a powerful analytical tool, but the current description does not provide sufficient information to evaluate the reliability of the results. This issue is particularly important because the authors conclude that tumbles account for less than 3% of escape events, which contrasts with previous paradigms. Automated tracking methods can be susceptible to artifacts, and therefore, rigorous validation of the tracking pipeline, supported by appropriate figures and benchmark data, is essential.

    4. Reviewer #3 (Public review):

      The manuscript by Bai et al presents a study of the effect of trapping on the efficiency of chemotactic spreading. While the overall impression of the study is positive, there are multiple drawbacks that accumulate and together make the statement of the paper not fully justifiable. Below, I provide some detailed comments in chronological order, and indicate those of particular importance.

      (1) On the first page of the Introduction, the authors use the following wording: "...how bacteria optimise their intrinsic motility parameters to maximise navigation efficiency". However, it is not shown or known whether they do. In the experiments, the authors fetch the bacteria at the far front and artificially select the ones with shorter run times. The ones at the front could be the effect of heterogeneity of the population rather than an adaptation. Moreover, the authors claim that the selective pressure is via trapping. But this can be due to a multitude of other factors that change with agar concentration, availability of nutrients, osmotic properties of water, etc.

      (2) At the beginning of the results section, the authors claim that for both agar concentrations, they observe a progressive increase in chemotactic navigation. I do not see how the data for 0.2 % agar would correspond to that. Migration speed remains flat.

      (3) (Important). The authors claim that the mean run speed remained constant. But this is definitely not true, as seen in the plots. The speed of modernity is increasing for both agar conditions. And here it is important to note that the chemotactic drift velocity is proportional to the square of run speed (which is not the case for the formulas in this paper, see comment below). Thus, even smaller changes in v_0 can result in a significant increase in the drift velocity.

      (4) (Important). Tumble bias is also significantly increasing in 0.3 agar concentration. While it is not clear from the paper what exactly the tumble bias is, if it is related to the persistence of the turning angle, this also has a linear effect on the chemotactic drift velocity.

      (5) (Important). When performing aTc dependence testing, the authors didn't report how other observables of swimming behaviour are changing.

      (6) (Very important). I'm not sure that by interfering with Che-Z expression, one does not affect the whole chemotactic circuit, for example, by changing G (in terms of the model) and thus the optimality occurs not due to the agar concentration/traps but due to the perturbations in the circuit. Also, the effect of different % seems to be much more minor compared to the overall induced changes in spreading speed.

      (7) (Very important). I was very confused by the statement of the authors about only 3% of traps being exited due to tumble. I don't think this is possible (in a way consistent with the suggested model). Mean free run times (Figure 1C) go down to 0.4 s. Duration of tumbles is 0.3s (Figure S2c), but the duration of traps is longer than tumbles (and a bit shorter than runs). So how can it be that a running cell gets into a trap and only in 3% cases it experiences a tumble? What would be the distribution of run durations if one combines pre-trap+trap_time+post_trap run time - would they still have a mean below 1s?? It really looks like the authors are not able to detect tumbles when bacteria are trapped. Or is there an active mechanism suppressing tumbles when in the trap?

      (8) It is not clear what it means that post-tumble angles were uniformly distributed. Does this refer to only trap-associated tumbles? It is known that in the freely swimming e.coli the tumbling angles are not isotropic but have a preference for the forward direction. Is it different in agar conditions?

      (9) (Very important) The authors assume an oversimplified model for the chemotactic drift based on biased random walks. As a result, the answer for chemotactic drift velocity has a wrong scaling with run speed. In the linear theory of chemotaxis by de Gennes, the scaling is v_0^2, while the authors use a linear relationship. Thus, the assumption of the simplified model is incorrect. The exact effect of the traps (where no tumbling is happening, and the directional memory is conserved) needs to be properly calculated, for example, in the same de Gennes framework. And I can't say what the result would be from the top of my head because the calculation is, in fact, not too trivial. Thus, the model used is oversimplified, and thus the fact that it shows a non-monotonous relationship with tau_f is of little predictive power.

      Taken together, you see that all the key points that are used in the chain of the argument about the optimality are not rock solid and allow for alternative explanations. I think all those either need to be tested explicitly or at least clearly discussed, and the respective conclusions of the paper need to be rephrased. In my view, this work needs major revision.

    1. eLife Assessment

      The nematode C. elegans is an ideal model in which to achieve the ambitious goal of having a genome-wide atlas of protein expression and localization. In this paper, the authors develop a rational strategy for at-scale tagging of all protein coding genes with fluorescent markers, providing solid evidence that it would be a feasible foundation for a community-based, genome-wide effort. This work should serve as an important springboard for discussions about how to achieve this worthy and impactful goal.

    2. Reviewer #1 (Public review):

      Summary:

      Eroglu and Hobert demonstrate that injecting CRISPR guides and repair constructs to target three genes at a time, tagging each with a different fluorescent protein, and selecting which gene to tag with which fluorophore based on genes' expression levels, can improve efficiency of gene tagging.

      Strengths:

      This manuscript demonstrates that three genes can be targeted efficiently with three different fluorophores. It also presents some practical considerations, like using the fluorophore least complicated by agar/worm autofluorescence for genes with low expression levels, and cost calculations if the same methods were used on all genes.

      Weaknesses:

      Eroglu has demonstrated in a previous publication that single-stranded DNA injection can increase efficiency of CRISPR in C. elegans, while inserting two fluorescent proteins and a co-CRISPR marker into three loci, and Paix et al 2015 demonstrated simultaneous insertion of two fluorescent tags. The current work is valuable and incremental advance. In general, I applaud the authors' willingness to strategize about how whole proteome tagging might be accomplished. I predict that the advance here will be one of many small advances that will get the field to that goal. The title oversells the advance presented, in my view, since seems like one among many key advances, and the first sentence of the Discussion seems a more apt summary of the key advance here.

      Some injections targeted genes on the same chromosome together, which will create unnecessary issues when doing crossing that will be useful for some future experiments. This made me wonder if injecting 3 together really is helpful vs targeting each gene separately, since only 5 worms need to be injected. It cuts time down by 2/3, but perhaps avoiding targeting the same chromosome with two tags would be useful.

      The limited utility of current blue fluorescent proteins makes me wonder if it's worth using at this stage, before there are better blue fluorescent proteins, or better yet, far red, to avoid issues with live imaging under phototoxic UV or near-UV illumination.

    3. Reviewer #2 (Public review):

      Overall, we found the responses to be quite recalcitrant.

      We have one remaining composite concern about the comparison between observed expression patterns with the new strains versus published data.

      First, the authors only report patterns for one stage while it should be not too much effort to image the different life stages. However, since this is a revision, we are not formally requesting they do this.

      Second, in the now provided Table (thank you) 'observed expression' (last column) is lacking for 9 of the 30 proteins, and for 6 of these the procedure was not successful. Why not report patterns for the other three? It is confusing also because on page 5, the authors say that "overall, 24 of 30 tags ...all of which were visible with fluorescence stereomicroscopy" - are we missing something? Also, they then said that they "obtained 6/9 of the originally failed tags"; why are the corresponding patterns not included in table 1, and are 9 proteins still labeled as "no" in the "success?" Column?

      Third, we strongly feel that the response to our comments about expression patterns is not adequate. On page 5 the authors say that "all proteins were expected to be ubiquitously expressed" and that "scRNA-seq indicated that transcript abundance was ubiquitous and without strong tissue-specific enrichment with few exceptions". However, in their rebuttal, the authors now argue for tissue-specific expression for proteins with paralogs, turning around their own argument! Moreover, their Table indicates that many genes show tissue-enriched expression by RNA-seq while many of their tagged proteins exhibit ubiquitous expression.

      Overall, this indicates that both the overall accomplishment of generating tagged protein strains and analyzing their expression is oversold.

    4. Reviewer #3 (Public review):

      Summary:

      The authors argue that establishing the expression pattern and sub-cellular localisation of an animal's proteome will highlight hypotheses for further study. This claim is probably accepted by many in the community. This manuscript seeks to confirm the feasibility of establishing such a resource, by using current transgenic methods to knock in DNA encoding different colored fluorescent tags into C. elegans genes.

      Strengths:

      The authors make the points above. For example, they provide evidence that the C. elegans germline harbors two populations of mitochondria that differ qualitatively in the proteins they express. They also confirm that labelling the whole proteome is an achievable goal with relatively limited resources and time.

      Weaknesses:

      The work is somewhat incremental in that it uses existing transgenic technology. Cell biology in C. elegans is challenging because of the small size of many of its cells, notably neurons. This can make establishing the sub-cellular localisation of a fluorescently tagged protein, or co-localizing it with another protein, tricky. The authors point out in their introduction that advances in light microscopy such as diSPIM, STED and ISM (a close relative of SIM), have increased the resolution of light microscopy. They also point out that recent advances in expansion microscopy can similarly help overcome the resolution limit. However, they do not use these technologies to characterize their transgenic strains.

    5. Reviewer #4 (Public review):

      Summary:

      Tagging the entire proteome of a metazoan would be a landmark achievement, providing a powerful complement and extension to existing "omic" catalogs in model systems. Here, Eroglu and Hobert argue that efficiently tagging multiple loci in a single "batch" would make the community-based achievement of this goal realistic. They provide rigorous evidence that such an approach is indeed feasible, exploring issues related to efficiency, design and screening strategies, disruption of gene function, and the potential for endogenously tagged alleles to reveal unexpected aspects of protein expression and localization. While the work has some minor gaps that are important to rigorously assess the feasibility of the proposed effort, the detailed and valuable insights that emerge should provide impetus to the community to coordinate efforts to make this ambitious goal a reality.

      Strengths:

      The work has numerous strengths. The authors provide compelling evidence that:

      - three distinct loci can be efficiently targeted with three distinct fluorescent tags in a single injection.

      - thoughtful targeting design can reduce the likelihood of disruption of function by the tag.

      - systematic design principles based on expression level and predicted localization/function can be used to optimize tagging strategies.

      - the resulting tags can provide unexpected insight into patterns of protein production and subcellular localization.

      Not all of these advances are novel in themselves, but taken together, they represent an important technical and conceptual advance. The most important strength comes from the exceptionally high value of the goal itself, in that the work is that it has the potential to spur a community-wide effort toward achieving the ambitious goal of proteome-wide tagging.

      Weaknesses:

      The work's shortcomings are minor.

      - One concern has to do with the feasibility of the proposed screening strategies. The experimental design cleverly coinjects tags for three loci in different gene expression 'zones'; this expression level determines which tag will be used. As the authors allude to, there is an important distinction between genes with the same overall FKPM value between those that are expressed broadly and those focally expressed in a specific tissue. The proposed strategy claims that there are a sufficient number of highly expressed genes "to be used as visible markers" for recovering successfully edited animals. It would be useful for the authors to discuss the issue of broad vs focused expression among this set of genes a bit more thoroughly, with an eye toward the issue of how likely it is that these genes could indeed consistently be used as visible markers, particularly for those at the low end of this limit.

      - What fraction of the proteome (on a per-gene basis) is secreted proteins? How difficult will it be to screen these for successful tags? Are there specific tags that would be more optimal for secreted proteins? (The authors mention the use of an SL2 or T2A cassette to label the cells in which these proteins are expressed but note that there are technical challenges associated with doing this at scale.)

      - For secreted and/or weakly expressed genes, it would be useful for the authors to estimate for what fraction of these would successful insertions need to be screened by PCR, and what resources (time and money) this would likely entail.

      - For how many genes would a single tag not capture all predicted isoforms?

      - Finally, some readers might object to the authors' assertion in the abstract that this work is "a first step in this direction" (presumably referring to designing a strategy for whole-proteome tagging). There is no concern that the authors are disregarding the extensive work of other groups, as they explicitly mention the contributions of other groups to the foundation that enables the present work. However, the spirit of the abstract could be misinterpreted by a well-intentioned reader.

    6. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      Summary:

      Eroglu and Hobert demonstrate that injecting CRISPR guides and repair constructs to target three genes at a time, tagging each with a different fluorescent protein, and selecting which gene to tag with which fluorophore based on genes' expression levels, can improve efficiency of gene tagging.

      Strengths:

      This manuscript demonstrates that three genes can be targeted efficiently with three different fluorophores. It also presents some practical considerations, like using the fluorophore least complicated by agar/worm autofluorescence for genes with low expression levels, and cost calculations if the same methods were used on all genes.

      Weaknesses:

      Eroglu has demonstrated in a previous publication that single-stranded DNA injection can increase efficiency of CRISPR in C. elegans, while inserting two fluorescent proteins and a co-CRISPR marker into three loci, and Paix et al 2015 demonstrated simultaneous insertion of two fluorescent tags. The current work is valuable and incremental advance. In general, I applaud the authors' willingness to strategize about how whole proteome tagging might be accomplished. I predict that the advance here will be one of many small advances that will get the field to that goal. The title oversells the advance presented, in my view, since seems like one among many key advances, and the first sentence of the Discussion seems a more apt summary of the key advance here.

      Some injections targeted genes on the same chromosome together, which will create unnecessary issues when doing crossing that will be useful for some future experiments. This made me wonder if injecting 3 together really is helpful vs targeting each gene separately, since only 5 worms need to be injected. It cuts time down by 2/3, but perhaps avoiding targeting the same chromosome with two tags would be useful.

      The limited utility of current blue fluorescent proteins makes me wonder if it's worth using at this stage, before there are better blue fluorescent proteins, or better yet, far red, to avoid issues with live imaging under phototoxic UV or near-UV illumination.

      These comments are a repeat of the original comments, and we refer the reader to our response to the original comments.

      Reviewer #2 (Public review):

      Original Review:

      The manuscript by Eroglu and Hobert presents a set of strains each harboring up to three fluorescently tagged endogenous proteins. While there is technically nothing wrong with the method and the images are beautiful, we struggled to appreciate the advance of this work - who is this paper for?

      As a technical method, the advance is minimal since the first author had already demonstrated that three mutations (fluorophore insertion and co-CRISPR marker) could be introduced simultaneously.

      As a pilot for creating genome-scale resources, it is not clear whether three different fluorophores in one animal, while elegantly designed and implemented, will be desired by the broader community.

      Finally, the interpretation of the patterns observed in the created lines leaves much to be desired. A Table with all the observations must be included and can replace the tedious (and often wrong) descriptions of the observations with the different lines. It would be too much to point out every mistaken expectation of protein expression. Two examples include:

      The expectation that ACDH-10 is enriched in the intestine and epidermal tissues (hypodermis) is naïve - there are multiple paralogs of this protein (look at WormPaths or WormFlux) that may share functions in different tissues. There is also no reason to assume that fatty acid metabolism does not occur in other tissues (including the germline). Finally, there are no published studies about this enzyme, so we really don't know for sure what it's doing.

      The expectation that HXK-1 is ubiquitously expressed is similarly naïve. There are three paralogous enzymes that are all associated with the same reaction, and we have shown that these three function redundantly in vivo, perhaps in different tissues (PMID: 40011787). Moreover, single cell RNA-seq data (PMID: 38816550) also shows enrichment of hxk-1 in gonadal sheath cells.

      The table should have at least the following information: gene/protein name - Wormbase ID - TPM levels of single cell data assigned to tissues for L2, L4 and adult (all published) - tissues in which expression is observed in the lines presented by the authors.

      Other points:

      (1) We would encourage the authors to provide systematic validation of the reported insertions. The manuscript reports that 24 of 30 tags were isolated and visible but does not clearly state whether each isolated line was confirmed by sequence‑level validation to be correctly in‑frame and free of unintended mutations at the target locus.

      (2) The manuscript presents aggregated success counts (e.g., 8/10 mTagBFP2 tags, 9/10 mStayGold, 7/10 mScarlet3) and useful narrative descriptions of injection outcomes. We suggest also to include per‑locus success rates.

      (3) For pools that required re‑injection after initial failures, we would like to see a description of the specific changes that were made to the injection mixes or procedures (e.g., new repair template prep, different Cas9 reagent lot, guide redesign). This will be useful troubleshooting information for others.

      (4) The authors states that the fluorophore sequences are codon-optimized for C. elegans. We suggest they provide the exact donor/tag sequences used specifically state whether the fluorophore sequences contain any synthetic/artificial introns or other sequence modifications (e.g., silent PAM‑disrupting mutations) were included in the donor templates.

      (5) Page 3: Include a reference for "The C. elegans genome encodes around 20,000 genes"

      We hope these comments are useful.

      Comments on Revised Version:

      Overall, we found the responses to be quite recalcitrant.

      We have one remaining composite concern about the comparison between observed expression patterns with the new strains versus published data.

      First, the authors only report patterns for one stage while it should be not too much effort to image the different life stages. However, since this is a revision, we are not formally requesting they do this.

      Second, in the now provided Table (thank you) 'observed expression' (last column) is lacking for 9 of the 30 proteins, and for 6 of these the procedure was not successful. Why not report patterns for the other three? It is confusing also because on page 5, the authors say that "overall, 24 of 30 tags ...all of which were visible with fluorescence stereomicroscopy" - are we missing something? Also, they then said that they "obtained 6/9 of the originally failed tags"; why are the corresponding patterns not included in table 1, and are 9 proteins still labeled as "no" in the "success?" Column?

      We appreciate the chance to clarify this matter: There are only 6 “no” in the “success” column. In two cases, HAT-1 and CBP-1, expression was dim at F1 but still sufficient to pick positive worms and quantify success rate at the locus. We noted these as “dim” on the table to indicate that if expression was lower, we likely would not have been able to isolate them at F1. In one case, COX-6B, expression was too dim at F1 to be isolated but was sufficient at F2 to be visualized and isolated from parents that were positive for the other two tags. We now clarified this distinction in the table and accompanying text: “Fluorescent signals of HAT-1::mScarlet3 and CBP-1::mScarlet3 in F1 progeny were dim but still sufficiently visible for quantification of knock-in efficiency, indicating that they are at the lower end of detectability for mScarlet3.”

      We imaged worms that had multiple tags as proof of principle and are happy to provide strains to those who would like to image/study them. At this point we are not convinced that imaging more worms would add to the conceptual framework.

      Third, we strongly feel that the response to our comments about expression patterns is not adequate. On page 5 the authors say that "all proteins were expected to be ubiquitously expressed" and that "scRNA-seq indicated that transcript abundance was ubiquitous and without strong tissue-specific enrichment with few exceptions". However, in their rebuttal, the authors now argue for tissue-specific expression for proteins with paralogs, turning around their own argument! Moreover, their Table indicates that many genes show tissue-enriched expression by RNA-seq while many of their tagged proteins exhibit ubiquitous expression.

      We respectfully disagree that there is contradiction. In our response, the discussion on paralogs was added as a clarification in response to the referee’s original comments (e.g., regarding ACDH-10): “There is also no reason to assume that fatty acid metabolism does not occur in other tissues (including the germline).” We wanted to make it clear that we were not concluding fatty acid metabolism (or other processes) does not occur in other tissues.

      We wish to stress that we never argued that paralogs could not fulfil the same essential function across tissues. The proteins were selected because their biological functions (e.g., glycolysis, fatty acid β-oxidation, translation) are broadly required, and that scRNA seq generally predicted broad expression with few exceptions as detailed in the text. Paralogs with similar activities (e.g., hxk-1, -2, -3) may overlap broadly in expression, or individual paralogs may carry out the process in different tissues provided one carries out the reaction in each tissue. For acdh-10 and hxk-1 specifically, both appear broadly expressed across tissues by scRNA-seq, with no consistent enrichment or depletion across datasets. So, our central point is that: for a specific gene involved in an essential process, transcript data alone are not sufficient to accurately predict tissue specific enrichment. Not that the processes do not occur in tissues where one paralog is absent. The possibility that a paralog may compensate for lack of expression is in no way contradictory with our conclusion.

      The table does not generally show tissue-enriched expression: it simply lists three tissues with the highest quantitative value in the respective dataset. For instance, taking the first gene from the list (Y82E9BR.3) and looking at the Ghaddar dataset, the top 3 tissues (log2(TPM)) are: pharyngeal muscle (13.4), gonadal sheath (12.9), marginal cells (12.9). The next 3 tissues are: body wall muscle (12.9), pharyngeal epithelium (12.8), and intestine (12.3). Even when there were apparent enrichments among the top 3 tissues, there were significant disagreements between datasets, and beyond top 3 even greater disagreements (the datasets agreed on the top tissue only 4 times over the 30 genes). These indicate that much of the variation is attributable to experimental noise rather than true predicted enrichment. The referee points to HXK-1 being correctly gonadal sheath enriched in one scRNA dataset; however, the other two datasets actually show different sites as being highest, and the same dataset misses effects in other cases. This is precisely why protein level data is needed.

      We further clarified this issue in the text: “We thus selected 30 genes across a variety of bulk transcript expression ranges which are generally predicted to be broadly expressed based on molecular function or, where molecular function was unknown (e.g., ZK632.9), single cell RNA sequencing (scRNA-seq) data (Table 1, Fig. 2A, B) (Gao et al., 2024; Ghaddar et al., 2023; Taylor et al., 2021).”

      Overall, this indicates that both the overall accomplishment of generating tagged protein strains and analyzing their expression is oversold.

      We have tried to make clear that our contribution is not a handful of new tagged strains added to the many that already exist. Rather, as stated in the abstract and elsewhere, we propose a strategy and provide proof-of-concept for scaling up tagging efforts. We believe the importance of this cannot be oversold.

      Reviewer #3 (Public review):

      Summary:

      The authors argue that establishing the expression pattern and sub-cellular localisation of an animal's proteome will highlight hypotheses for further study. This claim is probably accepted by many in the community. This manuscript seeks to confirm the feasibility of establishing such a resource, by using current transgenic methods to knock in DNA encoding different colored fluorescent tags into C. elegans genes.

      Strengths:

      The authors make the points above. For example, they provide evidence that the C. elegans germline harbors two populations of mitochondria that differ qualitatively in the proteins they express. They also confirm that labelling the whole proteome is an achievable goal with relatively limited resources and time.

      Weaknesses:

      The work is somewhat incremental in that it uses existing transgenic technology. Cell biology in C. elegans is challenging because of the small size of many of its cells, notably neurons. This can make establishing the sub-cellular localisation of a fluorescently tagged protein, or co-localizing it with another protein, tricky. The authors point out in their introduction that advances in light microscopy such as diSPIM, STED and ISM (a close relative of SIM), have increased the resolution of light microscopy. They also point out that recent advances in expansion microscopy can similarly help overcome the resolution limit. However, they do not use these technologies to characterize their transgenic strains.

      Reviewer #4 (Public review):

      Summary:

      Tagging the entire proteome of a metazoan would be a landmark achievement, providing a powerful complement and extension to existing "omic" catalogs in model systems. Here, Eroglu and Hobert argue that efficiently tagging multiple loci in a single "batch" would make the community-based achievement of this goal realistic. They provide rigorous evidence that such an approach is indeed feasible, exploring issues related to efficiency, design and screening strategies, disruption of gene function, and the potential for endogenously tagged alleles to reveal unexpected aspects of protein expression and localization. While the work has some minor gaps that are important to rigorously assess the feasibility of the proposed effort, the detailed and valuable insights that emerge should provide impetus to the community to coordinate efforts to make this ambitious goal a reality.

      Strengths:

      The work has numerous strengths. The authors provide compelling evidence that:

      Three distinct loci can be efficiently targeted with three distinct fluorescent tags in a single injection.

      Thoughtful targeting design can reduce the likelihood of disruption of function by the tag.

      Systematic design principles based on expression level and predicted localization/function can be used to optimize tagging strategies.

      The resulting tags can provide unexpected insight into patterns of protein production and subcellular localization.

      Not all of these advances are novel in themselves, but taken together, they represent an important technical and conceptual advance. The most important strength comes from the exceptionally high value of the goal itself, in that the work is that it has the potential to spur a community-wide effort toward achieving the ambitious goal of proteome-wide tagging.

      We appreciate the referee’s enthusiasm and hope that this will engage members of the community in a collective effort.

      Weaknesses:

      The work's shortcomings are minor.

      One concern has to do with the feasibility of the proposed screening strategies. The experimental design cleverly coinjects tags for three loci in different gene expression 'zones'; this expression level determines which tag will be used. As the authors allude to, there is an important distinction between genes with the same overall FKPM value between those that are expressed broadly and those focally expressed in a specific tissue. The proposed strategy claims that there are a sufficient number of highly expressed genes "to be used as visible markers" for recovering successfully edited animals. It would be useful for the authors to discuss the issue of broad vs focused expression among this set of genes a bit more thoroughly, with an eye toward the issue of how likely it is that these genes could indeed consistently be used as visible markers, particularly for those at the low end of this limit.

      To give two examples, this principle aided us with screening F54C8.1 and HAT-1. We added additional discussion on this to the first paragraph of the discussion: “For instance, we could clearly visualize F54C8.1::mScarlet3 in adult sperm by fluorescence stereomicroscopy despite a bulk FPKM of 16. Similarly, nuclear localized proteins will likely be easier to detect even at low expression levels, given the concentration of signal in small subcellular compartments. Indeed, this helped us detect HAT-1::mScarlet3 (56 bulk FPKM), which may have been too dim if distributed more broadly within cells.”

      What fraction of the proteome (on a per-gene basis) is secreted proteins? How difficult will it be to screen these for successful tags? Are there specific tags that would be more optimal for secreted proteins? (The authors mention the use of an SL2 or T2A cassette to label the cells in which these proteins are expressed but note that there are technical challenges associated with doing this at scale.)

      We added some of these points to the discussion: “Moreover, around 17% of the C. elegans genome (3,484 genes) may encode for secreted proteins (Suh and Hutter, 2012). Endogenous tagging of a substantial fraction of these proteins could reveal spatial patterns of secretion, distinguishing components that remain near their cell of origin from those that disperse to distal sites (Keeley et al., 2020). Tagging secreted proteins can also reveal sites of secretion – such as apical or basolateral membranes, or neurites – as has been observed for specific insulins (Sural et al., 2025) and for neuropeptides that localize selectively to synaptic regions (Toker et al., 2025).”

      Various tags have been used for secreted proteins including Venus, TagRFP, and mNeonGreen. The pH of secretory vesicles is ~5.0-5.5, so chosen FPs should have a pKa below this range to avoid denaturation. All 3 fluorophores used here (mStayGold, mScarlet3 and mTagBFP2) have pKa’s below this range and would likely be fluorescent within secretory vesicles.

      For secreted and/or weakly expressed genes, it would be useful for the authors to estimate for what fraction of these would successful insertions need to be screened by PCR, and what resources (time and money) this would likely entail. 

      We think that the bulk of ECM proteins would likely be visualizable without PCR due to their broad and stable expression, and as mentioned a good portion of these have been already tagged. However, it is likely that most of the secreted small peptides will have to be screened by PCR. We use homemade Taq, which makes material cost of the reagents minimal. A pair of genotyping primers costs ~$8 (~$27,872 for all secreted genes).

      Hands on time for lysis of 48-96 worms is approximately 20-30 minutes, with time to set up PCR around 5-10 minutes per target, and time to load a gel of 10 mins. In a given pool, 2/3 could be a putative secreted protein; thus, the same lysed population would enable screening for two targets at once. Collectively, around 40-60 mins of hands-on time would be required for two genes (around 20-30 mins per gene). Given 18 targets are injected per day, if 12 are screened by PCR, the screening could be done in 6 hours per day without affecting throughput. Most of the time spent on PCR would be replacing fluorescence screening time and would not overlap with the rate limiting injection step, performed by a separate specialist.

      For how many genes would a single tag not capture all predicted isoforms?

      Around 25% of C. elegans genes are thought to undergo alternative splicing (PMID: 21177968), with on average, ~2 isoforms per transcript. Among our selected genes, we only had one case where a single tag would not capture all isoforms (flad-1). We examined an additional 30 random genes and found no more examples by chance. So, in our view, this will be rare though we recognize in some cases a practical decision will need to be made, which could involve consideration of expression levels of each terminal exon.

      Finally, some readers might object to the authors' assertion in the abstract that this work is "a first step in this direction" (presumably referring to designing a strategy for whole-proteome tagging). There is no concern that the authors are disregarding the extensive work of other groups, as they explicitly mention the contributions of other groups to the foundation that enables the present work. However, the spirit of the abstract could be misinterpreted by a well-intentioned reader.

      We appreciate the referee’s perspective and have reworded this phrase in the abstract to: “As proof-of-principle for scalable pooled tagging, we undertook a pilot study in the nematode C. elegans, in which we set out to tag 30 different genetic loci with three different fluorophores, with 3 tags being introduced at a time.”

    1. eLife Assessment

      This study uses the yeast two-hybrid assay to identify proteins that may interact with yeast Set1 and other subunits of COMPASS/Set1C, the histone H3K4 methyltransferase, providing also some evidence for Set1 sumoylation and a role of SET1C methylating other factors in vitro. The results are valuable and they should contribute to understanding the functions of the conserved SET1C complex, as they suggest potential functional connections with RNA biogenesis, chromatin remodeling, and non-histone methylation whose implications would yet need to be explored. Nevertheless, apart from the fact that only a small subset of the Y2H interactions is further examined, the validating experiments are only partial or inconclusive, the strength of evidence being incomplete at this point, although with improvements over the previous version.

    2. Reviewer #1 (Public review):

      The manuscript has been improved in response to the reviewing. Although overinterpretation has been partially reduced compared to the previous version, the main concerns on the manuscript remain. The experiments have been conducted according to rigorous standards and the limitations of the results have been discussed to provide a comprehensive interpretation. However, this still represents an incomplete study in which the conclusions are insufficiently supported by the data provided.

    3. Reviewer #2 (Public review):

      Summary:

      This paper starts with a large-scale yeast two-hybrid (Y2H) screen using Set1 (full-length and smaller parts) and other Set1C/COMPASS subunits as bait. There are hundreds possible interactions identified, but only a small number are given any follow-up. While it's useful to document all the possible interactions, the unfocused and preliminary nature of the results makes the paper feel scattered and incomplete.

      Strengths:

      The Y2H screen was very comprehensive, producing lots of interesting possible leads for further experiments.

      Weaknesses:

      Most interactions were not further tested, and even in the case of those that were, the experiments are often inconclusive or incomplete.

    4. Reviewer #3 (Public review):

      The SET1C/COMPASS complex is the histone H3K4 methyltransferase in Saccharomyces cerevisiae, where it plays pivotal roles in transcriptional regulation, DNA repair, and chromatin dynamics. While its canonical function in histone methylation is well-established, its full interactome remains poorly defined. Moreover, whether SET1C methylates non-histone substrates has been an open question.

      In this study, Luciano et al. employ systematic yeast two-hybrid (Y2H) screening to uncover novel interactors and functions of SET1C. Their findings reveal potential functional connections to RNA biogenesis, chromatin remodeling, and non-histone methylation.

      The authors performed multiple Y2H screens using Set1 (full-length, N-terminal, and C-terminal fragments) and each of its seven subunits as baits. They identified high-confidence interactors that link SET1C to diverse cellular processes, including chromatin regulation (e.g., the SWI/SNF complex via Snf2), DNA replication (e.g., Mcm2, Orc6), RNA biogenesis (e.g., spliceosome components Prp8 and Prp22; polyadenylation factors Pta1 and Ref2), tRNA processing (e.g., Trm1, Trm732), and nuclear import/export (e.g., importins Kap104 and Kap123). Some of these interactions were further validated by immunoprecipitation or in vitro assays.

      Given the interaction of Set1 with Slx5 and Wss1-proteins involved in SUMO-dependent processes-the authors investigated and convincingly demonstrated that Set1 is sumoylated. This modification may influence the function and regulation of the SET1C complex.

      Finally, the authors provide evidence that SET1C methylates Snf2, the catalytic subunit of the SWI/SNF chromatin remodeling complex.

      One of the interactors, Nrm1, contains a domain resembling the H3K4-methylated sequence, which is also present in other proteins. Whether this H3K4-like domain is required for methylation remains to be demonstrated

      Strengths:

      This study offers valuable insights into the interactome of SET1C, suggesting potential links between the complex and a wide range of cellular processes. It also provides information on the possible regulation of Set1 by sumoylation. Finally, the finding that Snf2 is methylated in a Set1-dependent manner could significantly expand the known targets and functions of SET1C.

      Weaknesses:

      Many of the Y2H interactions remain to be validated and have to be considered as a starting point for further studies. Their functional significance remains to be explored. Several conclusions based on these 2HY data are speculative.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study uses the yeast two-hybrid assay to identify proteins that may interact with yeast Set1 and other subunits of COMPASS/Set1C, the histone H3K4 methyltransferase, providing also some evidence for Set1 sumoylation and a role of SET1C methylating other factors in vitro. The results are valuable, and they should contribute to understanding the functions of the conserved SET1C complex, as they suggest potential functional connections with RNA biogenesis, chromatin remodeling, and non-histone methylation, whose implications would yet need to be explored. Nevertheless, apart from the fact that only a small subset of the Y2H interactions is further examined, the validating experiments are only partial or inconclusive, the strength of evidence being at this point incomplete.

      We present a systematic SET1C interaction map that provides a structured resource for generating and testing new hypotheses on SET1C function. We emphasise that these interactions represent a hypothesis generating resource rather than a set of validated protein–protein interactions. To reflect this, the manuscript has been carefully revised to distinguish clearly between observation and interpretation, and to avoid overstatement of the data. Accordingly, we have revised the title and the abstract. Selected examples are explored further to illustrate how candidates from the dataset can be followed up, but the primary contribution of this work is to provide a structured framework and resource that can guide future mechanistic studies of SET1C function.

      We thank the reviewers for their thoughtful comments. We have followed their recommendations by modifying the structure of the manuscript, removing distracting results and relocating some figures to the supplementary materials to improve the readability of the manuscript. At the same time, the reviewers acknowledge that the dataset is extensive and that aspects of the validation work are valuable.

      The changes made to the manuscript's structure in accordance with the reviewers' recommendations are as follows:

      (1) Figure 1 is accompanied by a table (Table S2) with the raw data describing all the interactions from the ten 2H screens. This table also lists common interactors found in the independent screens. I'm afraid Table S2 was omitted from the initial submission of the manuscript

      (2) Figure 2 has been modified to include an AlphaFold modeling of a seven-subunit Set1C complex (Set1– Bre2–Sdc1<sub>2</sub>–Swd1–Swd3–Spp1) together with Kap104. Figure 2D has been moved to a new Figure S2

      (3) The initial figure S2, which was problematic, has been removed, along with the accompanying text.

      (4) Figure 3 of the original paper has been moved to the supplementary material and is now shown as a new Figure S3.

      (5) Figure 5 in the original paper becomes Figure 3 in the revised version

      (6) Figure S3 (Co-IP between Set1 and Prp22), which serves as validation data, has been moved to the main figures and is now presented as Figure 4.

      (7) Figure 6 in the original paper becomes Figure 5 in the revised version

      (8) Figure 4 from the original paper has been repositioned as the first figure (new Figure 6) of the biochemical characterization of the interaction between Snf2 and Set1C.

      (9) Figure 7 has been removed from the manuscript. We have kept the original Figure 7E as a new Figure S6.

      (10) Figures 8, 9, 10 become Figures 7, 8, 9.

      Public Reviews:

      Reviewer #1 (Public review):

      We thank Reviewer 1 for the careful and thoughtful evaluation of our manuscript. We fully agree that yeast two hybrid screening provides candidate interactions that require cautious interpretation, and we recognise that our original version did not always make this sufficiently explicit.

      In the revised manuscript, we have made substantial changes to address this central concern. All Y2H interactions are now consistently presented as candidate or potential interactions, and speculative statements have been either removed or explicitly framed as hypotheses. Our intention is that the reader can clearly separate the dataset itself from any proposed biological implications.

      Second, we have refocused the manuscript to better reflect its primary contribution. We now present the Y2H screens as a comprehensive resource that defines a set of candidate interactions for SET1C, rather than as a set of validated functional relationships. In line with this, we have reduced the emphasis on speculative models and removed sections where the connection to experimental evidence was not sufficiently strong. This includes the removal of Fig. S2 and Fig. 7 and the associated text, as well as the relocation of several figures to the supplementary material. Where appropriate, we have added statements highlighting the limitations of the approaches used and the need for future work to establish physiological relevance.

      More generally, we agree with the reviewer that the value of Y2H data lies in generating testable hypotheses rather than establishing conclusions. We have therefore revised the manuscript throughout to ensure that the interpretation remains proportionate to the strength of the evidence.

      We hope that these changes address the reviewer’s concerns and result in a clearer and more appropriately balanced presentation of the data.

      The manuscript by Luciano et al is a collection of experiments about the yeast histone 3 lysine 4 methyltransferase, Set1, starting with 10 yeast two-hybrid screens (Y2H). Y2H screens were briefly popular 20+ years ago, but the persistently unfavourable false-to-true positive ratios limited their utility, and the conclusion emerged that Y2H is an unreliable approach for gathering protein-protein interaction data. Y2H outcomes are candidate interaction lists at best, strongly contaminated by false positives. Here, the authors employed a company (Hybridomics) to perform the Y2H screens.

      The primary data is not presented, and the outcomes are summarized using the Hybridomics in-house quality scoring system in Figure 1A. It is not possible to evaluate these data, and the manuscript presents cartoon summaries that the reader must accept as valuable.

      Hybrigenics brings extensive experience from conducting numerous screens, enabling the team to recognize recurring false positives that commonly arise in screening assays. In their detailed analysis, Hybrigenics reports the number of clones recovered and the extent of overlap among interaction regions, both of which contribute to the confidence scores they assign. Table S2, provided in the revised version, more accurately reflects the raw data obtained by Hybrigenics. Nevertheless, we agree that false positives contaminate the list of potential interactors. Some interactions may also be indirect through a common interactor and do not reflect a physiological interaction.

      (1) Based on the extensive knowledge about Set1C/COMPASS acquired from genetics and biochemistry by many labs (including the Geli lab), the results presented here from the 10 Y2H screens are notably patchy. Of the 7 subunits of this complex, only one (Spp1) was identified using Set1 as bait. Conversely, as baits, Swd2, Spp1, Shg1, captured Set1, and the Bre2-Sdc1 interaction was reciprocally identified. These interactions were scored at the highest confidence level, which lends some confidence to the screens. However, the missing interactions, even at the third confidence level, indicate that any Y2H conclusions using these data must be qualified with caution. The authors do not appear to be cautious in their lengthy evaluations of these candidate interactions, which are illustrated with cartoons in Figures 2 and 3, with some support from the literature but almost without additional evidence. Snf2 is a particularly interesting candidate, which the authors support with pull-down experiments after mixing the two proteins in vitro (Figure 4). After Y2H, this is the least convincing evidence for a protein-protein interaction, and no further, more reliable evidence is supplied.

      We thank the reviewer for raising this important point regarding the strength of the evidence supporting the Set1– Snf2 interaction. We agree that the current data do not establish a definitive physiological interaction. In the discussion, we explicitly note the limitations of the current data.

      For Figure 2, as recommended by referee 2, we performed AlphaFold modeling of a seven-subunit Set1C complex (Set1–Bre2–Sdc1<sub>2</sub>–Swd1–Swd3–Spp1) together with Kap104. Consistent with the Y2H data, the model recapitulates binding of the Kap104 SID to the PY-NLS region of Set1 (residues 40–90).

      We have moved Figure 3 in the supplementary materials.

      (2) Figure 5 continues the cartoon summary of extrapolations from the Y2H screens, again without supporting evidence, except that the authors state.

      Figure 5 is now Figure 3. We have added the statement in the text: “It is not feasible to validate all of these interactions within the limits of this manuscript, and their validity should therefore be interpreted with caution. Nonetheless, these findings provide a useful basis for future research”.

      "We have refined the interaction region between Set1, Prp8 and Prp22, showing that Prp8 and Prp22 interact strongly with Set1-F4 (n-SET). Prp22 interacts in addition with Set1-F1 (Figure S2)." However, Figure S2 does not show this evidence and is incoherent.

      When we say that we have refined the interaction region between Set1, Prp8, and Prp22, we mean that we have restricted the interaction regions according to Y2H criteria. Indeed, we have not shown the spots illustrating the results. This statement has been deleted as well as Fig. S2

      The figure legends for Figure S2B and C do not correspond to the figure.

      (B) Expression of the F1-F5 fragments in yeast cells. Fusion proteins were detected with an anti-GAL4 monoclonal antibody. TOTO yeast cells (Hybrigenics) were transformed with the different pB66-Set1-F1 to F5 plasmids and subsequently with either P6, pP6-Snf2 762-968, pP6-Prp8 37-250, or pP6-Prp22 379-763 that were identified in the Y2H screens. Transformed cells were incubated 3 days at 30{degree sign}C on SD-LEU-TRP and then restreaked on SD-LEU-TRP-HIS with 3AT. Cell growth was monitored after 2 days at 30{degree sign}C.

      (C) Solid and dotted arrows indicate that transformed TOTO cells transformed with pB66-Set1-F1 to F5 and the indicated prey (Snf2, Prp8, and Prp22) are growing in the presence of 20 mM and 5 mM of AT, respectively.

      Figure S2D is two almost featureless dark grey panels accompanied by the figure legend D) Control experiment showing that TOTO cells transformed with p6 and pB66-Set1-F4 are not gowing (sic) in the presence of 5 mM or 20 mM AT.

      We agree that the legend for Figure S2 was unclear and does not accurately describe the panels shown in the figure. Fig; S2 has been deleted in the revised version. The results shown in the original Fig. S2 add limited information and may detract from the clarity of the main points.

      In the revised version, we have moved the CoIP analysis demonstrating the interaction between Set1 and Prp22 (previously shown in Figure S3) into the main figures (now Figure 4) to further support and validate the two-hybrid screening results presented there.

      Line 343. Interestingly, the two-hybrid screens reveal that Set1 1-754 interacted with Gag capsid-like proteins of Ty1 (Figure S5), raising the possibility that Set1 binding to Ty1 mRNA is linked to the interaction of Set1 1-754 with Gag.

      This is another example of the primary mistake repeatedly made by the authors -Y2H interactions are candidate results and not conclusive evidence.

      This statement is supported by our previous findings showing that Set1 binds Ty1 mRNA independently of its dRRM domain and represses Ty1 mobility at a post-transcriptional stage (Luciano et al., Cell Discovery, 2017; PMID: 29071121). One possible explanation for Set1 association with Ty1 mRNA is its interaction with the Gag capsidlike protein. In this context, the observed interaction between Set1(1–754) and Gag capsid-like proteins is consistent with this model.

      To further illustrate this point, the authors highlight the candidate interaction between Nis1 and 3 Set1C subunits.

      While we agree that the Nis1-Set1C interaction has not been demonstrated beyond doubt, we feel that our Y2H and in vitro binding experiments provide reasonable evidence that the interactions may be relevant. It is important to consider that any interaction assay can provide negative (and false positive) results, this includes Y2H, in vitro binding and mass-spec analysis of purified complexes from cells. We feel that it is not appropriate to only trust protein interactions that are strong and stable enough to be demonstrated via purified complexes. It is clear that some protein interactions do occur in transient and weak manner and therefore are not compatible with biochemical purification approach. This indeed is the strength of alternative methods like Y2H and in vitro binding assays, that interactions can be identified and tested even if the physiological context of the interaction may be more complex.

      (3) After multiple speculations based on the Y2H candidates, the authors changed to focus on sumoylation of Set1, which has previously reported to be sumoylated. Evidence identifying two sumoylation sites in Set1, in the N-SET and SET domains, is valuable and adds important progress to the role of sumoylation in the regulation of H3K4 methyltransferase, relevant for all eukaryotes. This illuminating part of the manuscript is only tenuously connected to the preceding Y2H screens and concomitant speculations.

      We thank Referee 1 for their comment. While it is true that there is only a modest connection between Set1 interactors involved in direct or indirect sumoylation and the characterization of Set1 SUMOylation sites, we believe that this does not constitute a weakness of the manuscript.

      (4) The manuscript then describes a red herring exercise involving Set1 methylation of Nrm1. In an already speculative and difficult manuscript, it is exasperating to read a paragraph about a failed idea. Apart from panel E, Figure 7 is a distraction, and I believe it should not be shared.

      (5) However, despite the failure with Nrm1, Line 443 - The H3K4-like domain in Nrm1 raised our attention to other yeast proteins that carry such sequences.

      This line of thinking is even less connected to the Y2H screens than the sumoylation work.

      However, the authors present a reasonable evaluation of the yeast proteome screened for six amino acids similar to the known H3K4 motif ARTKQT (Figure 7e).

      (6) However, this evaluation goes nowhere and has no connection with the next section of the manuscript, which is entirely speculation about the regulation of metabolism and stress responses based on the Y2H results and selected evidence from the literature.

      In response to comments 4 and 5, we have removed Fig. 7 and the paragraph titled “The transcriptional corepressor Nrm1 interacts with SET1C.” Part of this paragraph and the section describing the screen of the yeast proteome for six–amino acid sequences resembling the H3K4 motif (ARTKQT) has been kept as Fig. S6.

      In the abstract, we have removed the sentence: We demonstrate that the transcriptional corepressor Nrm1 is methylated by SET1C in vitro suggesting that H3K4-like domains may represent a class of non-histone substrates for SET1C.

      At the end of the introduction, we have deleted “the transcriptional corepressor Nrm1” in the sentence: In addition, we demonstrate that the transcriptional corepressor Nrm1 and the Snf2 AT-hook are both methylated by SET1C in vitro

      (7) The manuscript then describes more failed experiments regarding lysine methylation of Snf2 by Set1C, which unexpectedly reports arginine methylation rather than lysine. The manuscript does not currently meet the standard expected for this type of paper - the composition is somewhat incoherent and there are no previous reports of arginine methylation by SET domain proteins.

      We have integrated extensive in vitro reconstruction experiments with complementary in vivo studies, all conducted according to the rigorous standards expected by leading journals. These approaches have allowed us to reach the conclusions presented in this manuscript. While some of these findings are unexpected, they are supported by the data. We have carefully discussed the results and their limitations to provide a comprehensive interpretation.

      The manuscript presents a very experienced grasp of the literature and a sophisticated appreciation of the forefront issues, but a surprising failure to eliminate uninformative failures and peripheral distractions. The over interpretation of Y2H results is a dominating failure. There are some valuable parts within this manuscript, and hopefully, the authors can reformat to eliminate the defects and appropriately qualify the candidate data.

      We thank Referee 1 for these insightful comments. In the revised version, we have followed the advice to remove non-informative failures and peripheral distractions. Additionally, we exercise greater caution to avoid over-interpreting the Y2H results.

      Reviewer #2 (Public review):

      Summary:

      This paper starts with a large-scale yeast two-hybrid (Y2H) screen using Set1 (full-length and smaller parts) and other Set1C/COMPASS subunits as bait. There are hundreds of possible interactions identified, but only a small number are given any follow-up. While it's useful to document all the possible interactions, the unfocused and preliminary nature of the results makes the paper feel scattered and incomplete.

      Strengths:

      The Y2H screen was very comprehensive, producing lots of interesting possible leads for further experiments.

      Weaknesses:

      The results are useful but incomplete because only a small subset of the Y2H interactions is further examined. Even in the case of those that were further tested, the validating experiments are only partial or inconclusive.

      Referee 2’s comments align in some respects with those of Referee 1. In the revised version, we have followed the detailed Referee 2 suggestions to reduce the scattered nature of the manuscript. In addition, we include an AlphaFold model of the interaction between the Set1 N-term 1-754 with the SID domain of Kap104 that involves the proposed Set1 PY-NLS sequence.

      Reviewer #3 (Public review):

      The SET1C/COMPASS complex is the histone H3K4 methyltransferase in Saccharomyces cerevisiae, where it plays pivotal roles in transcriptional regulation, DNA repair, and chromatin dynamics. While its canonical function in histone methylation is well-established, its full interactome remains poorly defined. Moreover, whether SET1C methylates non-histone substrates has been an open question. In this study, Luciano et al. employ systematic yeast two-hybrid (Y2H) screening to uncover novel interactors and functions of SET1C. Their findings reveal potential functional connections to RNA biogenesis, chromatin remodeling, and non-histone methylation.

      The authors performed multiple Y2H screens using Set1 (full-length, N-terminal, and C-terminal fragments) and each of its seven subunits as baits. They identified high-confidence interactors that link SET1C to diverse cellular processes, including chromatin regulation (e.g., the SWI/SNF complex via Snf2), DNA replication (e.g., Mcm2, Orc6), RNA biogenesis (e.g., spliceosome components Prp8 and Prp22; polyadenylation factors Pta1 and Ref2), tRNA processing (e.g., Trm1, Trm732), and nuclear import/export (e.g., importins Kap104 and Kap123). Some of these interactions were further validated by immunoprecipitation or in vitro assays.

      Given the interaction of Set1 with Slx5 and Wss1 - proteins involved in SUMO-dependent processes - the authors investigated and convincingly demonstrated that Set1 is sumoylated. This modification may influence the function and regulation of the SET1C complex.

      Finally, the authors provide evidence that SET1C methylates proteins beyond histone H3K4, notably Nrm1, a transcriptional corepressor, and Snf2, the catalytic subunit of the SWI/SNF chromatin remodeling complex. Although Nrm1 contains a domain resembling the H3K4-methylated sequence (H3K4-like domain), this region does not appear to be required for its methylation. The search for other proteins containing similar domains as potential methylation candidates (p.12, first paragraph) seems less justified, given the lack of evidence supporting the requirement for the H3K4-like domain in methylation.

      This study offers valuable insights into the interactome of SET1C, suggesting potential links between the complex and a wide range of cellular processes. However, the functional implications of the Y2H interactions remain to be explored further. Additionally, the study provides intriguing information on the possible regulation of Set1 by sumoylation. The discovery of Nrm1 and Snf2 as methylation substrates could significantly expand the known targets and functions of SET1C.

      The results are supported by high-quality data.

      We thank referee 3 for their positive comments

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Restructure the manuscript into at least two papers.

      We thank the reviewer for this suggestion. In the revised manuscript, we have addressed this concern by substantially restructuring and streamlining the presentation. We consider the dataset, validation experiments, and functional observations to be closely integrated, and we believe that presenting them together provides the most coherent and impactful account of the work.

      Minor points

      There are several basic flaws in the manuscript that I feel indicate the co-authors have not proofread the manuscript sufficiently - 4 examples from early in the manuscript are listed below.

      (1) The reference for Hybridomics is (73) - obviously from an earlier version that used a different referencing system that has not been corrected.

      Thank you. This has been corrected.

      (2) Line 194 - 197. These screens have proven their power and effectiveness. In particular, they identified ...... the CTD of Rpb1 as an interactor of the N-terminal region of Set1 (Bae et al, 2020) (Figure S1). Rbp1 interaction is not identified in the screens presented here, and Figure S1 is a cartoon and not primary evidence.

      The interaction between the CTD of Rpb1 (Rpo21) and Set1 is reported in Table S2. The detailed characterization presented in Bae et al. (2020) was subsequently carried out as a direct follow-up to this screen.

      (3) Line 205-211. The highly confident interactors of the seven SET1C subunits are shown in Figure 1C-E. We found that Spp1, Shg1 and Swd2 interact alone with Set1 (Figure 1C). The minimum Set1 region for which an interaction is found for each of these 3 subunits is shown in Figure 1C. The high confidence interactors of the seven SET1C subunits are shown in Figure 1C-E. We found that Spp1, Shg1 and Swd2 display Y2H interactions with Set1 (Figure 1C). The high confidence interactors of Spp1, Shg1 and Swd2 are indicated in Figure 1D (see also Table S2).

      It is possible that Table S2 was omitted from the original submission, as it was requested during the production stage.

      (4) Line 335. We have classified all Set1 and subunit interactors according to these SET1C roles (Figure S5). However, this refers to Figure S4 - many further references to Figure S5 are also to Figure S4.

      Thank you. This has been corrected.

      Reviewer #2 (Recommendations for the authors):

      General recommendations:

      (1) Figures 1, 2, 3, and 5 and their associated main text are essentially just lists of interactors, put in graphic form and grouped to allow speculation about possible biological functions for the interactions. But almost none of the ideas are tested, so these sections take much more space than warranted. Having so much preliminary Y2H data actually distracts attention from the follow-up experiments that are shown. I would move most or all of this to the supplement, consolidating the Y2H results into fewer figures (or even just the Table).

      As mentioned earlier, the manuscript has been reorganized and Table S2 is provided.

      (2) The Snf2 interaction gets the most follow-up, so separating Figure 4 from Figures 8-10 broke the flow of that story. I would group these figures together since all are related to the Snf2 AT hook story.

      This was done accordingly.

      (3) I understand that it's impossible to validate all the possible interactions, particularly if resources are limited. However, at least for the interactions that get further attention, it could be very useful to try some AlphaFold multimer predictions. A high confidence AlphaFold score would provide a second orthogonal piece of evidence to support the Y2H results.

      We generated an AlphaFold model (Figure 2C) that recapitulates the key predictions for the Set1-Kap104 Y2H interaction.

      Comments on specific sections:

      (1) Y2H results. The text says Figure 1 shows all the high-confidence interactors. But the Set1 NTD interaction with the Rpb1 CTD is not shown here (it's in the supplement).

      In Table S2, an interaction is observed between full-length Set1 and the Rpb1-CTD (14 repeats), where Rpb1 is referred to as Rpo21.

      Figure 2 shows additional high-confidence interactors that do not appear in Figure 1, while others (like the Shg1Mog1 interaction) are shown in both Figures 1 and 2. It's confusing to scatter the data like this, which is why I recommend consolidating into a single figure or table.

      In Figure 2, the high-confidence interactors of Set1 (1–754) are highlighted in red and green (Snf2, Gbp2, and Kap104), and all are also present in Figure 1. Dbp1, identified as a high-confidence interactor of Spp1, likewise appears in Figure 1. Table S2 summarizes all of these interactions.

      (2) Line 219. How does a "high confidence" Set1-Kap104 Y2H interaction suggest the interaction is direct? Couldn't an indirect interaction also be tight and reproducible? This is an example where it would be worth seeing if AlphaFold also predicts an interaction and, if so, whether it involves the proposed NLS sequences.

      Y2H screening indicated that Kap104 binds to the N-terminal region (aa 1–754) of Set1 via its Set1 interaction domain (SID). To validate this, we used AlphaFold to model the seven-subunit Set1C complex (Set1-Bre2-Sdc1(x2) Swd1-Swd3-Spp1) with Kap104. The resulting model showed borderline confidence for the overall fold (pTM = 0.53) and low confidence in subunit positioning (ipTM = 0.5). Visualization in PyMOL confirmed Kap104 SID binding to Set1(1–754), consistent with Y2H results. The structure highlights Kap104 SID interaction with Set1’s PY-NLS at residues 40–90; the second PY-NLS is neither visible nor engaged in this model.

      (3) In the discussion of nuclear import interactors, what does it mean to say the Shg1-Mog1 interaction is "along the same line" as Set1-Kap104?

      We meant that the interaction between Shg1 and Mog1 represents another example of an interaction between a Set1C subunit and a protein involved in nuclear import. Along the same line has been deleted in the revised version.

      (4) To follow up on the Swd1-Nrm1 Y2H interaction, the paper shows that Nrm1 is methylated by Set1 in vitro (Figure 7), but it's not clear whether this has any biological significance. Without any in vivo follow-up, this figure is probably more appropriate for the Supplement.

      As noted above, Figure 7 has been removed, only panel E of Figure 7 is retained in the revised version.

      (5) Figures 6 and S8 show that Set1 is SUMOylated. Although it's not clear what this does to Set1 function or which E3 is responsible, the modification data looks convincing. The legend to Figures 6A and B says the Elutes samples are purified on nickel columns. Why are the Myc-Set1 and GB-Set1 proteins without the his-SUMO modification also binding to the nickel column? That's not happening in panels C and D. In the blots on the right for his-SUMO, is there any way to show that one of those bands is Set1? Maybe IP for MYC and then probe for the His tag?

      We thank the reviewer for this observation. His-SUMO purification using Nickel beads was used to purify HisSUMOylated proteins. Purified proteins were analyzed by Western blot using anti-MYC or anti-GAL4 antibodies to detect SET1-His-SUMO, as well as anti-His antibodies to confirm the presence of purified His-SUMOylated proteins. As mentioned by the reviewer, we detected unmodified MYC-Set1 and GAL4-Set1 in both the (-) and (+) His-SUMO eluates. This phenomenon is most likely due to the stickiness of unmodified Set1 to the beads. This is a commonly observed phenomenon in this type of biochemical assay, particularly when analyzing large proteins such as Set1 (124 kDa). This stickiness behavior has been observed in similar SUMOylation assays, e.g., for Hpr1 (88 kDa) (Bretes H, 2014. PMID: 24500206), Nup1 (114 kDa), and Nup2 (78 kDa) (Folz H, 2019. PMID: 30837289). This stickiness was not observed when using Set1 fragments (panels C and D), most likely because the fragments lost the stickiness to the beads, a characteristic belonging only to the full-length Set1. We mention this point in the legend of the new figure 5.

      (6) The Snf2 interaction gets the most follow-up. The GST pulldown validation of Set1 interaction with Snf2 AThook looks pretty good. However, the RGG repeats are necessary for the Set1 interaction with recombinant Snf2 proteins, but not for the co-IP of in vivo material. Again, AlphaFold could lend further support here.

      Thank you for this helpful suggestion. We agree that structural modelling could, in principle, provide an additional and orthogonal line of support for the Set1-Snf2 interaction. We did explore this using AlphaFold. However, both Set1 and Snf2 contain extensive intrinsically disordered regions, including the regions implicated in the interaction, and none of the models we obtained provided interpretable structural insight into the interaction interface. In particular, the predicted complexes showed low confidence in relative domain positioning, which limits their usefulness for supporting or refining the interaction model. One possible explanation is that additional components are required to stabilise a meaningful interaction in silico. While we modelled Set1 within a seven-subunit Set1C complex, Snf2 was necessarily included in isolation from its native context. Given that Snf2 functions as part of multiple, heterogeneous chromatin remodelling complexes, the absence of its physiological binding partners may prevent AlphaFold from resolving a relevant interaction interface. In light of these limitations, we have not included the AlphaFold models in the manuscript, as we felt they would not provide reliable or informative support. Instead, we have focused on the experimental evidence presented. We have clarified this point in the revised discussion to acknowledge both the potential and the current limitations of structural prediction approaches in this context.

      (7) The Snf2 methylation by Set1 is less convincing, and its biological significance is still unclear. I think it's pretty unlikely that Set1 could methylate arginine. The mass spectrometry is used for in vivo validation (mass spec), but mutating the lysines (Figure S11, S12) or Set1 deletion (Figure S14) doesn't seem to affect the signal. Could there be quantitative differences? Is there any way to quantitate the mass spec data to estimate the modified/unmodified ratio?

      We thank the reviewer for highlighting the unexpected nature of the methylation results. We agree that the observation of arginine methylation in this context is surprising, particularly given that SET domain proteins are classically associated with lysine methylation. This is why we performed multiple in vitro and in vivo experiments, and careful interpretation data that were clear led us to conclude that Set1C methylates the arginines within the ARTSTRGR motif of the AT-hook. We agree that the biological significance of this modification remains unclear. We obtained data showing that deletion of the SID domain of Snf2 impairs yeast growth on lactate, whereas this mutant grows normally on glucose and galactose, in contrast to the Snf2Δ mutant, which exhibits poor growth on both glucose and galactose. In comparison, deletion of the RG motif of Snf2 does not affect growth on lactate. These results provide insight into the interaction between Set1 and Snf2 but do not shed light on the potential importance of methylation of the RG motif. We therefore chose not to include them. In the discussion, we acknowledge the limitations of the current evidence. Our intention is to retain these findings as potentially interesting observations while ensuring that their interpretation remains appropriately cautious.

      Minor comments:

      (1) Lines 153 and 163: Stress response is listed twice, but with different references. Maybe these need to be further defined or else combined?

      We have deleted stress response line 163 and moved the references “Deshpande et al, 2022 and Nadal-Ribelles et al, 2015” line 153.

      (2) Line 193: better to say the proteins were fused to the C- or N-terminus (rather than upstream/downstream). It would be worth mentioning if there was a reason why Swd2 was fused to the N-terminus, unlike all the others.

      This has been done accordingly. In our hands, C-terminal fusions of Swd2 are not functional.

      (3) Is the scoring scheme (highest, high, good) that produces the colors in Figure 1 shown in the table? It doesn't say what the tan color (two of the Bre2 interactors) means.

      It is a mistake, Tea1 should be blue and Swi1 should not appear here. This has been fixed.

      (4) Line 206. It's not clear what it means to say that three of the subunits "interact alone with Set1". It can't mean they only interact with Set1, since other interactors are shown in Figure 1B. If it meant to say the interactions don't require other COMPASS subunits? I don't see how you can tell that from the Y2H assay. Please clarify.

      It means that these 3 subunits interact directly with Set1 without the need of another subunit, unlike of the other subunits.

      (5) Line 252. While discussing the Set1 - Snf2 interaction, the paper cites Hirschhorn et al. That paper talks about Swi-Snf, but doesn't mention Set1 anywhere. Maybe the authors meant to cite a different paper?

      We agree, this reference is not appropriated. It has been deleted.

      (6) Figure S2 panels A and C are redundant and could easily be combined.

      Figure S2 has been deleted.

      (7) Figure S4: Should the green category also include transcription? Ssl1 is a TFIIH subunit, which could be involved in either transcription initiation or NER. Sen1 and Nrd1 are transcription termination factors, although Sen1 may also function in R-loop resolution.

      We agree but it is already complicated as it is.

    1. eLife Assessment

      This study presents an important finding regarding the role of oxytocin neurons in thermogenesis and behavioral thermoregulation. The use of numerous converging methods, including behavior, fiber photometry, optogenetics, thermal recordings, metabolic analyses, and more, produces a multi-dimensional dataset delivering findings that provide solid support for the conclusions. The conclusions could be further strengthen by more extensive analyses of behavior and determining whether it is the release of oxytocin (rather than co-release of glutamate) from the PVN that is critical for the transition between behavioral states, nevertheless, the manuscript had many strengths, the findings are novel, and this work opens new doors for understanding the role of the PVT in thermoregulation. This work will be of strong interest to the thermoregulation, social behavior, and oxytocin signaling communities.

    2. Reviewer #1 (Public review):

      Summary:

      The authors identify and investigate a specific population of PVNOT neurons (oxytocin neurons of the paraventricular hypothalamus) that seem to be involved in both behavioral and autonomic thermoregulation. These cells are activated by social thermoregulatory behaviors, but can influence thermoregulation in both social and social contexts, specifically during transitions and when mice are at low core body temperature (Tb).

      Strengths:

      The manuscript has many strengths.

      This is a novel study, with a clear question that is addressed using an array of well-designed experiments employing integrative methods. Most of the Figures are well developed, and the analysis is generally rigorous and well detailed. The authors are clearly very experienced in this field, and indeed their scholarly introduction and discussion sections is in their credit.

      The link between thermoregulation and the oxytocin system is well established, as is the link between social behavioral and the same broad system. However, the link between these three things is novel, if it can be well substantiated. I am not persuaded that was achieved here, but I do think this manuscript has many novel and useful offerings.

      The authors use a cooling floor and only go town to 10 degrees Celsius. This is fine, but I would like to see the effects using ambient temperature also. This is not a crucial issue, as it is not necessary for the authors' interpretations, but it could improve measurement sensitivity.

      Through an elegant behavioral experiment in Fig. 1, the authors identify c-Fos patterns in the PVN that are activated by active social huddling, and they show that at the RNA level these cells overlap with oxytocin, indicating that they are oxytocin producing cells. But this is not well discussed or indeed quantified.

      The authors engage in deep analysis of fiber photometry experiments, first by observing PVNOT neuron overall activity during a variety of different behaviors in the context of three different temperatures. Activity was associated with nesting, quiescence, and both types of huddling (when social opportunities exist). Social situations did not strongly effect this, not did temperature conditions. These analyses indicate that the PVNOT neurons are involved in mediating specific behavioral outputs.

      With more detailed analysis, the authors investigated how PVNOT neuronal activity relate to behavioral state transition. They found that the probability of peak PVNOT neural activity strongly predicts the offset of quiescence or quiescent huddling and therefore can be argued to signal an increase in physical activity, and as such increased metabolism. However, the opposite pattern was observed for huddling and nesting (onset being associated with PVNOT activity), again arguing for increased thermogenesis as a function.

      What is particularly compelling is that these peaks of activity tend to occur during low Tb, again arguing for the function in increasing body warmth.

      The authors then employ an impressive set-up where they image brown dispose tissue (BAT) in tandem with DeepLabCut (DLC) based animal tracking. Crucially, BAT activity and surface temperature correlated with the calcium peak of PVNOT neurons.

      Lastly, optogenetic activation of PVNOT neurons increased Tb when it was in the lower range, but not when in the higher ranger. It also affected BAT and rump temperature, again at low Tb. However, there is no real affect on behavior, except a trend in activity.

      The authors do some interesting tracing work at the end, though this is not functionally explored. That's not a criticism as it does seem like this would be a follow-up whole study.

      Comments on revised version.

      As discussed before, the authors employ a wide range of techniques (FOS IHC, FP for fine scale PVN OXT population dynamics, behavioural analysis, core and surface temperature tracking, physiological recordings to assess AAV specificity, optogenetic activation of PVN OXT neurons, and projection tracing) to address a clear question. The outcomes of these techniques seem to drive the same conclusion that PVN OXT neurons signal transitions from rest to arousal (behavioural and thermogenic) in a state-dependent manner:

      - FOS data identifies PVN OXT population activity following behavioural onset

      - Ca activity in these cells peaks at behavioural and thermogenic state transitions

      - Rump temperature and BAT activity increase at state transition points

      - Optogenetic stimulation of these cells recapitulates the thermogenic effects seen during physiological state transitions (in low body temperature animals) with a trending increase in physical activity

      Despite the inconclusive IHC results when validating the specificity of their AAV, the virgin female/ lactation experiment is convincing that they are specifically targeting PVN OXT neurons. The rationale for this experiment is clearer in the revised manuscript.

      Generally, in terms of the revised manuscript, the authors give strong responses to reviewer comments, either incorporating feedback, or giving clear explanations for the choices they made in the original manuscript. The revised manuscript is clearer about the question the authors aim to address, the reasons for their choice of experiments, and the limitations of the techniques used.

      Criticisms:

      I appreciate and agree with the authors' point that this manuscript is more fundamental than simply social basis oxytocin neuron function. This is point is well made by their data, and in the revised text. However, I still believe more behavioural analysis would be welcome to any reader.

      They partly justify the lack of behavioural analysis in Figure 6 with the problem of "animal merging" on the SGBS images. However, in Figure 6C, they confirm that, in solo conditions, the SGBS readings are consistent with core body temperature readings. So why not stick to core body temperature, opto stimulate and analyse the social behaviour with DLC (with normal video recordings)?

      The lactation validation still seems out of place in manuscript order. It is a very valuable validation, but it feels more like supplementary data for Figure 1. I feel the authors wanted it as a main figure because of how much work it must have been. In that case, it still makes more sense to include it in Figure 1.

      Though their lactation experiment validates that they are targeting PVN OXT neurons, their optogenetic stimulation protocol may not be specifically inducing OXT release from these cells. PVN OXT neurons co-release glutamate but can also release glutamate independently of OXT following lower frequency tonic stimulation. OXT release from PVN neurons requires pulsatile stimulation at a higher frequency (Leithead et al., 2021; Piñol et al., 2014; Lincoln & Wakerley, 1975). In this paper, the authors use a low stimulation frequency (10Hz) and continuous pulse train (20s) to optogenetically manipulate the target PVN population which may bias the cells towards glutamate release over OXT. Therefore, though they find evidence that PVN OXT neurons are involved in driving the transition between states in their other experiments, their optogenetic stimulation may not necessarily involve OXT release/signalling. It may be valuable to separate this out to identify the signalling molecule underlying this behavioural/ thermogenic transition. This could be done by using an opto protocol that recapitulates physiological OXT release.<br /> The authors do however mention that isolating the specific contribution of OXT signalling compared to other co-transmitted molecules was not the aim of this study, so this is not an essential question for this manuscript.

      A loss of function experiment to test for sufficiency would be a nice addition to further confirm their claims, but the authors mention that there were technical limitations to their attempts at inhibiting PVN OXT neurons. I appreciate the authors declaring that the DREADDs attempt suffered from unfortunate confounds. But for optogenetic attempts, I don't think they need a closed-loop system to get some useful results. They still can shine the light at "random" moments (that will correspond to random body temperatures) and then separate the data per body temperature.

      Lastly, the mention of Raam et al. 2026 is insufficient. The authors just mention it regarding the potential differences with males, to be explored in future experiments. Even if not using males in the current study doesn't affect the stated conclusions, the fact that they chose females because "their thermo-behavioural states were readily discernible" is a considerable bias. Testing males in this very study might be out of scope, but more discussion is warranted.

      References

      Leithead, A. B., Tasker, J. G., & Harony-Nicolas, H. (2021). The interplay between glutamatergic circuits and oxytocin neurons in the hypothalamus and its relevance to neurodevelopmental disorders. Journal of neuroendocrinology, 33(12), e13061. https://doi.org/10.1111/jne.13061

      Lincoln, D. W., & Wakerley, J. B. (1975). Factors governing the periodic activation of supraoptic and paraventricular neurosecretory cells during suckling in the rat. The Journal of physiology, 250(2), 443-461. https://doi.org/10.1113/jphysiol.1975.sp011064

      Piñol, R. A., Jameson, H., Popratiloff, A., Lee, N. H., & Mendelowitz, D. (2014). Visualization of oxytocin release that mediates paired pulse facilitation in hypothalamic pathways to brainstem autonomic neurons. PloS one, 9(11), e112138. https://doi.org/10.1371/journal.pone.0112138

    3. Reviewer #2 (Public review):

      This is a very interesting study from Vandendoren and colleagues examining the role of PVN oxytocin neurons during thermoregulatory behaviors, in particular during thermoregulatory huddling. The findings are important and have implications for the thermoregulation field as well as the social/naturalistic behavior field. The findings are compelling and use a combination of state-of-the-art tools (photometry, optogenetics, automated behavior tracking, thermal imaging, and core body temperature measurement), often in combination with each other, to produce a rigorous and high-dimensional dataset.

      Comments on revised version.

      I appreciate the effort the authors have put into addressing all of my questions, and I have no remaining concerns.

    4. Reviewer #3 (Public review):

      Summary:

      This study investigates how the activity of hypothalamic paraventricular oxytocin (PVNOT) neurons relates to physiological states in female mice, with a particular focus on behavioral states and thermogenic sympathetic activity. To address this question, the authors combined automated video-based behavioral classification with calcium imaging of PVNOT neuron activity. Sympathetic thermogenesis was inferred from surface temperature changes measured by infrared thermography, and the authors have made their custom analysis scripts available. The authors report that strong, pulsatile activation of PVNOT neurons was "occasionally" observed immediately before transitions from resting to active states. This observation suggests that PVNOT neuronal activity may facilitate the transition from rest to activity. This phenomenon was observed in both pair-housed and individually housed animals. Taken together, these findings raise the possibility that the oxytocinergic system contributes to naturalistic behavior transitions even in the absence of social interactions. However, concerns regarding the selectivity of GCaMP expression in oxytocin-expressing neurons call into question the validity of the recorded PVNOT neuronal activity.

      Strengths:

      The oxytocinergic neural system is believed to subserve a wide range of physiological functions. Elucidating these roles requires monitoring PVNOT neuronal activity under diverse behavioral contexts, as well as manipulating this activity to establish causal relationships. In this study, the authors present a technically sound experimental framework that integrates behavioral tracking in both individually and group-housed mice with the monitoring and manipulation of PVNOT neuron activity. This setup represents a valuable methodological resource for researchers investigating the physiological functions of oxytocin.

      Weaknesses:

      (1) Immunohistochemical validation of selective GCaMP expression in oxytocin-expressing neurons showed that only 24-51% of GCaMP-positive neurons expressed oxytocin. As an alternative approach, the authors demonstrate that GCaMP-expressing PVN neurons in virgin females exhibit calcium peaks during rest-wake transitions with kinetics similar to those observed in PVNOT neurons during early lactation. However, this comparison is based solely on population-level peak profiles and does not provide direct evidence for cell-type specificity of GCaMP expression in oxytocin neurons. This limitation substantially undermines the validity of the optical calcium imaging data. In situ hybridization targeting oxytocin mRNA, rather than immunohistochemistry, may provide a more reliable assessment of expression specificity.

      (2) Although the authors' interpretation is generally consistent with the data presented, their main conclusions rely heavily on observational findings. Moreover, optogenetic stimulation of PVNOT neurons failed to robustly recapitulate behavioral state transitions (Figs. 6D and S5B). Further interventional experiments will be necessary to more rigorously test the authors' interpretation and to establish mechanistic insight into the causal relationship between PVNOT activity and rest-to-active transitions. In particular, loss-of-function approaches targeting the PVNOT system, such as OXTR antagonism, inhibitory DREADDs, or cell-type-specific ablation, will be essential to determine whether perturbation of this system alters behavioral state transitions These points should be addressed in future studies.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors identify and investigate a specific population of PVNOT neurons (oxytocin neurons of the paraventricular hypothalamus) that seem to be involved in both behavioral and autonomic thermoregulation. These cells are activated by social thermoregulatory behaviors, but can influence thermoregulation in both social and nonsocial contexts, specifically during transitions and when mice are at low core body temperature (Tb).

      Strengths:

      The manuscript has many strengths.

      This is a novel study, with a clear question that is addressed using an array of well-designed experiments employing integrative methods. Most of the figures are well-developed, and the analysis is generally rigorous and well-detailed. The authors are clearly very experienced in this field, and indeed, their scholarly introduction and discussion sections are to their credit.

      We are grateful for the reviewer’s careful reading and positive assessment, including their remarks on the clarity of the question, experimental design, and analysis.

      The link between thermoregulation and the oxytocin system is well established, as is the link between social behavior and the same broad system. However, the link between these three things is novel, if it can be well substantiated. I am not persuaded that was achieved here, but I do think this manuscript has many novel and useful offerings.

      We thank the reviewer for this thoughtful comment and for recognizing the novelty of the study. We wish to clarify the central goal of the manuscript: while social thermoregulation provided the initial influence for studying PVNOT neurons, our principal finding is that PVNOT activity during rest-to-arousal transitions is independent of social context. As stated in the manuscript, "To our surprise, these peaks were observed in both social and non-social contexts." Thus, our study demonstrates a broader role for PVNOT neurons in state-dependent thermoregulatory transitions—one that includes, but is not limited to, social contexts. We have revised the text to make this emphasis clearer throughout.

      We also added a short piece to the Discussion on this point. This is the fourth and final paragraph of the Discussion section called “State-dependent PVNOT activity during thermo-behavioral transitions.”

      The authors use a cooling floor, and only go down to 10 degrees Celsius. This is fine, but I would like to see the effects using ambient temperature also. This is not a crucial issue, as it is not necessary for the authors' interpretations, but it could improve measurement sensitivity.

      Both Reviewer 1 and Reviewer 2 raise important and related points: manipulating floor temperature provides a thermal stimulus that is distinct from manipulating whole-chamber ambient air temperature, and these modalities could engage partially different sensory pathways and circuits. (Note this response is copy-pasted to other relevant comments).

      We intentionally used floor cooling/heating because it provides a reliable, well-controlled stimulus that elicits thermoregulatory behaviors while keeping the experimental environment stable (e.g., avoiding changes in airflow/humidity that can accompany ambient cooling). To prevent conflation of these modalities, we revised the manuscript to consistently describe the manipulation as “floor temperature” (and not “ambient temperature”), and we added to the Discussion acknowledging that conductive floor temperature changes may differentially recruit peripheral thermoreceptors compared to ambient air temperature.

      While extending these experiments to whole-chamber ambient temperature changes could be informative in future work, it is not required for the central interpretations here, which focus on PVNOT activity dynamics during thermoregulatory behavior under controlled thermal conditions.

      Through an elegant behavioral experiment in Figure 1, the authors identify c-Fos patterns in the PVN that are activated by active social huddling, and they show that at the RNA level these cells overlap with oxytocin, indicating that they are oxytocin-producing cells. But this is not well discussed or indeed quantified.

      We thank the reviewer for catching this; Reviewer 2 made a similar comment. A typo in the figure legend led to this confusion. Figure 1I is in fact a quantification of the percent Oxytocin:Fos colocalized cells (not Fos:DAPI, as was written) in dorsal and ventral subregions of the PVN during active huddling and quiescent huddling. We have corrected the legend and clarified the quantification in the revised manuscript.

      The authors engage in a deep analysis of fiber photometry experiments, first by observing PVNOT neuron overall activity during a variety of different behaviors in the context of three different temperatures. Activity was associated with nesting, quiescence, and both types of huddling (when social opportunities exist). Social situations did not strongly affect this, nor did temperature conditions. These analyses indicate that the PVNOT neurons are involved in mediating specific behavioral outputs.

      With more detailed analysis, the authors investigated how PVNOT neuronal activity relates to behavioral state transition. They found that the probability of peak PVNOT neural activity strongly predicts the offset of quiescence or quiescent huddling, and therefore can be argued to signal an increase in physical activity, and as such, increased metabolism. However, the opposite pattern was observed for huddling and nesting (onset being associated with PVNOT activity), again arguing for increased thermogenesis as a function.

      What is particularly compelling is that these peaks of activity tend to occur during low Tb, again arguing for the function in increasing body warmth.

      The authors then employ an impressive setup where they image brown adipose tissue (BAT) in tandem with DeepLabCut (DLC) based animal tracking. Crucially, BAT activity and surface temperature correlated with the calcium peak of PVNOT neurons.

      Lastly, optogenetic activation of PVNOT neurons increased Tb when it was in the lower range, but not when in the higher range. It also affected BAT and rump temperature, again at low Tb. However, there is no real effect on behavior, except a trend in activity.

      The authors do some interesting tracing work at the end, though this is not functionally explored. That is not a criticism, as it does seem like this would be a whole follow-up study.

      Weaknesses:

      While novel and valuable, the manuscript feels incomplete in its current form.

      The main evidence lacking is a loss of function of the experiment. Ideally, the authors would chronically and/or acutely inhibit PVNOT neurons to establish their necessity. I know this seems obvious, but I think it is important.

      We agree with the reviewer that loss-of-function experiments are a valuable component of circuit mapping and we appreciate this suggestion. For transparency, we did attempt a chronic chemogenetic inhibition experiment using DREADDs in PVNOT neurons. However, the results were inconclusive, primarily owing to the confounding effects of pharmacological injections: both drug and vehicle-treated animals exhibited stress-induced hyperthermia following injection, and because inhibition could not be delivered while animals were asleep/resting the experimental conditions did not recapitulate the low-Tb quiescent state during which PVNOT peaks naturally occur. Given these confounds, we do not believe these data meet the standard required for inclusion in this manuscript.

      We did consider acute optogenetic inhibition. However, a clear prediction about inhibition was not as apparent in our model. Our photometry data identified a, testable hypothesis for activation: PVNOT peaks precede the exit from quiescence, therefore activation during quiescence should increase the transition, which it did (Figures 5 and 6).

      That said, new analyses of our data, driven by these reviews, have now uncovered what might be inhibition of PVNOT neurons during the approximate 60 seconds prior to entry to resting states (i.e., quiescence and quiescent huddling); see the new Fig. S3I-L. This raises the possibility that an appropriately timed photoinhibition of PVNOT neurons could facilitate the establishment of resting states. We believe that, in light of our chemogenetic and optogenetic activation experiments, for an inhibition experiment to be done appropriately would require a real-time, closed loop setup that is currently not available in our laboratory.

      We have added a caveat to the Discussion acknowledging the lack of LOF data as a limitation and have identified this as an important direction for future investigation.

      The relative lack of behavioral analysis following optogenetic activation of PVNOT neurons is puzzling. The authors must surely want to study what this intervention does to behavioral state transitions. I feel that the current level of analysis limits the overall conclusions of this study to a large extent.

      We appreciate this concern and wish to clarify two points.

      First, our decision to perform optogenetic activation in isolated (solo-housed) animals was driven by our initial finding that PVNOT activity profiles are mostly social-context independent during the transition from rest to arousal (Figures 2 and 3). By studying isolated animals, we could test the fundamental relationship between PVNOT activation and the rest-to-active transition without confounding social feedback. Additionally, we encountered technical challenges when using the SGBS thermographic model in paired contexts: the high thermal intensity at the point of contact between huddling mice created a thermal merging artifact that prevented accurate segmentation of individual body regions (BAT vs. rump).

      Second, we did examine the post-stimulation behaviors of solo-housed animals (Fig. S5B). While PVNOT activation significantly increased the probability of exiting quiescence, it did not trigger a singular, stereotyped behavioral output. Instead, it facilitated a generalized transition to an active state, within which animals engaged in various context-appropriate actions (nesting, grooming, locomotion). We note in the discussion that “Analysis of manually- annotated behaviors suggested that PVNOT stimulation did not activate a specific motor pattern output but instead resulted in combined increases in the time spent in nesting (linear mixed model estimate coefficient of ChR2+ stimulation: +38.0 sec), locomotion (+54.0 sec), and grooming (+14.5 sec), but not in eating/drinking (-0.4 sec) (Fig. S4B).”

      That photostimulation had relatively larger effects on nesting and locomotion is consistent with our model.

      Last, in the Discussion we acknowledge that future experiments should seek to disentangle the effects of PVNOT light simulation in the non-social vs social context (last paragraph of the Discussion section called “State-dependent PVNOT activity during thermo-behavioral transitions”).

      A broader criticism is that the social dimension of this manuscript seems overplayed. Naturally, oxytocin signaling can be implicated in social behavior based on a large literature. However, the focus on social thermogenesis seems like a crude integration of social behavior and thermogenesis. Given that the authors see their effects in both social and nonsocial cases of thermoregulation, I am not sure the attempts at integrating social functions and thermogenic functions of PVNOT neurons are warranted. That is, unless the authors have further experiments or analysis that can convincingly justify this link.

      We thank the reviewer for this comment. We understand the concern and wish to reframe our position. We argue that the equivalence of PVNOT signals across social and non-social contexts is itself a central finding. While the oxytocin system is widely regarded as a mediator of social bonding, and therefore a candidate mechanism underlying huddling, our data demonstrate that PVNOT neurons provide a signal for state-dependent thermoregulatory transitions that is unbiased by social context. Rather than overplaying the social dimension, we believe our study contextualizes the social function within a broader homeostatic role: PVNOT neurons facilitate transitions from rest to thermogenesis and arousal regardless of whether the resting state involves social huddling or solitary quiescence.

      While the thermoregulatory transitions are present in both contexts, we note that social context appears to modestly enhance some PVNOT downstream effects. Specifically, peak probability and frequency were slightly higher in the paired compared to solo context (Fig. 3F-I, Fig. S2D), and peaks were associated with a somewhat stronger increase in physical activity when a cagemate was present (Fig. 3B-E). Additionally, quiescent huddling (paired) bouts were associated with stronger body temperature regulation compared to solo quiescence (Fig. S3Q-V). This nuance supports that the social dimension is not overplayed but rather situated within a broader homeostatic function.

      We have revised the manuscript to ensure that this framing is consistent and clear. We emphasize that our goal was to uncover neural mechanisms underlying physiological transitions across behavioral and arousal states, using our social thermoregulation assay as a starting point (based on our previous publication). Counter to our initial hypothesis, the PVNOT signals generalized beyond the social setting.

      In addition, the analysis of virgin females and lactating mothers seems out of place in Figure 4.

      This point was echoed by Reviewers 1 and 3, and one we have taken several actions to address this. (Note this response is copy-pasted to the other reviewers).

      We agree with the reviewers that the rationale for the lactation data should be made more explicit. The primary purpose of this experiment was to validate the identity of oxytocinergic neurons of the PVN.

      Our efforts to use IHC to validate the identity of AAV-transfected cells were inconclusive, and we have now added new data to illustrate this point. We have added Fig. S4 that includes quantitative data on expression specificity. We observed significant variability in co-staining (OT+/GCaMP+) across brain slices, likely reflecting the dynamic nature of oxytocin peptide synthesis and storage, particularly with respect to processes lining the third ventricle. This finding is in accordance with other studies that are now cited in the text.

      We now emphasize that, because IHC provided variable co-localization, we employed the lactation model as an independent physiological validation of the identity of the recorded neurons.

      It is well established that PVNOT neurons undergo dramatic changes in firing dynamics and synchrony during lactation to support milk ejection (Yaguchi et al., 2023; Yukinaga et al., 2022). Conversely, AVP and CRF cell populations in the PVN do not appear to display synchronized pulsatile bursting during lactation (see response to Reviewer-2 comment-2 in ‘Recommendation for authors’ and our updated Discussion). Observing these characteristic changes in our recorded population provides high-confidence functional evidence that we are targeting oxytocin neurons. We have revised the text to clarify that Figure 4 serves primarily as a functional verification of genetic targeting.

      We also acknowledge in the Discussion the possibility that our Cre-line may capture a small percentage of nonoxytocinergic neurons, while noting that the dramatic shift in calcium dynamics during lactation (Figure 4I–L) strongly suggests the recorded population is dominated by oxytocin neurons.

      The c-Fos/oxytocin overlap needs to be quantified.

      We thank the reviewer for catching this; Reviewer 2 made a similar comment. A typo in the figure legend led to this confusion. Figure 1I is in fact a quantification of the percent Oxytocin:Fos colocalized cells (not Fos:DAPI, as was written) in dorsal and ventral subregions of the PVN during active huddling and quiescent huddling. We have corrected the legend and clarified the quantification in the revised manuscript. (Note this response is copy-pasted to other relevant comments).

      The methods section could be improved by explaining how the authors exclude animals that exhibit both types of huddling, if they occur within a 90-minute time window. This seems like it could cause significant confounds.

      We have clarified in the Methods that animals were not excluded if they exhibited both active and quiescent huddling during the recording session. Importantly, a prerequisite for inclusion in the FOS study was that animals had to be continually engaged in the target behavior for a minimum of 15 consecutive minutes from behavior onset, an established approach for behavior-driven immediate early gene mapping. The 90-minute window was then counted from that same onset for FOS IHC. Because active huddling frequently transitions directly into quiescent huddling (and vice versa), excluding such animals would have eliminated the majority of recordings. The heterogeneity of behavioral states within the FOS integration windows is precisely why we turned to fiber photometry, a technique with the temporal resolution necessary to dissociate neural signals associated with each behavioral state.

      The computer vision model is not well-explained. The authors need to be far more explicit here about how it was validated.

      We thank the reviewer for this comment and agree that the original manuscript did not sufficiently detail the validation framework. We have revised both the Methods and Results to explicitly detail how SGBS was evaluated.

      First, we now clearly describe model validation on a held-out dataset (20% of manually annotated images not used for training), reporting standard segmentation metrics (per-class IoU and Dice/F1) and directly comparing SGBS to an unmodified Mask R-CNN trained under identical conditions (same backbone initialization, dataset split, and training schedule). As shown in Fig. 5D, the skeleton-guided model converged more rapidly and achieved a lower final loss than the baseline network, demonstrating improved segmentation performance in occlusion-rich thermographic recordings.

      Second, we more explicitly describe an independent physiological validation step. SGBS-derived surface temperature trajectories were temporally aligned with simultaneously recorded implanted thermologger measurements, which were not used during model training. As shown in Fig. 5E, SGBS-derived signals strongly corresponded with core body temperature dynamics and reproduced expected thermophysiological relationships (e.g., BAT warming preceding core temperature rise). This establishes external validity beyond pixel-level segmentation metrics.

      The authors should cite and consider this preprint: https://www.biorxiv.org/content/10.1101/2024.09.17.613378v1

      We have cited this preprint (Raam et al., 2024) in the revised manuscript and integrated relevant findings into the Discussion, in the section called “Limitations and caveats”.

      Reviewer #2 (Public review):

      Summary:

      This is a very interesting study from Vandendoren and colleagues examining the role of PVN oxytocin neurons during thermoregulatory behaviors, in particular during thermoregulatory huddling. The findings are important and compelling, and have implications for the thermoregulation field as well as the social/naturalistic behavior field.

      Strengths:

      The study is very creative and tackles a challenging task to examine how natural and social behavior influences neural circuits for a homeostatic system such as thermoregulation. The authors use a combination of state-of-the-art tools (photometry, optogenetics, automated behavior tracking, thermal imaging, and core body temperature measurement), often in combination with each other, to produce a rigorous and high-dimensional dataset. Carrying out tightly temperature-controlled experiments and examining natural behavior, neural activity, and body physiology simultaneously is quite a feat. I applaud the authors for taking this on in a rigorous and detailed manner. This paper will be valuable for both the thermoregulation field as well as for researchers interested in naturalistic social behaviors. The conclusions are supported by the data.

      We appreciate the reviewer’s careful read and positive assessment of our integrated behavioral, neural, and physiological measurements and their relevance to both thermoregulation and social behavior.

      Weaknesses:

      I have a number of questions and suggestions for clarification that would help improve the interpretation of the findings.

      (1) Figure 1D-F: It would be helpful to include representative images of cFos expression in the PVN, LS, and DMH during both quiescent and solo huddling conditions, to better illustrate the reported differences.

      We have now addressed this in the revised manuscript. We had originally shown active huddle FOS expression in Fig. 1D-F and quiescent huddle in Fig. S1A-C. We have now added solo groom FOS expression to Fig. S1D-F.

      (2) Figure 1C: The data suggest a general suppression of neural activity during sleep-associated quiescent huddling, which somewhat complicates the interpretation of what specifically the active huddling cells are responding to. A more informative control might have been a comparison between huddling and a more generic form of social engagement (e.g., dyadic sniffing) to assess whether huddling-responsive neurons are broadly tuned to social stimuli. While it may not be feasible to add this experimentally at this time, a brief discussion of this limitation in the main text would be valuable.

      We thank the reviewer for this thoughtful suggestion. We agree that comparing huddling-responsive neurons with a more generic social engagement is an important consideration.

      We first note that the FOS study required animals to be continuously engaged in the target behavior for a minimum of 15 consecutive minutes, ensuring that FOS expression reflects sustained behavioral engagement rather than brief social contact. Furthermore, we believe the FOS association with active huddling in Figure 1C is likely driven by preceding bouts of quiescent huddling. Because these experiments were conducted during the light phase, active huddling bouts were almost always preceded by bouts of quiescent huddling.

      Given that FOS protein often integrates neural activity over ~60-90 minutes, the FOS signal during active huddling may reflect cumulative PVNOT activity during the quiescent to active transition, rather than active huddling by itself. This interpretation aligns with our fiber photometry data, which show that PVNOT peaks are concentrated at the offset of quiescent states and the onset of active states. Moreover, a broad-scale analysis of calcium data driven by these reviews, now shows there is a local minimum of PVNOT neurons during the transition into quiescent states and a local maximum of calcium activity during the offset of resting states and the onset of nesting and active huddling (Fig. S3I-L).

      To directly address whether PVNOT neurons are broadly tuned to social engagement or specifically associated with thermoregulatory state transitions, we examined neural activity during "Contact Initiated" (ConI) and "Contact Received" (ConR) events—brief social interactions (e.g., dyadic sniffing) that occur outside the context of huddling. These interactions, which typically last less than one second, did not trigger the large-amplitude calcium peaks observed during rest-to-arousal transitions. Specifically, there was no significant association between ConI or ConR events and PVNOT peak frequency or amplitude (Fig. S2H; Table S1; p = 0.505, p = 0.575, respectively). This reinforces our conclusion that PVNOT peaks are not a generic response to social stimuli but are specifically aligned with the coordinated autonomic and behavioral transitions required to exit a low-temperature quiescent state. We have added a clarifying paragraph to the Discussion.

      (3) Figure 2H-J vs. Figure 1: The fiber photometry data suggest increased PVN activity during quiescent huddling vs active huddling, which appears to contrast with the cFos results from Figure 1. It would be helpful for the authors to comment on possible reasons for this discrepancy-e.g., methodological differences, temporal resolution, or cell-type specificity.

      We agree that this apparent contrast deserves explicit discussion. The difference arises from the dramatically different temporal resolutions of the two techniques. Fiber photometry captures real-time neural dynamics at subsecond resolution, revealing that PVNOT neurons exhibit high-amplitude bursts primarily during the offset of quiescence (and to a lesser extent the onset of post-quiescence behaviors) (Figs. 3 and 5). Because these peaks occur while the animal is categorized as "quiescent," they appear as quiescence-associated activity in the photometry ethogram.

      Conversely, FOS integrates neural activity over ~30–90 minutes. In retrospect, and in light of our photometry data, an animal categorized as "Active Huddling" in the FOS study is one that has likely experienced PVNOT bursts and subsequently transitioned to an active state. The higher FOS signal in active animals therefore likely represents the cumulative activity of the transition itself and sustained activity in the active state.

      We have added a clarifying statement to the Discussion section, in the section called “State-dependent PVNOT activity during thermo-behavioral transitions”.

      (4) Figure 2O: A comparable linear regression for active huddling would be informative to assess whether the observed relationships extend across behavioral states.

      We agree. We have added linear regression analyses for active huddling and nesting to Fig. S2K-N including rsquared values, to complement the resting analyses in Figure 2O and 2L.

      This analysis shows that active huddling peak counts are also positively correlated with active huddle duration (but not nesting duration). The text has been updated accordingly.

      (5) Temperature manipulation: The use of floor temperature changes presents a distinct physiological and sensory experience from, for example, manipulation of ambient temperature. A discussion of how this choice may affect neural circuit engagement or interpretation of thermoregulatory responses would be beneficial.

      Both Reviewer 1 and Reviewer 2 raise important and related points: manipulating floor temperature provides a thermal stimulus that is distinct from manipulating whole-chamber ambient air temperature, and these modalities could engage partially different sensory pathways and circuits. (Note this response is copy-pasted to other relevant comments).

      We intentionally used floor cooling/heating because it provides a reliable, well-controlled stimulus that elicits thermoregulatory behaviors while keeping the experimental environment stable (e.g., avoiding changes in airflow/humidity that can accompany ambient cooling). To prevent conflation of these modalities, we revised the manuscript to consistently describe the manipulation as “floor temperature” (and not “ambient temperature”), and we added Discussion acknowledging that conductive floor temperature changes may differentially recruit peripheral thermoreceptors compared to ambient air temperature.

      While extending these experiments to whole-chamber ambient temperature changes could be informative in future work, it is not required for the central interpretations here, which focus on PVNOT activity dynamics during thermoregulatory behavior under controlled thermal conditions.

      (6) Correlations with behavior: Across the manuscript, it would be informative to see correlations between huddle duration and neural activity (e.g., cFos expression, calcium signal magnitude). Similarly, do longer huddles produce greater thermogenic effects?

      This is a great suggestion. The first point about huddle duration and neural activity echoes the Reviewer’s comment (4) above. For this point, we now show that the duration of active huddling is positively correlated with PVNOT peak count (Fig. S2K), which is similar to what we had shown for quiescence and quiescent huddling (Fig. 2K-P).

      Next, the point about huddle duration and thermogenic effects is also helpful. We have now added new analysis and panels to address this (Fig. S3M-R). We find that the duration of quiescent huddle bouts is negatively correlated with Tb (Fig. S3V). The other behaviors examined did not show correlations between duration and Tb. This finding supports our previous demonstration that quiescent huddling is an energy saving state in mice (Landen et al., 2024).

      Finally, we note that longitudinal correlations between bout length and peak counts are already reported in Fig. S3A-H.

      (7) Lactating vs. virgin mothers: The inclusion of maternal data is intriguing but feels somewhat disconnected from the central huddling-thermoregulation narrative. If these experiments are to remain, additional explanation of their rationale and how they fit into the broader story would help clarify their relevance.

      This point was echoed by Reviewers 1 and 3, and one we have taken several actions to address this.

      We agree with the reviewers that the rationale for the lactation data should be made more explicit. The primary purpose of this experiment was to validate the identity of oxytocinergic neurons of the PVN.

      Our efforts to use IHC to validate the identity of AAV-transfected cells were inconclusive, and we have now added new data to illustrate this point. We have added Fig. S4 that includes quantitative data on expression specificity. We observed significant variability in co-staining (OT+/GCaMP+) across brain slices, likely reflecting the dynamic nature of oxytocin peptide synthesis and storage, particularly with respect to processes lining the third ventricle. This finding is in accordance with other studies that are now cited in the text.

      We now emphasize that, because IHC provided variable co-localization, we employed the lactation model as an independent physiological validation of the identity of the recorded neurons.

      It is well established that PVNOT neurons undergo dramatic changes in firing dynamics and synchrony during lactation to support milk ejection (Yaguchi et al., 2023; Yukinaga et al., 2022). Conversely, AVP and CRF cell populations in the PVN do not appear to display synchronized pulsatile bursting during lactation (see response to Reviewer-2 comment-2 in ‘Recommendation for authors’ and our updated Discussion). Observing these characteristic changes in our recorded population provides high-confidence functional evidence that we are targeting oxytocin neurons. We have revised the text to clarify that Figure 4 serves primarily as a functional verification of genetic targeting.

      We also acknowledge in the Discussion the possibility that our Cre-line may capture a small percentage of non-oxytocinergic neurons, while noting that the dramatic shift in calcium dynamics during lactation (Figure 4I–L) strongly suggests the recorded population is dominated by oxytocin neurons.

      (8) Optogenetic manipulation: Have the authors tested the effect of PVN OT neuron stimulation or inhibition during huddling? Even a negative result would be of interest to the field. If these data exist (main or supplementary), I apologize for missing them. If not, the authors might consider including them or commenting briefly on any attempts or challenges in carrying out these experiments.

      We thank the reviewer for this question. We have not performed optogenetic manipulation during huddling. Our decision to perform optogenetic activation in solo-housed animals was driven by our fiber photometry finding that PVNOT activity profiles during the rest-to-arousal transition are social-context independent (Figures 2 and 3). Had the GCaMP data suggested that PVNOT peaks were specific to social huddling, optogenetic manipulation during huddling would have been the natural next experiment. However, because peaks aligned with thermoregulation broadly, rather than social behavior specifically, we designed our functional experiments to test the circuit's role in driving the autonomic and behavioral arousal transition.

      We also note that our experience with chemogenetic manipulation suggests that pharmacological approaches to study the rest-arousal transitions during huddling are not currently feasible. As described to our response to Reviewer 1, our DREADD inhibition experiments were confounded by stress-induced hyperthermia following injection, and because drug delivery could not occur while animals were asleep and resting, the experimental conditions failed to recapitulate the low-Tb quiescent state during which PVNOT peaks naturally occur. We share this experience because we believe it will be informative for others in the field considering similar approaches.

      Additionally, as described above (Reviewer 1, #5), the SGBS thermographic model encounters artifacts in paired contexts due to thermal merging between huddling mice. We have added a note in the Discussion addressing this, in the section called “Limitations and caveats”.

      Reviewer #3 (Public review):

      Summary:

      The authors aimed to elucidate the relationship between physiological state (i.e., behavioral status and thermogenic sympathetic activity) and the activity of hypothalamic paraventricular oxytocin (PVNOT) neurons in female mice. They studied this by combining automated classification of mouse behavior via video-based analysis with calcium imaging of PVNOT neuron activity. Sympathetic thermogenesis was inferred from surface temperature changes captured by infrared thermography, and the authors provided their custom analysis scripts in the manuscript. Notably, they found that a strong, pulsatile activation of PVNOT neurons was "occasionally" observed immediately before the animals transitioned from a resting to an active state. This pulsatile activity was observed in both pair-housed and individually housed animals. While PVNOT neurons are often associated with social behaviors, this finding suggests that the oxytocinergic system is also engaged during naturalistic behaviors, even in the absence of social interactions. If experiments were more convincingly performed and presented, the results would point to a broader physiological role of central oxytocin, including in the regulation of fundamental brain states and homeostatic processes, and offer a new perspective on the functional significance of central oxytocin signaling.

      Strengths:

      The oxytocinergic neural system is believed to subserve a wide range of physiological functions, and elucidating these roles requires monitoring PVNOT neuronal activity under various behavioral contexts, as well as manipulating this activity to establish causal links. In the present study, the authors show a technically sound experimental framework that integrates behavioral tracking in both individually and group-housed mice with the observation and manipulation of PVNOT neuron activity. This experimental setup represents a valuable methodological resource for researchers investigating the physiological functions of oxytocin.

      We thank the reviewer for the thoughtful review and for recognizing the value of our integrated framework for monitoring and manipulating PVNOT neuronal activity across behavioral contexts.

      Weaknesses:

      While this study successfully established a new experimental setup for simultaneous analyses of behavior and PVNOT neuronal activity, there are several concerns regarding the interpretation of the results and the robustness of the conclusions, which should be more thoroughly addressed.

      (1) The study relies on the assumption that calcium imaging and optogenetic manipulation were restricted only to PVNOT neurons. However, the specificity of AAV-mediated gene expression was not verified quantitatively. A fair number of cell bodies in the PVN expressed GCaMP8s, but not OT, indicating potential off-target expression (see Figure S2A, B). The lack of quantitative validation weakens confidence in the causal interpretation of the results.

      This point was echoed by Reviewers 1 and 3, and one we have taken several actions to address this.

      We agree with the reviewers that the rationale for the lactation data should be made more explicit. The primary purpose of this experiment was to validate the identity of oxytocinergic neurons of the PVN.

      Our efforts to use IHC to validate the identity of AAV-transfected cells were inconclusive, and we have now added new data to illustrate this point. We have added Fig. S4 that includes quantitative data on expression specificity. We observed significant variability in co-staining (OT+/GCaMP+) across brain slices, likely reflecting the dynamic nature of oxytocin peptide synthesis and storage, particularly with respect to processes lining the third ventricle. This finding is in accordance with other studies that are now cited in the text.

      We now emphasize that, because IHC provided variable co-localization, we employed the lactation model as an independent physiological validation of the identity of the recorded neurons.

      It is well established that PVNOT neurons undergo dramatic changes in firing dynamics and synchrony during lactation to support milk ejection (Yaguchi et al., 2023; Yukinaga et al., 2022). Conversely, AVP and CRF cell populations in the PVN do not appear to display synchronized pulsatile bursting during lactation (see response to Reviewer-2 comment-2 in ‘Recommendation for authors’ and our updated Discussion). Observing these characteristic changes in our recorded population provides high-confidence functional evidence that we are targeting oxytocin neurons. We have revised the text to clarify that Figure 4 serves primarily as a functional verification of genetic targeting.

      We also acknowledge in the Discussion the possibility that our Cre-line may capture a small percentage of nonoxytocinergic neurons, while noting that the dramatic shift in calcium dynamics during lactation (Figure 4I–L) strongly suggests the recorded population is dominated by oxytocin neurons.

      (Note, we have updated Figure S2A,B to more accurately reflect the extent of co-localization in this image).

      (2) The study focuses on the transition from rest to active states following pulsatile activity of PVNOT neurons. However, the physiological significance of this pulsatile activity remains unclear. According to the authors, pulsatile activity occurred with an approximately 20% probability within 100 seconds prior to the end of the resting state. This implies that, in the remaining 80% of rest-to-active transitions, pulsatile PVNOT activity did not occur, suggesting that it is not essential for initiating the transition. A comparative analysis of behavioral and thermogenic changes between transitions with and without pulsatile PVNOT activity would help to further clarify the functional relevance of this phenomenon and strengthen the authors' interpretation of the findings.

      These are excellent points, and here we address them separately.

      (1) probability of transitions.

      We agree that our wording could be misread and we have revised the text for clarity. The “~20%” value is not the fraction of rest-to-active transitions that exhibit pulsatile PVNOT activity within a 100-s window. Instead, Fig. 3F,H report an instantaneous (per-second) probability of observing a calcium peak as a function of time-to-bout offset (logistic regression). In other words, the probability of a peak increases sharply as the animal approaches rest offset (e.g., from ~2–3%/s near onset to ~14%/s for quiescence and ~25%/s for quiescent huddling near offset), indicating a strong state-dependent increase in peak likelihood rather than an all-or-none trigger.

      We further clarify in the Discussion that we do not claim PVNOT peaks are essential for initiating every transition; rather, PVNOT activity biases or enhances the probability of transition toward thermogenesis and behavioral arousal (added to section called “State-dependent PVNOT activity during thermo-behavioral transitions”).

      (2) the effect of peaks on transitions

      This is a very helpful suggestion and we agree that directly comparing transitions with vs. without pre-offset pulsatile PVNOT activity could strengthen interpretation of the functional relevance of these events. We have therefore added a new transition-aligned analysis of thermogenic dynamics at rest-to-active transitions (new Fig. 3P&S; and corresponding text in the Results and Statistics sections).

      Briefly, we extracted peri-transition body temperature (Tb) traces (−300 to +300 s) aligned to the offset of quiescence and quiescent-huddling bouts and classified each transition as Peak+ if it contained one or more calcium peaks in the 100 s preceding bout offset, and Peak− otherwise. To account for inter-individual differences in “balance point,” Tb was z-scored within mouse. We then quantified the post-offset thermogenic rise for each transition as the change in scaled temperature from a pre-offset baseline (−60 to 0 s) to the post-offset interval (0 to 300 s) and tested Peak+ vs Peak− differences using linear mixed-effects models. This revealed that Peak+ transitions exhibited significantly larger post-offset increases in scaled Tb than Peak− transitions for both quiescence offsets and quiescent-huddling offsets.

      Together, these results indicate that while pulsatile PVNOT activity is not present prior to every rest-to-active transition, when it occurs it is associated with a stronger thermogenic rise, consistent with a probabilistic modulatory role in promoting the transition rather than being strictly required to initiate it.

      We are grateful for this suggestion as this new data is very informative in the context of our model.

      (3) The study identifies a correlation between pulsatile activity of PVNOT neurons and rest-to-active transitions, and tests for a causal relationship using optogenetic stimulation. However, since PVNOT neurons are known to co-release other neurotransmitters such as glutamate, it remains unclear whether the observed effects are mediated specifically through oxytocin receptor signaling. To address this question, functional intervention experiments using oxytocin receptor antagonists or receptor knockout mice are necessary.

      We agree with the reviewer that PVNOT neurons co-release glutamate and that isolating the specific contribution of oxytocin signaling versus co-transmitted signals is an important question. However, our study was designed to identify the functional role of the PVNOT cell type during thermoregulatory state transitions, not to dissect the molecular mechanism of signaling at downstream targets. By demonstrating that the endogenous activity of this specific population aligns with the rest-arousal window and that their activation is sufficient to drive the phenotype, we provide an anatomical and functional framework for future mechanistic investigations.

      We also note that we provide anatomical evidence supporting a possible peptidergic mechanism: PVNOT neuron projections to the rostral medullary raphe (rMR), a key thermogenic control site, alongside oxytocin receptor mRNA expression in this region (Fig. S5). This anatomical link suggests a plausible pathway for oxytocinergic modulation of thermogenesis, but of course does not rule in/out glutamatergic signaling. We acknowledge this limitation in the Discussion and frame pharmacological and receptor knockout studies as important next steps.

      We address these points in the Discussion, in the section called “Limitations and caveats.”

      (4) The authors attempted to detect BAT thermogenesis and skin vasomotion using infrared thermography. This technique measures only skin hair temperatures (since the skin was not shaved), but does not measure "BAT temperature" or "vasomotor tone". As seen in Figure 5E, the temperatures of the body surface areas ("BAT", "Rump", and "Dorsal surface") mostly changed in parallel, indicating that these temperatures are strongly affected by body core temperature. Therefore, the thermographic measurements in this study did not provide convincing information on BAT thermogenesis or skin vasomotion. To avoid misleading reports, the authors need to use other techniques to directly measure temperatures, such as telemetry.

      We agree that infrared thermography measures surface radiance rather than internal tissue temperature. We have revised the manuscript to use more precise language (e.g., "surface temperature over the interscapular BAT region" rather than "BAT temperature"). However, surface measurements are not merely passive reflections of core temperature. Here we add background and explanation about our thermography data:

      Background on our approach

      Infrared thermography provides a non-invasive readout of heat emission over the interscapular region and has been validated as reporting UCP1-dependent BAT thermogenesis in mice under adrenergic stimulation (Crane et al., 2014). That said, there are known confounds (insulation/adiposity, blood flow, protocol variability) and standardized protocols are needed (Law et al., 2018). Direct telemetry or implanted thermocouples offer superior precision for measuring BAT temperature, so long as the probe is sutured to BAT itself or to Sulzer’s vein–a technical challenge because probes tend to drift over time (e.g., (Dodson et al., 2024)).

      Our BAT findings in context:

      Using SGBS, we demonstrate that the interscapular BAT region is significantly warmer than the adjacent rump surface (Fig. 5C). If surface temperature were purely a reflection of uniform core temperature, this consistent regional hotspot would not be observed.

      Our cross-correlation analysis from the photometry (Fig. 5E) shows the rise in BAT surface temperature precedes changes in other body regions by approximately 90 seconds, suggesting that BAT acts as a primary heat source during rest-to-arousal transitions rather than passively following core temperature. This finding is consistent with another study, using telemetric probes placed in BAT, finding that episodic onset of BAT temperature started to increase 3 minutes before body temperature (Ootsuka et al., 2009).

      Based on this Reviewer’s comment here and the subsequent one (5), we have now added a new analysis of the temporal patterning of arousal and thermogenesis in the optogenetic cohort of animals; see below for details.

      Vasomotor tone

      We agree that infrared thermography does not directly measure vasomotor tone. We have revised the text to remove language implying that our measurements directly quantify vasomotor tone, vasodilation or vasoconstriction.

      We note that the established approach for non-invasive assessment of vasomotion uses glabrous skin of the tail and ears (Garami et al., 2011; Meyer et al., 2017; Škop et al., 2020). Rump surface temperature measured over hairy, non-glabrous skin correlates more closely with core body temperature than with cutaneous vasomotor tone (Meyer et al., 2017; Škop et al., 2020) and is used in the literature as a reference point for calculating BAT thermogenesis.

      In our data, rump surface temperature decreased following PVNOT calcium peaks while BAT and dorsal surface temperatures increased (Fig. 5L-M). This pattern is consistent with sympathetically-driven thermogenesis in which peripheral heat loss is reduced while BAT drives core temperature upwards. We now acknowledge that our rump measurements do not isolate vasomotor contributions. We have revised the manuscript accordingly, replacing references to rump vasoconstriction with language describing the observed thermal pattern while avoiding attribution to a specific thermoeffector mechanism.

      Finally, we note that telemetry would strengthen deep-body temperature interpretation, but telemetry does not itself quantify vasomotor tone; the same distal heat-loss readouts described above would be required regardless of core Tb methodology.

      In sum, infrared thermography enables non-invasive, simultaneous tracking of multiple thermal features in freely moving, undisturbed animals—a requirement for studying the naturalistic state transitions central to this study. We have added a section to the Discussion acknowledging the limitations of surface infrared thermography.

      (5) Photostimulation of PVNOT neurons increased Tb after 400 sec (6.6 min) (Figure 5). This latency is too long to conclude that the neuronal stimulation elicited BAT thermogenesis. A more reasonable explanation is that the increase in Tb was caused by the induction of physical activity (Figure S4C), which slowly generates heat and contributes to the elevation of Tb. However, this view contradicts the authors' claim. To address this concern, the authors should directly measure BAT thermogenesis and compare it with the rate of Tb elevation. If BAT thermogenesis occurs, the rate at which the BAT temperature increases must exceed the rate at which Tb rises.

      We thank the reviewer for this thoughtful critique. With this response we first provide additional context about the timeline of temperature increases, and second add a new analysis addressing the relative contributions of activity and BAT-surface to Tb changes.

      (1) Additional context on the temporal progression

      First, the observed timescale does not, per se, rule out a contribution of BAT thermogenesis. While the kinetics of BAT activation and associated Tb increases can operate on a fast timescale in anesthetized animals, in vivo activation of BAT thermogenesis pathways can take several minutes to yield a statistically detectable difference. For example, activation of DMH→rMR glutamatergic signaling, a canonical thermogenic command pathway, takes several minutes to produce a significant increase in both Tb and BAT using telemetric temperature probes (Kataoka et al., 2014).

      This timescale could also be consistent with peptidergic neuromodulation by PVNOT neurons, which are more likely to be modulators (and not drivers) of the canonical thermogenic pathway. Oxytocin is known to act via volume transmission and metabotropic receptor signaling, which operate on slower timescales than ionotropic neurotransmission (Ludwig and Leng, 2006). Downstream recruitment of sympathetic outflow and BAT thermogenesis is likewise a multistep autonomic process, not an immediate synaptic event.

      Next, the thermal dynamics reported in Figure 5 and Figure S4 are not consistent with activity-induced heat production alone. Specifically:

      - Thermal increases were spatially localized to interscapular/dorsal regions corresponding to BAT depots before generalized surface warming.

      - Importantly, photostimulation-induced warming was observed even during behavioral states characterized by low baseline activity, suggesting that thermogenic activation was not simply a byproduct of movement.

      While we did not directly measure BAT sympathetic nerve activity, our surface thermography approach was designed specifically to resolve regional temperature dynamics over the interscapular BAT area. The spatial specificity and temporal profile of the warming are consistent with BAT thermogenesis rather than uniform musclegenerated heat.

      We acknowledge that direct measurement of BAT sympathetic activity or oxygen consumption would provide additional mechanistic resolution. However, given (i) the known role of PVN oxytocin neurons in autonomic regulation, (ii) the spatially localized dorsal temperature increase, and (iii) the temporal dissociation between stimulation onset and gradual systemic Tb rise, we conclude that BAT thermogenesis remains the most parsimonious explanation.

      We have revised the Discussion to more explicitly acknowledge these temporal dynamics by clarifying that photostimulation likely follows the timescales of peptidergic neuromodulation.

      (2) New analysis

      We have added a new analysis to address the relationship between Tb and BAT-surface temperature and locomotion the optogenetic cohort. In short, we show that across all mice changes in BAT typically precede changes in Tb, and that the effect of optogenetic stimulation on core Tb can’t be explained by physical activity (nor can it be explained by BAT-surface temperature).

      First, cross-correlation of derivatives suggested BAT surface temperature changes typically precede changes in dTb/dt across mice, whereas physical activity changes did not consistently precede dTb/dt. This result, now shown in Fig. S5G, is consistent with our cross-correlation analysis of the fiber-photometry cohort.

      Next, we used a lagged regression analysis to test whether photostimulation-evoked increases in core temperature are fully mediated by physical activity. Specifically, we modeled the derivative of core Tb (dTb/dt) using an impulseresponse representation of photostimulation, while controlling for distributed lags (0–120 s) of physical activity and BAT surface temperature derivative, with random effects for mouse and trial. Photostimulation remained a significant predictor of dTb/dt while controlling for activity and BAT-surface (likelihood ratio test, χ<sup>2</sup>=7.66, p=0.0056), indicating that the relationship between stimulation and Tb is not fully explained by activity.

      Recommendations for the authors:

      Editors note:

      We suggest including key statistical support for the claims in the main text (e.g., results or figure legends).

      We have added statistical support for key claims in the main text results. We have also added references to Table S1 where appropriate (e.g., where there is a long list of statistical results); we hope this aids the readability of the report.

      Reviewer #1 (Recommendations for the authors):

      See above - the authors should decide what to prioritize, but I only mention significant concerns above. The manuscript could be improved to 'Convincing' or even 'Compelling' with sufficient effort.

      Thanks for the careful reading of the manuscript. We’ve addressed many of these points, and feel the manuscript has been strengthened as a result.

      There were also some text errors here and there.

      Several text errors were identified and fixed. Thank you.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1I: The quantification shown here is a bit unclear from the figure and legend - are the authors reporting the percentage of cFos+ cells within the OXT+ population, or within the general DAPI+ population? If the latter, including a co-localization analysis to estimate the proportion of OXT+ cells activated would strengthen the interpretation.

      We thank the reviewer for catching this; Reviewer 1 made a similar comment. A typo in the figure legend led to this confusion. Figure 1I is in fact a quantification of the percent Oxytocin:Fos colocalized cells (not Fos:DAPI, as was written) in dorsal and ventral subregions of the PVN during active huddling and quiescent huddling. We have corrected the legend and clarified the quantification in the revised manuscript. (Note this response is copy-pasted to other relevant comments).

      (2) PVN cell types: It would be useful to briefly discuss the potential involvement of other PVN populations (e.g., CRF, AVP neurons) in huddling, given their known roles in social behavior, stress, and thermoregulation.

      Thank you for the insightful comment. We address these points in two parts.

      (1) PVN cell types and huddling

      Regarding the specific connection between these cell types and huddling: to our knowledge, no study has directly tested the effect of PVN CRF or PVN AVP neuron manipulation on huddling behavior. The most relevant data come from Bendesky et al. (Bendesky et al., 2017), who found that intracerebroventricular administration of AVP in Peromyscus inhibited nest building but had no effect on huddling, licking, or pup retrieval (though this pharmacological approach does not isolate PVN AVP neurons specifically). Their chemogenetic manipulation of PVN AVP neurons in Mus musculus confirmed the nest-building effect but did not assess huddling. For CRF, the available evidence suggests an opposing role to OT in social care contexts: chemogenetic activation of PVN CRF neurons impairs maternal behavior in postpartum mice (Melón et al., 2018), and intracerebroventricular CRF administration suppresses maternal care and can induce pup-killing in virgin rats (Pedersen et al., 1991).

      That said, PVN AVP neurons do promote wakefulness via lateral hypothalamic orexin neurons (Islam et al., 2022) and a recent preprint has implicated PVN AVP neurons in temperature-dependent maternal thermoregulatory behaviors, including co-nesting and shepherding, via projections to the central amygdala (Adahman et al., 2025). Notably, while that study focused on AVP neurons, their c-Fos data also revealed significant temperature-dependent modulation of PVNOT neurons (Fig. 3B), with suppressed activity at thermoneutrality relative to cooler conditions, a pattern suggesting that OT neurons are active under conditions where thermoregulatory effort is required. This data is consistent with our findings on PVNOT neuron involvement in rest-to-arousal transitions driven by thermoregulatory need.

      Additionally, Inada et al. (Inada et al., 2025) used an elegant series of viral-genetic experiments to demonstrate that PVN AVP neurons facilitate paternal caregiving behaviors via AVP to oxytocin receptor crosstalk in the preoptic area. Critically, their fiber photometry and circuit mapping data showed that chemogenetic activation of PVN AVP neurons did not recruit PVN OT neurons (Fig. 4), indicating that these populations operate independently in this context. We believe this finding is consistent with our interpretation that the thermoregulatory signals we observe reflect a cell-type specific property of PVNOT neurons. Future work examining how PVNOT, AVP, and CRF population interact during thermoregulatory state transitions would be valuable.

      (2) PVN cell types and stress and thermoregulation

      PVN CRF and AVP neurons have established roles in stress responses and social behavior, and future studies examining their involvement in huddling would be valuable. However, their direct roles in thermoregulation are limited. PVN CRF neurons are primarily stress-axis regulators whose thermoregulatory influence is mediated indirectly through downstream targets such as the DMH (reviewed in (Morrison and Nakamura, 2019)). AVP's thermoregulatory role is principally as an endogenous antipyretic acting via preoptic area neurons (Tabarean, 2021), rather than through PVN magnocellular AVP neurons.

      Importantly, the synchronized pulsatile bursting pattern that is characteristic of OT neurons during lactation (which serves as a key validation benchmark for our PVNOT calcium peaks), appears to be specific to OT neurons and does not generalize to other PVN populations. One study (Popescu et al., 2019) directly demonstrated that lactation-induced IPSC burst upregulation occurs selectively in OT magnocellular neurons, with no change in VP neurons within the same nucleus. VP neurons do exhibit phasic bursting, but these patterns are asynchronous, of longer duration, and serve antidiuretic rather than neuroendocrine-pulsatile functions (De Mota et al., 2004; Poulain et al., 1977; Wakerley et al., 1978). To our knowledge, no studies have reported synchronized burst activity in PVN CRF neurons during lactation or at rest. We have added a brief discussion of these points to the manuscript.

      (3) Figure 2B: Several behavioral abbreviations (e.g., LMA) are not intuitive and are missing from the legend. Spelling them out or including schematic illustrations would improve clarity.

      We have expanded the figure legends to define all behavioral abbreviations: LMA (Locomotor Activity), EaDr (Eating or Drinking), Groom (Grooming), Nest (Nesting or Nest Building), Quies (Quiescence), Sta (Stationary), ConI (Contact Initiated), ConR (Contact Received), AHud (Active Huddle), QHud (Quiescent Huddle).

      Reviewer #3 (Recommendations for the authors):

      (1) Figures 1D-F and S1A-C: The current magnification is insufficient to clearly resolve the distribution of FOS signals. FOS fluorescence is generally expected to be localized within cell nuclei. However, particularly in Figure 1F, the signals exhibit punctate or fibrous staining in addition to nuclear localization.

      This raises concerns about the quality of the tissue staining and the reliability of subsequent analyses. Including higher-magnification images would strengthen the credibility of the data presented.

      Thanks for the careful observation. We used a well-validated FOS protocol (see Methods; c-Fos (9F6) Rabbit mAb, Cell Signaling, 14609, 1:1000 dilution in block solution).

      To address this issue, in Figure 1 we have included better images of the regions of interest (DMH, LS, and PVN). We also show an inset with DAPI and the FOS IHC. These inset images show that the FOS signal does co-localize with nuclei.

      The reviewer notes that there is a fibrous staining in the PVN. We too noted this type of staining, due to clusters of bright dots in the PVN but not in other regions. This pattern was reproducible across several histological experiments. Fortunately, these bright dots were easily removed in our image processing routine using a selective median filter (pixel radius < 2.0 and and pixel intensity > 50).

      (2) Figures 2A, 4C, and 6A: As mentioned in the Public Review, the specificity of AAV-mediated gene expression is critical for the strength of the conclusions. Quantitative data demonstrating the expression specificity should be included.

      This point was echoed by Reviewers 1 and 3, and one we have taken several actions to address this. (Note this response is copy-pasted to the other reviewers).

      We agree with the reviewers that the rationale for the lactation data should be made more explicit. The primary purpose of this experiment was to validate the identity of oxytocinergic neurons of the PVN.

      Our efforts to use IHC to validate the identity of AAV-transfected cells were inconclusive, and we have now added new data to illustrate this point. We have added Fig. S4 that includes quantitative data on expression specificity. We observed significant variability in co-staining (OT+/GCaMP+) across brain slices, likely reflecting the dynamic nature of oxytocin peptide synthesis and storage, particularly with respect to processes lining the third ventricle. This finding is in accordance with other studies that are now cited in the text.

      We now emphasize that, because IHC provided variable co-localization, we employed the lactation model as an independent physiological validation of the identity of the recorded neurons.

      It is well established that PVNOT neurons undergo dramatic changes in firing dynamics and synchrony during lactation to support milk ejection (Yaguchi et al., 2023; Yukinaga et al., 2022). Conversely, AVP and CRF cell populations in the PVN do not appear to display synchronized pulsatile bursting during lactation (see response to Reviewer-2 comment-2 in ‘Recommendation for authors’ and our updated Discussion). Observing these characteristic changes in our recorded population provides high-confidence functional evidence that we are targeting oxytocin neurons. We have revised the text to clarify that Figure 4 serves primarily as a functional verification of genetic targeting.

      We also acknowledge in the Discussion the possibility that our Cre-line may capture a small percentage of non-oxytocinergic neurons, while noting that the dramatic shift in calcium dynamics during lactation (Figure 4I–L) strongly suggests the recorded population is dominated by oxytocin neurons.

      (3) Figure 2D: The authors should show an expanded view of a representative "PVNOT peak" from the spikes presented.

      We have added a representative peak to Fig. 2D.

      (4) Figure 2E-J: All the abbreviations of the behavioral states must be defined in the figure or legend.

      We added these abbreviations to the legend, and a text box reading “See legend for abbreviations” to the schematic.

      (5) Figure 2F, G, I, and J: The units on the y-axis should be indicated to facilitate interpretation.

      We have added these units. Thanks.

      (6) Figure 3A: Three large PVNOT peaks occurred between 01:30 and 02:00. However, these peaks did not cause an obvious transition in behavioral states or an increase in Tb within several minutes. Therefore, statements such as "PVNOT neurons predict transitions towards thermogenesis and behavioral arousal" in the text and subheading (pages 7 and 9) are questionable.

      We thank the reviewer for this careful observation. The three peaks between 01:30 and 02:00 that do not immediately lead to a behavioral transition illustrate a key aspect of our findings: the relationship between PVNOT activity and state transitions is probabilistic and state-dependent, not deterministic. Our logistic regression analysis (Fig. 3F, H, J, L) demonstrates that peaks increase the probability of a transition (up to ~20% per second) rather than acting as an obligatory "on switch." While individual variability exists in any single trace, the group-level analysis reveals a statistically significant increase in physical activity following PVNOT peaks (Fig. 3B–E).

      We therefore use ‘predict’ in a probabilistic sense: PVNOT peaks increase the conditional probability of impending state transitions in a manner that depends on behavioral context, rather than acting as an obligate trigger in every instance. We have taken care to not claim that PVNOT neurons are a necessary causal factor for transitions towards thermogenesis and arousal.

      We have updated the figure legend to clarify that Figure 3A shows an individual example trace, and revised the subheading on page 7 to more accurately reflect the probabilistic nature of this relationship: "PVNOT neurons predict increased likelihood of transitions towards thermogenesis and behavioral arousal in social and non-social contexts".

      We qualified the word “predicts” with “probabilistically” in the third paragraph of this section.

      Finally, this comment is related to the Reviewer’s comment-2 in the Public Reviews. To address that comment, we added a new analysis (now Fig. 3P&S) which shows that the presence of a peak in a bout of rest increases the thermogenic trajectory compared to bouts without a peak.

      (7) Figure 3F and H: If PVNOT peaks contribute to the initiation of transitions into the active state, the probability of peak occurrence should reach its maximum prior to the quiescence offset. However, the figures do not present the probability trajectory after the offset, which limits the ability to evaluate the authors' interpretation. Reanalysis extending to 150 seconds post-offset would be needed to clarify this issue.

      Thank you for this suggestion. We agree that examining PVNOT dynamics around the period following quiescence (and quiescent huddling) offset can further inform how PVNOT activity relates to rest-to-active transitions, and this has led to new insights within the manuscript.

      For background, in the original analysis (Fig. 3F,H), we used logistic regression to quantify how peak probability differs between bout onset versus near bout offset. We focused these analyses on the the timeframe of the bouts themselves (plus a small margin) because, in freely behaving animals, the pre-onset and post-offset period is heterogeneously composed of multiple potential subsequent behaviors (e.g., brief re-entry into quiescence, nesting, active huddling, locomotion, etc), which would confound a single post-offset probability trajectory (unless offsets are stratified by the identity of the subsequent behavioral state–beyond the scope of this paper).

      To address this concern, we now expand our peri-event baseline calcium analysis to include three minutes before and three minutes after both bout onset and bout offset for all four behaviors (new Fig. S3I–L). These extended traces show that for the two resting states (quiescence and quiescent huddling), baseline PVNOT calcium reaches a minimum near bout onset and a maximum near bout offset, whereas for the two active states (nesting and active huddling) baseline calcium shows the opposite pattern (maximum near onset, minimum near offset). Thus, the expanded post-offset analyses provide a more complete view of PVNOT calcium dynamics across the requested post-offset epoch and further support the conclusion that PVNOT activity is aligned with (and elevated around) behavioral transitions in a state-dependent manner. We have updated the Results text accordingly and now explicitly reference these new extended peri-event baseline analyses.

      (8) Figures 4H and I: Figure 4H shows that the waveform in the PPD2-7 group has a narrower FWHM than the Virgin group, which is the opposite of the group data in Figure 4I. Presenting scaled waveforms in parallel would allow for a clearer comparison across groups.

      Thank you for pointing out the inconsistency between the representative waveform in Fig. 4H and the group summary in Fig. 4I. You were correct: the PPD2–7 and Virgin waveforms in Fig. 4H had been mislabeled. We have corrected the labeling. (We verified that the underlying data are correct).

      As suggested, to enable visual comparison of waveform width across groups independent of amplitude differences, we derived peak-normalized average waveforms using a normalization procedure for every peak prior to averaging. Specifically, for each peak we (1) baseline-subtracted the trace by subtracting the mean fluorescence in a pre-peak baseline window, and then (2) divided the baseline-subtracted waveform by its own maximum value to scale the event amplitude to 1. We then computed the mean ± SEM of these peak-normalized waveforms across events within each group.

      We believe these changes resolve the discrepancy and improve the clarity of the figure, consistent with your suggestion.

      (9) Figure 5: In studies of thermoregulatory processes, tail blood flow or temperature is commonly used as an indicator of vasomotor responses. Is it feasible to track tail temperature using the SGBS system? If not, it may be helpful to acknowledge this as a technical limitation.

      We agree that tail temperature is a commonly used indicator of vasomotor responses. While SGBS could in principle be trained to segment the tail, the current model was optimized for dorsal body regions viewed from an overhead perspective. Reliable tail tracking presents substantial technical challenges in our configuration of homecage recordings. The tail’s thin geometry and rapid, multidirectional movement frequently result in partial or complete occlusion (e.g., beneath bedding or the animal’s body). In addition, during vasoconstriction the tail temperature approaches ambient floor temperature, reducing thermal contrast and making segmentation unreliable with the current thermal resolution limited by our camera. We have acknowledged this as a technical limitation in the Discussion, in the section called “Thermal tracking and validation of PVNOT recording specificity”.

      (10) Figure S5: Please describe the reason and histological background for the intravenous injection of FluoroGold.

      Intravenous injection of FluoroGold (FG) was used to histologically differentiate between magnocellular and parvicellular oxytocin neurons in the PVN. Because the posterior pituitary is located outside the blood-brain barrier,

      i.v. FG is selectively taken up by terminals of magnocellular neurons and retrogradely transported to their cell bodies. This allows us to infer the neuroanatomical identity (magno- vs. parvicellular) of the PVNOT neurons of interest. We have updated the Methods with a detailed description of the FG injection protocol as follows:

      “To distinguish between peripheral-projecting magnocellular and central-projecting parvicellular neurons, mice received 15 uL intravenous injection of 4% Fluoro-Gold (Fluorochrome) diluted in 100 uL of sterile saline. Prior to injection, mice were given an analgesic dose of carprofen (20 mg/kg, s.c.). Mice were briefly restrained using a modified 50 mL conical tube, in which holes were drilled to allow for proper air flow and respiration. Mouse tails were interposed between two heating pads to enhance visibility of the tail vein. Tails were wiped down with 70% ethanol and FG was administered via either right or left lateral tail vein using a 0.5 mL 28G syringe. Mice were sacrificed 24- 48 hours post-FG administration.”

      The following are minor points.

      (11) Figure 2E-G, Figure 3F,G, Figure S2G,I, Figure S3A: "quiesence" > "quiescence". This typo may appear elsewhere in the manuscript as well.

      Thanks. These edits have been made.

      (12) Page 7, line 14: Peaks were NOT significantly increased at 29{degree sign}C in Figure 2N.

      Thanks for the very careful read. By way of explanation: this difference had been significant in an earlier draft; however, when we added more replicates, the difference went away. We have corrected this sentence.

      (13) There are mislabeled figure numbers in the main text. The authors should carefully check this throughout the manuscript.

      We found mislabeled figure numbers and have corrected them.

      (14) Page 13, lines 1- 2: To make the description clearer, it might be better to rephrase the part that says, "some blue light stimulations occurred." As it stands, it could give the impression that the stimulations happened spontaneously. Using a phrase like "were delivered" would more clearly indicate that these were intentional, experimenter-controlled events.

      Agreed. Thanks. The edit has been made.

      Additional comments:

      The oxytocin system is thought to support a wide range of physiological and behavioral functions, and the circuits involving oxytocin neurons are likely to be regulated in complex and dynamic ways. As oxytocin research continues to expand, the growing body of evidence not only deepens our understanding but also highlights the system's complexity. In this context, the development of an approach that enables the observation of oxytocinergic neuron activity in parallel with naturalistic behavior represents a promising methodological contribution. It is likely that similar experimental frameworks will become increasingly common in future studies. While reading this manuscript, as a reader rather than a reviewer, I was wondering how OXT neurons detect or define the "rest balance-point," and how they might contribute to shifting the brain toward an "awake balance-point" (Figure 7). Given that eLife allows authors to include an "Ideas and Speculation" subsection within the Discussion, it would be appreciated - though not essential - if the authors could briefly share their perspective on this point. I believe such mechanistic insight would make the manuscript more intellectually stimulating.

      This is a great suggestion. We have added a new “Ideas and Speculation” section of the Discussion.

      References

      Adahman Z, Ooyama R, Gashi DB, Medik ZZ, Hollosi HK, Sahoo B, Akowuah ND, Riceberg JS, Carcea I. 2025. Hypothalamic Vasopressin Neurons Enable Maternal Thermoregulatory Behaviors. DOI: https://doi.org/10.1101/2025.01.23.634569

      Bendesky A, Kwon Y-M, Lassance J-M, Lewarch CL, Yao S, Peterson BK, He MX, Dulac C, Hoekstra HE. 2017. The genetic basis of parental care evolution in monogamous mice. Nature 544:434–439. DOI: https://doi.org/10.1038/nature22074

      Crane JD, Mottillo EP, Farncombe TH, Morrison KM, Steinberg GR. 2014. A standardized infrared imaging technique that specifically detects UCP1-mediated thermogenesis in vivo. Molecular Metabolism 3:490– 494. DOI: https://doi.org/10.1016/j.molmet.2014.04.007

      De Mota N, Reaux-Le Goazigo A, El Messari S, Chartrel N, Roesch D, Dujardin C, Kordon C, Vaudry H, Moos F, Llorens-Cortes C. 2004. Apelin, a potent diuretic neuropeptide counteracting vasopressin actions through inhibition of vasopressin neuron activity and vasopressin release. Proceedings of the National Academy of Sciences 101:10464–10469. DOI: https://doi.org/10.1073/pnas.0403518101

      Dodson AD, Herbertson AJ, Honeycutt MK, Vered R, Slattery JD, Goldberg M, Tsui E, Wolden-Hanson T, Graham JL, Wietecha TA, O’Brien KD, Havel PJ, Sikkema CL, Peskind ER, Mundinger TO, Taborsky GJ, Blevins JE. 2024. Sympathetic Innervation of Interscapular Brown Adipose Tissue Is Not a Predominant Mediator of Oxytocin-Induced Brown Adipose Tissue Thermogenesis in Female High Fat Diet-Fed Rats. Current Issues in Molecular Biology 46:11394–11424. DOI: https://doi.org/10.3390/cimb46100679

      Garami A, Pakai E, Oliveira DL, Steiner AA, Wanner SP, Almeida MC, Lesnikov VA, Gavva NR, Romanovsky AA. 2011. Thermoregulatory Phenotype of the Trpv1 Knockout Mouse: Thermoeffector Dysbalance with Hyperkinesis. The Journal of Neuroscience 31:1721–1733. DOI: https://doi.org/10.1523/JNEUROSCI.4671-10.2011

      Inada K, Hagihara M, Yaguchi K, Irie S, Inoue YU, Inoue T, Miyamichi K. 2025. Vasopressin-to-oxytocin receptor crosstalk in the preoptic area underlying parental behaviors in male mice. Nature Communications 16:10844. DOI: https://doi.org/10.1038/s41467-025-66908-0

      Islam MT, Rumpf F, Tsuno Y, Kodani S, Sakurai T, Matsui A, Maejima T, Mieda M. 2022. Vasopressin neurons in the paraventricular hypothalamus promote wakefulness via lateral hypothalamic orexin neurons. Current Biology 32:3871-3885.e4. DOI: https://doi.org/10.1016/j.cub.2022.07.020

      Kataoka N, Hioki H, Kaneko T, Nakamura K. 2014. Psychological Stress Activates a Dorsomedial HypothalamusMedullary Raphe Circuit Driving Brown Adipose Tissue Thermogenesis and Hyperthermia. Cell Metabolism 20:346–358. DOI: https://doi.org/10.1016/j.cmet.2014.05.018

      Landen JG, Vandendoren M, Killmer S, Bedford NL, Nelson AC. 2024. Huddling substates in mice facilitate dynamic changes in body temperature and are modulated by Shank3b and Trpm8 mutation. Communications Biology 7:1186. DOI: https://doi.org/10.1038/s42003-024-06781-7

      Law J, Chalmers J, Morris DE, Robinson L, Budge H, Symonds ME. 2018. The use of infrared thermography in the measurement and characterization of brown adipose tissue activation. Temperature 5:147–161. DOI: https://doi.org/10.1080/23328940.2017.1397085

      Ludwig M, Leng G. 2006. Dendritic peptide release and peptide-dependent behaviours. Nature Reviews Neuroscience 7:126–136. DOI: https://doi.org/10.1038/nrn1845

      Melón LC, Hooper A, Yang X, Moss SJ, Maguire J. 2018. Inability to suppress the stress-induced activation of the HPA axis during the peripartum period engenders deficits in postpartum behaviors in mice. Psychoneuroendocrinology 90:182–193. DOI: https://doi.org/10.1016/j.psyneuen.2017.12.003

      Meyer CW, Ootsuka Y, Romanovsky AA. 2017. Body Temperature Measurements for Metabolic Phenotyping in Mice. Frontiers in Physiology 8:520. DOI: https://doi.org/10.3389/fphys.2017.00520

      Morrison SF, Nakamura K. 2019. Central Mechanisms for Thermoregulation. Annual Review of Physiology 81:285– 308. DOI: https://doi.org/10.1146/annurev-physiol-020518-114546

      Ootsuka Y, de Menezes RC, Zaretsky DV, Alimoradian A, Hunt J, Stefanidis A, Oldfield BJ, Blessing WW. 2009. Brown adipose tissue thermogenesis heats brain and body as part of the brain-coordinated ultradian basic rest-activity cycle. Neuroscience 164:849–861. DOI: https://doi.org/10.1016/j.neuroscience.2009.08.013

      Pedersen CA, Caldwell JD, McGuire M, Evans DL. 1991. Corticotronpin-releasing hormone inhibits maternal behavior and induces pup-killing. Life Sciences 48:1537–1546. DOI: https://doi.org/10.1016/00243205(91)90278-J

      Popescu IR, Buraei Z, Haam J, Weng F, Tasker JG. 2019. Lactation induces increased IPSC bursting in oxytocinergic neurons. Physiological Reports 7:e14047. DOI: https://doi.org/10.14814/phy2.14047

      Poulain DA, Wakerley JB, Dyball REJ. 1977. Electrophysiological differentiation of oxytocin-and vasopressinsecreting neurones. Proceedings of the Royal Society of London. Series B. Biological Sciences 196:367– 384. DOI: https://doi.org/10.1098/rspb.1977.0046

      Škop V, Guo J, Liu N, Xiao C, Hall KD, Gavrilova O, Reitman ML. 2020. Mouse Thermoregulation: Introducing the Concept of the Thermoneutral Point. Cell Reports 31:107501. DOI: https://doi.org/10.1016/j.celrep.2020.03.065

      Tabarean IV. 2021. Activation of Preoptic Arginine Vasopressin Neurons Induces Hyperthermia in Male Mice. Endocrinology 162:bqaa217. DOI: https://doi.org/10.1210/endocr/bqaa217

      Wakerley JB, Poulain DA, Brown D. 1978. Comparison of firing patterns in oxytocin- and vasopressin-releasing neurones during progressive dehydration. Brain Research 148:425–440. DOI: https://doi.org/10.1016/00068993(78)90730-8

      Yaguchi K, Hagihara M, Konno A, Hirai H, Yukinaga H, Miyamichi K. 2023. Dynamic modulation of pulsatile activities of oxytocin neurons in lactating wild-type mice. PLOS ONE 18:e0285589. DOI: https://doi.org/10.1371/journal.pone.0285589

      Yukinaga H, Hagihara M, Tsujimoto K, Chiang H-L, Kato S, Kobayashi K, Miyamichi K. 2022. Recording and manipulation of the maternal oxytocin neural activities in mice. Current Biology 32:3821-3829.e6. DOI: https://doi.org/10.1016/j.cub.2022.06.083

    1. eLife Assessment

      In this study, the authors found that a species of aphid that is a known agricultural pest salivated longer and produced more honeydew when feeding at night. The authors identified aphid genes with diurnal expression patterns, including potential salivation-related genes. Silencing these genes reduced aphid performance only on plants and not on artificial diet, suggesting a specific role in plant feeding. This study is valuable for understanding plant-insect interactions in agriculture and presents solid evidence of diurnal rhythmicity in aphid activity and performance, although further research is needed to elucidate the function of the identified genes.

    2. Reviewer #2 (Public review):

      Summary:

      The authors conducted a time-course of whole-body transcriptional analysis of a pest aphid, Rhopalosiphum padi, and identified four major clusters of the genes that show diurnal rhythmicity in transcription. In addition, they have conducted the analysis of aphid feeding behaviour and showed that aphids salivate longer from the end of the day toward the beginning of the night while their phloem feeding time does not change throughout a day. The genes up-regulated at nighttime were enriched with the genes involved in metabolic activities, collaborating with the results showing higher number of honeydew excretion at night. The authors identified the list of candidate salivary genes that show diurnal rhythmicity in the transcription and silenced a salivary gene C002 and the candidate salivary gene E8696. Silencing of these genes reduced aphid fecundity and survival rate on the host plant but not on the artificial diet.

      Strengths:

      The time-course transcription study and its analysis will be of interest to researchers studying diurnal rhythms in insect biology. Also, the analysis of aphid feeding behaviour at different time of day is interesting. This study provides variable resources for those who study insect biology.

      Weaknesses:

      Without the knowledge of the functions of the salivary effectors, especially their targets, it is hard to conclude that the rhythmical expression is important for the aphid performance. In addition, it is not clear whether increase of gene expression is directly corelated with the increase of protein secretion into the saliva and the plant.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public Review):

      Summary:

      This study presents valuable data on diurnal patterns in aphid (Rhopalosiphum padi) feeding behavior and transcriptome profiles. The authors measured honeydew production by the aphids on plants and artificial diet during the day and night and conducted a comprehensive feeding behavior study using EPG with many biological replicates at 6 time-points in 24 hours. They also conducted transcriptome analyses of three samples of each 30 aphids at these time points. Differentially expressed transcripts were grouped into four clusters with distinct expression patterns. The expression of two genes found to be diurnally rhythmic was knocked down with RNAi and these aphids did less well, especially at night. They also analyzed the differential expression of candidate effector genes and found rhythmic ones to be enriched for more expression in aphid heads versus bodies - this pattern is expected given that effectors are most likely expressed in the salivary glands. Knockdown of a known effector (C002) that is diurnally rhythmic, and a novel effector gene, was found to alter aphid feeding dynamics and performance.

      Thank you for your thoughtful review and summary of our study. We would like to clarify one aspect of your summary regarding our clustering analysis. We did not cluster differentially expressed transcripts. Instead, our clustering analysis was performed exclusively on transcripts that were significantly diurnally rhythmic, as identified in our time-course transcriptomic analysis. This approach allowed us to reveal patterns of gene expression that exhibit robust rhythmicity over the 24-hour cycle, rather than grouping transcripts based solely on differential expression at individual time points.

      Strengths:

      The manuscript was highly accessible, with clear writing, and the figures provided were both comprehensive and of good quality. The datasets generated from this research are valuable to the research field, especially the findings for honeydew secretion, EPG analysis, and transcriptome experiments.

      The datasets generated in this study will be useful to scientists working on aphids and aphidplant interactions and will inform similar studies on other insect species.

      Weaknesses:

      The weaknesses mainly relate to the (depth of) analyses and interpretation of the data. Also, some methods require more explanation, as follows:

      In Figure 1, data show that aphids produce more honeydew at night than during the day. This suggests that the aphids ingest more phloem (E2 phase). However, in Figure 1d the duration of the E2 phase does not show obvious differences among the time points in the 24 hours. The authors contribute the explanation that the aphids may osmoregulate more during the night, leading to more honeydew secretion at night. This may be the case, but there could be other explanations. For example, the physiology, including regulation of water transport, of plants is known to change during night/day. The authors may focus this section more on the differences in the E1 phase, as this involves the delivery of aphid saliva and effectors into the plant phloem.

      Thank you for your constructive feedback. As noted, aphids excreted more honeydew at night, although the duration of the E2 phase did not differ significantly across time points. We agree that host plant physiology, particularly the composition of the phloem and its osmotic quality, also influences the observed osmoregulatory patterns in R. padi. However, a similar diurnal pattern of honeydew excretion was also observed on artificial diets (Fig. 1b), in which host-derived cues have been eliminated. This strongly suggests that increased nighttime honeydew excretion is primarily driven by enhanced aphid osmoregulation rather than plant factors alone. Nonetheless, we acknowledge that plant-derived factors may also contribute and cannot be entirely ruled out. We have revised the text in the discussion of the revised manuscript to reflect this broader interpretation. As suggested, we have also added further details to highlight the important role of the E1 phase in aphid salivation.

      Transcriptome data shown in Figure 2 (and the experimental procedure of Figure 5b) appears to be based on three biological replicates. However, these replicates appear to have been harvested at the same time in the experiment, and this makes them technical replicates, not biological replicates. The inclusion of true biological replicates that include samples from time series experiments done on different days should be considered.

      Thank you for your concern regarding the biological replication in our transcriptome analysis. Our experimental design included multiple independent pools of aphids collected at each time point. Specifically, each replicate consisted of a unique group of aphids collected from different leaf positions across multiple host plants, such that no individuals were shared among replicates. As a result, these samples represent biologically independent populations rather than technical replicates, which are defined as repeated measurements of the same biological sample used to estimate technical noise. Although all samples were collected within the same experimental time course, this approach is commonly used in time-series transcriptomic studies to minimize confounding variation associated with differences in insect age, entrainment history, or environmental conditions, all of which can obscure rhythmic gene expression patterns. By maintaining tightly controlled and consistent conditions across the sampling period, we aimed to ensure that observed transcriptional differences primarily reflected diurnal regulation rather than uncontrolled day-to-day variability.

      We acknowledge that conducting time-series experiments on different days could provide additional insight into biological variability. However, our approach aimed to reduce potential confounding effects caused by day-to-day environmental fluctuations – such as minor changes in temperature or humidity - which could significantly influence gene expression in insects. By maintaining consistent conditions, we sought to ensure that observed transcriptional differences were due to diurnal rhythms rather than uncontrolled variation. Similar designs or strategies have been employed in studies examining diurnal and circadian gene expression in both insects and plants. We have revised the Methods section to clarify our replication strategy.

      The authors conducted knockdown experiments targeting aquaporin 1 and gut sucrase 1 in aphids, resulting in reduced nymph production and decreased honeydew secretion. It is concluded that these results indicate significant roles of aquaporin 1 and gut sucrase 1 in diurnal regulation. However, it is essential to consider that these genes likely play crucial roles in aphid physiology beyond diurnal rhythms. Consequently, reduced expression would naturally impair aphid performance. The dsAQP1 and dsSUC1 aphids consistently produced less honeydew, regardless of the time of day, indicating a broader impact of gene knockdown. The observed increase of the phenotype at night may not be attributable to the specific roles of these genes in diurnal regulation but rather due to heightened aphid activity during that time (as evidenced by increased honeydew secretion) that could magnify the impact of the knockdown effect, making it easier to observe. Therefore, the knockdown of aquaporin 1 and gut sucrase 1 may exert a general negative influence on aphid fitness, independently of diurnal factors.

      We agree that these genes likely play fundamental roles in aphid physiology beyond diurnal rhythms, and that reduced expression may affect overall aphid performance. However, it is important to highlight that if the observed effects are solely due to general fitness impairments, we would more likely expect a comparable reduction across time points rather than a disproportionately stronger impact at night. We agree that the increased honeydew excretion at night is likely due to heightened aphid excretory activity. However, since this excretory behavior is downstream of osmoregulatory functions, such as water cycling to the midgut and digestion, polymerization, and excretion of oligosaccharides, the increased nighttime phenotype is likely a result of an increased nighttime regulation of osmoregulation in aphids. This hypothesis is further supported by our functional analysis, where the knockdown of AQP1 and SUC1 resulted in a loss of diurnal variation in honeydew production (Fig. 2h), indicating that the observed effects are not merely a general impact on aphid fitness but are likely tied to the genes' roles in regulating diurnal physiological processes. We have revised the discussion to clarify that our findings do not exclude general physiological roles for these genes but instead suggest that their functions intersect with diurnal rhythms to influence aphid feeding and excretion patterns.

      To analyze the roles of genes in diurnal regulation, additional controls should be incorporated. This could involve the knockdown of genes with essential functions that are not influenced by diurnal rhythms, providing a baseline comparison. Furthermore, consider including genes known to be involved in diurnal regulation in other insects, as documented in the existing literature, in the experimental design.

      We agree that incorporating appropriate controls in the RNAi experiments would provide a useful baseline for comparisons, helping to distinguish between general physiological effects and specific diurnal effects. Unfortunately, given the current limited knowledge on rhythmic genes in aphids, particularly in the context of aphid-plant interactions, it is challenging to identify appropriate rhythmic and non-rhythmic controls that can be definitively linked to or unaffected by diurnal regulation within aphids. We will ensure to consider this valuable suggestion in our future experiments.

      The same arguments as for aquaporin 1 and gut sucrase 1 above may be made for knockdown of effector genes (Figure 4). It has already been shown that knockdown of C002 impacts aphid performance, and the data herein may be explained by a general lower performance of aphids rather than a specific function of these effectors in diurnal regulation. It is also expected that knockdown of the effectors has less impact on aphids feeding from artificial diets. This does not necessarily indicate the role of the effectors in diurnal regulation.

      Our response to this comment mirrors that expressed in our earlier response regarding AQP1 and SUC1.

      In the abstract and elsewhere, the authors assert priority by stating, "...the first evidence of...". However, it's important to note that priority claims are often challenging to verify across many fields. Instead of relying solely on claims of precedence, the evidence presented in the research could stand on its own merit.

      We understand that priority claims can be difficult to substantiate across various fields, and we appreciate the importance of allowing the evidence to speak for itself. Considering this, we have revised the language in the abstract.

      Conclusion:

      The study presents intriguing new findings, particularly in the realms of honeydew analysis, EPG, and transcriptome analysis. However, the interpretation of subsequent studies employing gene knockdowns needs further consideration.

      We thank the reviewer for the thoughtful and constructive feedback. We appreciate the positive assessment of our findings on honeydew analysis, EPG, and transcriptome profiling. We have carefully revised the section on gene knockdown experiments to provide clearer interpretation and additional context, and we hope the concerns raised have now been appropriately addressed.

      Reviewer #2 (Public Review):

      Summary:

      The authors conducted a time-course of whole-body transcriptional analysis of a pest aphid, Rhopalosiphum padi, and identified four major clusters of the genes that show diurnal rhythmicity in transcription. In addition, they conducted the analysis of aphid feeding behaviour and showed that aphids salivate longer from the end of the day toward the beginning of the night while their phloem feeding time does not change throughout the day. The genes upregulated at night time were enriched with the genes involved in metabolic activities, collaborating with the results showing a higher number of honeydew excretion at night. The authors identified the list of candidate salivary genes that show diurnal rhythmicity in the transcription and silenced a salivary gene C002 and the candidate salivary gene E8696. Silencing of these genes reduced aphid fecundity and survival rate on the host plant but not on the artificial diet.

      Thank you for your thoughtful review and valuable comments on our study.

      Strengths:

      The time-course transcription study and its analysis will be of interest to researchers studying diurnal rhythms in insect biology. Also, the analysis of aphid feeding behaviour at different times of day is interesting. This study provides variable resources for those who study insect biology.

      Weaknesses:

      It is not clear to me which data was used to define the putative salivary effectors for R. padi, but the candidate salivary gene list made by Thorpe et al consists of the aphid genes encoding secreted proteins that are up-regulated in the head samples compared to the body samples. Although some proteins were confirmed to be secreted into the aphid saliva, many genes in the list are not confirmed to be expressed in the aphid salivary glands, and their products are not confirmed to be secreted into the saliva and the plant. Is E8696 expressed in the aphid salivary glands and secreted into its host plant? Without the data confirming the expression of the gene in the salivary glands and its secretion into the saliva and into the host plant, we cannot call the protein a salivary protein. Furthermore, without the observation that E8696 has some effect on plant biology, we cannot call it an aphid effector. Therefore, I cannot agree with the parts of the manuscript that refer to E8686 as an aphid salivary effector.

      We have revised the text in the Methods to clarify the database used for defining putative salivary effectors. We have also added a sentence in the discussion to indicate that these are putative effectors. The putative effector E8696 was confirmed to be expressed in the salivary glands; however, its secretion into saliva and the host plant remains undetermined due to the lack of E8696-specific antibodies. Over the past year and a half, we have been creating an antibody for E8696. However, the antibody we generated is non-specific, and as a result, we are still unable to demonstrate that E8696 is secreted into host tissue and functions as an effector. While our functional analysis provided strong evidence of E8696’s impact on aphid fecundity and mortality on host plants but not on artificial diets, we agree that without further confirmation of its secretion and effect on the host plant, E8696 should be considered only a putative salivary effector. We expect to address these important questions in future research. To prevent any confusion, we have revised our manuscript to reflect that E8696 is only a putative effector.

      It is interesting to know that some candidate salivary gene expression showed a diurnal rhythm. However, without the knowledge of the functions of the salivary effectors, especially their targets, it is not possible to conclude that the rhythmical expression is important for the aphid performance. In addition, I wonder whether the increase in gene expression is directly correlated with the increase of protein secretion into the saliva and the plant.

      The primary goal of this study was to determine whether aphid genes, particularly those associated with osmoregulation and salivary effectors, exhibit diurnal patterns of expression and whether disrupting these rhythms affects aphid performance. While we agree that the precise molecular targets of these effectors in host plants remain to be identified, our functional assays provide evidence that rhythmic expression is biologically relevant for aphid physiology. Our results demonstrate that silencing rhythmic effector genes resulted in increased aphid mortality, reduced fecundity on host plants, and, more importantly, the disruption of diurnal honeydew excretion patterns, especially for C002. As honeydew excretion is a critical physiological process for aphids, the alteration of this behavior suggests that the rhythmic expression of these genes is functionally important for aphid physiology. We believe our results provide compelling evidence that rhythmic expression plays a critical role in aphid biology. We agree that rhythmic transcript abundance does not necessarily imply proportional changes in protein secretion into saliva, and direct measurements of effector protein dynamics will be an important direction for future work. However, the observed physiological and performance consequences of disrupting rhythmic gene expression support the conclusion that temporal regulation of these salivary genes is functionally important for aphid biology, even in the absence of detailed target identification.

      Finally, the authors examined aphid survival, fecundity, and feeding behaviour. Those are important for overall aphid performance, but they do not "shape" aphid colonization. Aphid colonisation is shaped by the mechanisms by which aphids find and select their host plant and start to feed on it. Therefore, I do not agree with the title of this manuscript and some parts of the discussion.

      We agree with your perspective and have revised the title and discussion to more precisely reflect the scope of our findings, focusing on aphid performance rather than colonization. The revised title now reads “Diurnal rhythmicity in metabolism and salivary effector expression shapes aphid performance on host plants”.

      I would like the authors to develop how the knowledge of the diurnal rhythm of aphid feeding can contribute to optimise pest management. I see that there are some differences in aphid metabolism and feeding behaviour between day and night, but I would like to hear how such knowledge can optimise pest management strategies.

      We have expanded the Discussion to address how knowledge of diurnal rhythms in aphid physiology and feeding behavior could inform the optimization of pest management strategies. Specifically, we discuss how time-of-day variation in aphid feeding activity and metabolism may influence the efficacy of control measures and how chronobiological insights could be integrated into future pest management frameworks.

      Recommendations for the authors:

      Reviewing Editor:

      Based on comments from two reviewers, here are the six key areas that need to be addressed to improve the manuscript.

      Clarity and Specificity:

      (1) Salivary effectors: The manuscript defines "salivary effector" loosely. The reviewer argues for stricter criteria - a protein can only be called a salivary effector if it's confirmed to be produced in the salivary glands and/or secreted into the plant with saliva and function in or around the plant.

      We have addressed this comment and clarified the definition in the revised manuscript.

      (2) Diurnal rhythm: The paper finds a daily rhythm in aphid gene expression, but doesn't explain how these genes affect the plant. The reviewer argues that without understanding the function of these genes, the significance of the rhythm is unclear.

      We have addressed this comment and clarified that the scope of our study is to elucidate diurnal rhythmicity in aphid gene expression and to evaluate the functional importance of rhythmic genes for aphid performance. We agree that understanding how these genes interact with host plants is essential for fully elucidating their molecular functions under diurnal regulation; however, this is beyond the scope of the current study and will be pursued in future research.

      (3) Knockdown experiments: The reviewer suggests the observed effects of knocking down certain genes (aquaporin, sucrase, effectors) might be due to their general importance, not necessarily their role in the day-night cycle. They recommend including control genes and genes known to be involved in circadian rhythms for a more robust comparison.

      We have addressed this comment and clarified the interpretation of these experiments in the revised manuscript.

      Technical Issues:

      (4) Honeydew production: The explanation for nighttime honeydew production needs more exploration. Plant changes at night might also play a role, and the daytime saliva delivery phase deserves more attention in the analysis (Figure 1).

      We expanded the description and interpretation of the salivation phase by incorporating additional detail in the revised manuscript.

      (5) Gene expression data: The current data (Figure 2 & Figure 5b) lacks proper biological replicates. Replicates collected at different times are essential for stronger conclusions.

      We have addressed this comment and clarified the experimental design and replication strategy in the manuscript.

      (6) Priority claims: The reviewer advises against focusing on claiming novelty ("first evidence"). The research should be impactful based on its own merit, not just being the first to find something.

      We revised the sentences to avoid making claims of priority throughout the manuscript.

      Reviewer #2 (Recommendations For The Authors):

      Figures 2 f,g, and h : according to the legend, these experiments seemed to have a low number of replicates (n=3-5). However, Figure 2h has many data points. I understood that here n means the number of experimental replications, but it may be better to show the number of aphid samples examined.

      You are correct that the n refers to the number of experimental replicates, with each replicate comprising multiple individual aphids. Because our analyses were performed on replicate-level averages across multiple days, rather than on individual aphids, we believe this notation most accurately reflects the experimental design. To improve clarity, we have revised figure legends to explicitly state that each replicate includes several individual aphids.

      Are the orthologous proteins of E8696 expressed in aphid salivary glands or detected in saliva? Such data will strengthen the claim that E8696 is a salivary protein of R.padi.

      E8696 is expressed in aphid salivary glands, but it is not confirmed to be secreted into saliva or host plants due to the lack of specific antibodies. We have revised our manuscript to reflect that E8696 is only a putative effector. We will address this question in future research.

    1. eLife Assessment

      Huang and colleagues examined neural responses in mouse anterior cingulate cortex (ACC) during a discrimination-avoidance task. The authors present valuable findings that ACC neurons encode primarily "action content" over extended periods. The methodological approach is sound and the evidence in support of action state encoding is solid, though it is not conclusive to what extent ACC primarily encodes post-action events.

    2. Reviewer #1 (Public review):

      Huang et al. examined ACC response during a novel discrimination-avoid task. The authors concluded that ACC neurons primarily encode post-action variables over extended periods, reflecting the animal's preceding actions rather than the outcomes or values of those actions. The authors have made considerable revision to address the raised the concerns. However, it appears that some important issues remain unresolved.

      To what extent ACC neurons encode post action content remain as a major concern. This may be at least partially attributed by the analysis methods. If I understand it correctly, the authors compared pre- vs post-event neural activity and looked for significant changed. By default, this is to look for post-event changes, rather than pre-event. As a result, it would lead to the conclusion 'Our study also reveals that ACC neurons play a limited role in encoding pre-action variables associated with decision-making or planning, as evidenced by their minimal responses to auditory cues and the modest activity changes prior to shuttle initiation'.

      To determine whether ACC encode pre-action variables or planning, different time windows should be used in the analysis.

    3. Reviewer #2 (Public review):

      Summary:

      Huang et al recorded anterior cingulate cortex activity in mice while they performed a shuttle escape task. The task utilized two auditory cues, each of which informed the mice to stay or escape depending on which side they were on, and incorrect responses were punished by shock administration. Analyses focused on ACC neurons that fired when mice crossed the shuttle box in either direction (A-->B or B-->A), coined "action state", or when mice crossed in one direction but not the other, coined "action content". The authors characterized these populations, and ACC firing changes mostly occurred around the time of shuttle crossing. This work will likely be of broad interest to those who are interested in neocortical neurophysiology broadly, anterior cingulate cortex specifically, and their contributions to learning about actions. The task is well-designed and provides a nice background for neurophysiological recordings. The authors leveraged these strengths in characterizing the neural populations that fire to shuttle crossings in both directions vs one direction.

      Strengths:

      The factorial design nicely controls for sensory coding and value coding, since the same stimulus can signal different actions and values.

      The figures are well presented, labeled, and easy to read.

      Additional analyses, such as the 2.5/7.5s windows and place-field analysis, are nice to see and indicate that the authors were careful in their neural analyses.

      The n-trial + 1 analysis where ACC activity was higher on trials that preceded correct responses is a nice addition, since it shows that ACC activity predicts future behavior, well before it happens.

      The authors identified ACC neurons that fire to shuttle crossings in one direction or to crossings in both directions. This is very clear in the spike rasters and population scaled color images. While other factors such as place fields, sensory input, and their integration can account for this activity, the authors discuss this and provide additional supplemental analyses.

    4. Reviewer #3 (Public review):

      Summary:

      The authors record from the ACC during a task in which animals must switch contexts to avoid shock as instructed by a cue. As expected, they find neurons that encode context, with some encoding of actions prior to the context, and encoding of neurons post-action. The primary novelty is dynamic encoding of action-outcome in a discrimination-avoidance domain, while this is traditionally done using operant methods.

      Comments on revised version:

      I appreciate subsequent responses to my comments and other reviewers. My comments are addressed, and at this point, I think readers can judge the work appropriately in context.

    5. Author response:

      The following is the authors’ response to the previous reviews

      We thank the reviewers for their additional feedback. Below, we provide detailed responses to each reviewer’s major concerns. In addition, we identified an error in the previously submitted Fig. 6C and have corrected the X-axis labels accordingly.

      Public Reviews:

      Reviewer #1 (Public review):

      Motion-related signal in ACC: the new Fig. 2E looks good, but it is hard to visualize how it is just a reordering of the old Fig. 5C.

      We thank the reviewer for this feedback. Fig. 2E and the original Fig. 5C do bear resemblance. The primary difference is the temporal window and organization of the data. In the original Fig 5C, the time window was only ± 5 sec whereas Fig. 2E is ± 30 sec. The main objective we aim to highlight is that ACC shows both activation and inhibition in response to shuttle on an extremely prolonged order, up to 30 sec. Data is sorted to separate inhibition and activation to illustrate the sustained activity persists for both populations.

      All categories in the new Fig. 4D appear to respond to shuttle initiation, with less than 1s latency. For example, type 2a/2b consists of 40% of the population and their response to movement onset is apparent. Thus, it is not clear whether most neurons respond to shuttle crossing as described in the manuscript.

      We thank the reviewer for drawing attention to this discrepancy. It was not our intention to strike comparison between shuttle initiation versus shutting crossing responses across neurons, and we do not dispute that ACC responds to both events. While shuttle initiations and crossings provide a consistent temporal alignment point, they do not define the temporal focus of much of our analyses. Given that most shuttle responses terminate within ~2 sec, the extended windows analyzed (i.e. ± 5 sec; Fig. 4) largely reflect post-action ACC activity. Overall, although ACC neurons show mixed responses to initiations or crossings, the most consistent feature is prolonged modulation that persists beyond shuttle termination. We have revised the text to reflect this focus.

      Given this and the reviewer’s feedback, we further examined whether ACC activity is more strongly aligned with shuttle initiation, crossing, or termination. To determine which shuttle event (initiation, crossing, or termination) captured the most acute changes in ACC neuronal firing, we conducted an event-locked modulation analysis (Fig. S4). Our results showed that shuttle crossing was associated with the largest fraction of significantly modulated ACC neurons (Fig. S4). These findings suggest that shuttle crossing represents the most prominent event for ACC engagement during shuttle behaviors.

      Could the authors use relatively simple analysis, such as comparing spike rate before and after crossing, or before and after initiation, to quantify the response properties of each neuron? This could also help validate the classification analysis performed in Fig. 4.

      As mentioned above, we have added a new supplemental figure to directly address this question (Fig. S4).

      Reviewer #2 (Public review):

      I think the authors did a very admirable job revising the manuscript. It is much improved. However, I believe a formal analysis of action-state versus action-content neurons on A-->B versus B-->A crossing is still warranted. I appreciate the fact that this analysis may not be as reliable with smaller ensemble sizes, but with careful pseudo-ensemble and resampling approaches, such an analysis would go a long way towards increasing the strength of evidence.

      At present, we are not sure what the reviewer means as “formal analysis”. Below is our best effort in addressing this concern.

      Firstly, in our first revised manuscript, we implemented a generalized linear model-based classification of action-content and action-state neurons using direction specific regressors. Specifically, this analysis classified neurons as action-content or action-state based on coefficient contrasts (Δβ), with appropriate statistical testing and multiple comparison correction (see Methods; Fig. 7 C–E). Neurons were classified as action-content neurons if the corrected p-value for Δβ was significant and the absolute effect size exceeded a predefined threshold (|Δ β |> 0.5). Neurons were classified as action-state neurons if Δβ was not significant but both β1 and β2 were individually significant after correction. We believe our generalized linear model-based classification offers a sophisticated and formal classification of these two neurons classes.

      Subsequently, we performed an SVM decoder to distinguish A→B from B→A shuttles. Decoding accuracy depended on action-content neurons, as their removal drastically decreased decoding accuracy, whereas removal of non-action-content neurons had no effect, further strengthening the conclusion that these populations encode distinct information.

      In the updated revision, we performed an additional SVM decoding analysis while controlling for unequal neuronal population sizes between action-state and action-content neurons (Fig. S8). Specifically, we constructed pseudo-ensembles by randomly resampling neurons within each category and training SVM decoders on size-matched ensembles. Decoder performance was evaluated across repeated resamples to generate distributions of accuracy. We found that only decoders using action-content neuronal activity predicted shuttle content with high accuracy (>95%), whereas decoders trained using non-action-content neurons performed at chance levels (Fig. S8).

      Reviewer #3 (Public review):

      The only remaining comment that was not addressed pertains to anatomy and recording details. Some electrodes appear to be clearly in M2 (Fig 2A), and the tetrodes were driven each day. I would strongly suggest that this be included as a further limitation, particularly given the statement on line 178.

      We thank the reviewer for this feedback. In the previous revision, we added a supplemental figure showing tetrode locations for each mouse (Fig. S2) and described recording details in the Methods (Lines #481–488). We agree that this should also be noted as a limitation, and we have now added this to the Discussion (Lines #384–388).

    1. eLife Assessment

      This study presents valuable findings on the neuromodulatory underpinnings of adaptive learning in dynamic, probabilistic environments. Solid evidence for these claims comes from showing spatial correlations between model-derived fMRI responses and PET-based receptor density maps. The work will be of interest to cognitive and systems neuroscientists working on decision-making.

    2. Reviewer #1 (Public review):

      Summary:

      This study investigates whether the distribution of receptors and transporters of neurotransmitters accounts for the topography of cortical activity of confidence and surprise in probability learning. The authors first examined the invariance of functional correlates of confidence and surprises with multiple fMRI studies and then investigated whether 20 PET-derived receptor and transporter density maps account for this cortical invariant activity of confidence and surprise in probabilistic learning. Beyond these specific findings, the main novelty of this study lies in its attempt to bridge neuromodulatory systems and cognitive processes using neuroimaging data. This integrative approach is particularly valuable, as it showcases a framework to combine neurochemical architecture and cognitive computations.

      Strengths:

      This study attempts to link neuromodulatory systems with cognitive processes involved in probabilistic learning. Although the role of neuromodulatory systems in learning has been highlighted in several influential previous studies, it has not yet been widely investigated or systematically related to functional neuroimaging data so far. The authors used an efficient approach to address this question by combining group-averaged neurotransmitter maps with functional results from multiple fMRI studies using probabilistic learning tasks with similar structures. This approach provides informative insights into the relationship between the distribution of neuromodulatory systems and cognitive processes from neuroimaging data.

      Weaknesses:

      One limitation of the study stems from the unavoidable constraints of relying on pre-existing datasets rather than data specifically collected to address the present research question. Because the four fMRI studies differed in their measurements and task structures, the authors defined confidence and surprise on the basis of ideal observer behavior. Thus, "confidence" and "surprise" are not related to individual decision or subjective value, and the PET data is also from group-level data. Thus, it certainly has a limitation in linking with individual learning performance and brain activity. Also, "surprise" in this study does not seem to capture the nature of "surprise" in the learning process, which is a violation of expectation, as it was calculated with improbability. Moreover, the correlation of Study 1-4 for surprise was not consistent and not strong enough to argue for spatial invariance. Thus, these results may not yet be fully conclusive.

    3. Reviewer #2 (Public review):

      Summary:

      Learning in dynamic, stochastic environments is difficult, and neuromodulatory systems may shape where learning signals appear in the brain. Using fMRI from four probabilistic learning studies and a Bayesian ideal observer model, the authors examined latent variables driving learning, such as confidence and surprise. They found that brain activity related to confidence, and to a lesser degree surprise, is highly spatially invariant across tasks and modalities, suggesting a stable cortical organization. This invariant pattern aligns with PET-derived maps of receptors and transporters, implicating catecholamine and opioid systems, and supporting a neuromodulatory account of adaptive learning with receptor-level hypotheses.

      Strengths:

      (1) Elegant combination of computational modelling, functional magnetic resonance imaging (fMRI) and positron emission tomography (PET).

      (2) The authors describe results of four separate experiments, with very similar results, in effect providing internal replications.

      (3) Cross-validated results compared against a meaningful null model.

      Weaknesses:

      (1) Unclear rationale for using one-sided statistics (e.g., in Figure 3). One-sided tests appear to be invalid, given that the Introduction lacks a preregistered directional hypothesis at an operationalised level. This may have consequences for the following statement in the Discussion: "The associations between receptor architecture and functional topography were substantially weaker for the language network, which is not thought to rely strongly on neuromodulatory systems."

      (2) Limited computational modelling. Since learning rates probably differ across subjects, I wonder if they have considered fitting the "volatility" instead of using the generative one. Would that give more meaningful fMRI maps, and better explained variance when correlating these to the PET-based predictors? I was also wondering how their surprise measure relates to "change-point probability" (e.g., Murphy et al., Nat Neurosci, 2021). Finally, I think it would be helpful to show average time courses of surprise and confidence time-locked to state changes.

      (3) Lack of GLM validation. It would help to show that the model fits the data well. This is important given the many underlying assumptions (shape of the HRF, linear effects of variables, etc). For example, one could show average insula activity time-locked to state changes, as well as the model-predicted activity, and separately for three strata defined by how surprising the state change was (according to the ideal observer model). Related, the authors use a substantial number of predictors in their GLM, and the language in the Methods is a little casual. It would help to show part of a design matrix, and clearly describe the following: were (occasional) questions and responses modelled by separate stick functions? Which predictors (stimulus, questions, response) varied parametrically with which variables?

    4. Reviewer #3 (Public review):

      Summary:

      In this unusual paper, Hodapp and Meyniel relate the spatial topography of activity maps for confidence and surprise (from four learning tasks) to the spatial topography of receptor density maps from atlas data. They find that the brain maps for confidence and surprise are largely consistent across four studies using different stimuli/ task demands. They then use a general linear model to predict the spatial pattern of confidence/surprise-related activity from the spatial distribution of receptors (receptor types) for several neuromodulators. Further analyses test which neuromodulators are most important for predicting the functional maps.

      Strengths:

      The study gives an interesting new perspective on the brain networks for surprise and confidence, indicating that one reason for the involvement of different networks with these computational parameters is the neurochemical sensitivity of tissue within those networks.

      Weaknesses:

      I felt the paper was light on context.

      To what extent are the distributions of receptor types correlated with each other?

      What does the spatial topography of receptor density look like for the identified receptors (NET, MOR, 5HT1b)? Could these be displayed alongside the functional networks? I realise these are atlas data, but for me to interpret the result, I'd need to see the map, and I don't want to download the atlas.

      To what extent are the correlations with receptor maps network-wide, vs being driven by one big patch of activity in a single region with high receptor density? To me, this would be important - does this study demonstrate that distant regions united in a functional purpose by shared receptor profile (which would in my opinion be more intersting that the alternative, that there is a single region within each network driving the effect).

      Finally, I wasn't convinced by the spin test in this particular application. To my mind, permutation tests are valid when the permuted points are interchangeable under the null. The spin test, as used, preserves the distribution and spatial pattern of activity, but assumes that it could equally plausibly be relocated to any angle on the 3D surface (under the null). However, the brain has a lot of structure that is non-uniform across its surface (connectivity patterns and histological boundaries being important ones). The observed data probably follow this structure, but the 'spun' or permuted datasets probably overlay randomly on the connectivity structure (for example), so that one blob of activity has uniform connectivity in the real data, but overlaps the projections of multiple white matter tracts in the permuted data. But then the permuted data would likely be more heterogeneous in terms of both function and histology than the original data. Since connectivity, histology (layer structure) and receptor density are likely correlated, I think it must be impossible to find verticies that differ in one modality whilst being interchangeable in all others, therefore it may not be possible to use permutation logic to make a claim about (say) receptor density independently of connectivity and histology.

      I should add I'm not sure how one would carry out a permutation test that respects the underlying brain anatomy here, or whether this is even possible; that is a difficult question.

      I would add that I think the observation that the functional networks have different receptor profiles is interesting, even as a qualitative observation, but not convinced the statistical approach can be justified.

    1. eLife Assessment

      This valuable study uses naturalistic movie-viewing fMRI and stacked encoding models to investigate sensory feature representations in autistic and non-autistic youth, showing a relative shift toward low-level visual representations in higher-order social cortical regions in autism. The evidence is solid overall, supported by preregistration, a relatively large open dataset, and sophisticated encoding-model analyses, although several methodological and interpretive issues require further clarification and validation. The work will interest researchers in developmental cognitive neuroscience and naturalistic neuroimaging.

    2. Reviewer #1 (Public review):

      Summary:

      This study uses stacked encoding models to characterize differences in sensory (visual and auditory) processing between autistic and non-autistic children and adolescents. The authors found no significant enhancement of low-level feature encoding in either visual or auditory cortex, but reduced high-level visual representations and a relative shift toward low-level over high-level visual feature encoding in the posterior superior temporal sulcus (pSTS). The shift in pSTS correlated with social symptom severity (SRS scores). These findings support weak central coherence (WCC) theory over enhanced perceptual functioning (EPF) theory, suggesting an altered visual feature encoding in pSTS in autism.

      Strengths:

      This study uses sophisticated methodology and an open data set with a relatively large sample size. fMRI data are acquired during a naturalistic paradigm (i.e., movie watching), which promotes attention and engagement among participants, and provides greater ecological validity. The use of encoding models to explore population-level differences in neural representations of stimulus-computable features is novel. Overall, results provide somewhat modest yet still informative evidence for adjudicating between possible theories of altered sensory processing in autism.

      Weaknesses:

      Some important methodological details are missing and/or require justification. Some potential confounding factors or unconsidered differences between individuals and/or diagnostic groups should be explored and possibly addressed. Specific major and minor points are raised below.

      Major comments:

      (1) Unclear description of noise ceiling calculation (line 205-206, 632-634) and potential heterogeneity: it is not clear what data were "split" for the split-half correlation used to calculate noise ceilings. To our knowledge, each participant watched each movie once each, so there is no within-subject repetition available. Were these correlations across participants (i.e., ISC)? If so, does this across-subject metric provide a fair representation of the true noise ceiling, given that a) encoding models themselves are trained within subjects and b) autistic individuals are known to exhibit more idiosyncrasy in responses to naturalistic stimuli (e.g., Hasson et al., 2008)? Moreover, do noise ceilings differ between individual participants, diagnostic groups, and/or with age? If so, how might these differences affect the interpretation of results (e.g., R2 differences)?

      (2) Possibly underperforming visual model: given that the visual model in general performed worse than the audio model, the visual vs audio perceptual preference analyses (line 281-290) might be affected by the underlying mismatch between model performance. Though the visual and auditory regions showed similar noise ceilings (Figure 2 S1B), the stacked model performed better in auditory regions than in visual or multimodal regions (Figure 2 S1A). Supporting the same idea, the visual model in general showed lower fitting R2 than the audio model (Figure 2 S2A, Figure 2 S3A vs B). Instead of using mean motion (line 608-614), applying PCA on the raw features might help reduce noise inherent in the raw motion energy features (Malik et al., 2026), therefore improving model performance.

      (3) The clipping procedure for unique variance (lines 634-637) requires justification: the unique variance is defined by subtracting high-level R² from stacked R² with explicit clipping when high-level R² is negative or exceeds stacked R². However, in the original stacked regression framework (Lin et al., 2024), unique variance is defined by simple subtraction without such post-hoc adjustment, as the negative R2 is still meaningful, indicating the model performs worse than predicting using the mean value. This requires justification. How frequently does clipping occur, and in which brain regions? Is it an indicator of overfitting or poor model performance? How substantially do results change if clipping is removed? E.g., the hemisphere dominance comparison (line 271-280, Figure 6). Critically, does this procedure affect the key finding regarding SRS/sensory symptom severity correlations in pSTS?

      (4) The interpretation of the correlation between SRS with neural patterns is misleading (line 237-242, line 364-366): based on Figure 3, SRS and SSS showed more significant and robust relationship with unique variance of high-level visual feature, meaning that the decrement of high-level feature encoding in STSvp and STSdp, rather than the relative low-level preference, is likely driving the relationship with autism severity and sensory symptom.

      (5) Details are missing about how data from the two movie runs were combined. Were the time series concatenated without regard to which movie they originally came from, or was the distinction between movies taken into account for purposes of splitting data into train/test cross-validation folds? The results would be stronger if the authors could show that results replicate across the two movies when they are each analyzed independently, though we recognize that there is perhaps not enough data, especially in the shorter [~4min] movie, to do this. The authors discussed this in lines 412-417, but it would be helpful to provide a justification in the Methods section as well.

      (6) Potential feature weight differences across individuals and/or diagnostic categories: since the encoding models were trained for each subject, is there significant variability in feature weights across individuals and/or diagnostic categories (e.g., did the model predictions heavily rely on face for the non-ASD group but not for the ASD group)? If so, how does this change the interpretation of the R2 comparisons? The authors showed the results of stacked feature weight differences between diagnostic categories and their relationship with autism severity and sensory symptoms, but it might be informative to show the raw feature weightings before diving into stacked-weight differences.

    3. Reviewer #2 (Public review):

      Summary:

      This study by Mentch et al. uses naturalistic-movie fMRI and grayordinate-level stacked encoding models to test preregistered hypotheses about low/high-level and audio/visual feature encoding in autism and adolescence from openly available Healthy Brain Network data. Null results reported that autism was not linked to increased low-level encoding in primary sensory cortices. Exploratory analyses showed participants with autism showed reduced high-level visual encoding in social regions (pSTS, face areas), with the high-low feature shift tracking social responsiveness scale (SRS) scores. Age and laterality effects were also found.

      Strengths:

      (1) This study and hypotheses were preregistered.

      (2) The study utilised proper variance partitioning, split-half noise ceilings, FD-threshold sensitivity analyses, and an explicit modelling framework that recovers known sensory hierarchies in the aggregated sample. The developmental sampling adds to the interest.

      (3) The manuscript is written clearly, laying out the background and theories to be tested with encoding models. The analyses and reporting of results are clear.

      Weaknesses:

      (1) If I understand correctly, by only averaging the grayordinates that already passed a significance threshold, the resulting parcel value is guaranteed to look stronger than if all grayordinates had been included. This has been raised in neuroimaging (Kriegeskorte et al., 2009; Vul et al., 2009). Can the authors justify these choices?

      (2) I assume that the phrase "temporally permuting the order of observations" on Page 22 means random shuffling of time points. The details of this exact permutation are not specified. Both the fMRI BOLD signal and movie features have strong temporal autocorrelation, and random shuffling will destroy this structure. This is important as grayordinate-level survivors will propagate to parcel pools. Circular shifting or phase randomization preserving the autocorrelation spectrum is appropriate.

      (3) In the movie feature selection, the low-level visual model contains only two scalars: mean perceptual brightness and a single averaged value across 2,139 motion-energy filters. With only two low-level visual features, the low-level visual model potentially would underestimate low-level visual encoding. The H1.1 toward the null perhaps suggests to this. Principal components of the motion-energy outputs, as was done for the cochleagram, could be used.

      (4) The pilot sample composition is not described. Features were selected based on their performance on an independent set of 54 pilot subjects. Please provide age, sex, and diagnostic composition of the pilot sample. The main point being whether the selected features were optimised for a population that differs from the subject studied.

      (5) The authors acknowledge the lack of eye-tracking in theory study. I think this should be elaborated, especially why this modality is important for answering sensory and perceptual encoding. Face encoding may not be degraded, but just that faces are not being attended to.

      (6) I think a more nuanced distinction about the representational nature of encoding-model R² should be mentioned, especially when the interpretation of findings is related to perceptual functioning (EPF theory). R² measures how well a feature set predicts brain activity, not perceptual function or cognitive integration.

      (7) The literature also includes evidence for no Colavita effect, not just reverse Colavita in autism, and the framing should reflect this more even-handedly.

      (8) The 0.2 mm per-volume threshold is quite strict. The 40%/60%/80% sensitivity analyses partially address this, but a brief justification for the choice of 0.2 mm would strengthen the Methods.

      (9) Figure 1 seems confusing and would benefit from more information or text in the figure.

      (10) Figure 2 supplement has caption A labelled twice; please correct.

      (11) Acronyms. Please spell out MSI on first mention (page 2) and ISC/ISFC on first mention (page 4).

    4. Reviewer #3 (Public review):

      Summary:

      This study investigates the neural mechanisms underlying sensory-perceptual differences in autism through a naturalistic movie-viewing fMRI paradigm. By employing encoding models, the authors demonstrate that autistic children and adolescents exhibit a specific alteration in visual feature weighting, characterized by a shift toward low-level visual feature encoding in higher-order association regions, particularly the posterior Superior Temporal Sulcus (pSTS). This shift is linked to social symptom severity, providing empirical support for Weak Central Coherence accounts.

      Strengths:

      The study's primary strengths lie in its methodological rigor and innovative approach. The use of a pre-registered analysis plan ensures transparency and enhances the credibility of the findings, while the encoding models allow for a fine-grained dissociation of low-level versus high-level feature representations across the cortex. Overall, the writing is clear, the logic is sound, and the results offer a significant contribution to the field by refining our understanding of how sensory processing is differentially organized in autism.

      Weaknesses:

      While the study presents compelling findings regarding visual feature encoding in autism, several methodological and interpretive limitations warrant consideration. First, the Discussion focuses primarily on WCC and EPF theories, failing to explicitly address how the results intersect with other prominent frameworks mentioned in the Introduction, such as Bayesian predictive coding or E/I imbalance hypotheses. Second, the demographic characteristics and specific sample sizes of the ASD-ADHD and ASD+ADHD subgroups are not reported, limiting the interpretability of the stratified analyses; furthermore, the counterintuitive finding that the ASD+ADHD group resembles controls is not sufficiently discussed. Third, given the significant group difference in IQ and the known relationship between cognitive ability and neural processing, the potential confounding influence of IQ on the neuroimaging results requires more explicit acknowledgment, particularly since IQ was not included as a covariate in the primary models.

    1. eLife Assessment

      This valuable study uses an elegant visual-anagram approach to test whether perceived animacy structures visual working memory and attention while controlling for many low-level image properties. The evidence is solid, with converging results across seven preregistered experiments, but the central claim that animacy itself is represented independently of visual features should be tempered, as residual mid-level configural cues, ensemble or category structure, and broader semantic differences may also contribute to the effects. The work will be of interest to researchers studying high-level visual representation, attention, and working memory.

    2. Reviewer #1 (Public review):

      Summary:

      Evidence for visual representation of animacy.

      Strengths:

      This is a very cool paper that casts light on a persistent problem in the psychology and philosophy of visual representation: is there high-level perception? Every vision scientist agrees that low-level features such as shape, color, texture, motion and spatial frequency are represented in visual perception, but there is a great deal of controversy about the representation of high-level properties such as causation, faces, agency and animacy. Animacy is especially problematic because there are large differences in line curvature between stimuli that represent animate and inanimate items.

      This article uses a novel approach-visual "anagrams" that are exactly the same image, except one is rotated 90 degrees relative to the other. They found persistent differences in visual processing between animate and inanimate stimuli. (Of course, the stimuli aren't animate-they represent animate items.). For example, there were processing differences between changes between animate and inanimate items (rabbit to boot) that were not present in rabbit to dog. They also showed such differences in two kinds of visual search tasks.

      Of course, there are feature differences that exploit orientation. A classic example is the difference between a square and a diamond that is produced from the square by rotating it 45 degrees.

      They addressed an aspect of this challenge having to do with some features using silhouettes. There was no search advantage for silhouetted stimuli.

      Weaknesses:

      I thought this was an excellent submission. I have two suggestions for revision:

      (1) I thought that experiment 7 should have been described in more detail, with the upshot explained better. What exactly do the authors take it to show?

      (2) There should be a candid discussion of what the loose ends are and how they might be addressed. It would be good to have some examples like the square/diamond case with some indication of what would address such challenges.

    3. Reviewer #2 (Public review):

      Summary:

      The authors present a creative approach using visual anagrams matched on low-level image statistics to isolate animacy from low-level visual features and report consistent effects of animacy on visual working memory and attention. While this is a thoughtful design and is well executed across seven pre-registered experiments, it remains unclear whether the reported effect is truly driven by animacy, as opposed to broader differences in ensemble statistics or semantic structure across the "mixed animacy" versus "uniform animacy" conditions. As such, the interpretation of a "pure" animacy effect may be overstated.

      Strengths:

      (1) An important methodological advance in controlling low-level confounds that have historically complicated the study of animacy.

      (2) The converging effects across multiple experiments, together with the pre-registered design, strengthen the reliability of the reported findings.

      Weaknesses:

      (1) Specificity of the animacy effect vs. category-level ensemble structure

      The central claim is that animacy itself drives the observed effects. However, the key manipulation ("mixed animacy" versus "uniform animacy") also introduces differences in category-level ensemble structure. For example, in Experiments 1-2, cross-category change detection (e.g., dog to chair) may be easier not because of animacy per se, but because of a change in overall ensemble statistics (Brady & Alvarez, 2011, 2015). In addition, since each display contains five objects (two in one category and three in the other category), cross-category changes may also alter category balance in a way that further facilitates detection. In contrast, within-category changes preserve both ensemble structure and category composition, making them more difficult to detect.

      Brady, T. F., & Alvarez, G. A. (2011). Hierarchical encoding in visual working memory: Ensemble statistics bias memory for individual items. Psychological Science.

      Brady, T. F., & Alvarez, G. A. (2015). Contextual effects in visual working memory reveal hierarchically structured memory representations. Journal of Vision.

      (2) Limited stimulus set and potential learning effects

      The relatively small stimulus set (six anagram pairs) and repeated exposure raise the possibility of learning or familiarity effects. Does performance change over time? e.g., are there meaningful differences between early and late trials (e.g., first 10% vs. last 10%)? If such differences are present, they could suggest the development of task-specific strategies or increased efficiency with repeated exposure, rather than stable effects driven by the experimental manipulation itself.

      (3) Role of semantics

      Although the anagram paradigm effectively controls low-level visual features, it still relies on high-level semantics (e.g., "dog" vs. "boot"). These stimuli differ not only in animacy but also along other semantic dimensions such as natural versus manmade categories. From a semantic standpoint, it remains unclear whether the observed effects can be uniquely attributed to animacy or whether they reflect broader conceptual distinctions.

    4. Reviewer #3 (Public review):

      Summary:

      This study makes clever use of generative AI to create stimuli that are pixel-for-pixel identical but which have radically different meanings depending on their orientation, to investigate the perception of animacy while retaining control over low-level image features (so-called 'anagram' stimuli).

      The authors present seven elegantly designed experiments in a commendably compact format.

      Experiments 1 and 2 involved a working memory paradigm in which participants had to spot which of five objects in an array changed after a pause. Importantly, the changed object was an anagram stimulus that in one orientation matched the animacy/inanimacy of the changed object, and in the other orientation was the opposite (e.g., a rabbit is replaced by either a dog or a boot, where the dog and boot stimuli are actually identical, just rotated by 90 degrees). They found a difference in accuracy depending on whether the animacy of the objects matched.

      Experiments 3 and 4 used a visual search task in which the participants had to localize the target, and the distractors were anagrams that either matched the target in terms of animacy or did not. There was a significant cost in terms of response time when the animacy of the target was the same as that of the distractors. Experiments 5 and 6 also used a similar visual search design, except that the task was to determine if the target was present or absent from the display, and the distractors again either matched or differed from the target in terms of animacy. Again, the authors found slower responses when the distractor arrays matched the animacy of the target than when they differed.

      An obvious potential concern about the studies is addressed by Experiment 7. It is unclear if the observed effects are related to the specific orientations of the target and distractor stimuli selected in each condition. For example, it could be that all the animate versions of the anagrams involved tall and skinny shapes, while all the inanimate versions involved wide and short objects, due to the 90-degree rotational difference between the two versions of the stimuli. To control for this, the authors repeated the visual search experiment but with convex-hull silhouettes of each of the stimuli. In other words, all targets and distractors from each trial were replaced by a black splotch with approximately the same overall outline (envelope) as the corresponding stimulus. Importantly, in contrast to the anagram stimuli, the silhouettes had had no meaningful semantic interpretation, and their animacy did not change depending on their orientation.

      Strengths:

      The main strength is the elegant use of stimuli that control almost perfectly for low-level image features.

      Weaknesses:

      My only real concern about the study is whether the findings truly provide evidence for a high-level visual representation of animacy independent of the low-level stimulus characteristics, or whether, instead, the effects are essentially semantic priming, which is independent of visual processing per se. For example, if all the stimuli in the experiments were replaced with the verbal names of the depicted objects instead of pictures, would we expect different results? Words can also access semantic representations of the animacy of objects, and also don't suffer from low-level visual confounds. It would be helpful to add a discussion of this possibility to the article.

    5. Reviewer #4 (Public review):

      In this article, the authors investigate whether perceived animacy influences visual processing independently of lower-level visual features by using "visual anagrams." Across seven experiments, they test whether animacy, isolated from many lower-level visual properties, structures visual working memory and guides visual attention. The central claim is that the visual system may represent animacy itself, rather than animacy emerging solely from associations among low-level visual properties.

      I find this investigation compelling. The experiments described provide strong control over several lower-level visual features, including curvature, texture, and related image properties. However, the visual anagrams are not pixelwise-identical across orientations. Because the images are rotated, the retinal configuration of pixels and the spatial organization of some low- to mid-level shape features also change. As a result, the configural arrangement of mid-level visual features may still contribute to perceived animacy.

      I encourage the authors to discuss how independent perceived animacy is in this context from the contribution of mid-level visual features, such as configural shape cues that are diagnostic of animacy. This distinction would help sharpen the interpretation of the results and more precisely define the level of visual representation isolated by the visual-anagram approach.

      Additionally, previous studies have argued that low- and mid-level curvilinear features may contribute to animate/inanimate categorization, and may in some cases be sufficient to support such distinctions (e.g., PMID: 33798259; PMID: 28654965). I encourage the authors to clarify how these previous findings on curvilinearity and rectilinearity fit with the overarching claim of the current study, namely that the visual system may represent animacy itself rather than animacy emerging solely from associations among lower-level visual properties.

    1. eLife Assessment

      This useful study combines behavioral reports, EEG decoding, and computational modeling to address an interesting question: how delay-period distractors bias working-memory representations, and how these effects depend on target relevance, distractor location, and the strength of memory maintenance and distractor encoding. However, the supporting evidence is incomplete, as several key claims require clearer statistical tests, better integration of the behavioral and neural results, and more careful consideration of alternative explanations. Stronger engagement with prior literature would also substantially strengthen the manuscript and increase its potential interest to researchers in systems, cognitive, and computational neuroscience.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, Deepak V. Raya and colleagues combined behavioral measures with EEG recordings to investigate how distractors presented during the working memory delay influence memory representations. Using oriented gratings as stimuli and a continuous estimation task, the authors systematically manipulated factors that may modulate distractor interference, including the behavioral relevance of the WM item (cued vs. uncued) and the spatial relationship between the distractor and the WM item. By analyzing the relative orientation between the WM item and the distractor, the authors showed that distractors presented at the same location as the WM item induced an attractive bias (i.e., reported orientations biased toward that of the distractor), whereas distractors presented at the opposite location produced a weaker effect, with any systematic bias tending to be repulsive. Through a combination of behavioral analyses and EEG-based decoding, the authors further examined and revealed factors that modulate the magnitude of distractor interference, including cueing status, the strength of memory maintenance, distractor timing, and neural indices of distractor encoding and gating. Lastly, the authors propose a computational account of these effects by implementing a two-layer ring attractor model that captures several key behavioral patterns observed in the data.

      Strengths:

      The influence of distractors on working memory has been extensively studied both behaviorally and with neuroimaging. The present study advances this literature by providing a more comprehensive account that jointly manipulates and quantifies many key factors, including cueing (behavioral relevance), the spatial relationship between WM items and distractors, and distractor timing. This integrative approach enables a more systematic characterization of how different sources of interference interact. A particular strength of the study is the use of EEG combined with multivariate decoding to track the dynamics of memory and distractor representations. Compared to prior fMRI work, this approach provides a time-resolved view of how encoding, maintenance, and distractor processing unfold over time. This is especially valuable for dissociating memory maintenance and stimulus encoding, or gating contribute to behavioral interference, which is more difficult to achieve with fMRI.

      Behaviorally, while most previous studies have reported attractive biases by distractors, the current study identified a repulsive effect when distractors were in the opposite hemifield from the WM item. Overall, the study provides a rich investigation of distractor interference in working memory and will be of interest to researchers studying the neural and computational mechanisms that protect memory representations from distraction.

      Weaknesses:

      (1) In the paragraph starting around line 125, the authors reported a 2-way ANOVA (cue/uncued × same/opposite side) restricted to trials in which a distractor was present. However, the subsequent post-hoc analyses compared distractor-present trials (same or opposite side) with no-distractor trials, which were not included in the ANOVA. While both analyses were informative, presenting them together in this way was somewhat confusing, as the post-hoc tests extended beyond the factors and conditions analyzed by the ANOVA. I suggest presenting these analyses separately and clarifying their distinct purposes. Additionally, Figure 1C appeared to reflect only the pairwise comparisons; including a figure that directly visualizes the two-way ANOVA results would improve clarity.

      (2) In lines 138-150, the authors fitted von Mises functions to the distributions of memory error and reported that the effect of distractor location (same vs. opposite) was stronger in the uncued condition than in the cued condition. However, this result appears difficult to reconcile with the earlier 2-way ANOVA, which showed no interaction between cueing and distractor location. It is unclear whether this discrepancy arose from differences in the dependent measures (CSD vs. κ), statistical procedures, or other factors. Clarifying how these two sets of results should be interpreted together would improve the clarity of the findings.

      (3) For the analyses in Figures 1B and 1D, parametric functions were fitted to the distributions of memory error using aggregated data. Models of memory error distributions have been central to ongoing debates in the working memory literature (e.g., Schurgin, Wixted, & Brady, 2020; van den Berg, Awh, & Ma, 2014). Fitting functions/curves to aggregated data can be problematic, as it distorts the underlying distributions at the individual level. I suggest performing the fits on the individual data and analyzing the fitted parameters across participants using appropriate group-level statistical tests.

      (4) At the end of the first Results section (lines 234-235), the authors concluded that cued memoranda were "better shielded from interference" than uncued memoranda. However, I did not see a clear statistical test directly supporting this. This statement appeared to rely mainly on Figure 1D, which showed a stronger location effect (same vs. opposite) when the memory item was uncued. However, this analysis does not directly test whether distractors impair uncued items more than cued items overall. Supporting this broader claim would require a direct comparison of distractor effects (e.g., distractor vs. no-distractor) between cued and uncued conditions, or an interaction test involving cueing and distractor presence (e.g., either by pooling different distractor locations, or focusing on the same-location condition if opposite-location distractors show no significant effect).

      (5) While the attractive and repulsive biases are an interesting finding, it was demonstrated only at the behavioral level. It would be informative to examine whether the biases are reflected in the decoding results. For example, after deriving trial-wise orientation tuning functions, one could estimate decoded orientations (e.g., via vector averaging or the peak of the tuning curve) and assess bias at the neural level. Although EEG SNR may limit recovery of full function of the memory error (e.g., Figure 1F-G), grouping trials into fewer bins (even with just two bins) may still allow detection of the overall direction of the bias in the decoding results. This type of decoding bias has been reported in other contexts (GY Bae - NeuroImage, 2021).

      (6) The analysis P2/P3a requires more explanations. Typically, these components are extracted from trial-averaged ERP. The methods section also mentioned "averaged across channels and trials to obtain the ERP waveform." However, to split the trials, these components have to be identified at a single-trial level. More details are needed in the Methods.

      (7) Components such as P3a are often linked to attentional capture and orienting, which would predict increased, rather than decreased, distractor interference. The interpretation of this signal as reflecting gating appears to be inferred from the observed relationship between larger P3a amplitudes and weaker interference. The N2pc component is a well-established index of spatial attention allocation and may be particularly relevant (and useful) here, given the lateralized distractor. Have the authors tested whether distractor-evoked N2pc can be used to split trials and examine its relationship with the bias?

      (8) Line 676 in the Discussion states "possibly by error-correcting top-down control mechanisms." It is unclear which results provide support for this interpretation, except that there are stronger feedback connections at the cued location in the attractor ring model.

    3. Reviewer #2 (Public review):

      Summary:

      Understanding the factors and mechanisms underlying the deleterious effects of distraction, and protection from distraction, in working memory is an important question that has a long and rich history in psychology and neuroscience, and continues to be highly relevant. In this study, the authors recorded the EEG while subjects viewed the initial presentation of two oriented-grating stimuli, aligned on either side of fixation along the horizontal meridian (memory array), followed by a 70%-valid cue, then one of three distractor conditions (overlapping cued item (40%), opposite cued item (40%), no distraction (20%)), followed by recall ("delayed estimation"). The behavioral and EEG results from this procedure are complemented with computational modeling with a two-tier bump-attractor model.

      Weaknesses:

      Interpretation of the results is complicated by several factors. One is the non-consideration of a considerable amount of extant research that is highly relevant to the question of interest (these include seminal studies from Gi-Yeul Bae and from Tatiana Pasternak). Relatedly, the manuscript emphasizes biasing effects of distractors to the exclusion of a conceptually distinct effect: degradation of representational precision. (For example, the actual focus of the study of Wimmer et al. (2014) that the manuscript cites with reference to bias is the degradation of precision; one only has to read the title of this paper to know this.) Also relatedly, the authors are aware of the possibility of misbinding (a.k.a. "swap") errors, in which subjects mistakenly recall a high-fidelity representation of a foil (in this case, the distractor) rather than the target, but they (1) fail to cite any of the extensive literature on this topic and (2) seem to erroneously attribute what their analyses would seem to identify as misbinding errors as "antagonistic bias" exerted by the distractor on the target item.

      A second concern relates to the interpretation of patterns in the empirical results. In particular, Figure 1G is interpreted as displaying a pattern of repulsive bias exerted by the distractor on trials when the distractor appeared at the location opposite to the cued item. However, it is not clear that this is a repulsive bias. Rather, what the plot shows is that report error is attracted to "near" distractors with a positively signed offset but repelled by "near" distractors with a negatively signed offset. Stated another way, when one applies a model-free assessment of the influence of the distractor on the memorandum, there is no systematic bias: the AOC of positively signed offset values from 0 to +45 deg is roughly the same as the AOC negatively signed offset values from 0 to -45 deg. The same also seems to be true, albeit with a smaller magnitude, for trials featuring "stronger mnemonic neural representation" that are illustrated in Figure 2. And so it's unclear that the effect of the "Dist. Opp" distractor is indeed a repulsive bias, rather than a loss of precision.

      The third primary concern is that the results from simulations from the two-tier bump-attractor modeling are difficult to interpret due to several poorly motivated and seemingly "hand-coded" assumptions. These include the (seemingly arbitrary) strengthening of HCVC feedback connections by 20% for cued vs. uncued items; and the choice to "transiently block[ed] feedforward connections from the VC to the HC during the maintenance epoch" as a consequence of cuing. There is frankly no evidence that the latter phenomenon actually happens in primate brains performing comparable tasks, including in papers (such as from Xu and from Rademaker) that are cited in this manuscript. The current consensus is that priority-related rotations of representational geometry are the scheme employed by mammalian nervous systems to control the otherwise deleterious effects of distraction.

    1. eLife Assessment

      This revised manuscript retreats from the original claim of establishing a causal link between cardiolipin deficiency and the progression from steatotic liver disease to steatohepatitis and instead advances a more limited mechanistic conclusion: that cardiolipin deficiency perturbs electron transport and promotes electron leak from the mitochondrial respiratory chain. The experimental evidence supporting this revised claim is now solid, and the potential for increased electron leak to contribute to liver pathophysiology is demonstrated. However, absent evidence that cardiolipin deficiency is causally upstream of disease progression, the overall significance of the work remains limited. While the study provides a convincing analysis of mitochondrial bioenergetics, the narrowing of its central claim diminishes its impact relative to that proposed in the original submission.

    2. Joint Public Review:

      Cardiolipin, is a key lipid constituent of mitochondrial membranes. Perturbation of its abundance is thus poised to affect broad aspects of mitochondrial function. Given the important role of mitochondria, it is not surprising that cardiolipin deficiency would have pervasive effects on cell physiology.

      The original version of this paper advanced the idea that cardiolipin deficiency, and the attendant mitochondrial dysfunction, plays a causative role in the progression of fatty liver (a common feature in the human population) to a more pathogenic inflammatory state known as steatohepatitis. Given the prevalence of this form of liver disease in the human population this claim for discovery was deemed sufficiently interesting to merit peer review at eLife.

      Peer review reaffirmed the importance of the claim but also revealed important limitations in the experimental support provided. Specifically, the lack of experimental interventions that uncouple the correlation between progression in a mouse model and changes in cardiolipin abundance to test the causal relationship. The review process also recognised the utility of other aspects of the paper, namely the evidence implicating cardiolipin deficiency in altered properties of the mitochondrial membrane, its contribution to an electron leak and the potential for these features to contribute to pathology.

      The revised version of the manuscript now focuses on the importance of cardiolipin sufficiency to mitochondrial integrity and contains various improvements to the data supporting this aspect. At the same time the revised paper retreats from the most interesting claim of a causal role for cardiolipin deficiency in disease progression. We are left with a more convincing but less significant paper.

    3. Author response:

      The following is the authors’ response to the original reviews.

      As the reviewers noted, the evidence we provide is the strongest on the mechanistic link between hepatic cardiolipin deficiency and electron leak from the electron transport chain. This narrative is supported by our assessment of site-specific electron leak as well as reconstitution of exogenous cardiolipin in the small unilamellar vesicles deficient with CL. On the other hand, as pointed out by the Reviewer 2, the mechanistic link between cardiolipin to MASLD/MASH is less robust. At this moment, we have not experimentally demonstrated that the MASLD/MASH induced by CLS deletion can be rescued by replacement of mitochondrial CL in vivo. Taken together, our current narrative makes an incomplete loop between CL deficiency, electron leak, and MASLD/MASH. Nevertheless, as indicated by all the reviewers, this manuscript highlights a previously undescribed role that CL potentially plays in MASH pathology, particularly with the data that human MASH coincides with reduction in liver mitochondrial CL. We focused this revision primarily on additional descriptive experiments in CLS-LKO mice that were requested by the reviewers. Even though it is not a component of the current manuscript, we have recently successfully developed mice with hepatocyte-specific CLS overexpressing mice and began performing experiments to test causality of CL deficiency to MASLD/MASH which we hope to complete in a few years. We are hopeful that the MASLD/MASH research community will still find evidence on CL contained in this manuscript plausible, and that it provides critical information to our understanding of mechanisms for MASH pathogenesis.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Brothwell and colleagues describes a central role for hepatic cardiolipin deficiency in MASH. The authors identify cardiolipin as a mediator of two long-standing problems in the field: how dysregulated lipid metabolism relates to altered mitochondrial metabolism during MASLD, and what the innate changes are in the steatotic liver that cause the increased respiration. The authors identified reduced liver cardiolipin in humans with MASH and in a variety of mouse models with MASH. When they knocked out hepatic cardiolipin synthesis, mice developed steatosis and inflammation. These mice also recapitulated the elevated hepatic oxidative metabolism and oxidative stress found in obese humans with MASLD. Some of the in vivo functional data related to glucose homeostasis and substrate metabolism could be stronger, and interpretation of the in vitro flux data needs some clarification, but in both cases, the data are not essential to the main conclusions of the manuscript. Overall, the study offers compelling evidence that cardiolipin is reduced in MASLD and that impaired cardiolipin synthesis is sufficient to recapitulate many features of MASLD.

      We thank the reviewer 1 for the positive feedback emphasizing novel and important findings in our manuscript.

      Strengths:

      The main strengths of the study are:

      (1) The identification of reduced cardiolipin levels in the liver of humans with MASLD and in a variety of mouse models of MASLD.

      (2) The finding that loss of cardiolipin synthesis recapitulates steatosis and inflammation in MASH.

      (3) The finding that loss of cardiolipin increases mitochondrial respiration, ROS production, and fat oxidation (in a separate hepatocyte cell line), again recapitulates several previous studies in obese humans with MASLD.

      (4) Evidence, though less definitive, that cardiolipin deficiency promotes electron leak by disrupting respiratory supercomplexes and preventing CoQ reduction.

      Weaknesses:

      (1) Figure 3A-D tries to make the point that liver CLS KO causes defects in substrate handling in vivo, based on glucose and pyruvate tolerance tests. The KO mice have a blunted response to a glucose tolerance test, but the pyruvate tolerance test showed very little (almost no) effect on glucose levels in either WT or LKO mice. The small blunting of the response in the LKO is impossible to interpret (if it's real), since the ability to clear glucose is also increased, and no tracers were used. It might be useful to monitor pyruvate and lactate levels during the experiment. However, this reviewer doesn't think the data is essential to prove the authors' main points.

      Thank you for pointing this out. We have now revised our manuscript to correctly reflect our findings on GTT and PTT. In our initial submission, we failed to clearly articulate that CLS deletion appeared to increase systemic glucose handling, which is the opposite of what one might expect in liver with steatosis. We agree that additional experiments would be helpful to better understand the systemic substrate handling in the CLS-LKO mice. As the reviewer indicates, we decided to focus this particular manuscript on intracellular and mitochondrial metabolism because of cardiolipin’s known localization to mitochondria, and the central role that this organelle plays in the pathogenesis of MASLD.

      (2) After presenting convincing evidence that respiration is elevated in isolated mitochondria from CLS KO liver, the authors follow up the findings by investigating whether 13C-palmitate and 13C-glucose oxidation are altered by CLS knockdown in murine Hepa1-6 cells (Figure 4).

      A few comments are worth mentioning about Figure 4:

      (a) It is not clear why the authors chose to use a hepatoma cell line rather than primary hepatocytes from LKO mice. The latter would be more convincing, since there could be important differences in metabolism between hepatoma cells and hepatocytes (e.g., preference for fatty acids vs glucose). Nevertheless, I think the approach is sufficient to test the general effect of loss of CLS on substrate metabolism.

      We appreciate the sentiment and agree that primary hepatocytes would have been a better model. We simply have not had prior expertise to culture primary hepatocytes and do not have the system working. We completely agree that it’s important to discuss the limitation of hepa1-6 cells as a hepatoma cells and now discuss this in our manuscript.

      (b) The authors use the M+2 enrichments of TCA cycle intermediates to infer rates of oxidation of [U-13C] palmitate or [U-13C] glucose. It is important to note that this kind of data reports fractional carbon sources (i.e., substrate preference) rather than rates of oxidation. For example, data from the 13C-palmitate experiment indicates that the CLS KD cells increase the fractional contribution from 13C palmitate (compared to glucose, for example) to the TCA cycle, but the actual rate of palmitate oxidation is not implicit in the data. However, it is reasonable to suggest that, in combination with the increased rates of O2 consumption observed in isolated mitochondria, this data supports increased fat oxidation.

      We agree with the reviewer that the nuances are important: that M+2 enrichments from [U-13C] palmitate or [U-13C] glucose reflects the fractional contributions of labeled substrates to the TCA cycle rather than oxidation. We have now revised the text to clarify that the data represent carbon incorporation patterns.

      (c) I have some concern that the [U-13C] glucose experiment is more complicated to interpret than the description implies. I'm not sure what happens in this cell line, but in the liver, most labeling from pyruvate (i.e., originating from glucose in this case) enters the TCA cycle via pyruvate carboxylase, with smaller amounts entering via PDH (depending on the nutritional state). Since one could expect pyruvate carboxylase to contribute M+3 labeled TCA cycle intermediates initially, and M+2 on the first turn of the cycle, it's hard to conclude what the data indicates about glucose oxidation. The authors could generalize the conclusion by framing the TCA cycle enrichment data as the contribution of glucose carbons and noting in Figure 4A that pyruvate carbons can enter the TCA cycle via PDH or pyruvate carboxylase, without attempting to assign their relative contributions. There are better ways to do it, but it's a small nuance here since the authors aren't making a critical point about the pathways.

      This expert comment is much appreciated. We have revised the text to more broadly describe glucose carbon entry into TCA cycle through PDH and PC. We also revised the schematic to reflect this notion.

      Reviewer #2 (Public review):

      In this study, the authors show that alterations in the lipid composition of the inner mitochondrial membrane, particularly changes in cardiolipin (CL) content, lead to defects in electron transport, supercomplex formation, and oxidative stress. Using liver-specific CLS knockout mice, which are characterized by dysfunctional capacity for cardiolipin synthesis, the authors highlight an underappreciated role for CL in MASH pathology. Overall, this is an interesting study highlighting the importance of functional/physiological electron transport (and in this context, electron leakage) in MASH pathophysiology. Despite that, this manuscript has several weaknesses that require attention.

      We thank the reviewer 2 for the constructive criticisms and identifying areas of weakness were additional data or explanations can improve the manuscript.

      (1) For all LKO studies, it is stated that the decrease in hepatic CL is causal for the observed phenotype. However, it is evident that many other lipids are impacted by CLS KO, including a marked increase in hepatic PG. In this respect, the authors show no evidence that the observed metabolic phenotype is indeed due to the reduction in CL and not to other accompanying changes.

      Thanks for this comment. We agree that because deletion of CLS promotes changes in mitochondrial lipids other than CL, we cannot conclusively attribute phenotypes we observed to CL and not to other lipids such as PG. In our experience, rescuing mitochondrial phospholipids by exogenous supplementation is problematic as they most certainly are not exclusively destined to the tissue of interest, nor to the organelle of interest, and often metabolized to produce other lipids, etc, making it difficult to interpret the data. We now have mice that conditionally overexpress CLS, which could be used to address this question, but the study is in its early phase and are outside the scope of the current study.

      The one experiment we performed is the ex vivo CL supplementation by SUV fusion to mitochondria, which has an ability to rescue electron leak. While they do not demonstrate the role of CL in all phenotypes found in the CLS-LKO mice, we think that bioenergetic phenotype associated with CLS deletion is therefore likely due to the reduction in CL. We now provide these additional discussions in lines.

      (2) In the results, the authors highlight that 'MASLD has been shown to alter the total cellular lipidome in liver.' Given that this study focused on CL, it would be useful to include specific studies that pointed to changes in hepatic CL content in MASLD/MASH/fibrosis.

      We now provide citations for these studies (PMID: 30042157, PMID: 34257827).

      (3) The initial human mitochondrial lipidomics studies show a reduction in mitochondrial CL and PG content. What was the content/expression of CL synthase and PGP synthase in these samples? If this cannot be assessed, is there any association of CLS or PGPS expression and MASLD/fibrosis (etc) in publicly available databases (e.g, GEP liver) that may explain the reduction in mitochondrial PG and CL content?

      Thanks for this suggestion. Quantification of mitochondrial lipidome require a good amount of tissue, and we do not have sufficient biomaterials left to quantify gene expression. Upon our survey of publicly available database (including GepLiver), we did not find that human MASLD was associated with an increase in CLS or other enzymes of CL biosynthesis compared to healthy controls.

      (4) The validation of MASH in patients (Figure 1B) is not convincing (ie., no quantification/scoring provided). NAS /fibrosis scoring (according to Kleiner) would help to define if all patients have indeed MASH, and what subset has fibrosis. Could the reduction in CL/PG content be (also) associated with fibrosis? In addition, Masson's Trichrome should be added to Figure 1B.

      The diagnosis was based on obvious bridging fibrosis and/or regenerative nodules on H&E staining (see additional zoomed-out images in Figure 1 – figure supplement 1). Due to the severity of these cases, formal NAS scoring was not applied. We do not have the Trichrome staining available but all MASH samples had fibrosis. Thus, it is possible that reduced CL/PG is related to fibrosis. We now added more descriptions on this point.

      (5) In human lipidomics, the authors suggest that reductions are observed in tetralinoleoyl CL (Figure 1C). However, Figure 1C only shows the combined FA acyl chain length + unsaturation, therefore not allowing for FA-specific ID (unless such data are available from the LC/MS analysis).

      Thanks for pointing this out. Per lipidomic nomenclature guideline we assign combined FA acyl chain length + unsaturation when MS2 is not performed. We have validated that our 72:8 peak corresponds to TLCL, but we do not perform MS2 on every lipid species for every sample. We now clarify this point in our manuscripts.

      (6) Figures 1 J/K/I. It is obvious that the background in all murine immunoblotting analysis has been altered. The authors should provide unaltered images for these immunoblots.

      We apologizes with the confusion. In Figure 1J/K/L/M, each panel actually represents two western blots (not one, similar to Figure 3H). The above represents a western blot with OXPHOS antibody cocktail (CV, CIII, CIV, CII, and CI), while the bottom represents the second western blot with citrate synthase (CS). Thus, we had not manipulated parts of the western blot to look different. To clarify, we now place an outline in each of the western blot to clearly demarcate individual blots to avoid confusion (new Figure 1J-M).

      (7) For Figure 1, it is unclear what is meant by 'we performed all mitochondrial lipidomic analyses by quantifying lipids per mg of mitochondrial proteins'. Was the murine lipidomics carried out on fractionated mitochondria or whole liver? If whole liver, then how were the data corrected, particularly given that PG is not a mitochondria-specific lipid?

      The data are all from lipidomic analyses performed in isolated mitochondria.

      (8) While total CL content seems indeed decreased across the different mouse models, this is mostly due to 1-2 CL species showing a pronounced reduction, with the remainder being unaltered. This should at least be acknowledged in the results. This is similarly the case in the LKO livers.

      Thanks for pointing this out. We now provide additional clarification in the text.

      (9) Figure 2. A secondary biochemical analysis of changes in lipid content should be provided, e.g., total triglyceride content, particularly given that the histology analysis does not show any major changes in hepatic lipid droplets/steatosis. In addition, the Masson's Trichrome staining shows almost no collagen deposition.

      We now provide a quantification of triglycerides in Figure 2J.

      (10) Figure 3. 'CLS deletion modestly reduced glucose handling' should be reworded. The LKO mice show improved glucose tolerance (despite the MASH phenotype), which is not evident from the above wording.

      We modified our text accordingly.

      (11) Looking at the mechanism behind the increase in hepatic steatosis, the authors state that lipid accumulation can occur due to increased lipogenesis, or dysfunctional VLDL secretion or beta oxidation, and subsequently assessed the relevant proteins/pathways. What about fatty acid uptake, which is also one of the four major pathways impacted in MASLD? This should be included in this assessment in Figure 3.

      Thank you for this comment. We now provide data for genes involved in fatty acid uptake, which was not reduced with CLS deletion (Figure 3E).

      (12) For Figure 5A, it is simply stated 'CLS deletion promotes liver fibrosis in standard chow-fed condition', and it is unclear what is highlighted within the selected EM images and what the arrows refer to. The authors should clarify this within the text.

      We have modified the text accordingly.

      Reviewer #3 (Public review):

      Summary:

      Mitochondrial oxphos causes lipid accumulation, leading to MASH, although the mechanism has been poorly understood. In this study, Funai and colleagues identify that reductions in cardiolipin in the mitochondria cause disruptions in the electron transport chain. Knockout of cardiolipin synthase was sufficient to drive MASH phenotypes, increase respiratory capacity, and cause electron leak at complexes II and III. It is well established that loss of cardiolipin increases ROS. Studies to date have been performed on whole tissue lysates, but to rule out which changes in mitochondrial lipids are driven by changes in mitochondrial number versus lipid synthesis/turnover, the authors uniquely purified mitochondria from human and mouse livers in MASH and NASH models for this study. This study provides critical information to the field that will inevitably help us better understand the mechanisms underlying MASH and NASH onset. The evidence provided is both convincing and compelling. With further suggested revision experiments, this study has the potential to change our understanding of MASH and NASH pathogenesis.

      We would like to thank the reviewer 3 for the highly-encouraging feedback.

      Strengths:

      The authors use a unique approach of lipidomics on purified mitochondria. They also analyze many distinct MASH models and provide a unique resource for the field of comprehensive lipidomics analysis of the different ways in which MASH can be induced. The use of human tissue elevates the impact/significance of the findings.

      Weaknesses:

      The data on the super complexes was the least compelling, and frankly, I do not think the authors needed those data to make a compelling argument! The authors should shift their focus more to the compelling electron leak data they have collected. If possible, it would also strengthen the work to include cardiolipin rescues on more of the experiments. Finally, expanding their explanations of the model systems would be very helpful for the readership.

      Thank you for this comment. We have now revised our argument to highlight the electron leak data and less emphasis on the supercomplexes.

      Reviewer #4 (Public review):

      Summary:

      Here, the authors wish to shed light on factors that contribute to the development of liver disease in what used to be called 'the metabolic syndrome'. This is a human-health problem of considerable significance, and the insights they provide, namely the implication of a defect in mitochondrial cardiolipin (CL) content to the progression from metabolic dysfunction associated steatotic liver disease to steatohepatitis, are plausible.

      We would like to thank the reviewer 4 in an encouraging feedback.

      Strengths:

      The experimental evidence proffered is derived from the observation of lower levels of (CL) in mitochondria from the liver of patients undergoing liver transplant or resection due to endstage steatohepatitis compared with mitochondria derived from livers of patients with other conditions. This correlation is buttressed by observations made in mice with liver-selective compromise in CL synthesis and which suggest a pathological environment associated with mitochondrial dysfunction and enhanced oxidative stress, features deemed to play a role in the progression from steatotic liver disease to steatohepatitis.

      The paper is well written, and the findings are well explained and superficially convincing.

      Weaknesses:

      It is unclear how much can be learned from compromising a key enzyme that produces a key mitochondrial lipid in a busy metabolic organ like the liver - isn't the discovery of a mitochondrial defect in such a context rather trivial? And how reliably can these findings be related to the human observations? Most importantly, the chain of causality implied by the title is unproven: the key question of whether or not (somehow) preventing the drop in cardiolipin content affects the course of steatohepatitis remains unanswered.

      We agree with the reviewer that the current manuscript does not directly provide evidence that reduction in CL causes MASLD in humans, which as the reviewer describes, must be tested by rescuing CL content in the context of MASLD. We have now obtained mice with conditional overexpressor and have begun the experiments, but findings from these mice are beyond the scope of the current study. We have modified our title to “Cardiolipin deficiency disrupts electron transport chain AND drives steatohepatitis” to reduce the implication for causality.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The manuscript states that loss of mitochondrial respiration is expected in MASLD. Forexample, line 187 "MASLD is known to be associated with reduced mitochondrial oxidative capacity". A more accurate statement is that "MASH" is known to be associated with reduced mitochondrial oxidative capacity and increased ROS production in humans. As you correctly cite later for an ex vivo human mitochondrial respiration study, early MALSD, especially with obesity, is associated with elevated mitochondrial respiration (40). Since those measurements are maximal respiration rates, which might not reflect actual in vivo flux, you might also make readers aware that your data is consistent with in vivo human studies that found increased hepatic oxidative flux (TCA cycle flux) in obese subjects with moderate steatosis (PMID: 22152305), which appears to wane with severe steatosis and/or inflammation (PMID: 31012869, PMID: 40272888).

      Thank you for these suggestions. We have made the suggested changes to the text.

      Reviewer #3 (Recommendations for the authors):

      (1) Throughout the manuscript, the authors refer to the inner mitochondrial membrane, although they never perform assays to distinguish the inner vs outer mitochondrial membrane. It would be better to just refer to the cardiolipin being measured as "mitochondrial."

      Thank you. We made these changes.

      (2) In figures showing changes in cardiolipins, not all of them change; only a handful of them are reduced in NASH. Could the authors add commentary in the manuscript about what is known about these different cardiolipin species, and speculate as to why certain CLs are changing while others are not?

      Thank you. Reviewer #2 had similar comments and we provided additional discussions.

      (3) In the human tissues, what do the other mitochondrial inner membrane lipids (PC, PE, PI, PS, LPC, LPE) look like in the healthy vs NASH patients (Figure 1A-D)?

      Thank you for this request. We did not include these data in the manuscript as we have a separate ongoing study (the second author is the lead author on this paper) where we are following up on hepatic mitochondrial PS and PE, which we found to be decreased in human MASH samples compared to healthy livers. This turned out to be a convoluted story so we decided not to include it in the paper.

      (4) The descriptions of the different MASLD/MASH models are a little sparse. Especially needing more detail is the model for carbon tetrachloride injection, causing NASH. The authors should explain how each of these models typically induces MASLD/MASH.

      We now provide these details.

      (5) In figures 2E and F, total body mass is unchanged in CLS-LKO mice, but liver mass is decreased; yet on the chow diet, there appears to be lipid accumulation in the liver as well; I am wondering what the authors' reasoning is for this decreased liver mass.

      It is difficult to say conclusively, but we suspect it is due to cell death evidenced by fibrosis. It’s important to note that while there is lipid accumulation in the liver, steatosis is relatively mild and the increase in liver triglyceride is quite marginal (Figure 2J).

      (6) The lipidomics analysis and comparison of livers in these different models is a wonderful dataset that needs far more depth in terms of unpacking and describing the findings. For example, all the models of MASH show similar changes in most of the lipid species analyzed. NASH appears to be quite different than MASH. This, among other trends, is certainly worth highlighting as it will be of interest to the field.

      Thanks for this comment. We agree that while CL phenotype were common to mouse and human MASH samples, there were other changes that we observed in other lipids that may be biologically significant. As described above, we have an ongoing study pursuing mitochondrial PS in the liver.

      (7) Figure 2B - It is interesting that the CLS KO only impacts certain CLs. The 72:8 CL, which is regulated by CLS, is also a CL that appears to change in the human patient samples. The information on the specific CL that is changing seems critical to the mechanism of the role of the CL in the disease. Throughout the manuscript, it is important to specify which specific CL is being referred to, instead of broadly characterizing the changes to cardiolipins, especially since most of the cardiolipins shown do not change; only a handful of them do.

      Thank you for this suggestion. We have included additional discussions on 72:8 CL in the manuscript.

      (8) One potential non-specific mechanism whereby CLS knockout can cause MASH would be if the mice change their overall food consumption. It is an important control to test if the total food intake is different in WT vs KO mice to formally rule out this possibility.

      The food intake was not different between the group (Figure 2E).

      (9) To determine the extent to which de novo cardiolipin synthesis underlies the change in MASH/fatty liver observed in the HFD, GAN, and CCl4 models in Figure 1, the authors should also put the CLS KO mice on these diets and perform liver histology, analysis of inflammation markers, and analyze immune cell infiltration. Alternatively, the authors could try to rescue the CLS KO model by supplementing cardiolipin in the diet or by injection.

      Thank you. We have an ongoing experiment to examine the effect of hepatocyte-specific CLS overexpression on protection from GAN-induced MASLD.

      (10) Figure 3F shows a decrease in UQCRC2 by RNA but no change at the protein level in Figure 3H. The authors should comment a bit more on this disparity, and the data in Figure 3F don't mean much for the main point of the study if the levels of the proteins are unchanged.

      The reviewer is correct. We initially performed RNAseq in trying to broadly capture how CLS knockout influences liver health, which implicated that transcriptional program for mitochondrial proteins were downregulated. Nevertheless, gold standard measurements of mitochondrial content (mitochondrial protein or mtDNA) did not show change in the abundance with CLS deletion.

      (11) The increase in respiration and spare respiratory capacity upon CLS KO shown in Figure 3J is extremely interesting! The explanation of the experiment and its meaning should be significantly expanded upon.

      Thank you. We included additional discussion on this point.

      (12) Figure 4 - It is interesting that the fraction of the TCA cycle metabolites labeled is increasing with the palmitate tracer and decreasing with the glucose tracer. This implies a "fuel switch," such that more of the TCA cycle carbons originate from fatty acids than glucose upon loss of CLS. The authors should make note of this point. Also, to understand if the total molar quantity of labeling in the TCA cycle from palmitate and glucose is changing, the authors should also report the relative abundance (instead of just the fraction labeled) of the labeled metabolites and unlabeled metabolites.

      Thanks for this suggestion, we have now added this discussion.

      (13) In Figure 5C-F, the authors show that CLS deletion can activate the caspase pathway, but do not see any change in cytochrome c localization. Can the authors clarify if CLS deletion is sufficient to induce apoptosis?

      CLS deletion certainly causes cell death that induces tissue fibrosis. Activation of the caspase pathway suggests that the cell death may be due to apoptosis but we did not see changes in cytochrome c localization. Our lab is currently performing additional experience to test the possibility that CLS deletion may induce ferroptosis.

      (14) Figure 6A-C- The authors discuss the I + III2 + IV supercomplex substantially and consistently decreasing in the CLS-KO mice, however, the quantifications do not look statistically significant. Can the authors confirm if these changes are or are not significant and adjust the text accordingly?

      The reviewer is correct. Abundances of I+III2+IV supercomplexes are decreased in CLS-LKO mice compared to control mice when quantifying with supercomplex antibody cocktail or with UQCRSF1 (complex III subunit) antibody, but not with complex I antibodies. The discrepancy for these results are not entirely clear but it’s likely a combination of antibody sensitivity and a tricky nature to dissolve high molecular weight protein complexes.

      (15) The most compelling data to indicate electron leakage increasing upon CLS knockout is in Figures 7A-E. I would suggest the authors decrease their emphasis on the rearrangement of the supercomplexes and focus their discussion on the very compelling results of Figure 7.

      Thanks for this suggestion. We have modified our text.

      (16) Figure 7D shows that a major site of electron leak is from site II, and these results also fit with the profound succinate-induced respiration observed in earlier experiments. It would be nice if the authors could test the ability of cardiolipin to rescue these phenotypes, similar to the assay in Figure 5I. Assessing this rescue on the CoQ redox state would also strengthen the claims.

      Thank you for this comment. We are encouraged with your suggestions. We have thought about this quite extensively during the preparation of the manuscript but we refrained from making conclusive statements regarding complex II because the magnitude of the increase in electron leak is equally elevated at complex II and III. It’s true that CLS deletion increases succinate-induced respiration, but this might also be because succinate elicits the highest increase in respiration even in wildtype mice (see values in Figure 3K and L compared to other substrates). It would be intriguing to examine the influence of CLS deletion on complex II/III electron leak as well as succinate-induced respiration in tissues where succinate is not a preferred substrate. We have attempted cardiolipin rescue in SUV but unfortunately, we could not get this assay to work for site-specific electron leak measurements.

      (17) In Figure 7G-H, it would be nice to see a ratio of oxidized to reduced CoQ, in the CLS deletion mice and in human NASH livers, if samples are available.

      Thanks for this suggestion. Data shown (Figure 7- figure supplement 1P-S).

      (18) CoQH2 can also deliver electrons to complex II (via its reversal). Complex II shows a remarkable contribution to the electron leak phenotype (Figure 7D). Also, as the complex II monomer showed much larger changes in the native gels of Figure 6 than the complexes involving complex III. A more likely model is that oxidized CoQ accumulates in the CLS knockout model because of increased CoQH2 leak via complex II.

      Perhaps. We also thought about this but we are not sure if this fits with the observation that CLS deletion increases succinate-induced respiration, which suggests increased succinate to fumarate conversion, a notion that I am not sure can be congruent with increase CoQH2 reversal to complex II. Overall, I think we lack the tools or evidence to conclusively implicate whether CLS deletion primarily acts on complex II or III. Nevertheless, we appreciate the reviewer’s enthusiasm on these topics as we perform additional experiments on the mechanism of interactions between CL and the ETC.

    1. eLife Assessment

      This important study assesses the portability of epigenetic clocks across ancestries, including in the context of accelerated aging in Alzheimer's Disease patients. It provides convincing evidence for population differences in age estimation accuracy across a variety of epigenetic clocks, driven in large part by continuous variation in ancestry. Given the accelerating use of epigenetic clocks across fields, this study is likely to be of interest to researchers working on human genetic and epigenetic variation or who apply epigenetic clocks to diverse human populations.

    2. Reviewer #3 (Public review):

      The authors find that DNA methylation-based clocks are generally less accurate at predicting age in cohorts with large proportions of non-European (especially African) ancestry, compared to cohorts with high European ancestry proportions (which more closely reflects the genetic composition of individuals included in training sets). They provide evidence for this ancestry bias via ancestry-stratified analyses, and in analyses of continuous ancestry proportion effects on clock error. They then test two hypothesized underlying causes of ancestry bias: that ancestry-differentiated SNPs disrupt CpG sites preventing methylation, and that ancestry-differentiated SNPs influence DNA methylation levels. They find clear evidence especially for the second cause, in the form of meQTL that influence clock CpG sites and vary in frequency across ancestry groups. Finally, the authors provide key discussions of potential paths forward to alleviate bias and improve portability for future clock algorithms.

      The topic is timely due to the increasing popularity of DNA methylation-based clocks and the acknowledgment that many algorithms (e.g., polygenic risk scores) lack portability when applied to cohorts that substantially differ in ancestry or other characteristics from the training set. This has been discussed to some degree for DNA methylation-based clocks, but could of course use more discussion and empirical attention, which the authors nicely provide using an impressive and diverse collection of data. The inclusion of data from multiple cohorts, the analysis of ancestry as a continuous variable, and the attempts to address the underlying causes of ancestry-based differences in accuracy provide comprehensive evidence that genetic background influences clock portability.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      Cruz-Gonz´alez and colleagues draw on DNA methylation and paired genetic data from 621 participants (n=308 controls; n=313 participants with Alzheimer’s Disease). The authors generate a panel of epigenetic biomarkers of aging with a primary focus on the Horvath multi-tissue clock. The authors find weaker correlations between predicted epigenetic age and chronological age in subgroups with higher African ancestry than within a subgroup identified as White. The authors then examine genetic variation as a potential source for between-group differences in epigenetic clock performance. The authors draw on a large collection of publicly available methylation quantitative trait loci datasets and find evidence for substantial overlap between clock CpGs located within the Horvath clock and methQTLs. Going further, the authors show that methQTLs that overlap with Horvath clock CpGs show greater allelic variation in African ancestral groups pointing to a potential explanation for poorer clock performance within this group.

      Thank you for this summary.

      Strengths:

      This is an interesting dataset and an important research question. The authors cite issues of portability regarding polygenic risk scores as a motivation to examine between-group differences in the performance of a panel of epigenetic clocks. The authors benefit from a diverse cohort of individuals with paired genetic data and focus on a clinical phenotype, Alzheimer’s disease, of clear relevance for studies evaluating age-related biomarkers.

      Thank you.

      Weaknesses:

      While the authors tackle an important question using a diverse cohort the current manuscript is lacking some detail that may diminish the potential impact of this paper. For example:

      (1) Information on chronological ages across groups should be reported to ensure there are no systematic differences in ages or age ranges between groups (see point below).

      Thank you for pointing out this omission. The distributions are now presented in Supplementary Figure 1. While there is some variation in median age, the age ranges are similar across cohorts (median 73.1 to 79.3). The small differences do not explain the differences in accuracy between the cohorts, e.g., the median age of the African Americans (76.4) is lower than the median age for the White cohort (77.7).

      (2) The authors compare correlations between chronological age and epigenetic age in sub-groups within to correlations reported by Horvath (2013). Attempting to draw comparisons between these two datasets is problematic. The current study has a much smaller N (particularly for sub-group analyses) and has a more restricted age range (60-90yrs versus 0-100 yrs). Thus, is an alternative explanation simply that any weaker correlations observed in this study are driven by sample size and a restricted age range? Reporting the chronological ages (and ranges) across subgroups in the current study would help in this regard. Similarly, given the lack of association between AD status and epigenetic age (and very small effect in the white group), it may be of interest to examine the correlation between chronological age and epigenetic age in each group including the AD participants: would the between-group differences in correlations between chronological age and epigenetic be altered by increasing the sample size?

      Our conclusions about the reduced accuracy of the clock in admixed individuals are based on the comparison within the MAGENTA cohorts, not a comparison of MAGENTA to previously published studies. We find significantly reduced accuracy in the admixed cohorts compared to the White MAGENTA cohort. Further supporting this conclusion beyond he MAGENTA cohort, we analyzed three independent whole blood methylation datasets. Two focused on African American individuals—the Grady Trauma Project (n = 422) and the GENOA study (n = 1,394)—and one focused on White Swedish individuals (n = 729). As observed in MAGENTA, the Horvath clock had significantly lower accuracy for the African American cohorts (Figure 3 than for the White Swedish cohort.

      When comparing results across studies, the reviewer is correct that lower correlations are generally seen for older cohorts. Indeed, other studies applying the Horvath clock have seen similar correlations in older cohorts to those observed in MAGENTA (Marioni et al., 2015, Horvath 2013, and Shireby et al., 2020). We now also include the chronological age distributions of the cohorts in this study, along with their mean and standard deviations (Supplementary Figure 1). This shows that the distribution of chronological ages for White individuals is similar to the cohorts where the clocks did not perform as well. Finally, as suggested, we correlated chronological and epigenetic age with the inclusion of AD cases in each cohort for the Horvath clock. The significantly lower performance of the clock on Puerto Ricans and African Americans, relative to White individuals, remains even after including all individuals in each cohort. Thus, combining cases and controls did not qualitatively change the performance relationships for the African Americans and Puerto Ricans relative to the Whites (Supplementary Figure 3).

      (3) The correlation between chronological age and epigenetic age, while helpful is not the most informative estimate of accuracy. Median absolute error (and an analysis of MAE across subgroups) would be a helpful addition.

      We used correlation because it is commonly used to evaluate the performance of epigenetic age clocks, but we agree that other error quantification metrics provide a complementary perspective. We now include MAE and MSE comparisons across sub-groups in the revision (Supplementary Table 1). We find that across all accuracy metrics, the African American and Puerto Rican cohorts perform worse than the White and Peruvian cohorts. Interestingly, the Cubans show relatively high error despite a high correlation between predicted and chronological age. However, there are only 21 non-demented Cuban controls. In addition, we evaluated the same metrics in three replicate datasets (two African American cohorts and one for White Swedish individuals) and found the same patterns of lower accuracy across metrics in African ancestry individuals, albeit with some variation in accuracy between cohorts (Supplementary Table 2). Notably, as discussed above, this is not driven by differences in chronological age distributions: when we subset to older individuals (≥ 55 years old) in order to facilitate comparisons to MAGENTA study individuals, the median age for the White Swedish individuals (70 years old) is higher than that of the GENOA (62.7 years old) and Grady (58 years old) individuals. Despite the difference in median ages, the clock performs better on White Swedish individuals across all accuracy metrics than the African ancestry cohorts with younger individuals.

      (4) More information should be provided about how DNAm data were generated. Were samples from each ancestral group randomized across plates/slides to ensure ancestry and batch are not associated? How were batch effects considered? Given the relatively small sample sizes, it would be important to consider the impact of technical variation on measures of epigenetic age used in the current study. The use of principal Component-based versions of these clocks (Higgins Chen et al., 2023; Nature Aging https://doi.org/10.1038/s43587-022-00248-2) may help address concerns such concerns.

      Thank you for pointing out the need for additional context on data generation. We have added details to the Methods. All omics data from the MAGENTA study were generated using standard protocols that ensure minimal technical artifacts and batch effects. Samples were randomized across plates and chips to ensure that ancestry, age, and sex were not confounded with each batch. We also performed a principal components analysis of the normalized methylation data used as inputs for all MAGENTA analyses. We found that the samples did not stratify by sample plate, cohort, ethnicity, or ascertainment center along the principal components (Supplementary Figure 2).

      We also thank the reviewer for their suggestion to apply the principal component clock to account for potential technical variation. As outlined in the new section “Principal component versions of the methylation clocks also have lower age prediction accuracy for genetically admixed individuals,” using the principal component version of the Horvath clock did not result in consistent improvement in age prediction accuracy or generalization across MAGENTA cohorts (Supplementary Figures 4 and 5). The lower accuracy for age prediction in individuals with substantial African ancestry was present for the PC clock in the replication cohorts, just as in the MAGENTA cohorts (Supplementary Figure 6).

      (5) Marioni et al., (2015) found a very weak cross-sectional association between DNAm Age and cognitive function (r∼0.07) in a cohort of >900 participants. Given these effect sizes, I would not interpret the absence of an effect in the current study to reflect issues of portability of epigenetic biomarkers.

      We agree that previous links between DNAm Age and AD or cognitive function have been relatively small in magnitude. For example, the PhenoAge paper (Levine et al., 2018) and a study using the Horvath clock (Levine et al., 2015) found age acceleration of less than a year in AD patients relative to non-demented individuals. Similar results have also been observed in studies with smaller sample sizes (e.g., 700 for Levine et al. 2015 and 604 for Levine et al. 2018). Given these small effect sizes, we agree that accounting for statistical power is essential for interpretation of our results. We performed power calculations based on an effect of the size observed in previous studies (0.5 year acceleration). We have 86% power in the full MAGENTA data set to detect an effect of this size. Stratifying by cohorts, we have 75% power for the African Americans, 72% for the Puerto Ricans, 72% for the Whites, 65% for the Peruvians, and 47% for the Cubans. Thus, we believe we have high enough power that the consistent lack of association outside of the White cohort in MAGENTA is likely meaningful. Based on these calculations, there is only a 1% chance that we would not observe an effect in any of the other cohorts if the effect was present across cohorts. Nonetheless, we have added caveats about power and the small sample size to our suggestion that the reduced accuracy of the clocks contributes to the lack of AD association outside of Whites.

      (6) The methQTL analyses presented are suggestive of potential genetic influence on DNAm at some Horvath CpGs. Do authors see differences in DNAm across ancestral groups at these potentially affected CpGs? This seems to be a missing piece together (e.g., estimating the likely impact of methQTL on clock CpG DNAm).

      We agree. Thank you for this suggestion. We have added Figure 6 in the main text to address this gap. In short, we analyzed additional whole blood methylation data from inidividuals with African ancestry and found that a substantial proportion of the CpGs in methylation clocks are differentially methylated in African ancestry individuals relative to European ancestry individuals. In the case of the Horvath clock, we find that 84/353 (23.8%) of the clock CpGs are differentially methylated between ancestries. In parallel, we found that 56 of these differentially methylated clock CpGs are also affected by meQTL, many of which are at different frequencies between populations. We also investigated whether the meQTL-affected clock CpGs are associated with increased clock error in the MAGENTA individuals. We found 56 clock CpGs whose methylation levels associated with increased clock error, and 42 of these have at least one meQTL. Thus, while meQTL are not the only factor to affect the portability of methylation clocks across global populations, we suggest that they are a significant contributor, especially in the case of the Horvath clock.

      Reviewer #2 (Public review):

      Summary:

      This paper seeks to characterize the portability of methylation clocks across groups. Methylation clocks are trained to predict biological aging from DNA methylation but have largely been developed in datasets of individuals with primarily European ancestries. Given that genetic variation can influence DNA methylation, the authors hypothesize that methylation clocks might have reduced accuracy in non-European ancestries.

      Strengths:

      The authors evaluate five methylation clocks in 621 individuals from the MAGENTA study. This includes approximately 280 individuals sampled in Puerto Rico, Cuba, and Peru, as well as approximately 200 self-identified African American individuals sampled in the US. To understand how methylation clock accuracy varies with proportion of non-European ancestry, the authors inferred local ancestry for the Puerto Rican, Cuban, Peruvian, and African American cohorts. Overall, this paper presents solid evidence that methylation clocks have reduced accuracy in individuals with non-European ancestries, relative to individuals with primarily European ancestries. This should be of great interest to those researchers who seek to use methylation clocks as predictors of age-related, late-onset diseases and other health outcomes.

      Thank you for this summary.

      Weaknesses:

      One clear strength of this paper is the ability to do more sophisticated analyses using the local ancestry calls for the MAGENTA study. It would be valuable to capitalize on this strength and assess portability across the genetic ancestry spectrum, as was recently advocated by Ding et al. in Nature (2023). For example, the authors could regress non-European local ancestry fraction on measures of prediction accuracy. This could paint a clearer picture of the relationship between genetic ancestry and clock accuracy, compared to looking at overall correlations within each cohort.

      Thank you for this suggestion. To model portability across genetic ancestry as a spectrum, we regressed the Horvath clock error on the proportions of African ancestry in the genomes of the MAGENTA individuals, adjusting for chronological age. The proportion of African ancestry is significantly associated with increased Horvath clock error (p = 0.039), with the clock making less accurate age predictions by 1.46 years for individuals with full African ancestry compared to no African ancestry. We have added this new analysis to the Results.

      The authors present two possible reasons that methylation clocks might have reduced accuracy in individuals with non-European ancestries: genetic variants disrupting methylation sites (i.e., ”disruptive variants”) and genetic variants influencing methylation sites (i.e., meQTLs). The authors conclude disruptive variants do not contribute to poor methylation clock portability, but the evidence in support of this conclusion is incomplete. The site frequency spectrum of disruptive variants in Figure 4 is estimated from all gnomAD individuals, and gnomAD is comprised of primarily European individuals. Thus, the observation that disruptive variants are generally rare in gnomAD does not rule them out as a source of poor clock portability in admixed individuals with non-European ancestries.

      In the revision, we now additionally report ancestry-specific allele frequencies to demonstrate the rarity of CpGclock disrupting variants (Supplementary Figure 9). The global allele frequencies were so low that even if they all occurred in individuals of non-European ancestries, they would still be extremely rare.

      It is also unclear to what extent meQTLs impact methylation clock portability. The authors find that the frequency of meQTLs is higher in African ancestry populations, but this could reflect the fact that some of the analyzed meQTLs were ascertained in African Americans. The number of meQTL-affected methylation sites also varies widely between clocks, ranging from 6 to 271; thus, meQTLs likely impact the portability of different clocks in different ways. Overall, the paper would benefit from a more quantitative assessment of the extent to which meQTLs influence clock portability.

      We agree that the meQTL likely influence the clocks in different ways and that the ascertainment of the meQTLs in different populations makes direct comparisons challenging. To more directly link meQTL to clock performance, we identified 56 Horvath clock CpG sites whose methylation levels significantly associate with increased clock error in the MAGENTA study individuals. Of these, 42 (75%) are affected by an meQTL, including nine that are affected by an African ancestry-differentiated meQTL. As such, meQTL, and specifically meQTL that were likely not present in the training data of the Horvath clock, associated with both the methylation of CpG sites and clock error. However, as the reviewer suggests, determining causality among these factors is challenging. Given our incomplete knowledge of meQTL in different ancestries, we have added caveats to our conclusions about the effect of meQTL on clock portability.

      The paper implies that methylation clocks have an inferior ability to predict AD risk in admixed populations relative to white individuals, but the difference between white AD patients and controls is not significant when correcting for multiple testing. This nuance should be made more explicit.

      We agree that the signal is not strong in the white cohort; however, it is similar in magnitude to previous studies. As outlined in response to Reviewer 1’s Point 5, we have now added power calculations that indicate reasonable power (≥72%) to detect small effect sizes (0.5 year increase) in the white, Puerto Rican and African American cohorts. We now interpret the AD association tests in the context of these power calculations and multiple testing correction.

      Finally, this paper overlooks the possibility that environmental exposures co-vary with genetic ancestry and play a role in decreasing the accuracy of methylation clocks in genetically admixed individuals. Quantifying the impact of environmental factors is almost certainly outside of the scope of this paper. However, it is worth acknowledging the role of environmental factors to provide the field with a more comprehensive overview of factors influencing methylation clock portability. It is also essential to avoid the assumption that correlations with genetic ancestry necessarily arise from genetic causes.

      We entirely agree and have now clarified the scope of our analyses and importance of environmental factors in the revision. We intersected clock CpGs with enviromental-factor-associated CpGs from multiple epigenome-wide association studies (EWAS) and found overlaps that suggest an environemtnal contribution to differences in clock CpG methylation. However, given the lack of environmental data on the MAGENTA study individuals, as well as the lack of datasets for replication, we cannnot directly compare the environmental and genetic contributions to clock accuracy. Nevertheless, the new analyses in the revision highlight the contribution of both genetic and environmental factors to lack of portability for certain methylation clocks.

      Reviewer #2 (Recommendations for the authors):

      (1) Line 64: An association between methylation patterns and genetic ancestry does not presuppose that meQTLs vary in frequency between genetic ancestries; environmental factors could also play a role. It would be nice to comment on this further in the Introduction.

      We agree that environmental factors likely play a role in the decrease in methylation clock performance in admixed populations. We have added text highlighting this in the revised Discussion. Regarding meQTL, we agree that associations between methylation patterns and genetic ancestry do not necessarily imply that meQTL will vary in frequency between genetic ancestries. However, our new analyses in the revision find African-ancestry differentiated meQTL that associate with Horvath clock CpG methylation levels and overall clock error (Figure 6E-F and Supplementary Figure 13).

      (2) Line 116 implies Puerto Ricans have “substantial amounts of African ancestry” but the median ancestry is 15% (which is not much more than the Peruvian and Cuban cohorts).

      Thank you for pointing this out. We have clarified this statement in the text. While the median proportion of African ancestry in Puerto Ricans is 15% (vs. 6% and 2% for the Peruvian and Cuban individuals in MAGENTA), there are many individuals with substantially higher African ancestry. The upper quartile is >25% and several Puerto Ricans have >50% African ancestry.

      (3) In Figure 2B, Puerto Ricans have worse accuracy than Peruvians but a higher proportion of inferred CEU ancestry, which is interesting and defies intuition - is there any hypothesis for why this might be the case?

      In light of our new meQTL analyses, we hypothesize that the African ancestry differentiated meQTL that affect Horvath clock CpGs drive the increase in clock error for these individuals, despite having more European ancestry across their genome. Given that the Peruvians (and Cubans, for that matter) hold very little African ancestry, and also very few of the African-differentiated meQTL, this could explain some of the large difference in clock errors for the cohorts.

      (4) Figure 2C would be improved with confidence intervals.

      We thank the reviewer for this suggestion and have added confidence intervals for Figure 2C.

      (5) It’s interesting that the correlation with Cubans is positive in Figure 3B (for one clock, significantly so). Is there any rationale for this?

      We noticed this as well, but have not been able to come to a definitive conclusion. It is possible that environmental factors contribute. However, the Cuban cohort is the smallest in MAGENTA (22 cases and 21 controls) and the none of the differences are statistically significant, so more investigation in a large cohort is required.

      (6) Line 231: Which population(s) is allele frequency estimated in?

      This is the global frequency reported in gnomAD, which is calculated across all populations in gnomAD v3.0. As noted above, we now also report allele frequencies by gnomAD population (Supplementary Figure 9).

      (7) Were the meQTLs pruned? How many independent variants are there per methylation site? It would be nice to see a distribution for the sites in the Horvath clock.

      We now report the distribution of meQTL across clock CpG sites. The mean number of variants is 108; the median is 36; and the maximum is 1,699. We have now included a plot of the distribution for all 271 (out of 353) Horvath clock CpG sites (Supplementary Figure 14). We did not perform any pruning in these initial results for several reasons. First, we sought to demonstrate the great potential for meQTL to influence these CpGs and to compare the distributions of these common meQTL across populations (based on gnomAD data). Second, identifying the causal variant or variants is challenging. Given that many of these meQTLs likely reflect redundant signals, for the new analyses of African-differentiated meQTL, we restrict to a single variant per clock CpG site. We focus on the variant with the greatest absolute beta, as reported by the original meQTL study from which the variant originates.

      (8) Figure 5C might benefit from a geom density rather than overlapping bar plots; the trends are hard to see.

      We appreciate the reviwer’s suggestion and have now reworked the figure and based it on just the density curves so that readers may better appreciate the differences in allele frequencies.

      (9) Several figures would be more legible with larger font sizes.

      We appreciate this recommendations and have made the font sizes for all plots larger and more legible.

      Reviewer #3 (Public review):

      This manuscript examines the accuracy of DNA methylation-based epigenetic clocks across multiple cohorts of varying genetic ancestry. The authors find that clocks were generally less accurate at predicting age in cohorts with large proportions of non-European (especially African) ancestry, compared to cohorts with high European ancestry proportions. They suggest that some of this effect might be explained by meQTLs that occur near CpG sites included in clocks, because these variants may be at higher frequencies (or at least different frequencies) in cohorts with high proportions of non-European ancestry relative to the training set. They also provide discussions of potential paths forward to alleviate bias and improve portability for future clock algorithms.

      The topic is timely due to the increasing popularity of DNA methylation-based clocks and the acknowledgment that many algorithms (e.g., polygenic risk scores) lack portability when applied to cohorts that substantially differ in ancestry or other characteristics from the training set. This has been discussed to some degree for DNA methylationbased clocks, but could of course use more discussion and empirical attention which the authors nicely provide using an impressive and diverse collection of data.

      Thank you for this summary.

      The manuscript is clear and well-written, however, some key background was missing (e.g., what we know already about the ancestry composition of clock training sets) and most importantly several analyses would benefit from being taken one step further. For example, the main argument of the paper is that ancestry impacts clock predictions, but this is determined by subsetting the data by recruitment cohort rather than analyzing ancestry as a continuous variable. Extending some of the analyses could really help the authors nail down their hypothesized sources of lack of portability, which is critical for making recommendations to the community and understanding the best paths forward.

      Thank you for this suggestion. As noted in our response to Reviewer 2’s Point 1, we have analyzed ancestry as a continuous variable and found that the proportion of African ancestry in the genomes of the MAGENTA individuals significantly associates with increased difference in chronological and predicted age, even after controlling for chronological age (1.46 years more error for 100% vs. 0% African ancestry; p = 0.039). As outlined below, we have also added details on the training of previous clocks and the important additional previous work highlighted by the Reviewer.

      Reviewer #3 (Recommendations for the authors):

      Major comments

      There is previous literature addressing who is in the training set for methylation clocks. To my knowledge, this work has been primarily led by Nancy Krieger. It would be a valuable addition to discuss her work (and any similar work by other investigations) in the introduction. In other words, what do we currently know about the degree of bias in the training sets for methylation-based clocks? The assumption of the introduction is that the training sets are overwhelmingly European ancestry (which I assume is true) but I think some quantitative information about this would be helpful for understanding the source and magnitude of the problem.

      We thank the reviewer for bringing the work of Dr. Nancy Krieger to our attention. It directly supports the rationale for this study: the sociodemographic characteristics of the individuals used to train these clocks are poorly reported, limited to outdated population descriptors (for example, the use of “Caucasians” to describe some of the individuals used to train the Horvath and the Hannum clocks) or race and ethnicity labels. Moreover, where labels are available for training individuals, they tend to underrepresent the individuals of diverse backgrounds, as in the Horvath clock. We have incorporated Dr. Krieger’s work into the Introduction, including details of how this supports the rationale and purpose of our study.

      Related to the above comment, there has been pretty extensive previous work on the effects of race and ethnicity on epigenetic clock estimates (e.g., https://genomebiology.biomedcentral.com/articles/10.1186/s13059-016-1030-0), and that seems like it could be more explicitly weaved into the introduction and discussion.

      We thank the reviewer for highlighting this relevant article. We have added discussion of it into the Introduction. Several factors make direct comparison with our results challenging. First, the grouping of individuals based on race and ethnicity without consideration of genetic ancestry complicates comparisons. Race and ethnicity commonly do not match genetic ancestry components (see Gouveia et al., 2025 https://www.cell.com/ajhg/fulltext/S00029297(25)00173-9). Second, the study reports differences in epigenetic age accelerations (intrinsic and extrinsic) in individuals from various race and ethnic groups. It does not directly evaluate the accuracy of the epigenetic age predictions in these groups. Thus, it is challenging to interpret whether the differences in acceleration are driven by biological factors or biases in the performance of the clocks themselves.

      The main analysis that felt like it was missing was asking whether the age deviations are larger for individuals with greater proportions of African ancestry. The authors have the ability to analyze ancestry as a continuous variable, but instead performed analyses in various a priori subsets of the data; the subsets do have average differences in ancestry, but also there is heterogeneity within groups. Given that the authors calculated admixture proportions already, it seems like a missed opportunity not to use these estimates. This would also sidestep the issue of the problematic labels applied to the subsets, which mix ancestry, nationality, and race terms (note that I thought the legacy reasons why these labels are used were well-explained, but they are nevertheless problematic for biological explanations that center on ancestry/genetic information as the driver of bias).

      We appreciate the reviewer’s suggestion to investigate clock accuracy in the context of African ancestry proportions. As noted in the response to Reviewer 2’s Point 1, we modeled the clock error as a function of the fraction of African ancestry of each individual, adjusting for an individual’s chronological age. The proportion of African ancestry is significantly associated with increased Horvath clock error (p = 0.039), with the clock estimated to give less accurate age predictions by 1.46 years for individuals with 100% African ancestry compared to no African ancestry. We now report this in the Results.

      Another missed analysis opportunity occurs in lines 259-261, where the authors state “Thus, the clock with the largest decrease in performance in admixed cohorts (in terms of predicting chronological age and identifying age acceleration in AD) has the most and largest fraction of meQTLs influencing its CpGs.” This is another place where the authors make generalizations about a given cohort based on average ancestry rather than testing the claim empirically on an individual basis (e.g., by examining the number of meQTL variants a given individual is heterozygous for or has the non-European allele for).

      We thank the reviewer for this comment. This feedback motivated us to evaluate the relationship between differences in meQTL frequencies and methylation clock error. We found differences in meQTL frequency in the MAGENTA individuals, specifically many of the clock CpG affecting meQTL are most common in the African American cohort, consistent with our theory (Figure 6E,F). Nonetheless, there are 84 Horvath clock CpGs (24%) that are differentially methylated in AFR individuals, and 56 of these are affected by an meQTL, including 11 that are affected by an African ancestry-differentiated meQTL (Figure 6G). Finally, we find that 42 Horvath clock CpG sites in MAGENTA individuals with methylation levels that are significantly associated with increased clock error, and that are also affected by an meQTL (Figure 6B). However, at the individual level we do not find a clear relationship between the number of meQTL or ancestry-differentiated meQTL and methylation clock error. In light of these data, we have reframed our conclusions to state that meQTL likely contribute to clock error, while also being clear that they are not the sole cause.

      Can the authors explain or offer an investigation into why predicted age is often better in Cubans than Whites? They gave much attention to the opposite effect (of similar magnitude) in African Americans and Puerto Ricans but didn’t really discuss the surprisingly accurate prediction in Cubans.

      We did not focus on the results in the Cuban cohorts for several reasons. As discussed in response to Reviewer 2’s comment, the Cuban cohort had the smallest sample size (22 cases and 21 controls). Thus, while the correlation between methylation age and chronological age is similar to Whites, and in a few cases higher, the differences were not statistically significant. Second, looking at other error metrics, like mean absolute error, the clocks are comparatively less accurate in Cubans than on the White cohort (Supplementary Table 2). Finally, the clocks consistently find that Cubans with AD have lower predicted age than controls, though this is only significant for the ZhangEN clock. However, given these inconsisencies and the very small sample size, we caution against over-interpretation of these results. We clarify this in the manuscript and suggest that more work is needed on larger Cuban cohorts before any clear conclusions can be made.

      I was not a conceptual fan of the ensemble clock. The clocks are trained on very different things (e.g., chronological age versus clinical biomarkers) and are designed to capture different aspects of biology. Without more validation and motivation, I don’t think it makes sense to average values that are not designed to measure the same thing.

      We agree that combining the first and second-generation clocks for the task of age prediction is not sensible. However, for AD risk stratification, combining values from multiple clocks that capture different aspects of biology and aging could be beneficial. As mentioned in the main text, we took inspiration from approaches in polygenic risk scores, as well as the broader machine learning field, where ensembling often makes for better predictors. Nonetheless, consistent with the Reviewer’s intuition, we do not see improvement here.

      Minor comments

      (1) Typo in line 91.

      Thank you for bringing this to our attention. Fixed.

      (2) Lines 111-115, sample sizes would be helpful.

      We have added the sample sizes of the non-demented controls that were used to calculate these correlations in each cohort.

      (3) Line 137-138, the correlation stats would be helpful here. This is a common issue throughout the paper, more in-text statistics would help readers to evaluate the authors’ claims. For example, lines 249-251 as well. The authors refer the reader to Figure 5C, which itself has no statistics, this has two plots so it’s unclear which the authors are putting forward as the primary evidence.

      We have added more statistical details in the text and figures to address this comment. In this instance, we have removed the referenced figure.

      (4) Lines 258 and 261, I believe the authors report the same result in both these lines.

      Thank you for pointing out this lack of clarity. These lines report different, but related, results about the frequency of clock-affecting meQTL in different ancestral contexts. The first reports the frequency of clock CpGaffecting meQTL in individuals of African ancestry across all of gnomAD. The second result gives the frequency of those meQTL in different local ancestry backgrounds in admixed individuals. This is distinction is relevant since admixed individuals’ genomes are mosaics of multiple genetic ancestries. As such, a genetic variant might be present in haplotype whose ancestry is not in line with expectations based on global ancestry (e.g., an African American individual inherits a genetic variant within a European ancestry block). This local ancestry difference could modify the effect of the variant or obscure causal variants. Given the potential for confusion and similar results considering global and local ancestry context in this case, we have focused on the first result in the Main Text.

      (5) Somewhere, it would be helpful to provide the distribution/range of ages broken by cohort. Similarly, I didn’t see the breakdown of AD versus control cases within each cohort. Both of these features will impact power within a given cohort for certain analyses.

      We have added the distribution of ages by cohort in Supplementary Figure 1. Table 1 provides a breakdown of cases versus controls for each of the cohorts in the MAGENTA study.

      (6) Figure 3 is pretty hard to read. It would also be helpful if the authors put the white cohort in Figure 3A as a ’baseline’ comparison, as they use this as the baseline comparison in the text.

      We have made these changes to the figure and used larger text overall.

      (7) The various acronyms in the labels in Figure 5 are not explained. For Figure 5C - this is over-plotted and therefore hard to see.

      We have added the full population descriptors from gnomAD to the boxplots showing allele frequencies (Figure 6E). In addition, what used to be Figure 5C has been simplified and moved to Supplementary Figure 12.

      (8) The authors correct for cell type heterogeneity, which is known to vary across populations and can impact clock estimates. However, as far as I can tell, the cell type proportion estimates are coming from the DNA methylation data. The deconvolution algorithms for cell type proportions also have the same problem as the clocks of being trained on a very specific subset of human genetic and environmental diversity. Do the authors have any empirically derived estimates of cell type heterogeneity to sanity-check these deconvolution estimates? At the very least, it would be helpful to acknowledge this limitation.

      We thank the reviewer for commenting on this. There are no empirically derived estimates of cell type counts for the samples in the MAGENTA study. This is an inherent limitation of our study, and we have included text to make note of this.

      (9) There are very different sample sizes for each group, did the authors consider that their null results for the AD analyses in different cohorts are just a lack of power? This could be evaluated with power analyses or by comparing against sample sizes from similar studies in the literature.

      We agree that this is an important analysis and have added it to the manuscript. Given these small effect sizes, accounting for statistical power is essential for interpretation of our results. We performed power calculations based on an effect of the size observed in previous studies (0.5 year acceleration). Considering the full study, we have 86% power to detect an effect of this size. Stratifying by cohorts, we have 75% power for the African Americans, 72% for the Puerto Ricans, 72% for the Whites, 65% for the Peruvians, and 47% for the Cubans. Thus, we have high enough power that the consistent lack of association observed outside of the White cohort in MAGENTA is likely meaningful. Based on these calculations, there is only a 1% chance that we would not observe an effect in any of the other cohorts if the effect was present across cohorts. Nonetheless, we have added caveats about power and the small sample size to our suggestion that the reduced accuracy of the clocks contributes to the lack of association outside of Whites.

      (10) There has been a fair amount of discussion recently that single CpG-based clocks are much more variable than clocks that combine information across CpG sites, either using PC-based or window-based approaches. For example, the PC clock R package from the Levine Lab (https://github.com/MorganLevineLab/PC-Clocks) is very easily implemented and generally gives much less variable age estimations than site-level clocks. It would be nice to consider integrating or discussing these later-generation clocks as ways to improve clock performance in diverse human groups.

      We thank the reviewer for their suggestion to apply the principal component clock to account for potential technical variation. As outlined in the new section “Principal component versions of the methylation clocks also have lower age prediction accuracy for genetically admixed individuals,” using the principal component version of the Horvath clock did not result in consistent improvement in age prediction accuracy or generalization across MAGENTA cohorts (Supplementary Figures 4 and 5). The lower accuracy for age prediction in individuals with substantial African ancestry were present for the PC clock in the replication cohorts, just as in the MAGENTA cohorts (Supplementary Figure 6)

    1. eLife Assessment

      This study presents a comparison of the efficiency and precision of two prime editing methods to introduce single-nucleotide variants and longer exogenous DNA sequences into the zebrafish genome. Convincing data support the conclusion that the PE2 prime editor Nickase is more effective at introducing single-nucleotide variants, while the PEn prime editor nuclease is more effective at integrating sequences from 3 up to 46 base pairs, for both somatic and germline editing. The results will be valuable for the zebrafish community, in particular to model human disease variants in this model organism.

    2. Reviewer #1 (Public review):

      Ono et al., compared the activity of prime editor nickase PE2 and primer editor nuclease PEn in introducing SNPs and short exogenous DNA sequences into the zebrafish genome to model human disease variants. They find the nickase PE2 prime editor had a higher rate of precise integration for introducing single nucleotide substitutions, whereas the nuclease PEn prime editor showed improved precision of integration of short DNA sequences. In somatic tissue the percentage of SNP variant precision edits improved when using PE2 RNP injection instead of mRNA injection, but increased precision editing correlated with elevated indel formation. While PEn overall had higher rates of precision edits, the indel rate was also elevated. Similar rates were observed when introducing a 3 bp stop codon into the ror gene using a standard pegRNA with a 13-nucleotide homology arm, or a springRNA driving integration by NHEJ. Inclusion of an abasic sequence in the springRNA prevented imprecise edits caused by scaffold incorporation, but did not improve the overall percentage of precise edits in somatic tissue. Both PE2 and PEn showed higher frequency of 3 bp precision integration, compared to CRISPR HDR mediated knock-in using a single strand donor DNA template with short homology. Recovery of a germline ror-TGA integration allele using PEn with RNP was robust, resulting in 5 out of 10 founders transmitting a precise allele. The authors demonstrate PEn was effective at integration of a 30 bp nuclear localization signal into the 5' end of GFP in an existing muscle-specific reporter line. PEn-mediated integration of long sequences was further demonstrated by integration into the wls gene of a 46bp attP sequence for phiC31 integrase recombination. Additional analyses are needed to determine if the approach can be used to isolate stable germline alleles of variants that are potentially dominant negative or gain of function in nature.

      The conclusions of the paper are well supported, demonstrating PE2 increases precision, while PEn increases efficiency, for integrating short DNA sequences. Introducing longer sequences up to 46 bp wit PEn highlights the potential broad utility of this approach for insertion of functional motifs for protein modification and gene expression.

      (1) In Figure 3 the data indicates a significant increase in precise edits of the 3 bp TGA using PE2 RNP (11.5%) vs. PE2 mRNA (1.3%). At the adgrf3b locus both PE2 RNP, PE2 mRNA, PEn RNP and PEn mRNA were tested for introducing the 3 bp TGA and a longer 12 bp insertion. PEn RNP showed the highest rate of precision for integration of the longer 12 bp sequence. A comparison of somatic precision editing at additional loci, and analysis of germline transmission rates using PE2 vs. PEn, would support the conclusion that PEn is preferred for precise integration of longer templates, and recovery of germline integration alleles.

      (2) Figure 4 shows the results of introducing a TGA stop codon that is predicted to result in nonsense mediated decay. Testing the ability to also isolate different substitution mutations in the germline would be useful information for identifying the most effective approach for generating human disease variant models.

    3. Reviewer #2 (Public review):

      The manuscript by Ono et al compares two prime editing strategies in zebrafish, one based on a nickase and the other on a nuclease, and evaluates their performance for introducing substitutions, short insertions, and transmission to the next generation. The study aims to clarify the relative strengths of these approaches and to extend their use for inserting short DNA sequences in vivo.

      The study provides a useful and well-executed comparison of two editing strategies in a vertebrate model. In particular, the finding that the nuclease-based approach shows higher efficiency for short insertions is of practical interest for functional studies. The authors also present convincing evidence supporting their conclusions, including sequencing and phenotypic validation at selected loci. These results support the reliability of the approach in this system.

      The overall conceptual advance remains somewhat limited, as the general strategy of delivering prime editing components in zebrafish has been described previously. The present study extends this work by comparing two editing modes and exploring insertion efficiency, which represents a useful but incremental advance.

      Regarding the comparison between the two systems, the authors have made efforts to address concerns about generalizability by adding data from additional loci and by refining the scope of their conclusions. These additions strengthen the manuscript. However, the comparison is still based on a relatively small number of loci, and the conclusions may therefore remain somewhat context-dependent.

      Overall, the authors largely achieve their stated aims of comparing two editing strategies and demonstrating their applicability in zebrafish. The data generally support the conclusions, particularly within the tested loci. The work provides practical value to the community, especially for researchers seeking efficient strategies for short sequence insertion in this model system, although its broader impact is somewhat limited by its incremental nature.

    4. Reviewer #3 (Public review):

      The manuscript by Ono et al describes application of prime editors to introduce precise genetic changes in the zebrafish model system. Probably the most important observation is that compared to the "standard" PE2, prime editor with full nuclease activity appears to be more efficient at introducing insertions into the genome. Although many laboratories around the world have successfully used oligonucleotide-mediated HDR to insert short exogenous sequences such as epitope tags or loxP sites into the zebrafish genome, the method suffers from high frequency of indels at the edit site. Thus, additional tools are badly needed, making this manuscript very important.

      Comments on revised version.

      Thank you for thoroughly addressing my minor concerns.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Thank you very much for handling our revised manuscript and for the careful and constructive comments from the reviewers. We are grateful for the detailed feedback, which has helped us improve both the experimental presentation and the framing of the study. In response to the comments, we have substantially revised the manuscript, updated the figures and supplementary figures, and clarified several points in the text. We have also added new experimental analyses, which were essential to strengthen the manuscript.

      We would like to highlight the major changes in the revised version:

      Added the late phenotype analysis of the ror2 mutant, including loss of nasal and maxillary barbels and altered adult jaw morphology by microCT, strengthening the disease-model relevance.

      Added new data on a further target locus (wls) showing 46 bp attP insertion by PEn and comparison with HDR-mediated knock-in at the same site.

      Expanded the analysis of insertion performance at adgrf3b and clarified comparison with previously reported PE2 data.

      Added the analysis of HDR-mediated knock-in and prime editing substitution to generate ror2 W722X allele.

      Added comparative off-target analysis for PE2, PEn and HDR at three predicted off-target sites for the ror2 target.

      Resolved the cloning/NGS inconsistency for ror2 by increasing clone analysis

      We have also moderated several statements in the manuscript, for example, that editing efficiency is locus- and edit-dependent, and that broader comparison of germline transmission efficiencies between prime editing systems will require future work.

      A few reviewer suggestions would have required substantial additional experimental work that is technically demanding and beyond the immediate scope of the present methods-focused resubmission, for example, a direct side-by-side germline comparison of PE2 and PEn across several loci, or systematic cost benchmarking against HDR across multiple edit classes. Rather than overstate these points, we have acknowledged these limitations directly in the revised manuscript and narrowed our claims accordingly.

      Public Reviews:

      Reviewer #1 (Public review):

      From the work presented, it is unclear how prime editing could be used to transiently model human pathogenic variants, given the low frequency of precision edits in somatic tissue, or to isolate stable germline alleles of variants that are potentially dominant negative or gain-of-function in nature. Without a direct comparison with CRISPR/Cas9 nuclease HDR-based methods that use oligonucleotide templates to introduce edits, the advantage of prime editing is unclear. A cost comparison between prime editing and HDR methods would also be of interest, particularly for integration of longer DNA sequences

      We thank the reviewer for this important comment. In response, we added a direct comparison between PEn-mediated editing and HDR-mediated knock-in at the ror2 locus and the wls locus using insertion of a 46 bp attP sequence. This new dataset shows that PEn can achieve programmed insertion at a higher efficiency in ror2 and comparable efficiency in wls to HDR at the same target site, thereby providing a more direct benchmark within zebrafish embryos. We also revised the Discussion to better position prime editing as a practical donor DNA-free approach rather than as a universally superior method. We agree that a formal cost comparison would be informative; however, such an analysis would depend strongly on locus, edit size, optimisation burden, and local reagent production pipelines, and we believe this is beyond the scope of the present manuscript. Instead, we now discuss these practical considerations more cautiously in the revised Discussion.

      (1) In Figure 3, the data indicate a significant increase in precise edits of the 3 bp TGA using PE2 RNP (11.5%) vs. PE2 mRNA (1.3%). At the adgrf3b locus, only PEn mRNA was tested for introducing the 3 bp and 12 bp insertions. The previous study testing PE2 for 3 and 12 bp insertions was mentioned, but the frequency was not listed, and the study wasn't cited (lines 204 - 207). A comparison of germline transmission rates using PE2 vs. PEn would support the conclusion that PEn allows precise integration of longer templates and recovery of germline integration alleles.

      We appreciate this point. We revised the adgrf3b section to include the relevant reference and explicitly state the previously reported PE2 frequencies, allowing clearer comparison with our PEn data. We added our own experimental data to compare PE2 and PEn with mRNA or RNP form in adgrf3b locus (Figure 3i and j). We also refined the wording of our conclusions so that we do not imply a direct germline comparison between PE2 and PEn where such data are not available. In the revised manuscript, we now state that our germline transmission results apply to PEn-mediated insertions in the loci tested here. A full side-by-side germline comparison between PE2 and PEn across multiple loci would indeed be valuable, but this would require substantial additional animal work and time and is beyond the scope of the present resubmission.

      (2) Figure 4 shows the results of introducing a TGA stop codon that is predicted to result in nonsense-mediated decay. Testing the ability to also isolate different substitution mutations in the germline would be useful information for identifying the most effective approach for generating human disease variant models.

      We agree that this would be useful. In the present study, we focused experimentally on establishing stable lines for the insertion-based edits, while the substitution experiments were used to compare PE2 and PEn performance in somatic editing at the crbn locus. We also tested the generation of ror2 W722X allele by prime editing substitution (Supplementary Figure 3). We have therefore revised the manuscript to clarify the scope of the disease-modelling claim and now state more explicitly that our data support the generation of disease-relevant alleles in cases where short, programmed substitutions or insertions are sufficient.

      A comparison with the prime editing variant knock-in frequencies reported in the recent publication by Vanhooydonck et al., 2025, Lab Animal should be included in the Discussion.

      We have added this study to the revised manuscript and now discuss our findings in relation to the frequencies reported by Vanhooydonck et al. (2025).

      Reviewer #2 (Public review):

      The comparative analysis between PE2 and PEn systems suffers from limited evidentiary support. The comparison relies on single loci for substitutions (crbn) and insertions (ror2), raising concerns about generalizability. Additional validation across multiple loci is necessary to support broad conclusions about PE2/PEn performance

      We appreciate this concern. To strengthen the manuscript, we added new experimental data at an additional target locus, wls, where we tested insertion of a 46 bp attP sequence and compared PEn with HDR-mediated knock-in. We also included the adgrf3b insertion data more prominently. At the same time, we revised the wording throughout the manuscript so that our conclusions are more carefully limited to the loci tested here.

      Reviewer #3 (Public review):

      (1) The logic for introducing two nucleotide changes (at +3 and +10) to change a single amino acid (I378) should be explicitly explained in the main body of the manuscript. It is indeed self-explanatory when looking at Supplementary Figure 1. One way of doing it could be to include Supplementary Figure 1a in Figure 1.

      We thank the reviewer for pointing this out. We have now explained this directly in the main text. Specifically, we state that one nucleotide change introduces the desired missense mutation, whereas the second was included to reduce potential pegRNA misfolding caused by complementarity between the spacer and the PBS/RT template region.

      (2) It is not clear why a 3-nucleotide insertion was used to generate W722X. The human W720X is a single-nucleotide polymorphism, and it should be possible to make a corresponding zebrafish mutant by introducing two nucleotide changes.…

      We agree that this point and have now explained in the main text that the 3 bp stop-codon insertion was chosen as a proof-of-principle strategy for generating a precisely truncated protein through programmed insertion, a type of edit that can be broadly applied to target loci. We also tested the generation of ror2 W722X allele by prime editing substitution (Supplementary Figure 3). We also clarify that prime editing substitution was tested separately here.

      (3) Lines 137-138: T7 Endonuclease assay used in Figure 2d detects all polymorphisms, both precise changes and indels. Thus, if this assay were performed on embryos shown in Figure 1c-d, the overall percentage of modified alleles would be similarly higher for PEn over PE2 (add up precise prime edits and indels). The conclusion in the last sentence of the paragraph is, therefore, incorrect, I believe.

      We agreed with this point and revised the sentence accordingly. The text now states that no obvious cleavage was observed with the PE2/pegRNA condition, suggesting fewer editing events compared with PEn, rather than implying greater precision from the T7E1 result alone.

      (4) Use of terminology. "Germline transmission" is typically used to refer to the fraction of F0s transmitting desired changes (or transgenes) to their progeny, while "germline mosaicism" refers to the fraction of F1s with the desired change in the progeny of a given F0. "Germline transmission" in line 217 should be replaced with "germline mosaicism".

      We have replaced the terminology accordingly in the revised manuscript.

      (5) Lines 253-255: The fraction of injected embryos that had mosaic nuclear expression of GFP, indicative of NLS insertion, should be clarified. It should also be clarified whether embryos positive for nuclear GFP were preselected for amplicon sequencing and germline transmission analyses. This is extremely important for extrapolation to scenarios like epitope tagging, where preselection is not possible.

      We agree and have clarified this in the revised manuscript. We now state the fraction of injected embryos showing mosaic nuclear GFP expression, and we explicitly note that embryos were not preselected prior to sequencing or founder analysis. We further explain that preselection was not practical because the transgene is multicopy and individual fibres showed variable ratios of nuclear to cytoplasmic GFP, which made reliable scoring difficult.

      (6) Statistical analyses. It would be helpful to clarify why different statistical tests are sometimes used to assess seemingly very similar datasets (Figures 1c, 1d, 2b, 2c, 2f).

      We have clarified this in the Materials and Methods section and now state that the choice of statistical test depended on the normality and variance structure of the experimental data.

      (7) Discussion. Since authors suggest that PEn might be especially beneficial for insertion of additional sequences, it is important to stress locus-to-locus variability of success. While the precise +3 insertion was indeed tremendously efficient at both tested loci (ror2 and adgrf3b), +12 addition into adgrf3b was over 10 times less efficient. In contrast, +30 into smyhc:GFP using the shorter pegRNA was highly efficient again. Longer pegRNA did not work nearly as well. As dangerous as it is to extrapolate from small datasets, perhaps these observations indicate that optimization of RT template and PBS may be needed for each new locus in order to significantly outperform oligonucleotide-mediated HDR? If so, would the cost of ordering several pegRNAs and the effort needed to compare them factor in when deciding which method to use?

      We fully agree and have substantially revised the discussion to reflect this point. We now emphasise more clearly that editing efficiency is locus- and edit-dependent and likely influenced not only by insertion length but also by spacer sequence and pegRNA complexity. We cite the relevant literature on prime editing determinants and discuss that locus-specific optimisation may be required. We also softened our concluding claims so that the manuscript presents PEn as a practical donor DNA-free approach rather than as a universally high-efficiency solution.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Because this is a genome editing methods paper, including frequency or percentages of somatic and germline editing in the abstract, in comparison to previously published studies, it would be useful information for the intended audience

      We agree and revised the abstract to include concrete editing frequencies. We now indicate the strongest insertion efficiencies observed. We also retained the statement that edited alleles were transmitted to the next generation.

      Reviewer #2 (Recommendations for the authors):

      (2) Please include additional loci for substitutions and insertions to strengthen conclusions about PE2/PEn efficiencies.

      In response, we added further substitution data at the ror2 (Suppl. Data 3) and insertion at the wls locus (Suppl. Data 6) and strengthened the presentation of the adgrf3b insertion data: first, by adding new locus data where feasible; and second, by narrowing the wording of our conclusions so that they are explicitly limited to the loci tested here.

      (3) Please provide direct comparisons between zebrafish ror2 W722X phenotypes and human Robinow syndrome symptoms to support disease modeling claims.

      We addressed this by adding analysis of the late ror2 phenotype. In the revised manuscript, zygotic and maternal-zygotic mutants are reported to lack nasal and maxillary barbels, and one-year-old mutants show altered jaw morphology with a less protrusive lower jaw (Figure 4).

      (4) The substitution of two nucleotides (+3 G→C and +10 A→G) to target residue I378 of crbn is not justified. It is unclear why two substitutions were required to model thalidomide sensitivity or validate editing efficiency. Please explain why dual nucleotide substitutions were necessary in the crbn experiments and whether single substitutions would suffice.

      We now explain in the main text that the second substitution was introduced to reduce potential inhibitory intramolecular interactions within the pegRNA, while the primary substitution generated the intended amino-acid change. This clarification is now stated explicitly in the Results.

      (5) The reported 10.3% precise editing efficiency for PEn/pegRNA at ror2 conflicts with Supplementary Figure 2, where none of the 20 clones from PEn/pegRNA showed precise edits, while one clone from PEn/springRNA did. Please address the inconsistency between NGS and cloning results at ror2, possibly by increasing sample size or reanalyzing sequencing data.

      We addressed this directly by repeating and expanding the clone analysis. The revised Supplementary Figure 2 now includes the updated clone dataset, and the result is in much better agreement with the NGS-based frequency estimates.

      (6) Figure 3d highlights edits from PEn/springRNA but omits PEn/pegRNA results, despite the latter being described as superior. This creates ambiguity about the relative performance of pegRNA vs. springRNA. Please include PEn/pegRNA results in Figure 3d to fairly represent pegRNA performance.

      We agree. We therefore revised Figure 3e so that it now includes alignment data for PE2/pegRNA, PEn/pegRNA and PEn/springRNA, allowing more direct visual comparison of the editing outcomes.

      (7) The study does not specify the version of PEn used, or introduce some background of PE2 and springRNA. Comparisons to prior PE work in zebrafish, base editing, or HDR efficiencies are absent, obscuring the novelty of this approach. Please specify the PEn variant used, describe springRNA/PE2 structures, and compare results to prior zebrafish PE studies, BE, and HDR efficiencies for similar edits, contextualizing where PE2/PEn offers unique advantages.

      We thank the editors for this helpful suggestion. We have clarified the PEn and PE2 systems in the manuscript, specified the nuclease-based PEn used, and improved the background text introducing these editing strategies. We added the data to directly compare prime editing and HDR in the ror2 locus (Figure 3). We also expanded the Discussion to place the current findings in the context of prior zebrafish prime editing, HDR-based knock-in and base-editing work. We did not test all alternative systems experimentally in the current study, but we now discuss their relevance and clearly define the specific contribution of the present work.

      (8) The manuscript does not explore advanced PE variants (e.g., PE3, PEmax), codon optimization, or scaffold modifications to improve efficiency. Please discuss whether codon optimization, PE3/PEmax systems, or pegRNA modifications were tested or could improve outcomes.

      We agree that this should be discussed and we added recent work on zebrafish prime editing optimisation, codon optimisation, pegRNA engineering and related advances to the discussion, and explain that these are promising avenues for improving efficiency in future studies.

      (9) No data compares the off-target effects of PE2 and PEn, a critical consideration for evaluating specificity and safety. Please perform comparative off-target analyses for PE2 and PEn to assess specificity.

      In response, we performed comparative off-target analysis for the ror2 target and analysed three predicted off-target sites. These data are now included in Supplementary Figure 3 and show no significant increase in non-specific editing for the prime editing conditions tested.

    1. eLife Assessment

      This important study used five metrics to compare the cost-effectiveness of intramural and extramural research funded by the National Institutes of Health in the United States between 2009 and 2019. They found that each type of research had its own set of strengths: extramural research was more cost-effective in terms of publications, whereas intramural research was more cost-effective in terms of influencing clinical work. The evidence supporting these findings is solid.

    2. Reviewer #1 (Public review):

      Summary:

      This paper carefully compares intramural vs. extramural National Institutes of Health funded research during 2009-2019, according to a variety of bibliometric indices. They find that extramural awards more cost-effectively fund outputs commonly used for academic review such as number of publications and citations per dollar, while intramural awards are more cost-effective at generating work that influences future clinical work, more closely in line with agency health goals.

      Strengths:

      Great care was taken in selecting and cleaning the data, and in making sure that intramural vs. extramural projects were compared appropriately. The data has statistical validation. The trends are clear and convincing.

    3. Reviewer #2 (Public review):

      This article reports a cost-effectiveness comparison of intramural and extramural that NIH funded between 2009 and 2019. Using data obtained from NIH RePORTER, they linked total project costs to publication output, using robust validated metrics including Relative Citation Ratio (RCR), Approximate Potential to Translate (APT), and clinical citations. They find that after adjusting for confounders in regression and propensity-score analyses, extramural projects were generally more cost-effective, though intramural projects were more cost effective for generating clinical citations. They also describe differences in the topics of intramural- and extramural-funded publications, with intramural projects more likely to generate papers on viral infections and immunity or cancer metastases and survival, but less likely to generate papers on pregnancy and maternal health, brain connectivity and tasks, and adolescent experiences and depression. The authors aptly describe the different natures of the intramural and extramural funding models, including that extramural researchers spend much time writing grant applications and that the work described in extramural publications often receives funding from sources other than NIH grants.

      Strengths:

      The authors leveraged publicly available data (including RePORTER and the iCite repository) and used robust validated metrics (RCR, APT, clinical citations). They carefully considered a large number of confounders, including those related to the PI, and performed several well-described regression analyses.

    4. Reviewer #3 (Public review):

      This article demonstrates a comparative study on two funding mechanisms adopted by the National Institutes of Health (NIH). The authors adopted a quantitative approach and introduced five metrics to compare the output of intramural and extramural grants. These findings reveal the impacts of intramural and extramural grants on the scientific community, providing funders with insights into the future decisions of funding mechanisms they should take.

      Strengths:

      The authors clearly presented their methods for processing the NIH project data and classifying projects into either intramural or extramural categories. The limitations of the study are also well-addressed.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      Great care was taken in selecting and cleaning the data, and in making sure that intramural vs. extramural projects were compared appropriately. The data has statistical validation. The trends are clear and convincing.

      We thank the reviewer for highlighting the strengths of the manuscript.

      Weaknesses:

      The Discussion is too short and descriptive, and needs more perspective - why are the findings important and what do they mean? Without recommending policy, at least these should discuss possible implications for policy.

      The Discussion has been substantially expanded. We added several new paragraphs discussing: the 2024 Senate HELP Committee proposal for NIH reform; implications for portfolio management (positioning extramural for basic research, intramural for clinical translation); generalizability to other agencies (DoD, NSF FFRDCs, DoE national labs); and the extramural program's role in workforce training as a societal benefit distinct from research outputs.

      The biggest problem I have with this submission is Figure 3, which shows a big decrease in clinical-related parameters between 2014 and 2019 in both intramural and extramural research (panels C, D and E). There is no obvious explanation for this and I did not see any discussion of this trend, but it cries out for investigation. This might, for example, reflect global changes in funding policies which might also influence the observed closing gaps between intramural and extramural research.

      We added an explicit explanation in the Results: because the dataset is truncated at 2020, clinical citations naturally approach zero near the window's end, consistent with the ~7-year lag for clinical citations to accrue documented in prior work (Hutchins et al., 2019). The APT metric declines less steeply because it uses the forward citation network for predictions.

      Reviewer #2 (Public review):

      Strengths:

      The authors leveraged publicly available data (including RePORTER and the iCite repository) and used robust validated metrics (RCR, APT, clinical citations). They carefully considered a large number of confounders, including those related to the PI, and performed several well-described regression analyses.

      We thank the reviewer for highlighting these strengths of the manuscript

      Figure 3A shows intramural projects producing about 2.75 papers per year in 2009, whereas extramural projects are producing just over 1 paper per year. Extramural projects appear to catch up over the next five years. While the authors attempt to explain the difference in their figure legend, another explanation is that the intramural projects started well before 2009 but, as the authors state, intramural data only became available in 2009.

      We added a methodological note acknowledging that some intramural projects may have had start dates prior to 2009 that are not captured in the data, and that the ramp-up of new intramural projects is slower because they are more tied to new PI hiring. We also note the exclusion of projects matched in 2008 as possible continuations. However, the slow ramp-up of Intramural costs in Supplemental Figure 3 is consistent with hiring-associated lagged investment suggesting that our filtering of continuing projects was very successful. Nevertheless, because we cannot completely rule out some continuing projects made it through despite our efforts, we have made the caveats mentioned above in the “Comparison of research topics” section of the Results and the Data section of the Methods.

      As the authors note, funding information is often complex and difficult to characterize for an analysis like this. How did the authors handle: i) publications linked to multiple extramural grants; ii) publications linked to intramural and extramural grants; iii) publications linked NIH grants and non-NIH grants?

      I would think it necessary to somehow apportion credit, as otherwise it would appear that extramural projects are more productive than they truly are.

      We have now explicitly stated that papers with both intramural and extramural funding links were excluded, while papers with multiple links within the same funding type were retained. A new Supplemental Figure 6 was added showing the distribution of papers by number of funding sources for both extramural and intramural grants, demonstrating that the vast majority acknowledged only one project. These changes are in the Methods, Data section and Supplemental Figure 6

      Apportioning credit among a many-to-many graph like the ones used here is indeed a high value problem to solve, but one with many researcher-degrees-of-freedom about analytical design decisions that impact the results. We are working on a rigorous methodology for this, but the amount of time required to do this well is its own research project, and out of scope for manuscript revisions.

      Also, it is not clear if the authors took account of the indirect costs paid by the NIH to universities that have received extramural grants.

      We added explicit language clarifying that all cost comparisons use inflation-adjusted total costs (direct + indirect) for extramural grants. We also added a new sensitivity analysis (Supplemental Figure 4) inflating extramural indirect costs by 30% to approximate unrecovered university expenditures, with the finding that the fundamental pattern holds even under this adjustment. These are found in the “Comparison of funding” and “Comparison of cost effectiveness” sections of the Results, as well as Supplemental Figure 4.

      Reviewer #3 (Public review):

      Strengths:

      The authors clearly presented their methods for processing the NIH project data and classifying projects into either intramural or extramural categories. The limitations of the study are also well-addressed.

      We thank the reviewer for highlighting these strengths of the manuscript

      Weaknesses:

      The article would benefit from a more thorough discussion of the literature, a clearer presentation of the results (especially in the figure captions), and the inclusion of evidence to support some of the claims.

      The Introduction was updated with more specific framing of prior literature (e.g., explicit mention of risk management, funding disparities, and diminishing marginal returns as the focus of prior work). New references were added throughout, including Sampat (2012) on mission-oriented NIH research, Ioannidis et al. (2019) on grant competition inefficiencies, Drummond et al. (2005) on health economic evaluation methods, and the Cassidy (2024) Senate report, throughout the introduction and discussion.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The article would benefit from a more detailed analysis/discussion about the recovery of indirect costs for extramural research.

      I note that the authors are from the University of Wisconsin, which is part of the IRIS network (https://iris.isr.umich.edu/iris-members-map/). They could work with IRIS (also called UMETRICS) to get a better sense as to the true costs of extramural research for each project (e.g., all labor costs, all equipment costs). The IRIS data are extraordinarily robust. Here's an example of an IRIS / UMETRICS paper: https://www.science.org/doi/10.1126/sciadv.abb7348.

      They could, for example, re-do the analyses assuming that the recorded indirect cost covers only 70% of the true indirect costs. Thus, if they get $700,000 indirect costs from RePORTER, they should assume that the true indirect costs were $1,000,000. Similarly, they can add the costs of the time the PI spent writing the grant proposal, using the Bergstrom paper as a guide.

      Another option would be to conduct sensitivity analyses taking into account ~30% incomplete indirect cost recovery (see https://docs.house.gov/meetings/AP/AP07/20171024/106525/HHRG-115-AP07-Wstate-DroegemeierK-20171024.pdf) and lost efficiency due to excess time writing grant proposals (see https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.3000065).

      We conducted a sensitivity analysis as requested inflating extramural indirect costs by 30%, citing the Droegemeier (2017) Congressional testimony as the basis for this estimate. The cost of grant-writing time is now acknowledged in the Discussion as an unreimbursed hidden cost of the extramural system, citing Ioannidis et al. (2019). This narrowed the gap between extramural research and intramural research, but did not close it completely. In addition, our updated regression (Supplemental Figure 4) showed similar trends as our main Figure 4, but with the Intramural advantage heightened and the Extramural advantage diminished. Both remained significant. We have also added to the discussion that there are additional costs and benefits that may not be fully captured in an analysis such as ours.

      The authors appear to have used an agency-perspective for their cost-effectiveness analyses. Generally, it is preferable to use a wider societal perspective. While that may be difficult, the article would benefit from some discussion from the perspective of the government and universities.

      We added a new paragraph explicitly acknowledging the agency-centered perspective and its limitations, noting that it does not capture the full economic cost borne by universities (startup costs, philanthropy, endowments, state contributions, graduate student training, faculty retention, infrastructure). The extramural program's contribution to the US workforce pipeline is specifically highlighted as a societal benefit not captured by the cost-effectiveness metrics.

      Reviewer #3 (Recommendations for the authors):

      Line 84-87: "The overrepresentation of viral research is likely because of the outsize investment toward the intramural Vaccine Research Center, and the cancer/genetics overrepresentation due in part because National Cancer Institute intramural investigators conduct research at that institute as well as at the NIH Clinical Center for their human genetics work." What evidence is there to support this claim?

      A citation to the NCI Center for Cancer Research website was added to support the claim about NCI intramural investigators working at the Clinical Center and Center for Cancer Research, where vaccine research is extensively discussed.

      Lines 107-109. "Given that NIH funding for intramural research has remained relatively constant as a percent of total funding over the years, this indicates larger single awards for intramural research while extramural investigators may increasingly require multiple concurrent grants to sustain their labs." Authors may consider adding a panel to Figure 2 showing the percentage of total funding of intramural vs. extramural funding.

      Rather than adding a panel to Figure 2, we added a new Supplemental Figure 3 showing the cost breakdown and intramural percentage of total funding by year.

      Discussion section: Are any of the findings of this study relevant to other funding agencies in the US (such as the National Science Foundation, the Department of Energy, and the Department of Defense)?

      A new paragraph to the Discussion was added discussing implications for the Department of Defense (including the Congressionally Directed Medical Research Programs), NSF FFRDCs, and the Department of Energy's national labs and FFRDCs, arguing that the incentive-alignment logic likely generalizes across agencies.

      Methods section: Please add an explanation of the technique used for propensity score matching.

      A detailed step-by-step description of the PSM procedure was added, covering propensity score estimation, within-year matching, matched cohort construction, outcome regression on matched data, and visualization of results.

      Figure 1: Please clarify if the relative ratio of intramural projects is calculated from the numbers of grants (as suggested in lines 95-96 and 98-100) or the numbers of publications (as suggested in lines 82-83 and 97-98).

      Also, this figure would be more intuitive if, for each topic, it showed the relevant intramural number (as it currently does) and also the relevant extramural number.

      The caption and Methods were updated to clarify that clustering and ratio calculation are based on projects/grants, not publications. A formula was added to the Methods to make the ratio calculation explicit. The figure itself was not modified to add extramural bars, though the ratio calculation already implicitly encodes both.

      Figure 2: Please change "(red)" to "(blue)" in the caption, and remove the A as there is only one panel in this figure

      Figure 4: Please change "(red)" to "(blue)" in the caption.

      These changes have been made.

      Lines 19-21: I suggest rewriting this sentence as follows:

      "We find that extramural awards are more cost-effective for producing outputs commonly used for academic evaluation, such as publications and citations per dollar, while intramural awards are more cost-effective for generating research that influences future clinical work, more closely in line with agency's health goals."

      The sentence was rewritten substantially in line with the reviewer's suggestion, now reading more clearly with "per dollar" removed as a parenthetical and the structure of the comparison clarified.

      Lines 31-34: Please rewrite this sentence along the following lines to provide more context on previous research into the grant funding system:

      Certain aspects of the grant funding system have been the focus of research, such as AAAA (Azoulay et al., 2009), BBBB (Goldstein and Kearney, 2020), CCC (Hoppe et al., 2019), DDDD (Lauer et al., 2017), EEEE (Wahls, 2018a) and FFFF (Wahls, 2018b), but the relative merits of intramural and extramural funding have received little attention to date.

      The sentence was rewritten to name specific contributions of each cited paper (e.g., risk management, funding disparities, diminishing marginal returns), replacing the generic list of citations.

      Lines 41-44: Please explain "merit score" and please add a reference to an article or website that explains the review process at the NIH.

      "Merit score" was revised to "percentile ranking of overall impact merit score" and a citation to the NIH CSR website ("What happens to your application during and after review?," 2025) was added.

      Lines 53-54: Please change Intramural to intramural (two instances, and also in line 284), and Extramural to extramural.

      "Intramural" and "Extramural" were corrected to lowercase throughout.

      Line 65-67: This sentence ("Potential advantages of the intramural approach are that researchers in the NIH's own laboratories allow the NIH to hire researchers whose research agendas more closely align with its mission.") reads awkwardly. Please clarify.

      The sentence was rewritten to read more clearly: "An advantage of the intramural approach are that NIH has the direct ability to hire scientists whose research closely aligns with agency goals, and researchers do not need to devote time and effort on preparing and submitting grant applications."

      Line 95-97: Authors should consider including an equation to help explain the following sentence: "The relative ratio of intramural projects for each topic was calculated by taking a ratio of the proportions of total grants a topic represented in the intramural vs. extramural portfolios. A relative ratio >1 signifies a higher share of intramural project publications on that topic relative to their share across all topics."

      A formula was added to the Methods defining the topic-level ratio calculation explicitly.

      Line 143: The phrase "may reflect the extra attention intramural investigators are afforded" reads awkwardly - please reword.

      Reworded to "may reflect the extra time intramural investigators save because they do not have teaching and grant writing responsibilities."

      Lines 303-304: This sentence ("First, as the renewal of project contracts may alter the topic and arrangement of the projects, we dropped 70,297 projects with renewal records in our data.") reads awkwardly. Please clarify.

      Reworded to "Since the scientific focus of a study may drift over time, we dropped 70,297 projects with renewal records in our data."

      Line 378-379: Please specify the model of ChatGPT used.

      Done.

    1. eLife Assessment

      Du et al. present a valuable study examining neural activation in medial prefrontal cortex (mPFC) subpopulations projecting to the basolateral amygdala (BLA) and nucleus accumbens (NAc) during behavioral tasks assessing anxiety, social preference, and social dominance. The strength of the evidence linking in vivo neural physiology to behavioral outcomes was considered solid. Overall, the reviewers felt that the revised work provides insight into how distinct mPFC→BLA and mPFC→NAc pathways influence anxiety, exploration, and social behaviors.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      It is well known that neurons in the medial prefrontal cortex (mPFC) are involved in higher cognitive functions such as executive planning, motivational processing and internal state mediated decision-making. These internal states often correlate with the emotional states of the brain. While several studies point to the role of mPFC in regulating behavior based on such emotional states, the diversity of information processing in its sub-populations remains a less explored territory. In this study, the authors try to address this gap by identifying and characterizing some of these sub-populations in mice using a combination of projection-specific imaging, function-based tagging of neurons, multiple behavioral assays and ex-vivo patch clamp recordings.

      Strengths:

      The authors targeted mPFC projections to the nucleus accumbens (NAc) and basolateral amygdala (BLA). Using the open field task (OFT), the authors identified four relevant behavioral states as well as neurons active while the animal was in the center region ("center-ON neurons"). By characterizing single unit activity and using dimensionality reduction, the authors show differentiated coding of behavioral events at both the projection and functional levels. They further substantiate this effect by showing higher sensitivity of mPFC-BLA center-ON neurons during time spent in the open arms of the elevated plus maze (EPM). The authors then pivoted to the three-chamber social interaction (SI) assay to show the different subsets of neurons encode preference of social stimulus over non-social. This reveals an interesting diversity in the function of these sub-populations on multiple levels. Lastly, the authors used the tube test as a manipulation of the anxiety state of mice and compared behavioral differences before/after in the OFT and social interaction tasks. This experiment revealed that "losers" of the tube test spend less time in the center of the open field while "winners" show a stronger preference for the familiar mouse over the object. Using patch-clamp experiments, the authors also found that "winners" exhibit stronger synaptic transmission in the mPFC-NAc projection while "losers" exhibit stronger synaptic transmission in the mPFC-BLA projection. Given the popularity of the tube test assay in rank determination, this provides useful insights into possible effects on anxiety levels and synaptic plasticity. Overall, the many experiments performed by the authors reveal interesting differences in mPFC neurons relative to their involvement in high or low anxiety behaviors, social preference and social rank.

      Weaknesses:

      The authors have addressed all comments.

    3. Reviewer #2 (Public review):

      Summary:

      The goal of this proposal was to understand how two separate projection neurons from the medial prefrontal cortex, those innervating the basolateral amygdala (BLA) and nucleus accumbens (NAc), contribute to the encoding of emotional behaviors. The authors record the activity of these different neuron classes across three different behavioral environments. They propose that, although both populations are involved in emotional behavior, the two populations have diverging activity patterns in certain contexts. A subset of projections to the NAc appear particularly important for social behavior. They then attempt to link these changes to the emotional state of the animal and changes in synaptic connectivity.

      Strengths:

      The behavioral data builds on previous studies of these projection neurons supporting distinct roles in behavior and extend upon previous work by looking at the heterogeneity within different projection neurons across contexts, this is important to understand the "neural code" within the PFC that contributes to such behaviours and how it is relayed to other brain structures.

      Weaknesses:

      The diversity of neurons mediating these projections and their targeting within the BLA and NAc is not explored. These are not homogeneous structures and so one possibility is that some of the diversity within their findings may relate to targeting of different sub-structures within BLA or NAc or the diversity of projection neuron subtypes that mediate these pathways. This is an important future direction for this work but does not detract from the main finding as reported.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Weakness:

      The diversity of neurons mediating these projections and their targeting within the BLA and NAc is not explored. These are not homogeneous structures and so one possibility is that some of the diversity within their findings may relate to targeting of different sub-structures within BLA or NAc or the diversity of projection neuron subtypes that mediate these pathways. This is an important future direction for this work but does not detract from the main finding as reported. The electrophysiological data in Figure 7 have some experimental confounds that makes their interpretation challenging.

      We thank the reviewer for these thoughtful comments. We fully agree that targeting different substructures within the BLA or NAc, as well as the diversity of projection neuron subtypes mediating these pathways, represents an important direction for future investigation. We will certainly explore these possibilities in future studies.

      We have also removed the optogenetics and electrophysiology data, as they may introduce confounds. The removal of these data and figures does not affect our main conclusions.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (a) The authors have improved the manuscript somewhat by refining their description of the results. However, the normalized EPSC experiments still do not make much sense. If you have a higher light intensity or LED duration the curve of the EPSC response will saturate earlier. Similarly, if you are in a highly, or poorly labeled slice or subregion of a slice then you will see responses emerge at different intensities based on the number of synapses labelled. There is no standardization in the way these experiments were performed, so performing some arbitrary post hoc normalisation does not correct for this. Similarly, they also place the fibreoptic manually above the slice each time. This makes it much harder to determine the actual light intensity delivered to the slice on a cell by cell and group by group basis.

      I have reduced my public statement from significant experimental confounds, to some experimental confounds. But the way the experiments were performed does not allow the normalized data to really be interpretable. They still argue that normalized EPSCs are relatively larger. I don't even really understand what this means biologically.

      The subsequent rise/decay and other measures is now better described. However, they note that the decay constant is larger. This means that the kinetics are slower, not enhanced, as they describe.

      Again, we thank the reviewer for the careful advice. We recognize the limitations of the optogenetics and electrophysiology data and have therefore removed them to avoid potential confounds.

    1. eLife Assessment

      This multimodal neuroimaging study leverages fMRI, PET, and deep learning to predict memory performance. The authors introduce the brain-cognition gap to link these different imaging modalities to cognition and evaluate their results in two independent cohorts. The results are solid and provide an important contribution to the literature and will be of interest to neuroscientists working at the interface of cognition, neuroimaging.

    2. Reviewer #1 (Public review):

      Summary:

      The authors attempted to identify if a new deep learning model could be applied to both resting and task state fMRI data to predict cognition and dopaminergic signaling. They found that resting state and moving watching conditions best predict episodic memory, but only movie watching predicts both episodic and working memory. A negative 'brain gap' (where the model trained on brain connectivity predicts worse performance than what is actually observed) was associated with less physical activity, poorer cardiovascular function, and lower D1R availability.

      Strengths:

      The paper should be of broad interest to the journal's readership, with implications for cognitive neuroscience, psychiatry, and psychology fields. The paper is very well-written and clear. The authors use two independent datasets to validate their findings, including two of the largest databases of dopamine receptor availability to link brain functional connectivity/activity with neurochemical signaling.

      Weaknesses:

      The deep learning findings represent a relatively small extension/enhancement of knowledge in a very crowded field.

      It's unclear from these results how much utility the brain gaps provide above and beyond observed performance. It would be helpful to take a median split the dataset on observed performance, and plot aside the current Fig 3 results to see how the cardiovascular and physical activity measures differ based on actual performance. Could the authors perform additional analyses describing how much additional variance is explained in these measures by including brain gaps?

      Some of the imaging findings require deeper analysis. For figure 1f - Which default mode regions have high salience? DMN is a huge network with subregions having differing functions.

      Along the same lines, were the striatal D1R findings regionally specific at all? It would be informative to test whether the three nuclei (Accumbens, Caudate, Putamen) and/or voxelwise models would show something above and beyond what is achieved from averaging D1R across the striatum. What about cortical D1R, which are highly abundant, strongly associated with cognitive (especially WM) performance, and have much unique variance beyond striatal D1R? https://www.science.org/doi/full/10.1126/sciadv.1501672. The PET findings are one of the unique strengths of this paper and are underexplored. It's also unclear if the measure of brain entropy should simply be averaged across all regions.

      It is not clear from the text that the authors met the preconditions for mediation analysis (that is, demonstrating significant correlations between D1R and entropy, in addition to the correlation with brain gap. Could they please report this as well?

      Was age controlled for in the mediation analysis? I would not consider this result valid unless that is the case.

      The discussion is long, but the authors would do better to replace some less helpful sections (e.g., the paragraph on methodological tweaks to parcellations and model alignment) with a couple of other important points, including:

      (1) Discuss the 'sweet-spot' of movie watching for behavior prediction in the context of studies showing that task states 'quench' neural variability: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1007983. This may not be mutually exclusive of the discussion on dopamine and signal-to-noise ratio, but it would be helpful for the authors to discuss their potential overlap vs. unique contributions to the observed findings.

      (2) The argument that dopamine signaling increases signal-to-noise ratio is based on some preclinical data as well as correlational data using fMRI with pharmacological challenges. It is less clear how PET-derived estimates of D1R and D2R availability equate to 'dopamine signaling' as it is thought of in this context. Presumably, based on these data, higher D1R or D2R availability would be related to greater levels of tonic dopaminergic signaling. However, in the case of the COBRA dataset with D2R estimates, those are based on raclopride -- which competes with endogenous dopamine for the D2 receptor. Therefore, someone with higher levels of endogenous dopamine signaling should theoretically have lower raclopride binding and lower D2R estimates. I'm not arguing that the authors logic is flawed or that D1R and D2R are not good measures of dopamine signaling, but I'd ask the authors to dig into the literature and describe more direct potential links for how greater receptor availability might be associated with greater dopamine signaling (and hence lower entropy). Adding this to the discussion would be very valuable for PET research.

      Comments on revised version:

      I thank the authors for their extensive efforts to revise the manuscript. I have no further concerns.

    3. Reviewer #2 (Public review):

      The authors have made several corrections to the original manuscript. For example, they revised the bootstrapping analysis to avoid arbitrarily inflating the degrees of freedom. However, most substantive concerns remain inadequately addressed.

      (1) The primary issue is still the lack of baseline models against which to benchmark the predictive performance of the proposed DenseNet model. This concern was raised independently by two reviewers. Without such benchmarks, it is difficult to interpret the reported results in the context of prior work on MRI-based cognition prediction.

      Notably, the authors state: "While we compared our model with the connectome predictive modeling (CPM) approach and observed better performance with our deep learning framework, we did not conduct a comprehensive benchmark across all available machine learning methods, nor was this the aim of the present study."

      However, I could NOT find any discussion or results related to the CPM model in the manuscript. It is therefore unclear whether the DenseNet model was actually statistically compared with CPM, and, if so, how the comparison was conducted.

      Note that the statement, "While Vieira et al. show that the majority (76%) of prior studies used linear modeling approaches, including CPM and penalized regressions, these models are often vulnerable to overfitting, especially when applied to high-dimensional fMRI data," is not entirely accurate. Linear models typically have far fewer parameters than deep-learning models and are therefore often less prone to overfitting. In fact, it is well established that deep-learning models are particularly susceptible to overfitting and usually require substantially larger sample sizes to achieve stable and reliable performance. Although deep-learning models may outperform shallower models once sufficient data are available and training is well controlled, this does not justify the authors' claim as stated. I therefore disagree with the argument put forward by the authors.

      The authors further justify the absence of benchmarking by stating: "In this context, deep learning was employed as a flexible framework capable of modelling high-dimensional functional connectivity patterns across cognitive states, rather than as a claim of inherent methodological superiority. Thus, our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach." However, most shallow models can likewise be applied across different brain states and cognitive targets. This rationale does not establish deep learning as a uniquely appropriate or necessary choice. If deep learning is indeed a better approach in this context, the authors should demonstrate this empirically through appropriate benchmarking against established baseline models.

      (2) Additional analysis shows that "BCG is not significantly associated with cognition itself". This is the most perplexing result. This is like saying Brain Age Gap is not related to chronological Age. It is counterintuitive since the Brain Age Gap is calculated by chronological age minus actual age, and most research has shown a strong relationship between the Brain Age Gap and age.

      If the brain cognition gap is not related to cognition, is it possible that the results found are mainly due to the predictive model not fitting well with another dataset? Regardless, the lack of association between BCG and cognition deserves a discussion.

      (3) I still do not fully understand the rationale of the mediation analysis. The analysis and findings are still not related to aims 1 and 2, since DA and entropy are not part of the prediction models. But I appreciate the explanation that this part is related to the authors' previous work, and that the authors attempted to link to them somehow.

    4. Author response:

      The following is the authors’ response to the original reviews.

      In the revised version, our primary focus has been to more clearly demonstrate the unique contribution of the brain-cognitive gap (BCG) beyond what is captured by cognitive performance alone, and to show that the BCG is not trivially driven by the observed cognitive scores. Additional analyses now demonstrate that the BCG provides complementary and nuanced information regarding factors associated with cognitive resilience, above and beyond the cognitive measures themselves.

      In response to the comment regarding the inclusion of a baseline predictive model, we would like to clarify that the central aim of our study is to compare predictive utility across different cognitive states (resting state, movie watching, and n-back), rather than to establish a single universally optimal prediction model. Several previous studies have already systematically compared deep learning approaches with more traditional machine learning methods for functional connectome-based prediction. In contrast, the goal of the present study is to examine how brain state modulates the ability of AI-based functional connectome models to capture individual differences in working memory and episodic memory.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors attempted to identify whether a new deep-learning model could be applied to both resting and task state fMRI data to predict cognition and dopaminergic signaling. They found that resting state and moving watching conditions best predict episodic memory, but only movie watching predicts both episodic and working memory. A negative 'brain gap' (where the model trained on brain connectivity predicts worse performance than what is actually observed) was associated with less physical activity, poorer cardiovascular function, and lower D1R availability.

      Strengths:

      The paper should be of broad interest to the journal's readership, with implications for cognitive neuroscience, psychiatry, and psychology fields. The paper is very well-written and clear. The authors use two independent datasets to validate their findings, including two of the largest databases of dopamine receptor availability to link brain functional connectivity/activity with neurochemical signaling.

      Weaknesses:

      The deep learning findings represent a relatively small extension/enhancement of knowledge in a very crowded field.

      It's unclear from these results how much utility the brain gaps provide above and beyond observed performance. It would be helpful to take a median split of the dataset on observed performance and plot aside the current Figure 3 results to see how the cardiovascular and physical activity measures differ based on actual performance. Could the authors perform additional analyses describing how much additional variance is explained in these measures by including brain gaps?

      We thank the reviewer for raising this important point. In response to their request, we first examined the relationship between the BCG and the cognitive measure itself. We did not find any significant relationship in either the DyNAMiC sample (r =0.01, p =0.939) or the COBRA sample ((r =0.01, p=0.894) (see Author response image 1).

      Author response image 1.

      We then conducted additional analyses, splitting the sample into high and low EM performers, and compared their levels of physical activity and Framingham cardiovascular disease (CVD) risk scores. We found no significant difference in physical activity (DyNAMiC: p =0.56, 95% CI: –14.99 - 8.13; COBRA: p =0.29, 95% CI: –3.54 - 1.05) or Framingham CVD risk score (DyNAMiC: p =0.11, 95% CI: –1.08 - 10.72; COBRA: p =0.41, 95% CI: –1.86 - 4.58) between high and low EM perfprmers. Given the significant difference in physical activity and Framingham CVD risk score between positive and negative BCG groups, our results support that BCG provides unique information, beyond the observed cognitive measure (episodic memory score), regarding factors that contribute to cognitive resilience. These results have been added to Section 2.4, and Figure 3 has been updated.

      Some of the imaging findings require deeper analysis. For Figure 1f - Which default mode regions have high salience? DMN is a huge network with subregions having differing functions.

      Grad-CAM provides a coarse, gradient-based attribution that reflects how the learned feature maps contribute to the model output. It is not designed to produce specific input-level interpretations, such as symmetric edge-wise importance values. Therefore, the primary interpretation remains at the network level rather than at the level of individual FC edges.

      Along the same lines, were the striatal D1R findings regionally specific at all? It would be informative to test whether the three nuclei (Accumbens, Caudate, Putamen) and/or voxelwise models would show something above and beyond what is achieved from averaging D1R across the striatum. What about cortical D1R, which is highly abundant, strongly associated with cognitive (especially WM) performance, and has much unique variance beyond striatal D1R? https://www.science.org/doi/full/10.1126/sciadv.1501672. The PET findings are one of the unique strengths of this paper and are underexplored. It's also unclear if the measure of brain entropy should simply be averaged across all regions.

      In this study, we focused on D1DR/ D2DR averaged across the caudate and putamen, which has been reported in our previous work to be more strongly associated with cognitive functions (Johansson et al., 2023, Nyberg et al., 2016), compared to the nucleus Accumbens, which tends to show lower D1DR/D2DR levels and limited association with these cognitive domains. Following the Reviewer’s suggestion, we examined regional variations and found that while both caudate and putamen D1DR showed significant associations with BCG, there were no significant associations for D1DR in the nucleus accumbens or DLPFC with BCG. For D2DR, we observed a significant association between caudate/putamen D2DR and BCG.

      D1DR:

      Partial correlation between:

      Caudate_Bilateral vs. NegGap, (r =0.37, p =0.02

      Putamen_Bilateral vs. NegGap, r =0.34, p =0.03

      Accumbens_Bilateral vs. NegGap, r =0.07, p =0.69

      Mean (LRCaud, LRput, LRacc) vs NegGap, r =0.35, p =0.03

      DLPFC_Bilateral vs NegGap, r =0.21, p =0.21

      Striatum_Bilateral (Mean (LRCaud, LRput)) vs. NegGap, r =0.40, p =0.01

      Caudate_Bilateral vs. PosGap, r=–0.37, p=0.02

      Putamen_Bilateral vs. PosGap, r=–0.53, p=0.02

      Accumbens_Bilateral vs. PosGap, r=–0.25, p=0.31

      Mean (LRCaud, LRput, LRacc) vs PosGap, r=–0.41, p=0.08

      DLPFC_Bilateral vs. PosGap, r=–0.30, p=0.21

      Striatum_Bilateral (Mean (LRCaud, LRput)) vs. PosGap, r=–0.49, p=0.03

      Author response image 2.

      D2DR:

      Correlation between:

      Caudate_Bilateral vs. NegGap, r=0.36, p=0.0003

      Putamen_Bilateral vs. NegGap, r=0.22, p=0.03

      Accumbens_Bilateral vs. NegGap, r= –0.01, p=0.91

      Mean (LRCaud, LRput, LRacc) vs PosGap, r= –0.24, p=0.01

      Striatum_Bilateral vs. NegGap, r=0.39, p=0.0001

      Caudate_Bilateral vs. PosGap, r= –0.34, p=0.004

      Putamen_Bilateral vs. PosGap, r= –0.37, p=0.002

      Accumbens_Bilateral vs. PosGap, r= –0.21, p=0.09

      Mean (LRCaud, LRput, LRacc) vs PosGap, r= –0.38, p=0.001

      Striatum_Bilateral vs. PosGap, r= –0.49, p=0.0001

      We have added the following sentence to the Results section to highlight these regional differences in D1DR/D2DR in relation to BCG.

      “Both D1DR and D2DR availability in the striatum were associated with BCG, such that lower dopamine receptor availability was linked to a greater behavioral-cognitive gap. However, these associations varied by region. For D1DR, significant correlations with BCG were observed in the caudate (positive gap: r = –0.37, p =0.02; negative gap: r= 0.37, p =0.02) and putamen (positive gap: r = –0.53, p=0.02; negative gap:r=0.34, p=0.03), but not in the nucleus accumbens (positive gap: r= –0.25, p= 0.31; negative gap: r =0.07, p=0.69) or the DLPFC (positive gap: r = –0.30, p=0.21; negative gap: r =0.21, p=0.21). For D2DR, both caudate (positive gap: r = –0.34, p=0.004; negative gap: r =0.36, p=0.0003) and putamen (positive gap: r = –0.37, p=0.002; negative gap: r =0.22, p=0.03) showed significant associations with BCG.”

      Author response image 3.

      It is not clear from the text that the authors met the preconditions for mediation analysis (that is, demonstrating significant correlations between D1R and entropy, in addition to the correlation with brain gap. The authors should report this as well.

      This is a fair question. We recalculated entropy in the striatum, given that D1DR is more strongly expressed in this region and, therefore, reduced striatal D1DR may have a more pronounced impact on local entropy (as the reviewer suggested, it may not be appropriate to compute entropy across all brain regions). Our analyses showed that lower D1DR/D2DR levels were associated with higher entropy, which in turn was related to higher BCG.

      DyNAMiC; negative gap:

      Partial correlation between:

      Entropy and D1DR, r = –0.33, p=0.04.

      Entropy and NegGap, r = –0.36, p=0.03.

      DyNAMiC; positive gap:

      Partial correlation between:

      Entropy and D1DR, r = –0.56, p=0.01.

      Entropy and PosGap, r r =0.47, p=0.04.

      COBRA; negative gap:

      Correlation between:

      Entropy and D2DR, r = –0.22, p=0.03.

      Entropy and NegGap, r = –0.27, p=0.007.

      COBRA; positive gap:

      Correlation between:

      Entropy and D2DR, r = –0.26, p=0.03.

      Entropy and PosGap, r = 0.25, p=0.03.

      We have added these results under the result section 2.6. We have further updated Figure 4 in the revised manuscript, reporting these correlation results.

      Was age controlled for in the mediation analysis? I would not consider this result valid unless that is the case.

      We utilized the mediation package in R, and to control for a covariate age in the mediation analysis, we added age as a covariate in both the mediator model and the outcome model. The following information has been added in the method section in the revised version of the manuscript.

      “To assess the statistical significance of this mediation effect, we employed the bootstrapping method as outlined by Preacher and Hayes (145) and age has been controlled for in all statistical analysis.”

      The discussion section is long, but the authors would do better to replace some less helpful sections (e.g., the paragraph on methodological tweaks to parcellations and model alignment) with a couple of other important points, including:

      (1) Discuss the 'sweet-spot' of movie watching for behavior prediction in the context of studies showing that task states 'quench' neural variability: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1007983. This may not be mutually exclusive of the discussion on dopamine and signal-to-noise ratio, but it would be helpful for the authors to discuss their potential overlap vs. unique contributions to the observed findings.

      Thank you for the comment. We have now eliminated the section about methodological tweaks and extended the discussion on the sweet-spot of the task for behavioral prediction by referencing the paper that the reviewer suggested. Here comes the paragraph discussing this topic:

      “Additionally, previous research showed that movie-watching alters the propagation of activity across cortical pathways (105), particularly within and between regions involved in audiovisual processing and attention. These alterations lead to a less segregated and more integrated network organization (106). Similarly, the n-back task has been associated with increased integration of task-positive cortico-cortical connectivity (104, 107) and striato-cortical connectivity (102). Our findings also suggest that certain task contexts strike an optimal balance between reducing neural variability and maintaining sufficient richness to capture individual differences. Prior work shows that task states quench neural variability, leading to a more reliable and predictable neural signal (108). In this context, movie watching may represent such a sweet spot constraining neural dynamics through shared audiovisual stimulation, while simultaneously engaging a broad range of cognitive processes that preserve individual differences.”

      (2) The argument that dopamine signaling increases signal-to-noise ratio is based on some preclinical data as well as correlational data using fMRI with pharmacological challenges. It is less clear how PET-derived estimates of D1R and D2R availability equate to 'dopamine signaling' as it is thought of in this context. Presumably, based on these data, higher D1R or D2R availability would be related to greater levels of tonic dopaminergic signaling. However, in the case of the COBRA dataset with D2R estimates, those are based on raclopride -- which competes with endogenous dopamine for the D2 receptor. Therefore, someone with higher levels of endogenous dopamine signaling should theoretically have lower raclopride binding and lower D2R estimates. I'm not arguing that the authors' logic is flawed or that D1R and D2R are not good measures of dopamine signaling, but I'd ask the authors to dig into the literature and describe more direct potential links for how greater receptor availability might be associated with greater dopamine signaling (and hence lower entropy). Adding this to the discussion would be very valuable for PET research.

      Thank you for raising this important point. We agree that D1R and D2R availability should not be taken as direct proxies of dopamine signaling. However, prior work has suggested meaningful associations between pre- and post-synaptic markers. For instance, a well-powered study demonstrated a significant correlation between D2R availability and dopamine synthesis capacity measured by FMT (Berry et al., 2018). This finding supports the idea that postsynaptic receptor markers may, under certain conditions, serve as an indirect proxy for dopaminergic signaling. Moreover, the number of dopamine-producing neurons innervating the striatum during development has been proposed to shape the structural maturation and arborization of dendrites (McAllister, 2000; Whitford et al., 2002), potentially providing a structural and functional basis for observed associations between pre- and post-synaptic measures.

      At the same time, smaller-scale studies have yielded mixed findings, reporting either non-significant associations (Heinz et al., 2005; Kienast et al., 2008) or negative correlations (Ito et al., 2011). Importantly, the latter studies employed [18F]FDOPA to index dopamine synthesis, which has been argued to provide a less reliable estimate of synthesis capacity compared to FMT, as used in Berry et al. (2018). These inconsistencies underscore that the relationship between pre- and post-synaptic markers is not straightforward and requires further examination in larger, well-powered samples. The following paragraph has been added to the discussion.

      “An important caveat is that D1DR and D2DR availability do not provide a direct measure of dopamine signaling. Instead, they reflect receptor availability, which interacts with endogenous dopamine in a complex manner. PET measures of D1R and D2R availability reflect the density of unoccupied dopamine receptors and the degree to which endogenous dopamine competes with radioligand binding. D2R binding potential is sensitive to competition from synaptic dopamine, such that higher ambient dopamine generally reduces tracer binding; D1R binding, however, is less affected by endogenous dopamine under physiological conditions, reflecting more directly receptor expression levels. Previous studies demonstrated a significant association between D2R availability and dopamine synthesis capacity measured by FMT (117, 118), suggesting that postsynaptic receptor markers may, under certain conditions, serve as a proxy for dopaminergic signaling. Developmental factors, such as the number of dopamine-producing neurons innervating the striatum, may further influence the structural and functional relationship between pre- and post-synaptic markers. By contrast, smaller studies have reported non-significant (119, 120) or negative (121) associations, although these studies relied on [18F]FDOPA, which is considered a less precise index of dopamine synthesis than FMT. Taken together, these reports indicate that the relationship between pre- and post-synaptic markers is complex and not necessarily linear. Accordingly, our observation that lower receptor availability is associated with greater neural variability should not be interpreted as direct evidence of weaker dopaminergic signaling, but rather as reflecting the interplay between receptor density and endogenous dopamine occupancy, particularly in the case of D2DR.”

      Reviewer #2 (Public review):

      Summary:

      The authors developed a deep learning model based on a DenseNet CNN architecture to predict two cognitive functions: working memory and episodic memory, from functional connectivity matrices. These matrices were recorded under three conditions: during rest, a working memory task, and a movie, and were treated as images for the CNN algorithm. They tested their model's performance across different conditions and a separate dataset with a different age distribution (using the same MRI scanner, scanning configurations, and cognitive tests). They also calculated the "brain cognition gap" based on the model trained on resting functional connectivity to predict working memory. Extending from the commonly used index "brain age," the brain cognition gap was defined as the difference between the working memory score predicted by their model (predicted working memory) and the working memory score based on the working memory test itself (observed working memory). This brain cognition gap was found to be associated with physical activity, education, and cardiovascular risk. The authors also conducted additional mediation tests to examine whether regional functional variability mediated the relationship between PET-derived measures of dopamine and the brain cognition gap.

      Strengths:

      The major strength of this manuscript is the extensive effort the authors have put into creating a new 'biomarker' that links deep learning with fMRI, PET, physical activity, education, and cardiovascular risk across two studies. This effort is impressive.

      Weaknesses:

      There are several weaknesses in the current methods and results, making many of the claims unconvincing. These weaknesses include:

      (1) The lack of baseline models to benchmark the predictive performance of their DenseNet models.

      (2) The inappropriate calculation of the brain cognition gap due to the lack of control for regression-toward-the-mean and the influence of the working memory itself (a common practice in brain age studies).

      (3) The lack of benchmarking of the brain cognition gap against the 'corrected' brain age gap and the direct prediction of physical activity, education, and cardiovascular risk.

      (4) Minimal justification for their PET mediation analysis.

      We appreciate the reviewer’s constructive comments on the strengths and weaknesses of our study. In this revised version, we’ve addressed the concerns regarding the calculation of the brain-cognitive gap, clarified the unique variance that the brain-cognitive gap contributes beyond cognition itself, and provided additional justification for the PET mediation analysis. For the lack of a baseline model, it is important to highlight that our aim has never been to compare the predictive power of different deep learning or machine learning approaches. Therefore, the text in the introduction and discussion has been amended to avoid miscommunication on this topic.

      Regarding the impact of the work on the field and the utility of the methods and data to the community, I see its potential. However, addressing all the weaknesses listed above is crucial and likely to change the conclusions of the results.

      It is important to note that many statements in the manuscript are overstated, making the contribution of the manuscript seem exaggerated.

      We have run additional analysis based on the reviewer’s suggestions. The effect sizes and statistical values were adjusted due to the corrections; the overall conclusions remain largely consistent. The relationships between the brain-cognition gap and key factors such as physical activity, and cardiovascular risk persisted. We have updated the manuscript accordingly and revised the relevant sections to reflect these refinements and the resulting interpretations.

      For instance, the abstract claims "there is a lack of objective biomarkers to accurately predict cognitive function," and the discussion states, "across various studies, the correlation between predicted and actual fluid intelligence typically hovers around 0.25 (98-100)." However, a meta-analysis by Vieira and colleagues (2022 https://doi.org/10.1016/j.intell.2022.101654) found over 37 studies up to 2020 predicting cognitive abilities from fMRI with machine learning, with 24 studies published in 2019-20 alone. Since 2020, with the rise of machine learning and AI, even more studies have likely been published on this topic, all claiming to show objective biomarkers to accurately predict cognitive function. Vieira and colleagues also found an average performance of these objective biomarkers in predicting general cognition at r = .42, similar to what was found in this manuscript. Based on this alone, it is unclear how novel or superior their method is without a proper systematic benchmark.

      We appreciate the opportunity to clarify our study’s contribution relative to prior work. We have revised the introduction and discussion to highlight the contribution of other methods when it comes to biomarkers. As for the comment related to the work by Vieira and colleagues, Vieira et al. (2022) indeed present a comprehensive meta-analysis of studies predicting general and fluid intelligence using neuroimaging and machine learning. However, there are two critical differences between ours verus previous work:

      Target Cognitive Domains:

      Our study does not focus on general or fluid intelligence, but rather on comprehensive EM (3 tests) and WM (3 tests), two distinct cognitive domains that are critically important for aging research. These distinct abilities, in this context (measured by three independent tests to boost the reliability) are less frequently studied as predictive targets in the existing fMRI-ML literature, particularly using deep learning methods.

      Critically, our study explicitly compares predictive power across different cognitive states (rest, movie watching, n-back), with the aim of identifying the states that best capture individual differences across domains. Thus, our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach.

      Our primary objective is to test how brain state influences the ability of functional connectivity to predict domain-specific cognitive performance, using a deep learning framework. As now stated explicitly in the revised manuscript, this objective is operationalized through three clearly defined aims:

      (1) To compare the predictive utility of functional connectomes derived from different brain states (resting state, movie watching, and n-back task) for EM and WM;

      (2) To introduce and evaluate a brain-cognition gap as a marker of individual differences beyond chronological age; and

      (3) To examine the contribution of dopaminergic integrity to variability in connectome uniqueness and brain-cognition gaps.

      We have revised the manuscript text to make this focus clearer and to avoid any misinterpretation of our aims. Specifically, we removed statements in the Discussion that could be read as suggesting that our deep learning approach outperforms prior machine learning methods. While we compared our model with the connectome predictive modeling (CPM) approach and observed better performance with our deep learning framework for some of the prediction models, we did not conduct a comprehensive benchmark across all available machine learning methods nor was this the aim of the present study. Accordingly, we have adjusted the text to avoid implying methodological/biomarker superiority beyond the scope of our analyses.

      Modeling Approach:

      While Vieira et al. show that the majority (76%) of prior studies used linear modeling approaches, including CPM and penalized regressions, these models are often vulnerable to overfitting, especially when applied to high-dimensional fMRI data. Our use of a DenseNet-based CNN architecture is motivated by the need to leverage inductive biases suited to functional connectivity data, and we evaluate this approach across multiple cognitive tasks and independent datasets.

      Vieira and colleagues report that studies predicting general intelligence from fMRI (particularly from the HCP dataset) average around r =0.42, while those predicting fluid intelligence average around r =0.15. Our original claim about the correlation hovering around 0.25 is therefore not incorrect – and aligns with the Vieira meta-analysis. We have, however, nuanced this statement in the manuscript, now stating that correlations are higher for general intelligence than fluid intelligence.

      Altogether, we considered the reviewer’s comments and therefore conducted a careful revision of the manuscript text to moderate and clarify statements that may have come across as overstated. We have refined the language throughout the Introduction and Discussion sections to better align with the strength of the evidence and the scope of our contributions. A few examples are:

      “Our study explicitly compares predictive power across different cognitive states (rest, movie watching, n-back), with the aim of identifying the states that best capture individual differences across domains. The relative performance of deep learning and other non-linear approaches depends on multiple factors, including sample size, model architecture, feature representation, and domain-specific characteristics of the prediction target. In this context, deep learning was employed as a flexible framework capable of modeling high-dimensional functional connectivity patterns across cognitive states, rather than as a claim of inherent methodological superiority. Thus, our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach.”

      Also in page 14.

      “Our study introduces a deep neural network architecture that features dense connections and incorporates an attentional mechanism. While our findings demonstrate that a deep learning framework can provide reasonable predictive accuracy, it is important to note that other machine learning approaches (e.g., tree-based models) may offer comparable predictive power, as suggested by prior benchmarking work (29, 30).”

      Similarly, the authors claim superior performance of deep learning and mischaracterize machine learning algorithms: "In particular, deep neural networks (DNN) methods have been successfully applied to behavioral and disease prediction (24-26), and have been found to outperform other machine learning approaches (27-29)," and "Deep learning approaches overcome the limitation of predictive techniques that solely rely on linear associations between connectivity and behavioral phenotypes (17)." However, the superiority of deep learning is debatable. Studies show comparable performance between machine learning (such as kernel regression) and deep learning (such as fully-connected neural networks, BrainNetCNN, Graph CNN (GCNN), and temporal CNN), e.g., He and colleagues (2019) and Vieira and colleagues (2024) https://doi.org/10.1016/j.neuroimage.2019.116276 and Vieira and colleagues' https://doi.org/10.1101/2024.03.07.583858.

      We agree that the performance gap between traditional machine learning models and deep learning (which is a subcategory of machine learning) in neuroimaging is debatable and task-dependent. Indeed, both He et al. (2019) and Vieira et al. (2024) offer evidence that kernel regression can achieve performance on par with deep learning models, applied to appropriate datasets.

      We have therefore nuanced the statements in the revised version of the manuscript as follows:

      Introduction:

      “In particular, deep neural networks (DNN) methods have been successfully applied to behavioral and disease prediction (24-26), and were initially expected to outperform other machine learning approaches (27-29). However, this superiority remains debatable, as recent studies have reported comparable performance between DNNs and traditional methods (He et al.,2019; Vieira et al.,2024). Accordingly, the present study does not aim to benchmark deep learning against traditional machine learning approaches, but instead uses a consistent predictive framework to examine how brain state influences the utility of FC for cognitive prediction.”

      “Deep learning approaches offer a flexible modeling framework capable of capturing complex non-linear associations in high-dimensional data with potentially less sensitivity to training on a smaller subsample (Vieira et al., 2024)”.

      Discussion:

      We agree that traditional methods, such as kernel-based models, tree ensembles, and non-linear SVRs, can also effectively capture such relationships. The relative performance of our model and other non-linear approaches depends on several factors, including data size, model architecture, and domain-specific considerations. We have included additional explanations in the discussion to address this.

      Moreover, many non-deep learning predictive techniques are non-linear, e.g., XGBoost, CatBoost, random forest, kernel ridge, and support vector regression with non-linear kernels (such as RBF and polynomial). Thus, stating that machine learning can only model linear relationships is incorrect. Moreover, for the small amount of data the authors had, some might argue that a linear algorithm might be more appropriate to balance the bias-variance trade-off in prediction. Again, without a proper systematic benchmark, it is unclear how well their DenseNet algorithm performs compared to other algorithms.

      Thank you for bring this up. We have now removed statements implying that machine learning can only model linear relationship.

      Regarding the Brain Age literature, the authors also misinterpreted recent findings: "However, a recent study suggests that brain age predictions contribute minimally compared to chronological age for explaining cognitive decline (65), implying that cognitive predictions are more reliable." In this study, Tetereva and colleagues (2024) (https://doi.org/10.7554/eLife.87297.4) showed that non-deep-learning machine learning can make good predictions from MRI on both chronological age (with r up to .88) and fluid cognition (with r up to .627). Using the combination of functional connectivity matrices across rest and tasks to predict fluid cognition, they found performance at r = .565, comparable to what was found in the current manuscript with deep learning. Nonetheless, while brain age predicted chronological age well (and brain cognition predicted fluid cognition well), it was problematic to predict fluid cognition from brain age. They showed that, because brain age, by design, shared so much common variance with chronological age, brain age and chronological age captured the same variance of fluid cognition. When chronological age was controlled for in the prediction of fluid cognition, brain age no longer had high predictive ability. In the case of the current manuscript, the brain cognition gap is not appropriately controlled for cognition (to be more precise, a working memory score). I expect the performance in predicting physical activity, education, and cardiovascular risk will drop dramatically once cognition is controlled for. There are at least two ways to control cognition according to Tetereva and colleagues' study (see more in the recommendations).

      We thank the reviewer for breaking down the findings in the study by Tetereva and colleagues (2024). It was not our intention to suggest that Tetereva et al. showed brain age has little predictive value in general. Our understanding of the findings reported in that study is on par with the reviewers’ clarifications. We have now revised the introductions to avoid any misunderstanding:

      “A recent study demonstrated that while brain age can predict chronological age with high accuracy from MRI, its utility for predicting cognition is limited. Specifically, Tetereva and colleagues (2024) showed that brain age strongly tracks chronological age and that brain cognition (using functional connectivity) can predict fluid cognition. Yet, when used to predict cognition, brain age largely overlapped with chronological age, such that controlling for chronological age eliminated the predictive contribution of brain age. This finding suggests that brain-age models may provide little unique explanatory power for cognitive decline beyond what is already captured by chronological age. Building on this observation and extending the concept of a brain-age gap to a brain-cognition gap (BCG, defined as the discrepancy between predicted and observed cognitive performance), we propose that a BCG may serve as an informative marker of individual differences.”

      In addition, in response to the first comment from Reviewer 1, we have extended our results in the manuscript. We first showed that BCG is not significantly associated with cognition itself (see Author response image 1). Moreover, we conducted additional analyses, splitting the sample into high and low EM performers, and compared their levels of physical activity and Framingham cardiovascular risk scores. We found that no significant difference in physical activity (DyNAMiC: p =0.56, 95% CI: -14.99 – 8.13; COBRA: p =0.29, 95% CI: -3.54 – 1.05) or Framingham CVD risk score (DyNAMiC: p =0.11, 95% CI: -1.08 – 10.72; COBRA: p =0.41, 95% CI: -1.86 – 4.58) between high and low EM performers. Given the significant difference in physical activity and Framingham CVD risk score between positive and negative BCG groups, our results support that BCP provides unique information, beyond cognitive measures, regarding factors that contribute to cognitive resilience. This text has been added into the result section, and Figure 3 has been updated in the manuscript.

      The authors mentioned, "The third aim of the current study is to uncover the contribution of dopamine (DA) integrity to brain-cognition gaps." However, I fail to see how mediation analysis would test this. The authors also mentioned, "Insufficient DA modulation can affect neurocognitive functions detrimentally (69, 74, 76-78)." They should test if DA levels are related to working memory scores in their study, and if so, whether the relationship is mediated by the "corrected" brain-cognition gaps. Note see more on the recommendation for the calculation of the "corrected" brain-cognition gaps.

      Our mediation was not designed to test whether DA predicts episodic memory performance directly, nor whether BCG mediates such a relationship. Instead, we specifically investigated whether the effect of DA on BCG operates through functional variability, the theoretical framework emphasizing the role of DA on neuronal grain and signal-to-noise ratio (see our recent work in Korkki et al., 2025). We agree that future work could extend our approach by directly examining whether BCG mediates the link between DA and cognitive outcomes. However, in the present study, our primary focus was on testing the mechanistic pathway of DA → entropy → BCG.

      In line with this aim, we found that lower DA receptor availability was associated with larger BCGs (Figure 4). We then asked whether this relationship is mediated by functional signal variability, such that lower DA is linked to reduced signal-to-noise ratio (i.e., greater entropy), which in turn contributes to less reliable prediction of cognition and, consequently, larger BCGs. Our mediation analysis supports this pathway (please see also our reply to Reviewer 1, Comment 6).

      Reviewer #3 (Public review):

      Summary:

      This paper by Esmaeili and co-authors presents a connectome prediction study to predict episodic memory and relate prediction errors to other phonotypic variables.

      Strengths:

      (1) A primary and external validation dataset.

      (2) Novel use of prediction errors (i.e., brain-cognitive gap).

      (3) A wide range of data was investigated.

      Weaknesses:

      (1) Lack of comparisons to other methods for prediction.

      (2) Several different points are being investigated that don't allow any particular one to shine through.

      (3) Some choices of analysis are not well-motivated.

      (4) How do the n-back connectomes perform for prediction if the authors do not regress task activations from the n-back task?

      We thank the reviewer for raising these important points. For the lack of comparisons to other methods, it is important to highlight that our aim has never been to compare the predictive power of different deep learning or machine learning approaches. Rather, our primary objective was to test how brain state influences the ability of functional connectivity to predict domain-specific cognitive performance, using a deep learning framework.Therefore, the text in the introduction and discussion has been amended to avoid miscommunication on this topic.

      We chose to regress out task-evoked activations based on prior work demonstrating that failing to do so can produce spurious but systematic inflation of task functional connectivity estimates (Cole et al., 2019). In that study, as well as subsequent reports (e.g., Gao et al., 2020; Gonzalez-Castillo & Bandettini, 2018), connectomes derived without activation regression tended to capture task-evoked coactivations rather than background task functional interactions, which can artificially boost predictive performance but limit interpretability (whether it is co-activation or intrinsic connectivity during an entire goal-oriented task) and generalizability. For this reason, our analyses focused on the more conservative approach of regressing out task activations. Accordingly, we compared predictive performance only under this preprocessing strategy.

      We have added the following sentence to clarify this in the method: “To avoid spurious inflation of task functional connectivity by task-evoked activations, we regressed out task activation patterns from the n-back data prior to estimating functional connectivity, following recommendations by Cole et al. (2019) and related work.”

      (5) I am a little concerned about overfitting with the convolutional neural net. For example, the drop-off in prediction performance in the external sample is stark. How does the deep learning approach used here compare to something simpler, like a connectome-based predictive model or ridge regression?

      (6) It may be nice to try the other models in the validation dataset. This would also provide a sense of the overfitting that may be going on with overfitting.

      We thank the reviewer for raising this point. The prediction performance indeed dropped for episodic memory when models trained on the DyNAMiC sample were applied to the COBRA sample, whereas performance for working memory remained nearly identical across datasets. Moreover, our prediction power is on par with previous studies reporting reliable prediction of intelligence using deep learning approach (Vieira et al., 2021; Fan et al.,2020). While we compared our model with the connectome predictive modeling (CPM) approach and observed better performance with our deep learning framework, we did not conduct a comprehensive benchmark across all available machine learning methods nor was this the aim of the present study.

      We have revised the manuscript text to make this focus clearer and to avoid any misinterpretation of our aims. Specifically, we removed statements in the Discussion that could be read as suggesting that our deep learning approach outperforms prior machine learning methods. Finally, We have added the following paragraph to the discussion:

      “Our study used a deep neural network architecture that features dense connections and incorporates an attentional mechanism. While our findings demonstrate that a deep learning framework can provide reasonable predictive accuracy, it is important to note that other machine learning approaches (e.g., tree-based models) may offer comparable predictive power, as suggested by prior benchmarking work (29, 30). Our study explicitly compares predictive power across different cognitive states (rest, movie watching, n-back) to identify the states that best capture individual differences across domains. The relative performance of deep learning and other non-linear approaches depends on multiple factors, including sample size, model architecture, feature representation, and domain-specific characteristics of the prediction target. In this context, deep learning was employed as a flexible framework capable of modeling high-dimensional functional connectivity patterns across cognitive states, rather than as a claim of inherent methodological superiority. Thus, our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach.”

      (7) While predictive models increase the power over association studies, they still require large samples to prevent overfitting. Do the authors have a sense of the power their main and external validation sample sizes provide?

      We thank the reviewer for this important point. Our main sample size, together with the external validation in COBRA, is moderate for deep learning applications. To reduce the risk of overfitting, we employed several strategies, including external validation, early stopping, dropout, and regularization. As noted, performance for episodic memory decreased in the external sample, which we acknowledge, but key associations such as the link between BCG and resilient factors remained significant. Importantly, prediction of working memory was maintained across datasets, reducing the likelihood that the observed findings are driven by overfitting. We have added a statement in the Discussion to reflect on the limitations of sample size and the implications for generalizability.

      We added the following sentence to the discussion:

      “We acknowledge that our main and validation samples are moderate in size for deep learning, which constrains statistical power and generalizability. Although external validation, early stopping, dropout, and regularization help mitigate overfitting, larger samples will be needed in future work to fully establish the robustness of these predictive models.”

      (8) I am not sure that the Mann-Whitney is the correct test for comparing the distributions of prediction performances. The distributions are dependent on each other as they are each predicting the same outcomes. Using the typical degrees of freedom formula would overestimate the degrees of freedom.

      We appreciate the reviewer’s comment and agree that applying statistical tests directly to bootstrapped samples can lead to inflated or misleading p-values, as the degrees of freedom are determined by the number of bootstrap iterations rather than the actual number of independent observations.

      In our analysis, the Mann-Whitney U test was applied to 1000 bootstrapped correlation coefficients (r) for each model. While this number is relatively low and was chosen to limit overestimation of significance, we recognize that these bootstrapped samples are not independent, and thus the use of a Mann-Whitney U test can still be problematic. To address this concern, we have revised our statistical analysis. Rather than applying the Mann-Whitney U test to the bootstrapped r distributions, we now compute the difference in correlation coefficients (Δ r = r<sub>actual</sub> − r<sub>rest</sub>) for each bootstrap iteration. We then calculate a 95% confidence interval for Δr. If this interval does not include zero, we consider the difference statistically significant. This approach avoids artificially inflating the sample size and adheres more closely to proper statistical inference.

      We have updated the Methods (the following text) and Results sections accordingly and clearly stated the limitations regarding the degrees of freedom for all tests.

      “For the bootstrap-based comparison of model performance (bootstrap resampling with 1000 iterations), no test statistic with an associated degree of freedom is reported. Instead, statistical inference is based on the bootstrap distribution of the difference in correlation coefficients (Δr) and its 95% confidence interval. As bootstrap confidence-interval–based inference does not rely on an analytic sampling distribution, degrees of freedom are not defined for this procedure.” This has now been explicitly stated in the Methods section to avoid ambiguity.

      In the result section, we have reported with corresponding CI.

      (9) The brain cognition gap is interesting. It is very similar conceptually to the brain age gap. When associating the brain age gap with other phenotypes, typically age is regressed from the brain age gap and the other phenotype. In other words, age is typically associated with a brain age gap as individuals at the tail ages often show the largest gaps. Is the brain cognition gap correlated with episodic memory and do the group differences hold if episodic memory is controlled for?

      We thank the reviewer’s comment regarding the relationship between the brain cognition gap and episodic memory.

      Since this question was raised by all reviewers, we have conducted additional analyses. We did find that BCG is independent from the cognitive measure and provided additional information, beyond cognition alone, about factors contributing to resilience. Please visit our response to the first comment of Reviewer 1.

      (10) I have the same question for the dopamine results. Particularly, in the correlations that are divided by brain cognition gap sign. I could see these types of patterns arise due to a correlation with a third variable.

      For dopamine results, we explored whether age or cognition alone might confound the dopamine–brain cognition gap relationships. However, neither was significantly correlated with the brain cognition gap groups. The associations remained significant after controlling for age, suggesting that the observed patterns are not likely due to these potential third-variable confounder. This is also inline with our observation of significant associations between DA and GAP in an age-homogeneous COBRA sample. That said, we found that entropy, indeed, mediates the direct link between DA and BAG, suggesting that individuals with lower DA exhibit greater regional variability, and in turn larger BCG.

      These results have now been embedded into the manuscript. We also highlighted that age has been controlled for in reported correlation and mediation analyses.

      Recommendations for the authors:

      Reviewing Editor Comment:

      We particularly recommend that the authors: (a) compare the performance of their deep learning model with other baseline models, and (b) adjust for cognitive performance within the brain-cognition gap. These steps would strengthen the evidence base.

      We thank the editor for their comments. As for the first comments, our study explicitly compares predictive power across different cognitive states (rest, movie watching, n-back), with the aim of identifying the states that best capture individual differences across domains. Thus, our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach. We have revised the manuscript text to make this focus clearer and to avoid any misinterpretation of our aims. Specifically, we removed statements in the Discussion that could be read as suggesting that our deep learning approach outperforms prior machine learning methods. While we compared our model with the connectome predictive modeling (CPM) approach and observed better performance with our deep learning framework, we did not conduct a comprehensive benchmark across all available machine learning methods, nor was this the aim of the present study. Accordingly, we have adjusted the text to avoid implying methodological superiority beyond the scope of our analyses. Finally, we have added the following paragraph to the discussion:

      “Our study used a deep neural network architecture that features dense connections and incorporates an attentional mechanism. While our findings demonstrate that a deep learning framework can provide reasonable predictive accuracy, it is important to note that other machine learning approaches (e.g., tree-based models) may offer comparable predictive power, as suggested by prior benchmarking work (29, 30).

      Our study explicitly compares predictive power across different cognitive states (rest, movie watching, n-back) to identify the states that best capture individual differences across domains. The relative performance of deep learning and other non-linear approaches depends on multiple factors, including sample size, model architecture, feature representation, and domain-specific characteristics of the prediction target. In this context, deep learning was employed as a flexible framework capable of modeling high-dimensional functional connectivity patterns across cognitive states, rather than as a claim of inherent methodological superiority. Thus, our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach.”

      As for the second comment, we followed the instructions by Reviewer 1. In response to their request, we first examined the relationship between the Brain-Cognitive Gap (BCG) and the cognitive measure itself. Surprisingly, we did not find any significant relationship in either the DyNAMiC sample (r =0.01, p =0.939) or the COBRA sample (r =0.01, p =0.89) (see Author response image 1).

      We then conducted additional analyses, splitting the sample into high and low EM performers, and compared their levels of physical activity and Framingham cardiovascular disease (CVD) risk scores. We found no significant difference in physical activity (DyNAMiC: p =0.56, 95% CI: –14.99 - 8.13; COBRA: p =0.29, 95% CI: –3.54 - 1.05) or Framingham CVD risk score (DyNAMiC: p =0.11, 95% CI: –1.08 - 10.72; COBRA: p =0.41, 95% CI: –1.86 - 4.58) between high and low EM perfprmers. Given the significant difference in physical activity and Framingham CVD risk score between positive and negative BCG groups, our results support that BCG provides unique information, beyond the observed cognitive measure (episodic memory score), regarding factors that contribute to cognitive resilience. These results have been added to Section 2.4, and Figure 3 has been updated.

      Reviewer #1 (Recommendations for the authors):

      (1) The top and bottom triangles of the saliency maps, particularly in Figure 2, do not look symmetrical (this is most notable in the hotspot representing the between-network correlation of DMN and FPN). What is going on here? Was the image compressed or altered in some way, or is this a visual artifact of the interpolation method?

      We appreciate the reviewer’s insightful comment. Minor differences in the saliency maps between the upper and lower triangles of the FC matrix can arise due to several factors. For instance, Grad-CAM generates saliency maps at the resolution of the convolutional feature maps, which are then upsampled to match the input matrix dimensions. We initially used the default bilinear interpolation, which may have introduced slight asymmetries or blurring, resulting in interpolation artifacts. In response, we have reprocessed the saliency maps using spline interpolation in MATLAB. The updated saliency figures have been included in the revised version of the manuscript.

      (2) Pages 11-12. Please make it explicit in the text that the brain gap-education association was not significant in the COBRA dataset.

      Thanks for pointing this out. We added the following sentence to the discussion.

      “Note that the association with education was significant only in the DyNAMiC sample and did not reach significance in the COBRA dataset.“

      (3) Please overlay individual data points onto the boxplots in Figure 3 so that we can appropriately evaluate the data distributions.

      Figure 3 has now been updated.

      (4) Section 2.6: Was entropy calculated on movie-watching data, resting data, or all fMRI data? Please specify.

      We thank the reviewer for pointing this out. We have updated the text (Section 2.6) to clarify that entropy was calculated from the resting-state data. We intended to examine the mediating role of regional variability in the relationship between dopamine and the BCG of the winning model for episodic memory. Because resting state and movie-watching were the winning conditions for EM prediction, but movie-watching was not available in COBRA, we focused on entropy during rest, which exists in both datasets.

      (5) Was entropy during the resting state correlated with entropy during the task state, across individuals?

      We agree this is an interesting question. However, investigating the correlation of entropy between rest and task states goes beyond the scope of the present study. Our aim here was to test whether regional variability mediates the effect of dopamine on the BCG. Specifically, we examined whether individuals with lower striatal D1DR show higher local variability, which in turn relates to less accurate prediction and a larger gap. We assessed both the relationship between D1DR and entropy and the association between entropy and the gap, and these results have now been added to the manuscript (see also our response to Reviewer 1’s public comment).

      Reviewer #2 (Recommendation for authors):

      (1) The lack of baseline models to benchmark the predictive performance of their DenseNet models makes their results hard to interpret. This problem is quite common across ML literature. For instance, many DL-based algorithms were developed for tabular data without proper benchmarking against other ML algorithms. When they were properly tested, most weren't better than many tree-based ML algorithms (e.g., https://proceedings.neurips.cc/paper_files/paper/2022/file/0378c7692da36807bdec87ab043cdadc-Paper-Datasets_and_Benchmarks.pdf). I can see that a similar problem might happen here.

      For this particular manuscript, the authors made strong statements without doing a proper benchmark, e.g., from the discussion, "Indeed, the predictive power in the current study is stronger than for CPM-based predictions reported before." And "Unlike the BrainNet convolutional neural network, which focuses on staged transformations, our densely connected model promotes extensive feature reuse, possibly leading to more robust feature extraction." I hope to see the performance of the proposed algorithm against 1) other DL algorithms (e.g., fully-connected neural networks, BrainNetCNN, Graph CNN (GCNN), temporal CNN, GRU, and LSTM, see https://doi.org/10.1016/j.neuroimage.2019.116276 and https://doi.org/10.1002/hbm.26415), 2) ML algorithms (e.g., SVR with linear, RBF and polynomial kernels, Elastic Net, XGBoost, random forest, CPM), 3) data reduction algorithms (e.g., PCA regression, Partial Least Square). The results of this benchmark will substantiate the claims made by the authors.

      Our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach. We have revised the manuscript text to make this focus clearer and to avoid any misinterpretation of our aims. Specifically, we removed statements in the Discussion that could be read as suggesting that our deep learning approach outperforms prior machine learning methods. While we compared our model with the connectome predictive modeling (CPM) approach and observed better performance with our deep learning framework, we did not conduct a comprehensive benchmark across all available machine learning methods, nor was this the aim of the present study. Accordingly, we have adjusted the text to avoid implying methodological superiority beyond the scope of our analyses. Finally, we have added the following paragraph to the discussion:

      “Our study used a deep neural network architecture that features dense connections and incorporates an attentional mechanism. While our findings demonstrate that a deep learning framework can provide reasonable predictive accuracy, it is important to note that other machine learning approaches (e.g., tree-based models) may offer comparable predictive power, as suggested by prior benchmarking work (29, 30). Our study explicitly compares predictive power across different cognitive states (rest, movie watching, n-back) to identify the states that best capture individual differences across domains. The relative performance of deep learning and other non-linear approaches depends on multiple factors, including sample size, model architecture, feature representation, and domain-specific characteristics of the prediction target. In this context, deep learning was employed as a flexible framework capable of modeling high-dimensional functional connectivity patterns across cognitive states, rather than as a claim of inherent methodological superiority. Thus, our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach.”

      (2) From Figure 6b, it looks like the functional connectivity matrices were converted to different images, and each of the four images (in grey, blue, yellow, and red) was treated as a separate channel. What are these grey, blue, yellow, and red images?

      In our study, the inputs to the deep learning models were subject-specific FC matrices of size 273×273. To augment the data, we created different versions of each FC matrix by reordering specific brain networks within the matrix. To visualize that the inputs were augmented, we used different color codings (grey, blue, yellow, and red) in Figure 6b. These colors were intended solely to represent different augmented versions of the same subject’s FC matrix. They were not treated as separate channels in the model. To avoid any confusion or misinterpretation, we have revised this part of the figure and now use only grey coloring to represent the augmented FC matrices.

      (3) The differences in performance between within vs. outside studies might simply be due to the fact that the models trained from DyNAMiC captured the brain variation due to age, which is also related to cognitive abilities. I was wondering if age is controlled for, would performance be more similar across the studies? The authors should provide the performance of models that are controlled for age.

      We initially conducted partial correlation between FC features and cognitive measures while controlling for age. This is further supported by the fact that the model trained on the age-heterogeneous DyNAMiC sample provided a fairly reasonable prediction in the age-homogeneous COBRA dataset, particularly for working memory (see figure 2d). Moreover, in our post hoc analyses, we additionally controlled for age when examining associations, for example, between GAP and dopamine measures.

      (4) Related to point (3), from the discussion, "Validation outcomes thus affirm that the models, particularly those constructed from rest data, are robust to the particulars of the dataset." The performance dropped around half, so I am not sure if this conclusion is warranted.

      We thank the reviewer for raising this point. The prediction performance indeed dropped for episodic memory when models trained on the DyNAMiC sample were applied to the COBRA sample, whereas performance for working memory remained nearly identical across datasets. Although both EM and WM are sensitive to age, the divergence in cross-dataset performance suggests that factors beyond age alone may contribute to these differences. To address this, we have revised the discussion as follows:

      “Differences between the DyNAMiC and COBRA datasets make cross-dataset prediction a harder problem, as the age ranges of samples significantly vary, and prior studies highlight the importance of individual characteristics like age in predicting behavior from FC (33). In line with this, model performance decreased when predicting EM in the COBRA sample whereas prediction of WM remained largely unchanged. Thus, validation outcomes suggest that the models, particularly those predicting WM, show robustness across datasets, whereas the reduced EM performance highlights potential data-specific influences that limit generalizability.”

      (5) Please report the degree of freedom in all of the statistical analyses. Was the Mann-Whitney U test done on the bootstrapped r? If so, the degree of freedom was arbitrarily set by the number of bootstrapping, and hence the p-value can be higher or lower depending on the number of bootstrapping. This could lead to misleading conclusions.

      We appreciate the reviewer’s comment and agree that applying statistical tests directly to bootstrapped samples can lead to inflated or misleading p-values, as the degrees of freedom are determined by the number of bootstrap iterations rather than the actual number of independent observations.

      In our analysis, the Mann-Whitney U test was applied to 1000 bootstrapped correlation coefficients (r) for each model. While this number is relatively low and was chosen to limit overestimation of significance, we recognize that these bootstrapped samples are not independent, and thus the use of a Mann-Whitney U test can still be problematic. To address this concern, we have revised our statistical analysis. Rather than applying the Mann-Whitney U test to the bootstrapped r distributions, we now compute the difference in correlation coefficients (Δr = r<sub>actual</sub> − r<sub>rest</sub>) for each bootstrap iteration. We then calculate a 95% confidence interval for Δr. If this interval does not include zero, we consider the difference statistically significant. This approach avoids artificially inflating the sample size and adheres more closely to proper statistical inference.

      We have updated the Methods (the following text) and Results sections accordingly and clearly stated the limitations regarding the degrees of freedom for all tests.

      “For the bootstrap-based comparison of model performance (bootstrap resampling with 1000 iterations), no test statistic with an associated degree of freedom is reported. Instead, statistical inference is based on the bootstrap distribution of the difference in correlation coefficients (Δr) and its 95% confidence interval. As bootstrap confidence-interval–based inference does not rely on an analytic sampling distribution, degrees of freedom are not defined for this procedure.” This has now been explicitly stated in the Methods section to avoid ambiguity.

      In the result section, we have reported with corresponding CI.

      (6) For predictive performance, the correlation was reported in the table, while R<sup>2</sup> is reported in the text. This is confusing. Also, could you clarify if the R<sup>2</sup> is calculated using the sum square definition, not Pearson r squared? If Pearson r squared was used, then R<sup>2</sup> of a negative Pearson r would be positive, which is misleading (see 10.1001/jamapsychiatry.2019.3671). Also, other performance indices apart from Pearson r and R² should be reported (e.g., MSE and MAE, again see 10.1001/jamapsychiatry.2019.3671). This will allow a better understanding of the models' performance.

      We thank the reviewer for this helpful comment. We acknowledge the inconsistency in reporting predictive performance metrics and have revised the manuscript for clarity. In the text, we have reported the r value, whereas in the table, we have reported r<sup>2</sup> using the sum-of-squared definition. Specifically, we now consistently report Pearson correlation (r), mean squared error (MSE), and mean absolute error (MAE) across both the text and Tables 1 and 2.

      Regarding r<sup>2</sup>, we confirm that it was calculated using the sum-of-squares definition (i.e.,

      rather than as the square of the Pearson correlation coefficient. This ensures that negative correlations do not result in misleading positive R<sup>2</sup> values, as pointed out by the reviewer and discussed in Poldrack et al. (2020). All performance metrics (r, r<sup>2</sup>, MSE, and MAE) are now reported in Tables 1 and 2 to allow a more comprehensive and interpretable comparison of model performance.

      We have included a description of the method under section 4.9. Statistical significance analysis.

      (7) Could you clarify how data are standardized across training, validation, and tests (including Z-standardization for the cognitive tests)? This is to prevent data leakage.

      Thanks for the comments. We did standardization the cognitive test from both training and test, separately.

      We have added the following paragraph to the method section:

      “A composite score of performances across the three tests was calculated and used as the measure of the cognitive domain in question (i.e., episodic memory, working memory). For each of the three tests, scores were summarized across the total number of trials. The three resulting sum scores were z-standardized and averaged to form one composite score for each domain. The standardization has been carried out independently for the training (DyNAMiC) and test (COBRA) samples.”

      (8) There is really no ground truth to confirm that Grad-CAM provides actual feature importance used by the models. Perhaps the authors should compare that with Haufe transformation, which is commonly used in the predictive model for cognition (e.g., https://doi.org/10.1016/j.neuroimage.2021.118648 and https://doi.org/10.1016/j.neuroimage.2023.120115).

      We appreciate the reviewer’s comment and the suggested references. The Haufe transformation is primarily applied in traditional machine learning models, particularly in cognitive neuroscience, to interpret linear predictive models by mapping classifier weights back to the input space. However, its direct applicability to deep learning models, especially convolutional neural networks, remains an open research area with no widely established methodologies. Furthermore, the Haufe transformation does not provide feature importance in the same manner as Grad-CAM. Grad-CAM highlights spatial regions within an image that contribute to a model’s decision, making it particularly useful for interpreting convolutional networks in vision tasks. In contrast, the Haufe method offers a weight transformation that is more suited for understanding linear models and may not be as intuitive for feature attribution in complex hierarchical representations such as those learned by deep neural networks.

      While we acknowledge that Grad-CAM, like other interpretability methods, does not provide absolute ground truth validation for feature importance, it remains one of the most widely used and validated techniques for deep learning interpretability, particularly in medical imaging applications. Given its integration with frameworks such as Keras and TensorFlow and its ability to provide spatial attributions aligned with domain knowledge, we believe it is a suitable choice for our study. Future work may explore additional interpretability techniques, including adaptations of the Haufe transformation if applicable to deep learning architectures.

      We have added more details on Grad-CAM implementations in the Method.

      (9) Related to Grad-CAM, "These edges, indicated by a salience intensity of {greater than or equal to}.5, exert a significant influence on the model (Figure 1f)." What does 'significant' in this context mean? And how did the authors come up with the .5 threshold? Is it based on permutation or bootstrapping tests?

      We appreciate the reviewer’s comment and the opportunity to clarify our approach. In this context, the term "significant" refers to the regions' relative contribution to the model’s decision, as shown by the Grad-CAM saliency map. However, to avoid implying statistical testing, we will revise the term to "highly contributing."

      Regarding the 0.5 threshold, this value was selected empirically based on the normalized Grad-CAM activation values, where saliency scores range between 0 and 1. A threshold of 0.5 was used as a heuristic to highlight regions with relatively strong activation. However, this was not determined through statistical methods such as permutation or bootstrapping tests. We recognize the importance of rigorous threshold selection and will clarify this in the text. Future work could incorporate statistical methods to define thresholds more objectively.

      We have included the following text in the Method section:

      ”Grad-CAM saliency maps were interpreted qualitatively, with a heuristic threshold (≥ 0.5) applied to highlight regions with relatively higher contribution to the model’s predictions. These values do not reflect statistical significance and should therefore be interpreted descriptively.”

      (10) Still related to the saliency map, I believe the upper and lower triangles of the functional connectivity matrix are the same. If so, why are there some differences in saliency? While the difference is not prominent, this might affect the accuracy of Grad-CAM.

      Minor differences in the saliency maps between the upper and lower triangles of the FC matrix can arise due to several factors. For instance, Grad-CAM generates saliency maps at the resolution of the convolutional feature maps, which are then upsampled to match the input matrix dimensions. We initially used the default bilinear interpolation, which may have introduced slight asymmetries or blurring, resulting in interpolation artifacts. In response, we have reprocessed the saliency maps using spline interpolation in MATLAB. The updated saliency figures have been included in the revised version of the manuscript.

      (11) Why did the authors only report the cross-study for EM on rest, and for WM on n-back? This is a bit unexpected since COBRA has both rest and n-back. If there is no good justification, please report both.

      We focused on reporting cross-study results for EM using rest because rest was the winning condition for predicting EM in the DyNAMiC sample. Importantly, n-back did not significantly predict EM in DyNAMiC, and rest did not significantly predict WM. For this reason, we highlighted only the conditions that showed meaningful predictive power in the original analyses.

      (12) Are codes, trained models, and data available? To ensure transparency and reproducibility, I hope to see the code from preprocessing to modeling and statistical analyses.

      The analysis code is openly available on our GitHub page https://github.com/MorEsm/AI-based-Prediction-of-Cognitive-Function. Due to ethical considerations and GDPR restrictions in the European Union, we are not permitted to publicly share the raw data. However, we can provide detailed information about preprocessing steps and analysis pipelines to facilitate reproducibility.

      (13 &14) The authors did not appropriately control for regression-toward-the-mean and the influence of the working memory itself when calculating the brain cognition gap. This is commonly done to brain age (see https://doi.org/10.7554/eLife.87297.4https://doi.org/10.1002/hbm.25533https://doi.org/10.1016/j.nicl.2020.102229https://doi.org/10.3389/fnagi.2018.00317). Otherwise, the brain cognition gap still depends on the cognition/working memory score itself. Based on Tetereva et al., "If, for instance, Brain Age was based on prediction models with poor performance and made a prediction that everyone was 50 years old, individual differences in Brain Age Gap would then depend solely on chronological age (i.e., 50 minus chronological age)." Because of this, Tetereva and colleagues found that the 'uncorrected' brain age gap that predicted chronological age the worst became the best index to predict fluid cognitive abilities. This shows the pitfall of the 'uncorrected' brain age gap. You can apply the same logic to the brain cognition gap.

      (14) Additionally, another way to show the unique contribution of brain cognition, over and above cognition per se, is to add both brain cognition and cognition together to predict physical activity, education, and cardiovascular risk.

      We thank the Reviewer for raising this important point. In response to their request and also the request from Rev. 1, we first examined the relationship between the Brain-Cognitive Gap (BCG) and the cognitive measure itself. Surprisingly, we did not find any significant relationship in either the DyNAMiC sample (r =0.01, p =0.939) or the COBRA sample (r =0.01, p =0.894) (see Author response image 1).

      We then conducted additional analyses, splitting the sample into high and low EM performers, and compared their levels of physical activity and Framingham cardiovascular risk scores. We found that no significant difference in physical activity (DyNAMiC: p =0.56, CI: -14.99 – 8.13; COBRA: p =0.29, CI: -3.54 – 1.05) or Framingham CVD risk score (DyNAMiC: p =0.11, CI: -1.08 – 10.72; COBRA: p =0.41, CI: -1.86 – 4.58) between high and low EM perfprmers. Given the significant difference in physical activity and Framingham CVD risk score between positive and negative BCG groups, our results support that BCP provides unique information, beyond cognitive measure, regarding factors that contribute to cognitive resilience. These results have been added to Section 2.4, and Figure 3 has been updated.

      (15) Related to the brain age gap, the brain cognition gap is actually just another way to quantify how generalizable models are to another sample, similar to MAE or MSE. If the models built from DyNAMiC don't fit well with samples from COBRA, you will get a higher (i.e., wider) brain cognition gap, which means a poor fit. The authors should discuss this interpretation - should your biomarker's performance be due to a fit of the model?

      We appreciate this insightful comment. We agree that BCG can be interpreted not only as a marker of individual differences and resilience factors but also as a measure of model fit, analogous to error metrics, such as MAE or MSE. A higher gap may, in part, reflect poorer generalizability of models across samples. We have now revised the Discussion to explicitly acknowledge this alternative interpretation and to emphasize that BCG should be viewed both as a candidate biomarker and as a reflection of model performance.

      We added the following paragraph in the discussion:

      “An important caveat is that BCG can also be conceptualized as an error metric, similar to mean absolute error or mean square error, reflecting the extent to which models trained in one sample generalize to another. From this perspective, a larger gap may not only indicate individual differences related to resilience factors and dopaminergic function, but also reduced model fit or generalizability across datasets. Thus, BCG likely reflects a combination of meaningful biological variability and methodological variance.”

      (16) It is unclear why the authors binarized the brain cognition gap when predicting physical activity, education, and cardiovascular risk, and not doing so with the striatal D1DR. It is rarely a good idea to binarize a continuous variable (see 10.1136/bmj.332.7549.1080). In this case, people who had a bigger negative brain cognition gap were treated equally to people who had a smaller negative brain cognition gap. I also do not think it is necessary to separately analyze positive and negative gaps. Perhaps the authors should correlate the corrected brain cognition gap with physical activity, education, and cardiovascular risk and provide scatter plots and effect sizes.

      Following the reveiwer suggestion, we directly correlated BCG with physical activity and cardiovascular risk. Our results confirmed our initial analysis that individuals with a negative gap exhibited lower physical activity and higher Framingham CVD risk across both COBRA and DyNAMiC datasets. We have reported these results on page 10.

      Author response image 5.

      (17) Given that the motivation is to move away from brain age, the authors should benchmark the corrected brain cognition gap against the corrected brain age gap, as well as against the performance when directly predicting physical activity, education, and cardiovascular risk from the functional connectivity metrics.

      Author response image 6.

      We agree that benchmarking BCG against BAG in predicting lifestyle and vascular risk factors would be valuable. We have calculated adjusted BAG and related it to lifestyle and vascular risk factors. Interestingly, we did not find any significant association, suggesting that BCG might be more sensitive to cognitive resilience. However, this investigation was beyond the scope of the present study. Our aim was not to compare BCG with BAG, but rather to examine whether BCG provides information beyond cognition itself. We also note that introducing BAG would open a separate line of investigation, namely, which cognitive state (rest, movie-watching, n-back) best estimates biological age. While this is an interesting question in its own right, addressing it here would considerably broaden the scope and complexity of an already dense manuscript. To prevent misunderstanding, we have clarified this point in the Discussion and added a caveat noting that future work should explicitly benchmark these approaches. That said, if the Reviewer and/or the Editor incline to add these additional findings into the manuscript, we are open to doing so in a revision.

      We have added the following sentence to the Discussion.

      “While our focus was to investigate whether the brain–cognition gap provides information about factors contributing to cognitive resilience, we acknowledge that benchmarking BCG against the brain-age gap in predicting lifestyle and vascular risk factors would be valuable. However, addressing this question lies beyond the scope of the present study, and future work should systematically compare these approaches.”

      (18) Why was only the working memory score used to create brain cognition, and not episodic memory as well? Including both could provide a more comprehensive measure.

      We initially attempted to predict both episodic memory (EM) and working memory (WM). However, EM prediction was only reliable within and across samples for the resting state, whereas WM prediction generalized most strongly from the movie-watching condition. Because COBRA does not include a movie-watching paradigm, we could not evaluate WM prediction across datasets. For this reason, we focused on EM when examining the brain–cognition gap.

      (19) The PET mediation analysis seemed to come out of the blue. Is there existing literature showing the relationship between striatal D1DR and cognition? If so, did the authors find a similar relationship in the current data? I also suggest rewriting this section to strengthen the justification for the PET mediation analysis.

      We have previously conducted studies in which DA found to be associated with memory (Johansson et al., 2023, Nyberg et al., 2016).

      The third aim of our study was to examine whether DA integrity is implicated in brain–cognition gaps (BCG), which we propose as a marker of cognitive resilience. In line with this aim, we found that lower DA receptor availability was associated with larger BCGs (Figure 4). We then asked whether this relationship is mediated by functional signal variability, such that lower DA is linked to reduced signal-to-noise ratio (i.e., greater entropy in functional connectivity), which in turn contributes to less reliable prediction of cognition and, consequently, larger BCGs. Our mediation analysis supports this pathway (see also our reply to Reviewer 1, Comment 6).

      Thus, our mediation was not designed to test whether DA predicts episodic memory performance directly, nor whether BCG mediates such a relationship. Instead, we specifically investigated whether the effect of DA on BCG operates through functional variability. We agree that future work could extend our approach by directly examining whether BCG mediates the link between DA and cognitive outcomes. However, in the present study, our primary focus was on testing the mechanistic pathway of DA → entropy → BCG.

      Minor recommendations:

      (1) Task-based connections are not truly task-based, as they are around 70-80% related to the resting state, capturing non-task-specific functional connectivity. Task-based connections should refer to techniques that derive task-related connectivity, such as psychophysiological interaction and beta-series correlation. Perhaps use terms like "functional connectivity during tasks."

      Thank you. This has been corrected throughout the manuscript.

      (2) Are there really two studies? The same MRI was used with the same configurations, and participants were from the same city. The only difference is the age range. It may be more appropriate to refer to this as "across age groups" rather than "cross-datasets."

      Thank you for this comment. While the two samples share some similarities, there are also several marked differences beyond age range. For example, Movie-watching was administered in DyNAMiC but not collected in COBRA. The resting-state fMRI sequence was 12 minutes in DyNAMiC but only 6 minutes in COBRA. Moreover, DyNAMiC included dopamine D1-receptor PET, whereas COBRA assessed dopamine D2-receptor availability. Even the questionnaires used to measure physical activity differed between the two studies. Given these methodological and measurement differences, we believe that referring to them as “cross-datasets” rather than “across age groups” more accurately captures the distinction.

      (3) What kind of movie is "Cockpit"? Can you explain? Different movies may elicit different patterns of connectivity.

      We apologize for not providing information about the movie, which has been presented in our recent work (Johansson et al., 2023).

      The participants’ reactions to the content of the movie were not monitored, but the clips were selected to be as neutral in their content as possible. The content of the movie: Following his termination as a pilot and the end of his marriage, Valle embarks on a quest to secure new employment. Faced with desperation in the job market, he resorts to disguising himself as a woman with the intention of obtaining a position at a company specially seeking a female pilot.

      This information is added to the method section.

      “During the fMRI session, participants viewed a 12-minute segment from the Swedish comedy film Cockpit (2012). We did not monitor participants’ responses to the movie, and the chosen clips were selected to be relatively neutral in emotional content. The storyline follows Valle, a recently fired pilot whose marriage has ended, as he struggles to find new employment. In a desperate attempt to secure a job at an airline specifically recruiting a female pilot, he presents himself as a woman.”

      (4) There is a typo in the equation numbering (i.e., two equations are designated as #1).

      We have now corrected the typo.

      (5) From the discussion: "Importantly, this prediction generalizes across conditions." This is not surprising given the similarity between conditions, with around 70-80% variance.

      We agree with the reviewer that the high similarity of FC across states likely increases the chance of cross-condition generalizability. However, this generalization is not guaranteed for all models. For example, the model trained on FC during movie-watching successfully predicted episodic memory during rest, but it did not generalize to episodic memory during the n-back condition, although movie-watching and n-back FC patterns are themselves highly correlated. Thus, the observed generalization is meaningful in demonstrating that not all models transfer equally well across states.

      That said, we have added the following sentence to the Discussion:

      “Importantly, this prediction generalizes across conditions and datasets, suggesting that features derived from resting state FC serve as a relatively stable marker of individual differences in EM, though with reduced strength in COBRA. While such generalization is partly facilitated by the similarity of functional connectivity across states, it is not a trivial outcome. For instance, the model trained on movie-watching data generalized to EM prediction during rest but failed to do so for the n-back condition, even though movie-watching and n-back connectivity patterns are themselves highly correlated. This indicates that successful generalization depends not only on shared variance across states but also on the cognitive processes most relevant to the target behavior.”

      (6) It might be helpful to include some figures for the cognitive tasks used. The description is a bit hard to follow without visual aids.

      Thanks for the comment. We have had a figure describing this in the initial paper about DyNAMiC (Nordin et al., 2022). We have added the Supplementary Figure (Fig S3) in the manuscript.

      Fig S3. Overview of the cognitive tests included in the DyNAMiC study. Adopted from Nordin et al. with permission.

      (7) It may not be appropriate to use the term "cross-validation" here, as one dataset was used for testing and the other for training, but not vice versa (so no "cross" per se).

      We thank the reviewer for pointing this out. We agree that the term “cross-validation” is not precise in this context, since we trained the model in one dataset and tested it in another without performing the reverse. We have revised the manuscript to use the term “external validation” instead of “cross-validation” to more accurately describe our cross-dataset approach.

      (8) I don't have access to the supplementary materials or code/data, so all of the comments here are based on the main text.

      We have added the supplementary materials and inserted the GitHub link to the code.<br />

      Reviewer #3 (Recommendations for the authors):

      I suggest benchmarking against other simpler algorithms and controlling for memory in the brain cognition gap analyses.

      The authors might also want to simplify some aspects of the paper. There is a lot going on, which leaves less space to go into enough details for some analyses to warrant claims in the discussion. For example, the authors only compare the deep net to CPM and kernel ridge based on the literature. Direct comparisons would be needed.

      Thanks for the comment. We have made an attempt to address the concerns outlined in the public recommendation. Our study explicitly compares predictive power across different cognitive states (rest, movie watching, n-back), with the aim of identifying the states that best capture individual differences across domains. Thus, our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach. We have revised the manuscript text to make this focus clearer and to avoid any misinterpretation of our aims. Specifically, we removed statements in the Discussion that could be read as suggesting that our deep learning approach outperforms prior machine learning methods. While we compared our model with the connectome predictive modeling (CPM) approach and observed better performance with our deep learning framework, we did not conduct a comprehensive benchmark across all available machine learning methods, nor was this the aim of the present study. Accordingly, we have adjusted the text to avoid implying methodological superiority beyond the scope of our analyses. Furthermore, we have controlled for memory as suggested by the reviewer and outlined in response to reviewer 1.

    1. eLife Assessment

      This important study used whole-genome data to investigate Beefalo ancestry for the first time, providing insight into the genetics of Beefalo cattle and challenging the long-held claim of 37.5% bison ancestry reported by the American Beefalo Association. Despite some limitations regarding sequencing depth and sampling, the expert use of a comprehensive set of population-genomic methods allowed the authors to demonstrate convincingly that Beefalo and bison hybrid ancestry profiles are consistent with repeated backcrossing to either parental species. The work will be of significant interest to evolutionary biologists, population geneticists, animal breeders, and those involved in the conservation genetics of bovine species.

    2. Reviewer #1 (Public review):

      Summary:

      This study used whole genome data to investigate Beefalo ancestry for the first time, filling the gap in the field of Beefalo ancestry. The authors used preserved semen samples to generate genomic data on 47 registered Beefalo and 3 bison hybrids, further questioning the ABA's stated goal of ⅜ bison ancestry. In addition, the authors also show that ancestry profiles of Beefalo and bison hybrid genomes are consistent with repeated backcrossing to either parental species, demonstrate the value of genomic information in examining gene flow between species in the genus Bison. Overall, these data thus demonstrate the utility of genomic information in validating specific breeding claims for a more complete understanding of gene flow and genetic variation among bovine species. This is an interesting study, but there are still some major weaknesses that exist.

      Strengths:

      Numerous genetic analysis methods such as PCA, ADMIXTURE, F4 ratios, and local ancestry inference techniques revealed that no single Beefalo set meets the ancestry requirements set by the American Beefalo Association (ABA) and some beefalo had detectable indicine cattle ancestry.

      Comments on revised version:

      The authors have made further revisions in the revised manuscript, and these revisions have undoubtedly helped improve the article. No further comments.

    3. Reviewer #2 (Public review):

      Summary:

      Shapiro et al. set out to verify the American Beefalo Association's claim that Beefalo cattle possess 37.5% bison ancestry. They employ a comprehensive range of well-established population genomics methods to estimate ancestry in these hybrid populations, including PCA, ADMIXTURE, D and F statistics, and local ancestry inference. Their findings conclusively demonstrate that most Beefalo lack the claimed bison ancestry, with only 8 out of 47 samples showing any detectable bison ancestry, ranging from 2-18%.

      Strengths:

      The primary strength of this analysis lies in the comprehensive dataset available to the authors, which includes important foundational Beefalo individuals and various reference populations. The rigorous and multi-faceted methodological approach employs several well-established techniques in population genomics for detecting and measuring admixture. Each method used has a firm basis in the field, providing consistent and robust results. The authors' approach of using PCA to initially assess the data within a global context, followed by more specific analyses using ADMIXTURE and D-statistics, provides a clear and logical progression of evidence. The presentation of these results in figures is particularly effective, clearly illustrating the key findings of the study. Additionally, the examination of both autosomal and sex chromosome ancestry offers a more complete understanding of Beefalo genetic composition and the mechanics of bison-cattle hybridisation.

      Weaknesses:

      One limitation of this analysis is the relatively low coverage (~2x) of many Beefalo samples. However, the authors have taken steps to mitigate biases that may arise from this, and their downsampling experiment demonstrates that this level of coverage is appropriate for summarising species-level ancestry across Bos. Another potential weakness is the limited sampling of contemporary Beefalo populations, as the study focuses primarily on historical samples. The authors have justified this choice on the grounds that contemporary Beefalo breeding involves no further bison input, so founder-era individuals are the most informative samples for addressing the study's central question.

      Appraisal:

      The authors have clearly achieved their primary aim using a rigorous and comprehensive methodology. Their extensive dataset and multi-faceted analytical approach provide strong support for their conclusions. The study not only addresses its main research question but also reveals unexpected insights into Beefalo genetics, particularly the presence of zebu ancestry, predominantly from Brahman cattle.

      Discussion:

      This study is valuable for several reasons beyond its primary findings. First, it definitively addresses and refutes the claim of 37.5% bison ancestry in Beefalo, providing crucial information for those studying these interspecies hybrids and the viability of their offspring. Second, it reveals the unexpected presence of zebu ancestry, predominantly from Brahman cattle, in many Beefalo, raising intriguing questions about the breed's development and the potential role of zebu cattle in achieving desired traits. This finding suggests that the distinctive appearance of Beefalo may be due in part to zebu admixture rather than bison ancestry. Third, the study highlights the significant barriers to admixture between bison and cattle, both in controlled breeding programs and potentially in wild populations. This has important implications for conservation genetics and our understanding of gene flow between these species. Lastly, the study demonstrates the power of genomic analysis in verifying breed claims and understanding the complex history of domestic animal breeds. These findings open new avenues for research in bovine genomics, breed development, and the dynamics of interspecies hybridisation.

      Comments on revised version:

      Thanks for the responses, which address my comments in full. I have no further concerns.

    4. Reviewer #3 (Public review):

      Summary:

      The American beefalo cattle breed was developed as a mixture of 5/8 domestic cattle and 3/8 (or 37.5%) bison ancestry. The authors sequenced 50 genomes from bison and hybrids (historical and present-day). They found that most animals did not carry any detectable bison ancestry, with only a few between 2-18%, while other beefalo had taurine/zebu cattle ancestry, which may explain morphological traits. Breeding design was likely each time to a parental instead of to other admixtures.

      The authors utilize whole genome sequence data to explore the ancestry of beefalo with respect to expected and possible contributions from cattle lineages. Using molecular and analytical methods central to questions exploring genomic ancestry and identity, the authors very nicely show evidence that calls into question ability of ancestry to be deduced from breed club documentation without considering reproductive challenges that are known in hybridization between cattle lineages.

      Comments on revised version:

      The authors have addressed all my comments to help improve presentation of specific details, results, and readability. Thank you!

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study used whole genome data to investigate Beefalo ancestry for the first time, filling the gap in the field of Beefalo ancestry. The authors used preserved semen samples to generate genomic data on 47 registered Beefalo and 3 bison hybrids, further questioning the ABA's stated goal of ⅜ bison ancestry. In addition, the authors also show that ancestry profiles of Beefalo and bison hybrid genomes are consistent with repeated backcrossing to either parental species, demonstrating the value of genomic information in examining gene flow between species in the genus Bison. This is an interesting study that still has some major weaknesses that exist, but overall, the work demonstrates the utility of genomic information in validating specific breeding claims for a more complete understanding of gene flow and genetic variation among bovine species.

      We thank the reviewer for their thoughtful assessment of our work.

      Strengths:

      Numerous genetic analysis methods such as PCA, ADMIXTURE, F4 ratios, and local ancestry inference techniques revealed that no single Beefalo set meets the ancestry requirements set by the American Beefalo Association (ABA) and some beefalo had detectable indicine cattle ancestry.

      Weaknesses:

      While this study contributes to our knowledge of Beefalo ancestry, there are some key issues that need to be addressed in terms of analysing the specific results as well as writing the article.

      We have followed the reviewer’s suggestions for improving our study in detail (specified below), and appreciate their close reading of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      Shapiro et al. set out to verify the American Beefalo Association's claim that Beefalo cattle possess 37.5% bison ancestry. They employ a comprehensive range of well-established population genomics methods to estimate ancestry in these hybrid populations, including PCA, ADMIXTURE, D and F statistics, and local ancestry inference. Their findings conclusively demonstrate that most Beefalo lack the claimed bison ancestry, with only 8 out of 47 samples showing any detectable bison ancestry, ranging from 2 - 18%.

      We thank the reviewer for their thoughtful assessment of our work.

      Strengths:

      The primary strength of this analysis lies in the comprehensive dataset available to the authors, which includes important foundational Beefalo individuals and various reference populations. The rigorous and multi-faceted methodological approach employs several well-established techniques in population genomics for detecting and measuring admixture. Each method used has a firm basis in the field, providing consistent and robust results. The authors' approach of using PCA to initially assess the data within a global context, followed by more specific analyses using ADMIXTURE and D-statistics, provides a clear and logical progression of evidence. The presentation of these results in figures is particularly effective, clearly illustrating the key findings of the study. Additionally, the examination of both autosomal and sex chromosome ancestry offers a more complete understanding of Beefalo genetic composition and the mechanics of bison-cattle hybridisation.

      Weaknesses:

      One limitation of this analysis is the relatively low coverage (~2x) of many Beefalo samples. However, the authors have taken steps to mitigate biases that may arise from this. Another weakness is the limited sampling of contemporary Beefalo populations, as the study focuses primarily on historical samples. This may limit our understanding of how Beefalo genetics may have changed over time.

      The reviewer is correct that the low coverage obtained for many Beefalo is one potential limitation, although we believe that the downsampling experiment we performed (Fig. S4) shows that this level of coverage is appropriate for summarizing species-level ancestry across Bos, as the reviewer notes.

      Sampling contemporary Beefalo individuals would be valuable, though as the focus of our study was to understand the origins of bison ancestry in Beefalo, we prioritized sampling individuals which played an important role in establishing the breed. We also note that contemporary Beefalo breeding involves crossing between Beefalo individuals or backcrossing to cattle, with no additional bison ancestry input since the formation of the Beefalo. As such, sampling individuals that existed close to the breed’s founding should provide the most insight into bison ancestry in Beefalo.

      Appraisal:

      The authors have clearly achieved their primary aim using a rigorous and comprehensive methodology. Their extensive dataset and multi-faceted analytical approach provide strong support for their conclusions. The study not only addresses its main research question but also reveals unexpected insights into Beefalo genetics, particularly the presence of zebu ancestry.

      Discussion:

      This study is valuable for several reasons beyond its primary findings. First, it definitively addresses and refutes the claim of 37.5% bison ancestry in Beefalo, providing crucial information for those studying these interspecies hybrids and the viability of their offspring. Second, it reveals the unexpected presence of zebu ancestry in many Beefalo, raising intriguing questions about the breed's development and the potential role of zebu cattle in achieving desired traits. This finding suggests that the distinctive appearance of Beefalo may be due in part to zebu admixture rather than bison ancestry. Third, the study highlights the significant barriers to admixture between bison and cattle, both in controlled breeding programs and potentially in wild populations. This has important implications for conservation genetics and our understanding of gene flow between these species. Lastly, the study demonstrates the power of genomic analysis in verifying breed claims and understanding the complex history of domestic animal breeds. These findings open new avenues for research in bovine genomics, breed development, and the dynamics of interspecies hybridisation.

      Reviewer #3 (Public review):

      Summary:

      I really like this topic and study. But I think much can be more focused and tightened up. All the components are here - just some more refining to really make the storyline clear, the journey of discovery, and the impact of such knowledge.

      We thank the reviewer for their thoughtful assessment of our work.

      Strengths:

      The authors dive directly into the question of genomic ancestry as compared to the breed club's reported ancestry with heavy, quantitative data and critical analytical methods. The questioning line is direct and does not meander. The reader learns about the challenges of breeding associations, and values of understood ancestry, and presents a clear need of re-evaluating the breed standards and expectations of beefalo (if ancestry is indeed the primary goal instead of a phenotype-driven breed mission).

      Weaknesses:

      Much of the quantitative results are only referred to in the main text with qualitative language. Please incorporate more written quantitative results to highlight evidence that underlines the study narrative because it is quite an interesting study!

      The reviewer highlights an important point, and we agree that the qualitative language used to describe the results was generally lacking. We have now described the results quantitatively throughout the manuscript where possible.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) This study is not the first to question claims surrounding bison ancestry in the breed and is the sample size too small to be representative of the entire genetic structure of Beefalo?

      The reviewer correctly points out that this study is not the first to address uncertainty in the amount of bison ancestry present across beefalo. All earlier studies, to our knowledge, have been highlighted in the introduction and discussion (Lenoir and Lichtenberger, 1978 and Stormont et al, 1986). However, these studies examined a narrow range of Beefalo sources and used older methods (karyotyping and blood typing), such that comprehensive statements about the proportion of bison ancestry in Beefalo could not be made.

      We also agree that an appropriate sampling scheme is crucial for making definitive statements about Beefalo ancestry across the breed. As Beefalo breeding typically involves breeding select “full-blood” individuals with cattle, the ancestry across contemporary Beefalo is likely complex, with the cattle component coming from a wide range of breeds. Therefore, our sampling emphasized “full-blood” representatives, especially those that were involved in the founding of the breed and from which later Beefalo descend. This involved an exhaustive survey of the Beefalo individuals contained within the USDA’s National Animal Germplasm Program. Although we did not extensively evaluate current Beefalo diversity, we believe this approach is most suited for characterizing bison ancestry within Beefalo, as bison ancestry is maintained primarily through the continued use of genetic material from these “full-blood” individuals rather than repeated hybridization between bison and cattle.

      (2) Although genomic information is important for breeding research, this requires quality of data. The coverage of the data used in this study was mainly ~2X, and although multiple methods of analysis gave similar results, the ability to identify rare variants (e.g. insertions or deletions of long segments of the genome) may be limited at low coverage, affecting the confidence of the results.

      This is an important consideration, and we agree with the reviewer that the sequencing depth obtained for most individuals in our study precludes accurate genotype calling. Therefore, we did not attempt to perform traditional genotype calling. Rather, we used a pseudohaploid calling approach in which a random base was selected to represent the genotype at each position for each individual, using a pre-ascertained set of variants discovered in gaur, a closely related outgroup to bison and cattle. This pseudohaploid approach is common in other situations where coverage is low, for example in analyzing ancient DNA.

      Furthermore, our ancestry analyses focused on biallelic SNPs which were discovered in gaur and we did not attempt to call structural variants, given the limitations in coverage. As this outgroup ascertainment approach seeks to target SNPs which were polymorphic in the ancestor of both bison and cattle, which should yield unbiased results in population genetic analyses, we were less interested in discovering rare variation within the species and populations we examined here.

      Finally, we performed downsampling experiments comparing low coverage read data to genotypes called from high coverage data, and obtained consistent results between low and high coverage analyses using read-level data and called genotypes (Fig. S7).

      (3) Missing from the conclusions is the very important presentation of the results of genomic calling, the basics of what these data look like, coverage histograms, number of SNPs, categorization, annotations, and so on. These are necessary prerequisites for subsequent population analysis.

      The reference to “5.29M” on page 14 has been replaced with the exact number of SNPs used in analyses (5,291,534). The average sequencing depth for each sample is also included in Table S1.

      (4) The manuscript mentions "most" in a number of places, but can the authors give an accurate number based on the current data? "Most" is not a rigorous description. Based on the simulations of genomic data, how many Beefalo cattle were not detected as hybridized? This may be related to both sample size and where the authors sampled.

      We thank the reviewer for this important suggestion. We have now replaced vague summaries of results with precise numbers. However, we are unsure what “simulations” means in this context, as all results were obtained by analyzing empirical data from Beefalo, bison, cattle, and other bovines, rather than simulations.

      (5) The information in the third and fourth paragraphs of the Introduction is not sufficiently coherent and could be further consolidated into a more logical presentation.

      We have now condensed these paragraphs and edited them for clarity.

      (6) "For some analyses we also incorporated published genomes from outgroups". The description here is unclear as to what criteria were used to select these data, and it is possible that the choice of outgroups could lead to different conclusions from the analyses. In addition, ancient DNA data from cattle may be useful for this study and the authors are encouraged to consider it.

      Outgroup choice can certainly have a large impact on population genetic analyses. For the species examined in our study, we considered other Bos species, including yak, gaur, and banteng, as suitable outgroups, along with water buffalo, which is the closest outgroup outside of Bos. We have added comparisons of D-statistics using yak as an outgroup as a supplementary figure (Fig. S4), in addition to those using water buffalo as the outgroup which were presented in Figure 2.

      As we were examining species-level ancestry, and given the high level of divergence between bison and cattle, relative to that between published ancient and modern cattle genomes, we believed that it was most appropriate to use high quality modern cattle data, rather than poorer quality ancient cattle genomes, for analyses. Additionally, as any hybridization which took place between bison and cattle in the formation of Beefalo would have occurred within the past ~50 years, modern cattle are likely to be the most appropriate proxy for the cattle ancestry in Beefalo, especially given the lack of published historical North American cattle genomes.

      (7) The coordinates of the PCA plot need to be further supported by providing values.

      We have now updated axis labels for the PCA in Fig. 1A to include the proportion of variance explained for the first two components.

      (8) In Figure 1, Beefalo has one individual, NAGP9109, which belongs exclusively to the indicine group. For this individual, wouldn't it be nicer to label it separately in the PCA and ADMIXTURE plots, like Joe's Pride (JP), to make the presentation of the results clearer?

      This individual was one which was determined to be mislabeled as Beefalo within the NAGP and is actually a Brahman cattle. Therefore, we have relabeled it as zebu, rather than Beefalo, throughout the figures.

      (9) As the sex chromosome data do not fully support the authors' claims, some caution may be needed in describing the results.

      We interpret the sex chromosomal results as being fully consistent with patterns seen in the autosomes. However, they do shed some light on the dynamics of bison-cattle hybridization, and suggest male-mediated gene flow in which bison ancestry in Beefalo was introduced primarily through bison bulls.

      (10) Would it be appropriate to analyse the results at K = 3 only? The admixture analysis of all bison, cattle, bison hybrids, and buffalo individuals at different K values should further refine the results.

      We now also show ADMIXTURE results at K=2 and K=4 (Fig. S2) and present the cross-validation results from ADMIXTURE (Fig. S3).

      (11) The conclusions of this article about bison ancestry in Beefalo individuals are completely inconsistent with the American Beefalo Association, and should a description of possible reasons for this discrepancy be added to the discussion?

      Our analyses make it clear that there was much less hybridization between bison and cattle leading to the formation of the Beefalo that was previously believed. As the genetic data does not provide insight into exactly why this might be the case, we can only speculate on the precise reasons bison-cattle hybridization did not take place, which we have avoided here.

      Reviewer #2 (Recommendations for the authors):

      The manuscript is well written, the figures are easily understandable, and the claims made are justified by the results obtained.

      It is need to clarify cattle breeding terminology, particularly concerning breeds like the Brahman. While often described as zebu-taurine hybrids, Brahman cattle typically show over 90% zebu ancestry when analysed using ADMIXTURE against panels including European Bos taurus, African Bos taurus, and Bos indicus animals. This context would help explain why "NAGP9109" clusters with the Zebu group.

      We thank the reviewer for this useful context, and agree that most Brahman cattle have a high proportion of zebu ancestry. In fact, the zebu group we included primarily consists of Brahman individuals, which we have now clarified in the text, which now reads:

      “The reported pedigree in the NAGP for this animal lists its composition as 1/2 Brahman, 1/4 Charolais, 1/8 bison, 1/16th Hereford, and 1/16th Shorthorn, but the American Brahman Breeders Association records this animal (#309519) as purebred Brahman, which is a zebu breed (5 of the other 6 zebu individuals analyzed here are Brahman cattle).”

      I suggest three other improvements:

      (1) Standardise terminology: The manuscript alternates between "zebu" and "indicine" when referring to these cattle. While both terms are correctly defined in the introduction as "indicine (zebu; Bos indicus)" using one term consistently throughout would improve readability. I prefer "zebu" but leave this choice to the authors.

      We agree that this mixed terminology was confusing and have replaced all instances of “indicine” with “zebu.”

      (2) Add PCA metrics, including the percentage of variance explained by each principal component would demonstrate the genetic distinctiveness between bison and cattle, and between Taurus and zebu cattle. This would also support the selection of K=3 for the ADMIXTURE analysis.

      The axis labels for the PCA have been updated to include the proportion of variance explained for each component. We now also show ADMIXTURE results at K=2 and K=4 (Fig. S2) and present the cross-validation results from ADMIXTURE (Fig. S3).

      (3) Improve quantitative precision: The authors could improve precision by replacing qualitative statements with exact counts. For example "39 of 47 Beefalo showed no detectable bison ancestry." The same suggestion applies when describing how many Beefalo had zebu ancestry.

      We thank the reviewer for this useful suggestion, and agree that the manuscript used imprecise language in describing the results of certain analyses. We have now added quantitative detail throughout the Results section.

      Reviewer #3 (Recommendations for the authors):

      (1) Introduction

      The introduction sets a tone that is heavily focused on the genetic revelation that the economics of beefalo are somewhat of a facade. Beefalo are indeed not part-buffalo (bison). It is unclear to me if the introduction also could benefit from motivating this with more of a theoretical framework based on evolution, inheritance, or trait transmission. If this is really meant to be an economics-focused article, then lean more heavily into that. As it stands, it straddles a bit of economics, a bit of legacies that appear false (beefalo are not part bison at all!), and a bit of admixture genetics theory.

      We intended the focus of this study to be on documenting the species-level ancestry of Beefalo, and concentrated the information presented in the Introduction on this topic. Given that less hybridization between bison and cattle appears to have taken place to form the Beefalo breed than was previously described, we believe that broader theoretical statements about admixture are less relevant here, beyond highlighting examples of successful and failed interspecies hybridization in Bos. We also avoided speculating on the history of the establishment of the breed beyond what could be understood from the genetic data.

      Can the authors give a bit more details about beefalo breeding? Did the breeders select for any quantitative traits and is there a targeted phenotype for beefalo they used as a standard?

      Limited information exists about the precise origins of Beefalo, which were never publicly shared—possibly in part for reasons this manuscript addresses. The only criteria defining Beefalo is the proportion of bison ancestry, and so no quantitative traits or specific phenotypes are related to breed standard.

      Can the authors provide a few examples of what is known about the incompatibilities and reproductive challenges? What is known from past research or from the Beefalo Association documenting the breeding history?

      We provided a general summary of hybridization and incompatibility across Bos, but unfortunately cannot provide details about incompatibilities in Beefalo specifically. Though there is a long history of challenges interbreeding bison and cattle (referenced in the third paragraph of the Introduction), to our knowledge no examination has been carried out of Beefalo specifically and little is known about Beefalo pedigrees (again, perhaps for reasons related to information presented in this study).

      (2) Results Section Sequencing Beefalo genomes

      Please report the number of polymorphic sites to accompany the genomic read depth averages. It seems the authors could include a larger summary of the genomic data that was used for downstream analyses (like the PCA in the next section). Also, does this dataset include the sex chromosomes? How many variants that are retained for analyses are autosomal, sex-linked, or haploid? Please provide more characteristics of the data that was generated after QC and filtering.

      We have now replaced “5.29M” on page 14 with the exact number of SNPs (5,291,534) and added a description of genotype calling to the Results section. We have also included the number of SNPs used for sex chromosomal analyses.

      (3) Results section Estimating bison ancestry in beefalo

      What is a "foundational" individual? Is this a beefalo pedigree founder, a common sire, or an individual with remarkably high bison content? I see in the introduction Joe's Pride was the "most expensive cattle" but there are surely other aspects of "foundational" that the reader should understand as the results are presented.

      We agree that this terminology was imprecise, and have now clarified that we use foundational to mean an early individual that was important in the founding of the Beefalo breed, such as those that were first bred by Bud Basolo.

      For the sentence "The reported pedigree in the NAGP for this animal [NAGP9109] lists its composition as 1⁄2 Brahman, 1⁄4 Charolais, 1⁄8 bison, 1/16th Hereford, and 1/16th Shorthorn, but the American Brahman Breeders Association records this animal (#309519) as purebred Brahman.", this is difficult for a reader with limited cattle breed knowledge to infer significance of this. What is the origination of Brahman breed cattle? Does Brahman ancestry come from another mixed origin that could explain this discrepancy? Does the PCA have references to resolve the origin of Brahman? I realize this may sound extraneous but if membership to a breed that is recently formed from several other lineages or breeds, could you be seeing the deeper parts that compose Brahman cattle? How could one validate that the contributors erroneously labeled this individual as a beefalo?

      We have now noted that the Brahman breed has primarily zebu ancestry. The placement of this individual in the PCA supports the American Brahman Breeders Association metadata, and suggests that the NAGP labeling is incorrect:

      “The reported pedigree in the NAGP for this animal lists its composition as 1/2 Brahman, 1/4 Charolais, 1/8 bison, 1/16th Hereford, and 1/16th Shorthorn, but the American Brahman Breeders Association records this animal (#309519) as purebred Brahman, which is a zebu breed (5 of the other 6 zebu individuals analyzed here are Brahman cattle). We believe NAGP9109 was erroneously labeled as Beefalo by the contributors.”

      Figure 1A: Please add % explained by each PC.

      We have now updated axis labels for the PCA to include the proportion of variance explained for each component.

      Figures 1B and 1C are identical except for the Y axis. Please combine them into a graph with 2 Y-axes (one for PC1 and one for ADMIXTURE). Also, please include the bison in this panel as well.

      We have now updated these panels to include bison, although have kept the labeling so that they may be referenced separately in the text.

      I see that the authors did both unsupervised and supervised. Can the main text have the supervised graphical result instead of the unsurprised? That is more relevant for ancestry proportions via an assignment probability to ancestry groups. Or, if possible, could the authors consider STRUCTURE to also obtain the probability of assignment to a prior defied parental up to 2-generations back? This is by far the best way to leverage the ancestry information of the cattle and bison parental references in addition to the known F1/bison hybrids. Swap the Supplementary Figure 1 with Figure 1D!

      The supervised and unsupervised ADMIXTURE results are highly consistent, as could be expected given the high levels of divergence between species. We prefer to show the unsupervised results in the main text, as this makes the fewest assumptions about the ancestry of the examined individuals, and so also shows that the panels used to represent each species (taurine cattle, zebu cattle, and bison) do not contain individuals which were themselves highly admixed, which could have influenced the supervised ADMIXTURE analyses.

      For the unsupervised ADMIXTURE analyses, what were the cross-validation values per K value tested? How did the authors decide that K=3 was the best one to show?

      We now also show ADMIXTURE results at K=2 and K=4 (Fig. S2) and present the cross-validation results from ADMIXTURE (Fig. S3).

      Regarding "D-statistics ..... are consistent with 0 for most individual Beefalos....", I have two comments. First, by "consistent with", do you mean "are not significantly different from 0", indicating that (explain what this means in your words). Next, "most individual beefalos" means how many? Please provide numbers and values to highlight points or specific findings.

      The interpretation of the D-statistics has been clarified and Z-scores and numbers of individuals to quantitatively describe these results have been added. The text now reads:

      D-statistics of the form D (taurus, Beefalo; bison, water buffalo), which test whether Beefalo share more alleles with bison than taurine cattle, again show 39 Beefalo have no excess affinity with bison compared to taurine cattle (-13.04 < Z < 3.14), although the same eight Beefalo identified in PCA and ADMIXTURE as having bison ancestry also have an excess of bison alleles (6.16 < Z < 34.86), confirming their bison ancestry (Fig. 2A).”

      "In Beefalo with bison ancestry, that ancestry tends to be present in large contiguous blocks, often tens of megabases in size, indicative of recent admixture (Figure 3A, B)". Please display the quantitative results (mean, max, range, standard deviation, etc.) in the main text and point the reader to the table that contains the values for each individual. The rest of this paragraph also uses the words "most' or "always" - please provide numbers. Is most 30/46 beefalo? Is it always exactly all 47 beefalo? Readers want to see numbers!

      The reviewer is correct that this section lacked specificity. We have now provided the exact number of individuals identified with bison and zebu ancestry.

      The section starting "Several lines of evidence attest to the efficacy of using these source panels..." could realistically come first in the Results section and before beefalo results are presented. This would build confidence for the reader that this panel of samples passes a QC and will indeed be able to resolve ancestry-based questions.

      This section specifically refers to the local ancestry analyses, which we have now clarified in the text.

      Figure 3A-C: Please include on each of these figure panels the documented (breeder association) ancestry percentage and the percentage of bison ancestry you obtained from your genomic analyses. Moving it from the legend to the figure is more immediately powerful for the reader. If the authors dated the admixture events as well, please include the meta-data of the association pedigree reporting when bison entered the target individual's genome versus the genome-estimated number of generations since admixture.

      Figure 3 has now been updated to include the reported bison ancestry. No attempt was made to date the admixture event or compare with reported pedigrees, as documented Beefalo pedigrees are typically very sparse (and may be unreliable, as our results suggest).

      Figure 3 legend: Move the following text from the figure legend to the Results section: "Three bison hybrids are inferred to have ~75% bison ancestry, while eight Beefalo have detectable bison ancestry, ranging from 2-18%. Indicine ancestry is detected in most Beefalo at variable levels, ranging from 2-38%, with most Beefalo having between 2-18%.".

      This sentence has been removed from the legend and is now worked into the main text. The corresponding paragraph in the results now reads:

      “Local ancestry inference across individual Beefalo and bison-cattle hybrid genomes provides similar estimates of overall Beefalo ancestry, inferring an absence of bison ancestry across the 37 Beefalo that lacked evidence for such ancestry in previous analyses (Fig. 3). Three bison hybrids are inferred to have ~75% bison ancestry, while eight Beefalo have detectable bison ancestry, ranging from 2-18%. Zebu ancestry is detected in 38 Beefalo at variable levels, ranging from 2-38%, with all but two of Beefalo having between 2-18%.”

      (4) Results section Beefalo sex chromosome ancestry

      Check that the authors do not reference Figure 4B before Figure 4A.

      Thank you to the reviewer for noticing this, it has now been corrected.

      Figure 4A: Could this panel be considered to merge with the autosomal admixture plot? It helps with comparison. Not a firm request - but it is nice to see what is consistent versus what is discordant.

      To avoid cluttering the figure with two highly similar plots, we preferred to separate the autosomal and sex chromosomal results.

      Figure 4C: Could this panel be merged with the autosomal ancestry bar graph to help the reader with visual comparisons?

      We thank the reviewer for this suggestion, but do not understand exactly which figures they are suggesting to be merged.

      (5) Materials and Methods: Modeling Beefalo ancestry:

      The language used in this sentence "This approach allows for directly understanding the ancestry of Beefalo individuals relative to these three groups while mitigating the effects of the low sequencing depth obtained for many Beefalo." conflicts with a sentence later in this paragraph which called PCA a model-free analysis. Please correct.

      Unfortunately, we are unsure what the reviewer refers to here and believe that this sentence does not conflict with the characterization of PCA as a model-free analytical approach.

    1. eLife Assessment

      This study provides a detailed anatomical and functional framework for understanding CO₂ processing and behavioral flexibility in Drosophila. The significance of the work is important, as it identifies how specific neural circuits, such as LN23, modulate innately aversive signals across different contexts. The strength of the evidence is convincing, supported by a robust combination of connectomics, anatomical reconstructions, and targeted behavioral manipulations.

    2. Reviewer #1 (Public review):

      Summary:

      The authors set out to better understand how Drosophila responses to CO2 can be aversive or attractive depending on context (especially presence of food odors, temperature, humidity). While some aspects of this circuit had been previously identified, the authors uncovered additional, critical aspects of the circuit to more fully explain these phenomena. One important discovery was the identification of the LN23 interneuron, which receives input from the V glomerulus. LN23 relays sensory input via an extraglomerular CO2 pathway, and manipulation of LN23 activity revealed a dominant role in CO2-induced avoidance behavior.

      Through a careful series of experiments, the authors demonstrate important aspects of these parallel (and sometimes converging) circuits - differential sensitivity to CO2 concentration changes, synaptic plasticity, circuit connectivity, developmental origins, and the effect of chemo and optogenetic manipulations on behavior. Together, they piece together a complex and interconnected circuit diagram for CO2-dependent behaviors that can be modulated by external factors. This finding will be impactful not only for the fly olfactory/gustatory field but also for many others in the sensory neuroscience community who are very interested in understanding state-dependent modulation of sensory circuits.

      Strengths:

      The experiments were well described and controlled. The addition of the developmental trajectory of the LN23 neurons was interesting. The inclusion of multiple levels of analysis from synaptic contacts and activity-dependent labeling of synapses, circuit analysis guided by connectomes, and detailed behavior analysis for each part of the circuit were all strengths.

      Weaknesses:

      The circuit is very complex and interconnected. This is important for its function, but it makes reading through the manuscript a challenge. The diagrams are helpful, but still somewhat confusing, and some of the experimental findings do not completely support the model outlined in the final figure.

      The main difficulty is visualizing the "default/predominant aversive" LN23 circuit - in the final diagram, there is no "stop" sign on that side, although it's depicted as an inhibition of a "go".<br /> Also, importantly, the findings shown in Figure 5 demonstrate pretty convincingly that LN23 inhibition reduces CO2 avoidance "almost entirely". Also supporting a central role for LN23 is the opposite effect of silencing LN23, with chronic CO2 inducing attraction. If this is the case, then where is the contribution of the other canonical aversive pathway? How does the silencing of LN23 override the PNvbi/uni pathways to aversion? Incorporating this into the figure more prominently would improve the understanding of this contribution to the circuit.

      A minor weakness is that CO2 levels were not reduced below ambient air. For the first part of the paper addressing the activation of these circuits, there seemed to be a ceiling effect for the LN23 neurons at ambient CO2 levels. It would be interesting to see if there would be some change to the activity labeling experiments if CO2 were reduced or eliminated from the air.

    3. Reviewer #2 (Public review):

      Summary

      The authors investigate how parallel olfactory pathways contribute to CO₂ valence processing in Drosophila. By combining multiple approaches, the study identifies LN23 as a previously unrecognized component of the CO₂ circuit and proposes a model in which distinct downstream pathways contribute to aversive and attractive behavioral responses. More broadly, the work aims to connect circuit organization with context-dependent sensory processing and behavioral valence.

      Strengths

      A major strength of the study is the integration of multiple complementary approaches spanning anatomy, circuit analysis, and behavior. This combination provides a rich and valuable framework for understanding how CO₂ information may be processed across different levels of the olfactory system. The identification of LN23 as an important component of the CO₂ pathway is particularly interesting and will likely be useful for future studies investigating olfactory processing, behavioral state modulation, and valence coding. The connectomic and anatomical analyses also provide a valuable resource for the community.

      Another strength of the manuscript is its conceptual ambition. The work moves beyond a simple labeled-line view of olfactory processing and proposes that flexible behavioral responses may emerge from interactions between parallel downstream pathways and multimodal integration centers. The behavioral manipulations further support an important role for LN23 in CO₂-related behaviors.

      Weaknesses

      Several aspects of the conceptual interpretation would benefit from additional clarification or more cautious framing relative to the current experimental evidence. In particular, the distinction between atmospheric versus experimentally elevated CO₂ conditions, as well as the interpretation of chronic exposure in terms of habituation, remains somewhat unclear throughout the manuscript.

      Some conclusions regarding valence coding and multimodal integration also appear more inferential than directly demonstrated experimentally, especially when moving from anatomical connectivity to functional interpretation.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Javorski and colleagues investigate how CO2 valence is processed in the Drosophila olfactory system. Although CO2 is classically associated with an aversive labeled‑line pathway, its behavioral significance can be modulated by environmental context, such as the presence of food‑related cues. The circuit‑level mechanisms underlying this flexibility remain incompletely understood. The authors address this gap by examining how CO2 sensory information diverges at early stages of olfactory processing and how distinct neural pathways contribute to opposing behavioral outcomes. By identifying the local interneuron LN23 as a relay for CO2‑induced aversion, the study suggests that CO2 valence processing may begin to diverge at the level of the antennal lobe, prior to synaptic integration in higher‑order brain regions such as the lateral horn.

      Strengths:

      A major strength of this study is its comprehensive, multi-level experimental design that effectively links neuronal identity, synaptic organization, and behavior. The authors combine calcium‑based anatomical mapping, activity‑dependent reporters, optogenetic and thermogenetic manipulations, and connectomic analyses with behavioral readouts under genetically defined neuronal activation or silencing conditions. Specifically, the identification of LN23 as a component of the CO2 avoidance pathway is supported by anatomical, genetic, and behavioral evidence. Both silencing and activation experiments indicate that LN23 plays an important role in mediating CO2‑induced aversive responses. In contrast, manipulation of the projection neurons (PNv bi and PNv uni) produces more modest behavioral effects, suggesting a degree of specificity for LN23‑associated circuitry within the avoidance pathway. Moreover, the use of previous reported connectome to identify downstream third‑order neurons strengthens the proposed circuit model and provides anatomical support for early divergence of CO2 valence processing.

      Weaknesses:

      While the study provides a strong mechanistic framework for CO2 aversion, some aspects of context‑dependent valence modulation are less directly addressed and may benefit from further experimental exploration.

    1. eLife Assessment

      This study presents analyses of single neuron activity in the subthalamic nucleus (STN) of monkeys performing a decision-making task that manipulates both perceptual evidence and reward. The study shows convincing evidence of distinct subpopulations of neurons in STN that differ in their representations of key quantities related to decision formation. These findings reveal important functional heterogeneity within the STN that helps provide new insights into its contributions to decision processing.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript offers a careful and technically impressive dissection of how subpopulations within the subthalamic nucleus (STN) support reward-biased perceptual decision-making. The authors recorded STN neurons in monkeys performing an asymmetric-reward visual motion discrimination task, then combined single-unit analyses, regression modeling, and drift-diffusion model (DDM) fitting to identify functionally distinct neuronal clusters. Each subpopulation shows unique relationships to computational decision variables - evidence accumulation rate, decision bound, and non-decision time - as well as to post-decision evaluative signals including choice accuracy and reward expectation. The revised manuscript substantially strengthens the original submission by improving both the objectivity of neuron selection and the robustness of the clustering solution.

      Strengths:

      The asymmetric-reward paradigm cleanly separates perceptual and motivational contributions to STN activity, allowing the authors to characterize how neurons blend these distinct sources of information. The dataset is extensive and well-controlled, and the behavioral and neural analyses are tightly integrated. Relating cluster-specific activity to DDM parameters provides an interpretable computational link between population signals and behavior. The clustering solution is now validated across two algorithms, two monkeys, and subsets of trials - establishing that the three-cluster structure is robust. The new Figure 9 offers a conceptually useful, if necessarily speculative, synthesis connecting the identified subpopulations to distinct basal-ganglia pathways (hyperdirect versus indirect). The new Figure 8 documenting the anatomical intermingling of subpopulations is also important, as it directly informs the interpretation of prior and future STN stimulation studies.

      Weaknesses:

      The inferred relationships between neural clusters and DDM parameters remain correlational - the authors now appropriately flag this throughout, and the causal inference gap is acknowledged in the Discussion with concrete proposals for future targeted perturbation strategies. While a generative multi-cluster model would further strengthen mechanistic interpretation, the conceptual framework in Figure 9 provides a reasonable intermediate step given the scope of the study and the absence of simultaneous population recordings, which preclude direct inter-cluster covariation analyses. These remaining limitations are inherent to the experimental design rather than analytical oversights.

    3. Reviewer #2 (Public review):

      This study uses monkey single-unit recordings to examine the role of the STN in combining noisy sensory information with reward bias during decision-making between saccade directions. Using multiple linear regressions and clustering approaches, the authors overall show that a highly heterogeneous activity in the STN reflects almost all aspects of the task, including choice direction, stimulus coherence, reward context and expectation, choice evaluation, and their interactions. The authors report in particular how three classes of neurons map to different decision processes evaluated via the fitting of a drift-diffusion model. Overall, the study provides evidence for functionally diverse and anatomically intermingled populations of STN neurons, supporting multiple roles in perceptual and reward-based decision-making.

      This study follows up on work conducted in previous years by the same team and complements it. Extracellular recordings in monkeys trained to perform a complex decision-making task remain a remarkable achievement, particularly in brain structures that are difficult to target, such as the sub-thalamic nucleus. The authors conducted numerous analyses of STN activities, using sophisticated statistical approaches and functional computational modeling.

      One criticism that I would still make in the revised version of the paper concerns the description of the behavior of the two monkeys which is still minimal, while acknowledging differences in their choice and RT performance that reflect "individual differences in sensitivity to motion stimulus and a common heuristic-based satisficing strategy". This sentence is not clear to me. Moreover, the potential consequences of these differences on neuronal activity are only considered in the cluster analysis done for each of the two animals separately and for which it turns out there is no notable difference.

      Compared to the first version of the paper, the cluster analysis in this revised version yields three distinct populations instead of the previous four. While the authors suggest that these subpopulations play important roles in encoding different aspects of decision-making, the identification of three rather than four subpopulations seems to me an important update that warrants discussion.

      Finally, I think it would have been interesting to identify the level of collinearity in the model proposed by the authors (equation 7). Indeed, one can expect significant collinearity between some of the proposed explanatory factors of neuronal activity, such as choice and coherence level, for example. Similarly, for the analysis relating neuron activity to decision evaluation signals (p 16), firing rates calculated using sliding averages with 1-ms steps are compared, but the method does not specify controls for multiple comparisons or for non-independent data.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      (1) The inferred relationships between neural clusters and specific drift‑diffusion parameters (e.g., bound height, scaling factor, non‑decision time) are intriguing but inherently correlational. The authors should clarify that these associations do not necessarily establish distinct computational mechanisms.

      We agree and have revised the text to avoid any mention of a causal relationship.

      (2) While the k‑means approach is well described, it remains somewhat heuristic. Including additional cross‑validation (e.g., cluster reproducibility across monkeys or sessions) would strengthen confidence in the four‑cluster interpretation.

      We took several steps to increase our confidence in the clustering results. First, we made improvements in how we used the k-means method, primarily by using activity vectors with finer time resolution and filtering out “outlier” neurons (details in Methods) that were dissimilar to other neurons to reduce spurious clustering results. Second, we performed a new set of clustering procedures based on the linkage method, in addition to the k-means method that we originally used. The two clustering methods generated very similar neuron groupings, with a Rand index of 0.93. We now present k-means results in the main figures and linkage results as supplements (e.g., compare Fig 5 and Fig 5-S2). Third, following the reviewer’s suggestion, we performed clustering based on the two monkeys’ data both combined and separately (new Fig 5-S3). Clustering of data from both monkeys combined, compared to each monkey considered separately, had rand index values of 0.94 and 1 for monkeys C and F, respectively (i.e., neurons from one monkey tended to be assigned to the same cluster regardless of whether the clustering was based on data from that monkey alone or both monkeys together), indicating comparable cluster boundaries for the two monkeys. Lastly, we performed clustering based on pseudo-vectors derived from sampling a subset of trials for each neuron and found that the clustering results were stable and robust based on as low as 40% of the trials (new Fig 5-S4).

      Because most neurons were recorded in separate sessions, we cannot perform session-based cross validation.

      (3) The functional dissociations across clusters are clearly described, but how these subgroups interact within the STN or through downstream basal‑ganglia circuits remains speculative.

      We agree and have made sure any speculative claims we make are clearly described as such.

      (4) A natural next step would be to construct a generative multi‑cluster model of STN activity, in which each cluster is treated as a computational node (e.g., evidence integrator, bound controller, urgency or evaluative signal).

      (5) Such a low‑dimensional, coupled model could reproduce the observed diversity of firing patterns and predict how interactions among clusters shape decision variables and behavior.

      (6) Population‑level modeling of this kind would move the interpretation beyond correlational mapping and serve as an intermediate framework between single‑unit analysis and in‑vivo perturbation.

      We agree that such a model would be extremely useful. However, given that designing, implementing, and testing a model like that would require a good deal of speculation about functional and anatomical interactions that we did not measure, it is also well outside the scope of the current study.

      That said, we appreciate the suggestions, which spurred us to go further in terms of providing a summary of our findings (new Figure 9) with a bit of informed speculation about how the different functionally defined subgroups of STN neurons that we characterized might relate to not only different computations but also different pathways through the basal ganglia (i.e., the hyperdirect versus indirect pathway, both of which include the STN). We hope that this summary, along with our more detailed findings, will inform new modeling studies by us and others.

      (7) Causal inference gap - Without perturbation data, it is difficult to determine whether the identified neural modulations are necessary or sufficient for the observed behavioral effects. A brief discussion of this limitation - and how future causal manipulations could test these cluster functions - would be valuable.

      As suggested, we have added the following to the Discussion (line 365): “The exact contributions of these subpopulations are challenging to elucidate, as their intermingled localization make common perturbation techniques, such as electrical microstimulation or optogenetic manipulations, not suitable. It would be interesting to examine if these subpopulations differ in molecular or connectivity properties (e.g., as we speculated above) that can be capitalized to precisely target each subpopulation.”

      Reviewer #1 (Recommendations for the authors):

      (1) Develop or outline a generative multi‑cluster model:

      Consider constructing, even at a conceptual level, a generative network model in which the identified STN clusters serve as interacting computational nodes (e.g., evidence integration, bound modulation, urgency, or evaluative nodes).

      Such a framework could reproduce the simultaneous presence of ramping, transient, and context-sensitive activity patterns observed across clusters.

      Even a simulated or schematic implementation - showing how parameter coupling among these clusters gives rise to the reported firing diversity and behavioral effects - would help clarify the mechanistic implications of your findings.

      As noted above, we believe that a full modeling study is well outside the scope of the present work. However, we have provided a conceptual framework, shown in Figure 9, summarizing our findings and providing some informed speculation about how different subgroups of STN neurons could provide different functions along distinct anatomical pathways.

      (2) Strengthen the link between cluster activity and computation:

      Use cross‑validated or hierarchical regression models to verify the robustness of correlations between cluster‑specific firing measures and fitted drift‑diffusion parameters. This would make the mapping between neural activity and model components more statistically grounded.

      We appreciate the suggestion and thought hard about how we might implement it but ultimately decided our approach is most appropriate, given the strengths and limitations of our dataset. The fundamental issue is that it takes many trials to obtain reliable estimates of DDM parameters. Our approach of creating twelve “pseudo-sessions” for each neuron (half of those for trials with high firing rates, half for trials with low firing rates) balances our ability to obtain those estimates while testing for relationships with firing rate. Any further subdivision of the data for cross validation yields unreliable parameter estimates (i.e., with big error bars). We also chose not to use a hierarchical model and instead took a more unbiased approach by considering how all of the DDM parameters relate to firing rate.

      Despite the simplicity of our approach, we believe that these results are statistically grounded. It is possible that more complex regression models may reveal additional (e.g., non-linear) relationships, but those results would also be less intuitive to interpret. We therefore decided to retain our analysis choice.

      (3) Assess cluster reproducibility:

      Report or include in the supplement the degree of correspondence of cluster identities across monkeys or across independent subsets of trials. Cluster stability metrics (e.g., bootstrap or split‑half analysis) would reassure readers that cluster structure is not dataset‑specific.

      Please see our response above to the main comment #2 regarding the robustness and stability of clustering results.

      (4) Explore population interactions directly:

      You could analyze pairwise or population‑level covariations (e.g., principal components or canonical correlation analysis) to test whether inter‑cluster interactions correspond to model‑predicted dynamics such as competition or normalization.

      Because most of the neurons were recorded in separate sessions and not simultaneously, the suggested population analyses are not feasible.

      Discuss briefly how the proposed generative or dynamical multi‑cluster model could be empirically tested-e.g., using selective perturbation (microstimulation, optogenetic, or pharmacological) in future studies-to evaluate interactions inferred from the current dataset. If feasible, mention how this framework might generalize to other decision contexts beyond oculomotor tasks, such as effort‑reward tradeoffs or inhibitory control, reinforcing the broad relevance of STN computations.

      As suggested, we have added the following to the Discussion (line 366): “The exact contributions of these subpopulations are challenging to elucidate, as their intermingled localization make common perturbation techniques, such as electrical microstimulation or optogenetic manipulations, not suitable. It would be interesting to examine if these subpopulations differ in molecular or connectivity properties (e.g., as we speculated above) that can be capitalized to precisely target each subpopulation.”

      Reviewer #2 (Public review):

      One criticism I would make is that the authors sometimes seem to assume that readers are familiar with their previous work. Indeed, the motivation and choices behind some analyses are not clearly explained. It might be interesting to provide a little more context and insight into these methodological choices. The same is true for the description of certain results, such as the behavioral results, which I find insufficiently detailed, especially since the two animals do not perform exactly the same way in the task.

      We apologize for the lack of detail regarding the behavioral results and analysis choices. To address this issue, we substantially revised the text, particularly in Results and Methods.

      The differences in behavior for the two monkeys were the subject of an entire published study (Fan Y, Gold JI, Ding L, 2018, Ongoing, rational calibration of reward-driven perceptual biases. Elife 7: e36018.). That study showed that these differences most likely arose from the monkeys’ individual sensitivity to the motion stimulus, combined with a heuristic-based strategy to gain satisficing rewards that they all seem to use. We revised the text to acknowledge the individual differences and refer readers to our previous study (line 78): “Both monkeys showed consistent biases toward the large-reward choice (Figure 1B, C). The individual differences in their choice and RT performance reflect individual differences in sensitivity to motion stimulus and a common heuristic-based satisficing strategy, as we demonstrated in a previous study (Fan et al., 2018).”

      Another criticism is the difficulty in following and absorbing all the presented results, given their heterogeneity. This heterogeneity stems from analytical choices that include defining multiple time windows over which activities are studied, multiple task-related or monkey behavioral factors that can influence them, multiple parameters underlying the decision-making phenomena to be captured, and all this without any a priori hypotheses. The overall impression is of an exploratory description that is sometimes difficult to digest, from which it is hard to extract precise information beyond the very general message that multiple subpopulations of neurons exist and therefore that the STN is probably involved in multiple roles during decision-making.

      In response to the three reviewers’ comments on data inclusion and the clustering analysis we presented, we have substantially improved the objectivity and robustness of our approaches, by: 1) applying a data-driven criterion for identifying neurons with robust task-relevant modulation (Figure 4C), 2) removing “outlier” neurons that appear not to share activity profiles with any other neurons in our sample (note that these outlier neurons would be at the outskirts in the cluster space instead of between clusters), 3) increasing the temporal resolution for generating firing rate vectors, and 4) comparing clustering results based on two methods (k-means and linkage). These improvements both sharpened the cluster boundaries and allowed us to observe more robust and distinctive subpopulation-specific relationships between neural activity and computational components in the DDM framework (new Figures 5–7 and their supplementary figures). We believe these updated results clearly demonstrate that: 1) there are different STN subpopulations, and 2) each of the subpopulations encodes a distinct set of functions.

      It would also have been interesting to have information regarding the location of the different identified subpopulations of neurons in the STN and their level of segregation within this nucleus. Indeed, since the STN is one of the preferred targets of electrical stimulation aimed at improving the condition of patients suffering from various neurological disorders, it would be interesting to know whether a particular stimulation location could preferentially affect a specific subpopulation of neurons, with the associated specific behavioral consequences.

      We have added a new Figure 8 to show the localization of neurons with and without task modulation and of neurons from different subpopulations. Consistent with our previous demonstration of intermingled distribution of STN subpopulations, we did not observe any activity pattern-based segregation.

      To relate the activity patterns to previously reported stimulation effects, we added the following to the Discussion (line 307): “This functional diversity, along with a lack of clear anatomical organization, is consistent with the multiple effects of STN stimulation in patient populations on decision-making and out previous results in monkeys, including reductions in response times, a weaker dependence on evidence, and changes in the maximal value and trajectories of the decision bound (Frank et al., 2007; Cavanagh et al., 2011; Coulthard et al., 2012; Green et al., 2013; Zavala et al., 2014; Herz et al., 2016; Pote et al., 2016; Branam et al., 2024).”

      Therefore, this paper is interesting because it complements other work from the same team and other studies that demonstrate the likely important role of the STN in decision-making. This will be of interest to the decision-making neuroscience community, but it may leave a sense of incompleteness due to the difficulty in connecting the conclusions of these different studies. For example, in the discussion section, the authors attempt to relate the different neuronal populations identified in their study and describe some relatively consistent results, but others less so.

      We hope that our revised Results and Discussion clarify the conclusions that can be drawn from this and other related studies.

      Reviewer #2 (Recommendations for the authors):

      (1) Introduction, l. 47-48: It would be interesting to provide more details on these three populations in order to better understand why we need additional experiments to more comprehensively define their roles.

      We now give more details in the Introduction about the remaining questions we aimed to address in this study (line 50): “However, the specific computational roles that these different subpopulations play in decision-making and other cognitive functions remain not well understood. For example, two of the subpopulation had overall activity patterns that were consistent with two different models in which the STN modulated the decision bound (Ratcliff and Frank, 2012; Wei et al., 2015), but the exact nature of this modulation is not known. The other subpopulation’s general activity patterns were consistent with a model of STN mediating evidence accumulation (Bogacz and Gurney, 2007), but it is unclear if and how this activity contributes to how evidence is weighed, biased, or accumulated.”

      Our previous attempt to distinguish these alternatives using electrical microstimulation was unsuccessful because that manipulation likely affected highly intermingled subpopulations with different functions.”

      (2) Results, l. 71-73: A slightly more detailed description of the behavioral results would be appreciated, especially since the two monkeys do not behave exactly the same way in the task, particularly in terms of reaction times (Figure 1B top-right versus bottom-right).

      We revised the text to acknowledge the individual differences and refer readers to our previous study (line 78): “Both monkeys showed consistent biases toward the large-reward choice (Figure 1B, C). The individual differences in their choice and RT performance reflect individual differences in sensitivity to motion stimulus and a common heuristic-based satisficing strategy, as we demonstrated in a previous study (Fan et al., 2018).”

      (3) Figure 2G-I: Were the multiple linear regressions performed only in the asymmetric reward condition?

      Yes. We added in Methods (line 487): “All analyses were performed on activity from the asymmetric-reward task.”

      (4) Very often in the text, the authors use terms that refer to concepts or methods that are difficult to grasp on the first reading, especially if we are not familiar with the team's previous publications. This is the case, for example, with "joint modulation," "reward context," "reward expectation," "k-means clustering," "tSNE," "Silhouette score for neurons," "Rand index," etc. All the explanations are minimal, and it would be helpful to clearly define these terms and provide some justification and insight to support the use of the analyses and the resulting variables, all of which would facilitate the reading of the manuscript.

      We now define these terms explicitly in the text (emphasis added here for clarity):

      (Results, line 129): “Using a previous definition of “joint modulation” (Doi et al., 2020), including modulation separately by motion coherence and reward context or reward size and modulation by the interaction of motion coherence and reward size, we found that ~40% of the neurons showed joint modulation during motion viewing.”

      (Results, line 71): “… for which we separately manipulated the noisy evidence (motion direction and strength) and reward context (a larger juice reward for a correct choice associated with one of the two directions).”

      (Results, line 250): “Choice accuracy describes the probability that a choice is correct given the evidence. Reward expectation describes the the expected reward given a choice.”

      (Methods, line 550): “To quantify the consistency between two runs of clustering, we computed the Rand index as the number of neuron pairs with consistent grouping (i.e., they were placed in the same cluster for both runs or they were placed in different clusters for both runs), normalized by the total number of possible neuron pairs. A value of 1 indicates that the two clustering runs produce identical results, and a value of 0 indicates that the two runs do not agree on any pairs of neurons.”

      To quantify the separation of clusters, we computed silhouette scores as the difference between mean intra-cluster distance and the mean nearest-cluster distance, normalized by the maximum of the two values. A positive score indicates that the member is closer to its same-cluster neighbors than different-cluster neighbors. Clustering runs with high mean silhouette score were considered to have better cluster separation.

      We no longer use tSNE visualization.

      (5) Figure 5A, caption: A quick description of the parameters would be useful.

      We added the description of DDM parameters in the caption of new Figure 4.

      (6) Results l. 222: Why does the analysis only concern epoch 5? I suggest justifying this choice. Also, the text indicates a "trend" but Figure 5C shows a significant result (p=0.0129).

      These statements have been removed from the updated manuscript.

      (7) Methods, l. 443: The authors should report more details about how they decided that neurons were task-related or not. "Visual inspection" sounds like a very vague and subjective criterion.

      We now apply a more objective criterion for identifying neurons with task-relevant modulation:

      (Results, line 145): “To focus on neurons with the most robust task-relevant activity, we measured firing rates during a baseline period (300 ms before motion onset) and sliding 100 ms windows from motion onset to 150 ms after saccade onset in 50 ms steps. We identified the maximal and minimal z-scores, representing the peak activation and suppression, respectively, for each neuron across all trial conditions (Figure 4C). We applied a threshold of z-score >1.5 for either activation or suppression and focused further analyses on the 87 neurons that met this selection criterion (n = 62 and 25 for monkeys C and F, respectively).”

      (8) A map of the location of the different STN neuron clusters found in this study within the structure would be very interesting.

      We have added a new Figure 8 to show the localization of neurons with/without task modulation and of neurons from different subpopulations.

      (9) Unless I am mistaken, there is no mention of data availability in this manuscript.

      The data availability statement was/is on the submission form.

      Data Availability: All electrophysiological data and the code for the analyses presented in the paper will be deposited in a publicly accessible domain when the paper is published.

      Previously Published Datasets: Source data for Figure 3-S2 in eLife paper:

      https://doi.org/10.7554/eLife.60535.: Fan, Doi, Gold, Ding, 2020,

      https://cdn.elifesciences.org/articles/60535/elife-60535-fig3-data1-v1.csv,

      https://cdn.elifesciences.org/articles/60535/elife-60535-fig3-data1-v1.csv

      Reviewer #3 (Public review):

      The primary weakness of the paper lies in the claim that STN contains multiple sub-populations with distinct involvements in decision making, which is inadequately supported by the paper's methods and analyses.

      First, while it is clear that the ~150 recorded neurons across 2 monkeys (91, 59 respectively) display substantial heterogeneity in their activity profiles across time and across stimulus/reward conditions, the claim of sub-populations largely rests on clustering a *subset of less than half the population - 66 neurons (48, 15 respectively) - chosen manually by visual inspection*. The full population seems to contain far more decision-modulated neurons, whose response profiles seem to interpolate between clusters. Moreover, it is unclear if the 4 clusters hold for each of the 2 monkeys, and the choice of 4-5 clusters does not seem well supported by metrics such as silhouette score, etc, that peak at 3 (1 or 2 were not attempted). From the data, it is easier to draw the conclusion that the STN population contains neurons with heterogeneous response profiles that smoothly vary in their tuning to different decision variables, rather than distinct sub-populations.

      In response to the three reviewers’ comments on data inclusion and the clustering analysis we presented, we have substantially improved the objectivity and robustness of our approaches, by: 1) applying a data-driven criterion for identifying neurons with robust task-relevant modulation (Figure 4C), 2) removing “outlier” neurons that appear not to share activity profiles with any other neurons in our sample (note that these outlier neurons would be at the outskirts in the cluster space instead of between clusters), 3) increasing the temporal resolution for generating firing rate vectors, and 4) comparing clustering results based on two methods (K-means and linkage). These improvements both sharpened the cluster boundaries and allowed us to observe more robust and distinctive subpopulation-specific relationships between neural activity and computational components in the DDM framework (new Figures 5–7 and their supplementary figures). We believe these updated results clearly demonstrate that: 1) there are different STN subpopulations, and 2) each of the subpopulations encodes a distinct set of functions.

      We performed additional analysis to assess the robustness of the clustering results. First, following the reviewer’s suggestion, we performed clustering based on the two monkeys’ data both combined and separately (new Fig 5-S3). Clustering of data from both monkeys combined compared to each monkey considered separately had rand index values of 0.94 and 1 for monkeys C and F, respectively (i.e., neurons from one monkey were assigned to the same cluster regardless of whether the clustering was based on data from that monkey alone or both monkeys together), indicating comparable cluster boundaries for the two monkeys. Second, we performed clustering based on pseudo-vectors derived from sampling a subset of trials for each neuron and found that the clustering results were stable and robust based on as low as 40% of the trials (new Fig 5-S4). Third, we generated a new figure (Figure 5-S1), using dendrograms to visualize how the neurons relate to each other. The dendrogram in Figure 5-S2 is more consistent with (at least) three distinct subpopulations of neurons than with the null hypothesis of a continuous distribution with smoothly-varying response profiles.

      Second, assuming the existence of sub-populations, it is unclear how their time- and condition-varying relationship with DDM parameters is to be interpreted. These relationships are inferred by splitting trials based on individual neurons' firing rates in different task epochs and reward contexts, and regressing onto the parameters of separate DDMs fit to those subsets of trials. The result is that different sub-populations show heterogeneous relationships to different DDM parameters over time - a result that, while interesting, leaves the computational involvement of these sub-populations/implementation of the decision process unclear.

      The improvements we made of the clustering procedure both sharpened the cluster boundaries and allowed us to observe more robust and distinctive subpopulation-specific relationships between neural activity and computational components in the DDM framework (new Figures 5-7 and their supplementary figures). These updated results demonstrate that: 1) there are different STN subpopulations, and 2) each of the subpopulations encodes a particular set of functions.

    1. eLife Assessment

      This valuable study raises the intriguing possibility that crickets use bat-associated odors as cues of predation risk, extending the classic bat-insect arms race beyond its usual acoustic framework. The authors combine fecal metabarcoding, behavioral assays, electrophysiology, chemical analyses, and field observations to show that Loxoblemmus equestris avoids the odor of the insectivorous bat Scotophilus kuhlii, and that synthetic (-)-limonene can elicit antennal responses, avoidance in the laboratory, and reduced calling activity in the field. However, the evidence is currently incomplete because the identity, biological source, natural concentration, and ecological specificity of limonene as a bat-derived predator cue require stronger support, including clearer quantification, contamination controls, individual-level odor data, and evidence that crickets can distinguish bat-associated limonene from common environmental sources. The work will be of interest to researchers in sensory ecology, chemical ecology, predator-prey interactions, and bat-insect coevolution.

    2. Reviewer #1 (Public review):

      The manuscript examines whether insects can use bat odor as a cue of predation risk. The authors focus on the insectivorous bat Scotophilus kuhlii and the cricket Loxoblemmus equestris. They first use fecal DNA metabarcoding to show that crickets are part of the bat's diet, and field surveys to show that L. equestris is abundant at local foraging sites. In laboratory Y-tube assays, the authors show that crickets strongly avoid air carrying bat body odor. Gas chromatography coupled with electroantennographic detection showed that cricket antennae respond to components of bat odor. Chemical analyses identified several volatile compounds, with 2,2-dimethylheptane and (−)-limonene associated with antennal responses. Further analyses suggested that snout secretions are likely to contribute to the bat's body odor. The authors then tested individual compounds. Among the commercially available candidates, (−)-limonene elicited a strong antennal response and was sufficient to cause avoidance in the olfactometer. In field plots, spraying (−)-limonene reduced cricket calling activity relative to pre-exposure levels, whereas calling increased in control plots treated with hexane. Overall, the study argues that crickets can detect a vertebrate predator through olfactory cues and that a single bat-associated volatile can trigger antipredator behavior.

      This is an interesting and enjoyable study that addresses an understudied aspect of predator-prey interactions. The manuscript is clearly written, the experiments are presented in a logical sequence, and the figures are crisp and easy to follow. I really appreciated the combination of behavioral assays, electrophysiology, chemical analysis, and field observations.

      My main issue concerns the identity and biological origin of the proposed bat odor cue, (−)-limonene. Limonene seems like an unusual compound to be emitted endogenously by a mammal, particularly by an insectivorous bat. It would be helpful if the authors could clarify whether mammals are known to synthesize this compound de novo, and, if not, what the likely source of this plant-associated terpene would be in S. kuhlii. Possible sources could include environmental exposure, diet, roosting material, handling, or temporary housing conditions.

      I do not doubt that crickets avoid synthetic (−)-limonene. Indeed, this result is quite plausible given that limonene is widely used in insect repellent or repellent-associated fragrance products. However, this also makes contamination an important issue to address explicitly. How did the authors exclude the possibility that limonene entered the samples from human-associated sources, such as insect repellents, soaps, cleaning products, field equipment, cloth bags, cages, gloves, or other materials used while handling wild-caught bats? It would strengthen the manuscript to report limonene levels for individual bat odor collections, all relevant blanks, and any handling or housing controls.

      More broadly, given the common occurrence of limonene in plants and human-associated products, I am not yet convinced that it would function as a reliable "keystone kairomone" as suggested around line 253. How would crickets distinguish bat-associated limonene from limonene emitted by a mint leaf, citrus peel, pine material, or other non-threatening environmental sources? The authors may wish to soften this interpretation or provide additional evidence that crickets respond to limonene in a bat-specific context, perhaps through concentration, temporal patterning, co-occurring volatiles, or enantiomeric composition.

    3. Reviewer #2 (Public review):

      Summary:

      Many insects possess extremely sensitive olfactory systems that can detect chemical signals from distances of several kilometers. For decades, the arms race between bats and insects has served as a prime example of acoustic co-evolution. The auditory adaptations of insects to echolocation have been well documented. Cricket has a multi-sensory predator recognition system with keen olfactory, tactile, and auditory senses. However, whether crickets can use the scent of bats to avoid them remains unknown at present. The authors hypothesized that cricket prey (Loxoblemmus equestris) might eavesdrop on predator bat (Scotophilus kuhlii) VOCs as an early warning. L. equestris is one of the prey species of S. kuhlii, and the authors demonstrated that the body odor of the insectivorous bat S. kuhlii triggers robust avoidance and electrophysiological responses in the cricket L. equestris, and that a single compound, (-)-limonene, is sufficient to elicit this avoidance in the laboratory and suppress calling in the field. Overall, this paper has a complete chain of evidence and should be a highly praised study.

      Comments:

      (1) Olfactory eavesdropping can transcend the evolutionary divide between vertebrate predators and invertebrate prey, enabling invertebrates to trigger defensive avoidance behaviors in response to predator-derived volatile odors. This phenomenon is empirically well-documented and requires no excessive emphasis.

      (2) Without quantitative analysis and without knowing the relative content of this key substance limonene, I don't quite understand how to determine the concentration of limonene standard for EAD, as well as the concentration in field experiments. How is the concentration of limonene determined in field spraying, and is this actually the case in the wild environment?

      (3) Figures 1C and D should compare the GC-EAD response of L. equestris to the odor of bat body and the odor of bat nasal secretions. It should not be compared with the air control group. Figure 1D has the same problem.

    4. Author response:

      eLife Assessment

      This valuable study raises the intriguing possibility that crickets use bat-associated odors as cues of predation risk, extending the classic bat-insect arms race beyond its usual acoustic framework. The authors combine fecal metabarcoding, behavioral assays, electrophysiology, chemical analyses, and field observations to show that Loxoblemmus equestris avoids the odor of the insectivorous bat Scotophilus kuhlii, and that synthetic (-)-limonene can elicit antennal responses, avoidance in the laboratory, and reduced calling activity in the field. However, the evidence is currently incomplete because the identity, biological source, natural concentration, and ecological specificity of limonene as a bat-derived predator cue require stronger support, including clearer quantification, contamination controls, individual-level odor data, and evidence that crickets can distinguish bat-associated limonene from common environmental sources. The work will be of interest to researchers in sensory ecology, chemical ecology, predator-prey interactions, and bat-insect coevolution.

      We thank the editors for recognizing the novelty and significance of our work.

      The central aim and contribution of this study is to provide direct evidence that an insect can detect a phylogenetically distant vertebrate predator, an insectivorous bat, via olfaction and initiate avoidance behavior. Our dietary analysis of bats and survey of potential prey in foraging habitats established a predator–prey relationship between the Asiatic lesser yellow house bat (Scotophilus kuhlii) and the cricket Loxoblemmus equestris. In addition, behavioral assays showed that the crickets strongly avoid air carrying bat body odor, and electrophysiological recordings using GC-EAD confirmed that volatiles from S. kuhlii body odor are detected by L. equestris antennae. Together, these results provide strong evidence that this insect can perceive and avoid the body odor of its predator, S. kuhlii. We are grateful that the editors and reviewers acknowledged the main conclusions.

      We also investigated the sources of bat body odor, its major volatile components, and the behavioral, physiological, and ecological responses of the crickets to limonene. The purpose of these studies was to test the hypothesis that elemental perception—detection of a single compound—could serve as a mechanism by which crickets perceive bat odor. We found that limonene was present in bat odor, elicited antennal responses in crickets, induced avoidance behavior in olfactometer assays, and reduced calling activity in the field. Together, these results support the idea that elemental perception is a plausible and efficient strategy for initiating anti-predator behavior against bats.

      We appreciate the editors’ constructive comments. The editors and Reviewer #1 suggested that limonene, as a bat-derived predator cue, requires more evidence, mainly for two reasons:

      (1) Limonene is common in plants but rare in mammals; the reviewer raised the possibility that the limonene identified in our study may have been introduced as exogenous contamination during handling.

      (2) The ability of crickets to distinguish bat-associated limonene from limonene originating from common environmental sources (e.g., plants) remains unclear.

      Below we address these points and describe the revisions we will make to strengthen the manuscript.

      On the potential contamination origin of limonene

      We agree that limonene is common in plants and human-associated products, and we carefully considered this possibility. However, several lines of evidence suggest that contamination is highly unlikely. First, we followed strict experimental protocols: all instruments were cleaned with ethanol and oven-dried before each use; bats were held in stainless-steel cages and cloth bags made of degreased bleached cotton washed with purified water. Second, limonene was not detected in any blank controls (empty-chamber air samples for whole-body odor collection, nor blank cotton swabs for secretion analysis), whereas it was consistently identified in multiple bat snout-secretion samples. Third, previous studies have independently reported limonene in the secretions of several bat species (Faulkes et al., 2019, PeerJ; Zhang et al., 2022, Ann. N.Y. Acad. Sci.). Moreover, recent work suggests that skin-associated microorganisms can contribute to bat volatile profiles (Sun et al., 2026, BMC Biology), and some microbes possess enzymes involved in limonene biosynthesis. Therefore, we are confident that the limonene we detected originates from the bats themselves (either endogenously or via their microbiota), not from exogenous contamination.

      On how crickets might distinguish bat-derived limonene from environmental sources

      This is an insightful question. As discussed in our original manuscript (Discussion section), crickets may not rely exclusively on limonene as a standalone cue. First, our GC-EAD analyses showed that cricket antennae respond to multiple bat volatiles beyond limonene, suggesting that additional compounds, either alone or in synergistic blends, may contribute to predator recognition. Elemental perception via (–)-limonene therefore likely represents one effective strategy within a broader olfactory toolkit, rather than the sole mechanism. Second, under natural conditions, crickets could also integrate olfactory information with non-chemical ecological signals, such as temporal patterning (bats are active at night) and spatial patterning (specific foraging habitats), to further reduce false alarms.

      However, fully testing these hypotheses would require substantial additional work. It would be necessary to quantify natural limonene concentrations in bat odors versus various plant sources, conduct choice experiments with ecologically relevant concentrations and blends, and perhaps manipulate the olfactory landscape in the field. It would also be necessary to examine how other volatile compounds in bat body odor interact with limonene, alone or together, to shape cricket recognition. After all, bat body odor contains dozens of compounds, and it is challenging to determine the necessity and sufficiency of each. These kinds of difficulties are not unique to our study; they are widespread in chemical ecology. Problems like how animals distinguish identical compounds from different biological sources are common in chemical ecology, and they are rarely solved in a single study. These lines of investigation, from quantifying natural concentrations to examining compound interactions, are important, but they are not the focus of the present study. So we have put this forward as an important direction for future research.

      Revisions we will make:

      (1) In the Methods section, we will add detailed descriptions of contamination controls and report blank-control results to demonstrate that exogenous contamination is very unlikely.

      (2) In the Discussion section, we will expand the discussion of the possible biological sources of limonene (including microbiota) in light of our results and the literature.

      (3) In the Discussion and Conclusion, we will state more cautiously the role of limonene as a bat-derived cue, acknowledging that while it is sufficient to trigger avoidance, additional work is needed to establish its ecological specificity.

      We believe these revisions will address the editors’ and the reviewers’ concerns while preserving the main conclusion that olfaction can mediate bat detection by crickets.

      Reviewer #1 (Public review):

      The manuscript examines whether insects can use bat odor as a cue of predation risk. The authors focus on the insectivorous bat Scotophilus kuhlii and the cricket Loxoblemmus equestris. They first use fecal DNA metabarcoding to show that crickets are part of the bat's diet, and field surveys to show that L. equestris is abundant at local foraging sites. In laboratory Y-tube assays, the authors show that crickets strongly avoid air carrying bat body odor. Gas chromatography coupled with electroantennographic detection showed that cricket antennae respond to components of bat odor. Chemical analyses identified several volatile compounds, with 2,2-dimethylheptane and (−)-limonene associated with antennal responses. Further analyses suggested that snout secretions are likely to contribute to the bat's body odor. The authors then tested individual compounds. Among the commercially available candidates, (−)-limonene elicited a strong antennal response and was sufficient to cause avoidance in the olfactometer. In field plots, spraying (−)-limonene reduced cricket calling activity relative to pre-exposure levels, whereas calling increased in control plots treated with hexane. Overall, the study argues that crickets can detect a vertebrate predator through olfactory cues and that a single bat-associated volatile can trigger antipredator behavior.

      This is an interesting and enjoyable study that addresses an understudied aspect of predator-prey interactions. The manuscript is clearly written, the experiments are presented in a logical sequence, and the figures are crisp and easy to follow. I really appreciated the combination of behavioral assays, electrophysiology, chemical analysis, and field observations.

      My main issue concerns the identity and biological origin of the proposed bat odor cue, (−)-limonene. Limonene seems like an unusual compound to be emitted endogenously by a mammal, particularly by an insectivorous bat. It would be helpful if the authors could clarify whether mammals are known to synthesize this compound de novo, and, if not, what the likely source of this plant-associated terpene would be in S. kuhlii. Possible sources could include environmental exposure, diet, roosting material, handling, or temporary housing conditions.

      I do not doubt that crickets avoid synthetic (−)-limonene. Indeed, this result is quite plausible given that limonene is widely used in insect repellent or repellent-associated fragrance products. However, this also makes contamination an important issue to address explicitly. How did the authors exclude the possibility that limonene entered the samples from human-associated sources, such as insect repellents, soaps, cleaning products, field equipment, cloth bags, cages, gloves, or other materials used while handling wild-caught bats? It would strengthen the manuscript to report limonene levels for individual bat odor collections, all relevant blanks, and any handling or housing controls.

      More broadly, given the common occurrence of limonene in plants and human-associated products, I am not yet convinced that it would function as a reliable "keystone kairomone" as suggested around line 253. How would crickets distinguish bat-associated limonene from limonene emitted by a mint leaf, citrus peel, pine material, or other non-threatening environmental sources? The authors may wish to soften this interpretation or provide additional evidence that crickets respond to limonene in a bat-specific context, perhaps through concentration, temporal patterning, co-occurring volatiles, or enantiomeric composition.

      We thank Reviewer #1 for the positive evaluation and for recognizing the study as “interesting and enjoyable.” We greatly appreciate the endorsement of our integrative approach combining behavioral assays, electrophysiology, chemical analysis, and field observations. The core conclusion that crickets can detect and avoid bats via olfaction is well supported by our data, and we are pleased that the reviewer has recognized this central finding.

      We are grateful for the reviewer’s constructive comments on the biological source and ecological specificity of limonene. In our response to the editor above, we have already responded to both aspects in detail; here we will briefly restate the key points.

      On the biological origin of limonene and potential contamination

      We agree that limonene is common in plants and human-made products, but relatively unusual for a mammal to emit endogenously. We have carefully examined the possibility of contamination and believe it is highly unlikely for the following reasons:

      (1) Strict experimental protocols: All experiments were conducted in a dedicated space. Instruments were cleaned with ethanol and oven-dried before and after each use. Cloth bags used to hold bats were made of degreased bleached cotton and washed with purified water; holding cages were stainless steel.

      (2) Blank controls: Limonene was not detected in any blank control samples, neither in the empty-chamber air controls for whole-body odor collection nor in the blank cotton swabs used for secretion analysis. In contrast, limonene was consistently identified in multiple bat snout-secretion samples.

      (3) Independent reports: Limonene has been previously identified in the secretions of several bat species (Faulkes et al., 2019, PeerJ; Zhang et al., 2022, Ann. N.Y. Acad. Sci.), indicating that its presence is not unique to our study or handling conditions.

      (4) Potential microbial origin: Even if bats do not synthesize limonene de novo (a capacity for which there is currently no evidence), recent work shows that skin-associated microorganisms can substantially shape bat volatile odors (Sun et al., 2026, BMC Biology). Some of these microbes possess enzymes involved in limonene biosynthesis, making bat-associated microbiota a plausible biological source of this compound.

      (5) Thus, the limonene we detected is highly likely to originate from the bats themselves (directly or via their microbes) rather than from contamination.

      On how crickets distinguish bat-associated limonene from environmental sources

      This is an excellent and important question. As we briefly discussed in the original manuscript, crickets may not rely exclusively on limonene as a bat-specific cue. Under natural field conditions, they could integrate olfactory information with other ecological cues, for example, temporal and spatial patterning (bats are active at night in specific foraging habitats), co-occurrence with other bat-specific volatiles (the full odor blend contains many compounds), or even concentration thresholds that differ between bat emissions and plant sources. We hypothesize that such context-specific integration could minimize false alarms.

      However, we also recognize that fully testing these hypotheses would require substantial additional work: quantify natural limonene concentrations in bat odors versus various plant sources, conduct choice experiments with ecologically relevant concentrations and blends, and perhaps manipulate the olfactory landscape in the field. These are important questions, but they are not the central focus of the present study, whose primary aim is to provide evidence that olfaction—and elemental perception of a single compound—can function in this predator-prey system. We have therefore framed this as an important direction for future research.

      Revisions we will make:

      (1) In the Methods section, we will add detailed descriptions of contamination controls and present blank-control results to demonstrate that exogenous contamination is very unlikely.

      (2) In the Discussion section, we will expand the discussion of limonene’s biological sources (including microbial contributions) and explicitly acknowledge the need for future work on how crickets might discriminate bat-derived from plant-derived limonene.

      (3) In the Conclusion, we will more cautiously characterize limonene’s ecological role, emphasizing that it is sufficient to trigger avoidance but that its natural specificity requires further investigation.

      We thank the reviewer again for these insightful comments, which will help us improve the manuscript.

      Reviewer #2 (Public review):

      Summary:

      Many insects possess extremely sensitive olfactory systems that can detect chemical signals from distances of several kilometers. For decades, the arms race between bats and insects has served as a prime example of acoustic co-evolution. The auditory adaptations of insects to echolocation have been well documented. Cricket has a multi-sensory predator recognition system with keen olfactory, tactile, and auditory senses. However, whether crickets can use the scent of bats to avoid them remains unknown at present. The authors hypothesized that cricket prey (Loxoblemmus equestris) might eavesdrop on predator bat (Scotophilus kuhlii) VOCs as an early warning. L. equestris is one of the prey species of S. kuhlii, and the authors demonstrated that the body odor of the insectivorous bat S. kuhlii triggers robust avoidance and electrophysiological responses in the cricket L. equestris, and that a single compound, (-)-limonene, is sufficient to elicit this avoidance in the laboratory and suppress calling in the field. Overall, this paper has a complete chain of evidence and should be a highly praised study.

      Comments:

      (1) Olfactory eavesdropping can transcend the evolutionary divide between vertebrate predators and invertebrate prey, enabling invertebrates to trigger defensive avoidance behaviors in response to predator-derived volatile odors. This phenomenon is empirically well-documented and requires no excessive emphasis.

      (2) Without quantitative analysis and without knowing the relative content of this key substance limonene, I don't quite understand how to determine the concentration of limonene standard for EAD, as well as the concentration in field experiments. How is the concentration of limonene determined in field spraying, and is this actually the case in the wild environment?

      (3) Figures 1C and D should compare the GC-EAD response of L. equestris to the odor of bat body and the odor of bat nasal secretions. It should not be compared with the air control group. Figure 1D has the same problem.

      We sincerely thank Reviewer #2 for the high praise (“complete chain of evidence,” “highly praised study”) and for the constructive suggestions to further improve the manuscript.

      On the novelty of olfactory eavesdropping across the vertebrate–invertebrate divide

      We agree with the reviewer that “olfactory eavesdropping can transcend the evolutionary gap between vertebrate predators and invertebrate prey” and that such phenomena have been documented. However, we would like to note that empirical examples remain relatively scarce, especially those that combine chemical identification, electrophysiology, behavioral assays, and field validation within a confirmed predator–prey relationship. We will adjust the wording in the Introduction and Discussion to more accurately reflect this current state of knowledge, acknowledging prior work while clarifying the added value of our study.

      On quantitative analysis and concentration choices for limonene in EAG and field experiments.

      EAG concentration gradients: The concentrations used in our EAG experiments (including the 1% and 10% v/v dilutions of (−)-limonene) were selected based on standard practices in insect chemical ecology and on previous studies investigating dose-dependent antennal responses to volatile compounds (e.g., Tang et al., 2024, Int. J. Biol. Macromol.). The goal was to determine whether L. equestris antennae are capable of detecting limonene across a range of concentrations, not to precisely match natural emission levels or to determine behavioral thresholds. Our data clearly show concentration-dependent antennal responses, establishing physiological sensitivity.

      Field spray concentration: We acknowledge that the concentration used in the field experiment (10% v/v limonene sprayed over 25 m²) does not represent the exact amount of limonene naturally emitted by bats. Natural odor plumes are highly complex; the diffusion, dilution, and persistence of volatiles depend on multiple factors (airflow, turbulence, temperature, humidity, vegetation structure, etc.). Accurately reconstructing such dynamics would require detailed quantitative measurements and possibly fluid-dynamic modeling, which were beyond the scope of this study. The aim of the field experiment was functional: to test whether limonene, as a single bat-associated volatile, could alter cricket calling behavior under semi-natural conditions, not to establish the concentration threshold for this effect. Therefore, we did not design experiments to determine the exact concentration at which crickets begin to respond. The positive result supports the ecological relevance of limonene as an avoidance cue, but we do not claim that the applied concentration matches natural levels. We will clarify this point in the revised Methods and Discussion sections and acknowledge that quantitative characterization of natural bat-odor compositions and their diffusion dynamics is an important direction for future research.

      On Figures 1C and 1D comparing bat body odor with air control rather than with snout secretions.

      We thank the reviewer for this suggestion. The comparison between bat body odor and snout secretions is indeed novel and informative, and we agree that it could help identify anatomical sources of active volatiles. However, the purpose of Figures 1C and 1D in the current manuscript is to answer a more fundamental question: whether bat body odor (as a whole) contains volatile components that elicit antennal responses in crickets, compared to an odor-free control. This establishes the basic phenomenon of olfactory detection. The identification of snout secretions as the primary source of body odor is addressed separately in Figure 2, using HS-SPME-GC-MS and PCA. In the revised manuscript, we will clarify this rationale in the Methods and Results sections to avoid confusion. We also note that the reviewer’s idea, directly comparing GC-EAD responses to snout secretions versus whole-body odor, is an excellent suggestion for future experiments and would further strengthen the source attribution.

      Revisions we will make:

      (1) In the Introduction and Discussion, we will adjust the wording to more accurately reflect the current state of knowledge on olfactory eavesdropping across the vertebrate-invertebrate divide, acknowledging prior work while clarifying the added value of our study.

      (2) In the Methods and Discussion, we will clarify the rationale for our concentration choices in the EAG and field experiments, acknowledging that our aim was functional (testing sufficiency) rather than determining quantitative thresholds.

      (3) In the Methods and Results, we will clarify the rationale for comparing bat body odor with air controls in Figures 1C and 1D, and note that the reviewer’s suggestion of comparing with snout secretions is an excellent direction for future work.

      We thank Reviewer #2 again for the thoughtful comments, which have helped us improve the manuscript.

    1. eLife Assessment

      The authors use convincing methodology to investigate the detachment and reattachment kinetics of kinesin-1, 2 and 3 motors against loads oriented parallel to the microtubule. The conclusions drawn from the valuable experiments as well as the overall interpretation of the results are fully supported by the presented data.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      Noell et al have presented a careful study of the dissociation kinetics of Kinesin (1,2,3) classes of motors moving in-vitro on a microtubule. These motors move against the opposing force from a ~1 micron DNA strand (DNA tensiometer) that is tethered to the microtubule and also bound to the motor via specific linkages (Fig 1A). Authors compare the time for which motors remain attached to the microtubule when they are tethered to the DNA, versus when they are not. If the former is longer, the intepretation is that the force on the motor from the stretched DNA (presumed to be working solely along the length of the microtubule) causes the motor's detachment rate from the microtubule to be reduced. Thus, the specific motor exhibits "catch-bond" like behaviour.

      Strengths:

      The motivation is good - to understand how kinesin competes against dynein through the possible activation of a catch bond. Experiments are well done and there is an effort to model the results theoretically.

      Weaknesses from original round of review:

      The motivation of these studies is to understand how kinesin (1/2/3) motors would behave when they are pitted in a tug of war against dynein motors as they transport cargo in bidirectional manner on microtubules. Earlier work on dynein and kinesin motors using optical tweezers has suggested that dynein shows catch bond phenomenon, whereas such signatures were not seen for kinesin. Based on their data with DNA tensiometer, the authors would like to claim that (i) Kinesin1 and kinesin2 also show catch-bonding and (ii) The earlier results using optical traps suffer from vertical forces, which complicates the catch-bond interpretation.

    3. Reviewer #2 (Public review):

      Summary:

      To investigate the detachment and reattachment kinetics of kinesin-1, 2 and 3 motors against loads oriented parallel to the microtubule, the authors used a DNA tensiometer approach comprising a DNA entropic spring attached to the microtubule on one end and a motor on the other. They found that for kinesin-1 and kinesin-2 the dissociation rates at stall were smaller than the detachment rates during unloaded runs. With regard to the complex reattachment kinetics found in the experiments, the authors argue that these findings were consistent with a weakly-bound 'slip' state preceding motor dissociation from the microtubule. The behavior of kinesin-3 was different and (by the definition of the authors) only showed prolonged "detachment" rates when disregarding some of the slip events. The authors performed stochastic simulations which recapitulate the load-dependent detachment and reattachment kinetics for all three motors. They argue that the presented results provide insight into how kinesin-1, -2 and -3 families transport cargo in complex cellular geometries and compete against dynein during bidirectional transport.

      Strengths:

      The present study is timely, as significant concerns have been raised previously about studying motor kinetics in optical (single-bead) traps where significant vertical forces are present. Moreover, the obtained data are of high quality and the experimental procedures are clearly described.

    4. Reviewer #3 (Public review):

      Summary:

      Several recent findings indicate that forces perpendicular to the microtubule accelerate kinesin unbinding, where perpendicular and axial forces were analyzed using the geometry in a single-bead optical trapping assay (Khataee and Howard, 2019), comparison between single-bead and dumbbell assay measurements (Pyrpassopoulos et al., 2020), and comparison of single-bead optical trap measurements with and without a DNA tether (Hensley and Yildiz, 2025).

      Here, the authors devise an assay to exert forces along the microtubule axis by tethering kinesin to the microtubule via a dsDNA tether. They compared the behavior of kinesin-1, -2, and -3 when pulling against the DNA tether. In line with previous optical trapping measurements, kinesin unbinding is less sensitive forces when the forces are aligned with the microtubule axis. Surprisingly, the authors find that both kinesin-1 and -2 detach from the microtubule more slowly when stalled against the DNA tether than in unloaded conditions, indicating that these motors act as catch bonds in response to axial loads. Axial loads accelerate kinesin-3 detachment. However, kinesin-3 reattaches quickly to maintain forces. For all three kinesins, the authors observe weakly-attached states where the motor briefly slips along the microtubule before continuing a processive run.

      Strengths:

      These observations suggest that the conventional view that kinesins act as slip bonds under load, as concluded from single-bead optical trapping measurements where perpendicular loads are present due to the force being exerted on the centroid of a large (relative to the kinesin) bead, need to be reconsidered. Understanding the effect of force on the association kinetics of kinesin has important implications for intracellular transport, where the force-dependent detachment governs how kinesins interact with other kinesins and opposing dynein motors (Muller et al., 2008; Kunwar et al., 2011; Ohashi et al., 2018; Gicking et al., 2022) on vesicular cargoes.

    5. Author response:

      The following is the authors’ response to the current reviews.

      Reviewer #1 (Public review):

      I am not fully convinced about the responses from authors, so I would like to retain my original assessment of the paper. The same may be made available for public viewing, along with the responses of the authors. Readers can go through both and form their opinion.

      Unfortunately, this response from Reviewer 1 impacted the Assessment Statement but did not provide specific points for us to address. In the first round, the concerns of Reviewer 1 were: 1) the validity of the WLC prediction; 2) the claim that catch-bond measurements are generally made with superstall loads; 3) the role of vertical forces for dynein and a question about the orientation of the forces for kinesin; and 4) a request that we repeat the study using dynein. In rereading our responses to points 2-4 following our first revision, we felt that there were no unresolved issues around those points that affect our conclusions in any way. However, for point 1 regarding the validity of the WLC prediction, we had responded only in the reviewer response letter, and both reviewer 2 and the editors felt that there were points that we had addressed in the response letter that should be incorporated into the revised manuscript. Therefore, to clarify Reviewer 1’s question, we revised the text to address why we were justified to approximate the dsDNA force-extension curve using a WLC model with a 50 nm persistence length and why the precise shape of the force-extension curve has no impact on our conclusions.

      Reviewer #2 (Public review):

      The authors extensively entered into a scientific debate with the reviewers in their Response Letter. This led to a few changes and some (limited) new data in the manuscript. This is great and did improve the manuscript.

      However, in the view of this reviewer, (i) a significant number of responses fall short of actually addressing the concerns of the three reviewers (e.g. wrt using the same kinesin-1 neck-coil domains for all motors) and or (ii) a significant number of arguments now only occur in the response letter but not in the manuscript. The authors may check themselves critically for both. In principle, each longer discussion in the response letter warrants mentioning the appropriate facts and arguments in the main text of the manuscript.

      Based on this feedback, the first change we made was to rewrite the section justifying our choice of using a common coiled-coil dimerization domain for the three motors. Secondly, we went through our responses to all three reviewers to identify any instances where we either didn’t fully address the reviewer concerns or we provided arguments in the response letter but did not add corresponding text in the manuscript.

      Reviewer #3 (Public review):

      The authors attribute the differences in the behaviour of kinesins when pulling against a DNA tether compared to an optical trap to the differences in the perpendicular forces. However, the compliance is also much different in these two experiments. The optical trap acts like a ~ linear spring with stiffness ~ 0.05 pN/nm. The dsDNA tether is an entropic spring, with negligible stiffness at low extensions and very high compliance once the tether is extended to its contour length (Fig. 1B). The effect of the compliance on the results is not fully considered in the manuscript.

      In our first revision we added a paragraph in the ‘Geometry Calculations section of the Supplementary Methods addressing the dsDNA stiffness and comparing it to an optical trap. We considered moving this paragraph to the main text but decided against it because we felt it interrupted the flow of the Discussion. Instead, we expanded and clarified this paragraph to more specifically address the stiffness question. The paragraph with revised text now reads as follows:

      “Another consideration when comparing the DNA tensiometer to optical trap measurements is the relative stiffness of the trap and dsDNA. Optical traps stiffnesses are generally in the range of 0.05 pN/nm [13,14]. To calculate the predicted stiffness of the dsDNA spring, we computed the slope of theoretical force-extension curve in Fig. 1B. The stiffness is highly nonlinear and is <0.001 pN/nM below 650 nm extension. We compare motor performance under this low stiffness regime to the unloaded case in Fig. 3. In contrast, at the predicted stall force of 6 pN (960 nm extension), the dsDNA stiffness is ~0.2 pN/nm, which is stiffer than most optical traps, but it is similar to the estimated 0.3 pN/nm stiffness of kinesin motors themselves [13,14]. An 8 nm step at the 0.2 pN/nm stiffness of the dsDNA leads to a 1.6 pN jump in force and at the 0.05 pN/nm stiffness of an optical trap leads to a 0.4 pN jump in force; this is important because it means that in both cases the motors are likely dynamically stepping at stall. Because both experimental approaches allow for dynamic stepping at stall and because the stiffnesses of the instrument in both cases are less than the motor stiffness, there is no reason to expect that differences in stiffness between optical traps and the dsDNA spring lead to different motor detachment kinetics.”

      In the main text, we now address this compliance point in the ‘Comparison to previous work’ section of the Discussion:

      “stiffness differences are an unlikely explanation because at stall the stiffness of the DNA tether (~4 fold stiffer than optical tweezer) is still sufficiently low to allow for dynamic motor stepping at stall, and in any case it is still below the estimated motor stiffness (see Geometry Calculations in Supplementary methods).”.

      There were two points the reviewer felt we had sufficiently addressed. They were presented in the second review as a reiteration of the first review comments with a sentence appended, and are reproduced here. We added no new text based on these two points:

      In the single-molecule extension traces (Fig. 1F-H; S3), the kinesin-2 traces often show jumps in position at the beginning of runs (e.g. the four runs from ~4-13 s in Fig. 1G). These jumps are not apparent in the kinesin-1 and -3 traces. What is the explanation? Is kinesin-2 binding accelerated by resisting loads more strongly than kinesin-1 and -3? In their response, the authors provide an explanation of the appearance of jumps due to limited imaging speeds. The authors state that the qualitative difference in the kinesin-2 traces compared to the kinesin-1 and -3 traces may be due to the specific rebinding kinetics of kinesin-2.

      When comparing the durations of unloaded and stall events (Fig. 2), there is a potential for bias in the measurement, where very long unloaded runs cannot be observed due to the limited length of the microtubule (Thompson, Hoeprich, and Berger, 2013), while the duration of tethered runs is only limited by photobleaching. Was the possible censoring of the results addressed in the analysis? The authors addressed this concern by applying a Markov model to estimate the duration parameter.

      There was one final point from Reviewer 3 in the first round of reviews that we had addressed in the reviewer response (and that the reviewer was satisfied with), but we did not incorporate into the manuscript. Based on the suggestion from Reviewer 2 and the editors that we incorporate more from our responses to reviewers into the manuscript, we added new text on this point. That point (with the new sentence in the second review underlined), our response from first revision, and our response for this second revision are given below:

      The mathematical model is helpful in interpreting the data. To assess how the "slip" state contributes to the association kinetics, it would be helpful to compare the proposed model with a similar model with no slip state. Could the slips be explained by fast reattachments from the detached state? In their response, the authors addressed this question by explaining that a three-state model is required to model the recovery time distributions.

      In the model, the slip state and the detached states are conceptually similar; they only differ in the sequence (slip to detached) and the transition rates into and out of them. The simple answer is: yes, the slips could be explained by fast reattachments from the detached state. In that case, the slip state and recovery could be called a “detached state with fast reattachment kinetics”. However, the key data for defining the kinetics of the slip and detached states is the distribution of Recovery times shown in Fig. 4D-F, which required a triple exponential to account for all of the data. If we simplified the model by eliminating the slip state and incorporating fast reattachment from a single detached state, then the distribution of Recovery times would be a single-exponential with a time constant equivalent to t<sub>1</sub>, which would be a poor fit to the experimental distributions in Fig. 4D-F.

      Reviewer 3 noted that they were satisfied with our explanation of this point. However, based on Reviewer 2’s suggestion that we incorporate more of our responses into the text of the manuscript, we added the following clarification point in the model section of the Results:

      “We note that recapitulating the tri-exponential restart time distribution in Figure 4D-F required this slip/detached formulation and that lumping all events into a single detached state resulted in a single-exponential distribution of recovery times.”

    1. eLife Assessment

      This study characterizes a potentially targetable mechanism by which phosphate scarcity drives polymyxin B resistance in Enterobacteriaceae. The findings are important. While some aspects of the approach are very strong, particularly the diversity of techniques, it is recommended to include genetic controls and antibiotic resistance experiments in order to strengthen the evidence, which is currently solid. The clarity and presentation of the findings could also be improved.

    2. Reviewer #1 (Public review):

      This manuscript by Zhang et al addresses how Pi scarcity/depletion drives PMB resistance in Enterobacteriaceae, because it proposes a mechanistically distinct pathway from the better-known PhoBR-linked phospholipid-remodeling responses in other Gram-negatives. The authors also suggest an intervention strategy based on Mg repletion or Fe chelation. The results are substantial and include genetic analyses, mass spectrometry, reporter assays, phospho-signaling readouts, metal quantification, and comparative analyses across enterobacterial species.

      The paper reads well with the emphasis on the Mg loss followed by Fe mobilization during Pi depletion that induces PmrAB TCS activation for lipid A modification through transcriptional activation of ugd and arn genes. However, PmrAB is a well-known TCS responsible for PMB resistance through lipid A modification in the extensive studies by the Groisman lab. PmrA is a well-known transcriptional regulator to activate the transcription of the ugd gene in Salmonella and Yersinia by Mg depletion and Fe mobilization. Therefore, the current paper should focus more on the upstream signaling to connect the dots between Pi depletion and Mg loss. This is important because Ugd gene expression is not affected by PmrAB in Pi depletion. It should also be considered that Mg loss is temporally associated with Fe mobilization, but the manuscript does not quantitatively show that Mg dissociation/redistribution is sufficient to trigger Fe mobilization in the absence of Pi depletion, considering that Mg is a macronutrient, whereas Fe is a micronutrient.

      Second, the relationship between arn and ugd regulation needs a clearer mechanistic resolution to orchestrate the synthesis of the L-Ara4N during Pi depletion, because the manuscript shows that arn activation is PmrAB-dependent, whereas ugd is only partially PhoBR-dependent and not dependent on PmrAB. Yet the current model and narrative treat the system as a unified "ugd-arn" output. This should be carefully addressed, given that Pi depletion and Mg depletion might trigger different signaling modules.

      Third, the manuscript argues that this is a "conserved" circuit in Enterobacteriaceae. The evidence for conservation is presently strongest in E. coli MG1655 and includes supportive observations in E. coli O157, one UTI strain by lipid A MS, several UTI isolates by killing assay, and S. Typhimurium for key phenotypes. No direct mechanistic validation is shown in other important genera belonging to Enterobacteriaceae, which include Klebsiella, Enterobacter, Citrobacter, Yersinia, Serratia, or other clinically important Enterobacteriaceae.

      Fourth, the reversal and translational claims are a bit stronger than the current evidence supports. The title and Abstract state that identifying and targeting the circuit reverses Pi depletion-driven PMB, and the manuscript suggests a pharmacological intervention framework based on Mg supplementation or Fe chelation. The actual intervention evidence is limited to in vitro killing assays under acute Pi-depleted minimal-medium conditions in E. coli and S. Typhimurium, without in vivo testing, in that the experiments are performed under an acute 3-hour starvation in MOPS medium, not in host-mimicking or infection-relevant environments. The reversal needs to be shown not only at the level of survival curves, but also by the quantitative MIC/MBC measurements.

      More importantly, the authors demonstrated that the signaling module upon Pi limitation in Enterobacteria differs from that in other Gram-negative bacteria such as Pseudomonads. However, they did not discuss why this difference would impact the life of Enterobacteria. The authors should consider the glycolytic pathways (i.e., EMP pathway for enterobacteria vs ED pathway for pseudomonads), in that the ED pathway requires less Pi, whereas the EMP pathway requires more Pi. It should be noted that Pi supply is highly limited in the natural environment for the free-living bacteria, rather than in the host environment for the commensals.

    3. Reviewer #2 (Public review):

      Summary:

      Using E. coli K-12 as a model system, the authors investigated how phosphate (Pi) depletion induces polymyxin resistance in Enterobacteriaceae, which notably lack the canonical phospholipid remodeling pathways commonly associated with phosphate starvation responses. They demonstrated that low-phosphate conditions promote L-Ara4N modification of lipid A, thereby enhancing polymyxin resistance. Proteomic analyses revealed significant upregulation of the arn operon and ugd under phosphate-limited conditions, and promoter activity assays further confirmed that both promoters are strongly induced during Pi depletion. Through gene deletion experiments, the authors showed that arn expression is regulated by the PmrAB two-component system, whereas ugd is controlled by PhoBR under low-phosphate conditions. Using ICP-MS analysis, they further found that phosphate limitation increases cell-associated Fe levels, and that reducing Fe availability abolishes PmrAB-dependent activation of the arn operon. Finally, the study demonstrated that Mg supplementation and Fe chelation can suppress polymyxin resistance, highlighting the critical role of metal homeostasis in phosphate depletion-induced antimicrobial resistance.

      Strengths:

      Overall, I found this study to be well conducted, with convincing results that strongly support the proposed model. Through comprehensive genetic analyses and detailed characterization of metal ion homeostasis and membrane lipid modifications, the authors uncovered a novel regulatory connection among Mg²⁺, Fe³⁺, and the PmrAB pathway, a key driver of polymyxin resistance. These findings are highly interesting and have important implications for understanding the evolution of the Fe-sensing PmrAB system, as well as the broader role of nutrient availability in shaping antibiotic resistance.

      Weaknesses:

      I did not identify any particular weaknesses.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript examines how phosphate limitation primes E. coli and Salmonella for defense against polymyxin antibiotics. Other environmental signals, such as altered levels of extracellular Mg or Fe, were previously shown to induce polymyxin resistance in Enterobacteriaceae, and phosphate limitation was known to augment polymyxin resistance in other organisms such as A. baumannii and P. aeruginosa; however, whether phosphate limitation boosted polymyxin resistance in Enterobacteriaceae was not known. This study shows that this indeed occurs, and the mechanism is distinct from that in A. baumannii and P. aeruginosa. The model proposed is: (1) low phosphate causes bacteria to jettison Mg to balance cellular P/Mg ratio, (2) extracellular Fe3+ associates with the cell envelope to replace Mg as LPS-bridging cation, and (3) envelope Fe3+ activates PmrAB, which mediates a transcriptional response leading to L-Ara4N modification of lipid A and protection from polymyxin B. Flooding with Mg or chelating the surface Fe3+ blocks the protective response to low phosphate in E. coli and Salmonella but not in P. aeruginosa despite Fe still mobilizing in the latter. The differential response between Enterobacteriaceae and P. aeruginosa is connected to the presence/absence of Fe-sensing motifs in the PmrB periplasmic domain.

      Strengths:

      The strengths of the study are the wide array of approaches used and the thorough characterization of a novel stress-response mechanism involving metal mobilization. Combined with the analysis of multiple bacterial families, the results clarify how different strategies have evolved to defend against polymyxins during phosphate starvation.

      Weaknesses:

      Controls are needed in some of the genetic experiments, namely complementation, to verify linkage of defective survival phenotypes to the genes mutated and to rule out protein stability defects for the PmrB variants tested. In addition, the generalizability of the metal mobilization feature of the model would be strengthened by examining media with differing metal composition. Claims about antibiotic resistance would be strengthened by data examining bacterial growth in the presence of an antibiotic.

    1. eLife Assessment

      This study used pupillometry to provide an objective assessment of a form of synesthesia in which people see additional color when reading numbers. It provides convincing evidence that subjective color ratings are matched by changes in pupil size that recapitulate brightness-mediated changes when exposed to the real color. The work provides a valuable contribution to the literature on both synesthetic perception and the use of pupillometry to probe perception and related psychological processes.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      Knowing that small pupil-size variations accompany brightness variations (even when these are illusory), the authors asked whether pupil constrictions would accompany the synesthetic perception of a brighter color (compared with a darker one), induced by the presentation of a black-white character. This grapheme-colour synesthesia is only experienced by few participants, sixteen of whom were enrolled in this study. The results reliably showed that a relative pupil constriction would "betray" the perception of a brighter color in these participants, while no such effect would be observed in control participants who were asked to report a color in association with each grapheme, even though they did not perceive any.

      Strengths:

      The main strength of the study lays in its combination of psychophysics (brightness ratings) and pupillometry, which allowed for showing clear-cut results.

      Impact:

      This work is likely to improve our understanding of synesthesia, providing a new tool to quantify the subjective sensations; an interesting potential extension would be using pupillometry for tracking changes over time of the synesthetic experiences, opening up the possibility to evaluate the importance of learning for this peculiar experience.

    3. Reviewer #2 (Public review):

      Synesthesia is a neurological condition where stimulation of one sensory channel leads to involuntary, automatic, and consistent experience of another, unrelated percept. For example, Sir Francis Galton (1880, Nature) famously described the robust tendency of some individual (synesthetes) to associate numerals with a distinct color. Ever since, synesthesia keeps attracting a broad interest in the cognitive neurosciences in light of its implications for the study of domains such as perception, consciousness, and brain connectivity, among others.

      Strauch, Leenaars, and Rouw measured pupil size in a group of 16 grapheme-color synesthetes and two matched control groups. The participants were presented with gray digits - that is, visual stimuli having identical physical properties in terms of brightness. Each participant subsequently rated the corresponding evoked color and brightness: unlike controls, synesthetes did so in a very consistent and reliable fashion. Accordingly, this was also shown in their pupils: despite the same objective luminance, digits associated with brighter percepts caused their pupils to constrict and digits associated with darker percepts caused their pupils to dilate more than controls. These results highlight how crossmodal correspondences are deeply rooted in synesthetes, and puts forward pupillometry as a particularly appealing biomarker for some phenomenological experience (at least those grounded in "brightness").

      Further strengths of the technique are its temporal resolution and its responsiveness to several constructs. Across several tasks, the authors show for example that responses to synesthetic light are somewhat slower than responses to real light (i.e., they are likely mediated), but at the same time faster than responses to mental imagery. The role of mental imagery can also be reasonably dismissed when considering the second feature of pupil size: its responsiveness to mental effort and cognitive load. The pupils tend to dilate with demanding, challenging tasks, and this was the case when control participants were asked to report the color of a digit for which they did not consistently experience a synesthetic association. The same task was, instead, seemingly effortless for synesthetes, again speaking in favor of the automaticity of number-color correspondences in their case.

      Overall, the findings by Strauch, Leenaars, and Rouw are highly significant for the field and likely to be impactful. The strength of their evidence, when accounting for the relatively small sample size and the inherent variability of both phenomenology (color perception and subjective reporting) and physiology (pupil size), is adequate and sufficiently convincing.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      The pupil traces in Figure3 (main results) are heavily pre-processed (per-participant demeaned), loosing any feature besides the effect of interest. As I argued in my first review, I worry that this format gives unrealistic expectations about the effect (the perception of dark/bright colors do not generate a net dilation/constriction of the pupil; perception-related modulations of pupil size are always relative and generally small compared to the numerous other effects registered in pupil size; these include a pupil dilation that is more prominent in the controls and that gets analyzed later on in the manuscript; I do not think that eliminating one of the effects of interests from a main results figure helps the reader understand the results). In the revised manuscript, the authors addressed this concern by adding a Supplementary Figure 4, where a more complete representation of the results is shown (traces from individual trials are baseline corrected and averaged, resulting in more informative timecourses). I would strongly recommend that Supplementary Figure 4 is brought to the main text (Figure 3 could be presented in Supplementary).

      We agree that it is important to counter unrealistic interpretations of the effect. However, figures in the main article are the ones that are depicting the effects. Instead, it seems that additional clarification on these effects is needed. First and foremost, Figure 3 in the main manuscript visualizes the core effect: pupil size reveals that synesthesia is a sensory process and the phenomenology of the synesthetic experience can be measured physiologically. Secondly, this allows to advance synesthesia (and phenomenology) research as a new and powerful method.

      No doubt, our effect is relative in nature (as almost any pupillometry, fmri, eeg effect etc.). Including variation that is unrelated to the effect would increase rather than decrease confusion, as individual differences (i.e., how the pupil of an individual responds irrespective of the synesthetic experience) are unmeaningful to the question we set out to answer. Individual variations in pupil response shape irrespective of synesthetic color brightness are removed in Figure 3 but still present in Supplementary Figure 4. Thus, Figure 3 is better suited to illustrate our core effect than Supplementary Figure 4, as individual average responses (illustrated on the right) cannot be meaningfully related to the core effect anymore, only the difference can be.

      At the same time, the reviewer is correct that this may, not so much among researchers as among a general audience, create the expectation that the pupil will always net dilate when experiencing a dark synesthetic percept. This is clearly not the case, but only over its counterfactual (i.e., not seeing that dark synesthetic percept). We now counter such an unrealistic expectation:

      “Note that the effects here are visualized as counterfactuals. So while the pupil dilated for dark relative to bright experienced colors in synesthetes, this does not mean that the pupil net dilates and constricts to dark and bright experienced colors relative to baseline, but only relative to the counterfactual (see Supplementary Figure 4 for net pupil size changes).”

      We updated the caption of Supplementary Figure 4 as follows:

      "Supplementary Figure 4: Pupil size change to graphemes, split by 0.5 reported color lightness (dark gray = low lightness; light gray = high lightness) without demeaning (i.e., removing the average pupil response shape in the 4s stimulus interval per individual irrespective of brightness perception). (…)"

      Responses to physical brightness modulations were only measured in the synesthethes group, not in controls. The authors point out that pupillary light responses have been thoroughly characterized in previous studies, and conclude that synesthethes' responses were in line with the expectations both in terms of amplitude and latency. However, as we are not dealing with standardized measurements, subtle differences in pupil reactivity across the two populations remain a possibility. I recommend that this possibility is mentioned in the discussion.

      We agree with the reviewer, if there were any differences in the PLR between the two groups, they must be minor given that the responses follow those reported in the literature so closely. Yet, subtle differences cannot be ruled out fully unless tested and it doesn’t hurt mentioning this in the discussion, which we now do as follows:

      Finally, pupil light responses in Block 2 were only assessed in synesthetes. While these closely match such of control populations [50,51], subtle between-group differences cannot be excluded and could ideally be assessed in future and replication work.

      Reviewer #2 (Public review):

      Synesthesia is a neurological condition where stimulation of one sensory channel leads to involuntary, automatic, and consistent experience of another, unrelated percept. For example, Sir Francis Galton (1880, Nature) famously described the robust tendency of some individual (synesthetes) to associate numerals with a distinct color. Ever since, synesthesia keeps attracting a broad interest in the cognitive neurosciences in light of its implications for the study of domains such as perception, consciousness, and brain connectivity, among others.

      Strauch, Leenaars, and Rouw measured pupil size in a group of 16 grapheme-color synesthetes and two matched control groups. The participants were presented with gray digits - that is, visual stimuli having identical physical properties in terms of brightness. Each participant subsequently rated the corresponding evoked color and brightness: unlike controls, synesthetes did so in a very consistent and reliable fashion. Accordingly, this was also shown in their pupils: despite the same objective luminance, digits associated with brighter percepts caused their pupils to constrict and digits associated with darker percepts caused their pupils to dilate more than controls. These results highlight how crossmodal correspondences are deeply rooted in synesthetes, and puts forward pupillometry as a particularly appealing biomarker for some phenomenological experience (at least those grounded in "brightness").

      Further strengths of the technique are its temporal resolution and its responsiveness to several constructs. Across several tasks, the authors show for example that responses to synesthetic light are somewhat slower than responses to real light (i.e., they are likely mediated), but at the same time faster than responses to mental imagery. The role of mental imagery can also be reasonably dismissed when considering the second feature of pupil size: its responsiveness to mental effort and cognitive load. The pupils tend to dilate with demanding, challenging tasks, and this was the case when control participants were asked to report the color of a digit for which they did not consistently experience a synesthetic association. The same task was, instead, seemingly effortless for synesthetes, again speaking in favor of the automaticity of number-color correspondences in their case.

      Overall, the findings by Strauch, Leenaars, and Rouw are highly significant for the field and likely to be impactful. The strength of their evidence, when accounting for the relatively small sample size and the inherent variability of both phenomenology (color perception and subjective reporting) and physiology (pupil size), is adequate and sufficiently convincing.

      Comments on revisions:

      I thank the authors for addressing all my comments in a satisfactory way. I think that the paper has improved, especially in terms of transparency of the reporting and clarity of the results.

      We thank R1, R2, and R3 for their very useful input to improve our manuscript.

    1. eLife Assessment

      This study presents valuable findings on the high prevalence of pain in women with polycystic ovary syndrome and its association with distinct future health risks across different racial groups. The evidence supporting the conclusions is compelling, utilizing a massive global dataset and rigorous propensity score matching to identify pain as a critical, yet underexplored, clinical marker. The work will be of interest to reproductive endocrinologists, medical biologists, and clinicians involved in the diagnosis and management of polycystic ovary syndrome.

    2. Reviewer #1 (Public review):

      Summary:

      This retrospective study provides a new data regarding the prevalence of pain in women with PCOS and its relationship with health outcomes. Using data from electronic health records (EHR), the authors found a significantly higher prevalence of pain among women with PCOS compared to those without the condition: 19.21% of women with PCOS versus 15.8% in non-PCOS women. The highest prevalence of pain was conducted among Black or African American (32.11%) and White (30.75%) populations. Besides, women with PCOS and pain have at least a 2-fold increased prevalence of obesity (34.68%) at baseline compared to women with PCOS in general (16.11%). Also, women with PCOS had the highest risk for infertility and T2D, but women with PCOS and pain had higher risks for ovarian cysts and liver disease. Regarding these results, authors suggested the critical need to address pain in the diagnosis and management of PCOS due to its significant impact on patient health outcomes.

      Strengths:

      The problem of pain assessment in PCOS patients is well described and authors provided a clear rationale selection of the retrospective design to investigate this problem.

      A large number of analyzed patient's records (76,859,666 women) and its uniformity increases the power of the study. Using the Propensity Score Matching makes possible to reduce the heterogeneity of the compared cohorts and influence of comorbid conditions.

      Analysis in different ethnic cohorts provides actual and necessary data regarding the prevalence of pain and its relationship with different health conditions that will be helpful for clinicians to make a diagnosis and manage the PCOS in women of different ethnicity.

      Assessment of risk of different health conditions as including PCOS-associated pathology as other common groups of diseases in PCOS women with or without pain allows to differentiate the risk of comorbid conditions depending on the presence of one symptom (pelvic or abdominal pain, dysmenorrhea).

      Weaknesses:

      The significant weakness of the study is the absence of Latin American cohort. Probably the White cohort includes Latin Americans or others, but results of the study cannot be extrapolated to particular White ethnicities.

      Comments on revised version:

      At present, I have no questions or recommendations for the authors, as they have exhaustively addressed the previous comments and incorporated the necessary corrections.

    3. Reviewer #2 (Public review):

      Summary:

      The study offers a thorough analysis of the prevalence of pain in women with polycystic ovary syndrome (PCOS) and its associations with health outcomes across various racial groups. Furthermore, the research investigates the prevalence of PCOS and pain among different racial demographics, as well as the increased risk of developing various conditions in comparison to individuals who have PCOS alone.

      Strengths:

      The study emphasizes pain as a significant comorbidity of PCOS, an area that is critically underexplored in existing literature. The findings regarding the increased prevalence of some of the diseases in the PCOS + pain group provide valuable direction for future research and clinical care. I believe physicians should incorporate pain score assessments into their clinical practice to improve patients' quality of life and raise awareness about pain management. If future research focuses on the mechanisms of pain, it would provide a better understanding of pain and allow for a focus on the underlying causes rather than just symptomatic management. The study also highlights the association between PCOS+pain and various comorbidities, such as obesity, hypertension, and type 2 diabetes, as well as conditions like infertility and ovarian cysts, offering a holistic view of the burden of PCOS.

      Weaknesses:

      Due to the nature of retrospective design, some data may not be readily available in the EHR system. Diagnosis of PCOS, pain is based on ICD codes, which may lead to misclassification and may not capture symptom severity or patient-reported experiences.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This retrospective study provides new data regarding the prevalence of pain in women with PCOS and its relationship with health outcomes. Using data from electronic health records (EHR), the authors found a significantly higher prevalence of pain among women with PCOS compared to those without the condition: 19.21% of women with PCOS versus 15.8% in non-PCOS women. The highest prevalence of pain was conducted among Black or African American (32.11%) and White (30.75%) populations. Besides, women with PCOS and pain have at least a 2-fold increased prevalence of obesity (34.68%) at baseline compared to women with PCOS in general (16.11%). Also, women with PCOS had the highest risk for infertility and T2D, but women with PCOS and pain had higher risks for ovarian cysts and liver disease. Regarding these results, the authors suggested the critical need to address pain in the diagnosis and management of PCOS due to its significant impact on patient health outcomes.

      Strengths:

      (1) The problem of pain assessment in PCOS patients is well described and the authors provided a clear rationale selection of the retrospective design to investigate this problem.

      (2) A large number of analyzed patient records (76,859,666 women) and their uniformity increases the power of the study. Using the Propensity Score Matching makes it possible to reduce the heterogeneity of the compared cohorts and the influence of comorbid conditions.

      (3) Analysis in different ethnic cohorts provides actual and necessary data regarding the prevalence of pain and its relationship with different health conditions that will be helpful for clinicians to make a diagnosis and manage PCOS in women of different ethnicities.

      (4) Assessment of the risk of different health conditions including PCOS-associated pathology as other common groups of diseases in PCOS women with or without pain allows to differentiate the risk of comorbid conditions depending on the presence of one symptom (pelvic or abdominal pain, dysmenorrhea).

      We would like to thank the Reviewer for their positive feedback on this manuscript. Pain assessment in women with PCOS is of paramount interest and because of a gap in this research area, we are trying to address it.

      Weaknesses:

      (1) Although the paper has strengths in methodology and data analysis, it also has some weaknesses. The lack of a hypothesis doesn't allow us to evaluate the aim and significance of this study.

      We would like to thank the Reviewer for their valuable feedback regarding the hypothesis of this study. We understand that the hypothesis may not have been written clearly under the objectives and we have corrected this in the formal revision.

      The primary hypothesis of this study is that women with PCOS experience a higher prevalence to pain (including dysmenorrhea, abdominal pain and pelvic pain) compared to women without PCOS, and this prevalence varies by racial groups. Our hypothesis aims to explore the relationship between PCOS and pain, the associated health risks, and the potential racial disparities in pain prevalence and long-term health outcomes. Additionally, we seek to assess the effect of treatment on reducing pain symptoms in women with PCOS. This study not only examines the immediate burden of pain but also investigates its long-term consequences, including risks of infertility, obesity, and type 2 diabetes.

      To enhance clarity for readers, we explicitly stated this hypothesis in the revised manuscript and have ensured that its connection to the study’s objectives is clearly articulated. We appreciate the Reviewer’s insights and have incorporated these refinements to strengthen the manuscript.

      (2) The exclusion criteria don't include conditions, that can lead to symptoms similar to PCOS: thyroid diseases, hyperprolactinemia, and congenital adrenal hyperplasia. Thyroid status is not being taken into account in the criteria for matching. All these conditions could occur as on prevalence results as on risk assessment.

      We would like to thank the Reviewer for highlighting the need to include these additional conditions that mimic PCOS. After excluding hypothyroidism, hyperprolactinemia, and adrenal hyperplasia from the PCOS and PCOS and pain cohorts, we observed that 7,690 patients (1.65%) with PCOS and 1,854 patients (1.36%) with PCOS were removed. Based on this observation, we added these three conditions to our exclusion criteria and reran all our analysis for disease for our resubmission. The manuscript, figures, and tables have been updated to reflect these exclusions. Additionally, we have added rationale for excluding these conditions to the Discussion. With these major changes to the analysis, we aim to improve transparency and provide more accurate results and precise interpretations of our findings to the field.

      (3) The significant weakness of the study is the absence of a Latin American cohort. Probably the White cohort includes Latin Americans or others, but the results of the study cannot be extrapolated to particular White ethnicities.

      We appreciate the Reviewer’s suggestion to include Latin American cohorts in this study. The TriNetX platform has both self-reported race and ethnicity demographic information. In Table 3 - Figure Supplement 5 and Table 4 - Figure Supplement 6 we include baseline demographic information for both race (Asian, Black or African American, Native Hawaiian or Other Pacific Islander, Other, White, and Unknown Race) and ethnicity (Not Hispanic or Latino, Unknown, and Hispanic or Latino). In this paper we focused our future health outcome sub-analysis on four self-reported race groups: Asian, Black or African American, Other (Native Hawaiian or Other Pacific Islander, Other, Unknown Race), and White. We agree that including Latin American cohorts in the analysis is essential to better understand the health disparities affecting this population. Future work to better define Latin American cohorts in EHR data would significantly aid our ability to investigate this further.

      (4) The authors didn't provide sufficient rationale for future health outcomes and this list didn't include diseases of the digestive system or disorders of thyroid glands, which can also cause abdominal pain.

      We appreciate the Reviewer comment and concern regarding additional rationale for future health outcomes. We originally chose to investigate general future health outcomes like disease of the digestive system, circulatory system, etc. These disease groups were selected based on being general and having high prevalence as future health outcomes for patients with PCOS and Pain.

      Our initial results highlight the prevalence of disorders of the digestive system (Figure 2). However, after considering the Reviewers comments and to further strengthen our analysis, we included the most prevalent digestive system disorder in our relative risk (RR) analysis. Gastro-esophageal reflux disease (GERD) was identified as the most prevalent future digestive condition for women with PCOS and Pain (13.5%). There was also a 10.5% prevalence in women with PCOS overall.

      We were not able to include the same analysis for thyroid dysfunctions as this condition is a part of our exclusion criterion. These updates have been incorporated into the revised manuscript to ensure clarity and completeness.

      Reviewer #2 (Public review):

      Summary:

      The study offers a thorough analysis of the prevalence of pain in women with polycystic ovary syndrome (PCOS) and its associations with health outcomes across various racial groups. Furthermore, the research investigates the prevalence of PCOS and pain among different racial demographics, as well as the increased risk of developing various conditions in comparison to individuals who have PCOS alone.

      Strengths:

      The study emphasizes pain as a significant comorbidity of PCOS, an area that is critically underexplored in existing literature. The findings regarding the increased prevalence of some of the diseases in the PCOS + pain group provide valuable direction for future research and clinical care. I believe physicians should incorporate pain score assessments into their clinical practice to improve patient's quality of life and raise awareness about pain management. If future research focuses on the mechanisms of pain, it would provide a better understanding of pain and allow for a focus on the underlying causes rather than just symptomatic management. The study also highlights the association between PCOS+pain and various comorbidities, such as obesity, hypertension, and type 2 diabetes, as well as conditions like infertility and ovarian cysts, offering a holistic view of the burden of PCOS.

      We sincerely appreciate the Reviewer’s insightful comments. We hope that our findings will encourage further research on the occurrence of pain in women with PCOS and that others will replicate our results to strengthen the evidence in this area. As noted in our introduction, there are currently no standardized abdominal pain score assessments specifically for women with PCOS. We hope that the findings from this study will contribute to efforts toward developing a standardized pain assessment for the PCOS community. In the meantime, further research across more diverse populations will be essential to build a more comprehensive understanding of this issue.

      Weaknesses:

      Due to the nature of the retrospective study, some data may not be readily available in the system. Instead of simply categorizing participants based on whether they experience pain, it would be more useful to employ a pain scale or questionnaire to better understand the severity and type of patients' pain. This approach would allow for a more thorough analysis of pain improvement following treatment with the three widely used medications for PCOS. Additionally, it would be beneficial for the authors to specify subtypes of the disease rather than generalizing conditions, such as mentioning specific digestive system disorders or mental health disorders. The lack of detailed analysis of specific disorders limits the depth of the findings. This may cause authors to make incorrect conclusions.

      We appreciate the Reviewer for highlighting the importance of categorizing pain levels experienced by women with PCOS.  However, there is currently no standardized pain assessment for abdominal pain, and therefore more research is required before such a classification can be made. Additionally, the electronic health record data we leveraged via the TriNextX platform does not include any pain scale data from unstructured notes. Despite these limitations, this study is an important step toward recognizing abdominal and pelvic pain in women with PCOS. Our findings indicate that women with PCOS report abdominal pain independent of digestive conditions such as irritable bowel syndrome— a condition often associated with pain in this population.

      We would like to thank the Reviewer for their thoughtful comment with respect to subtyping future health outcomes. To get at the most impactful future health outcomes affecting women with PCOS and Pain, we have included the top 5 most prevalent health outcomes associated with PCOS and Pain. Specifically, we included analysis for anxiety disorder, depressive episodes, essential hypertension, Gastro-esophageal reflux disease (GERD), and acute pharyngitis. We observed that 17.1%, 11.5%, 10.5%, 10.0% of patients with PCOS and 20.1%, 13.7%, 13.5%, 13.3% of patients with PCOS and Pain were at risk of developing anxiety, depression, acute pharyngitis, and GERD respectively. For our revision, we have included these 5 conditions in our PCOS, PCOS and Pain and self-reported race-stratified future health outcome relative risk (RR) analyses. The revised manuscript, figures, and tables all reflect these changes.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) I highly recommend checking all papers and supplements for misprints. There are a lot of missing spaces in the Introduction.

      We would like to thank the Reviewer for bringing this to our attention. We have carefully reviewed the manuscript and all supplementary materials and corrected formatting issues, including missing spaces and typographical errors throughout the Introduction and the rest of the document.

      (2) Supplementary Table 3: numbers from the first line in "%No PCOS" should be in "No PCOS"?

      We thank the Reviewer for bringing this error to our attention. We have identified the source of the problem and values have been added to the appropriate column.

      (3) Why for the matching authors use the categorical data for overweight/obesity and not the entire values? There are different stages of obesity that can be predominant in different cohorts and contribute to the results.

      We would like to thank the Reviewer for their insightful question. While TriNetX does have some BMI values for patient participants, this data is not included for all patients. For example, only 29-30% of women in the PCOS control and case cohorts have BMI recorded. Therefore, we focused on ICD codes for obesity instead to include as much data as possible.

      (4) What criteria were being used to determine hyperlipidemia and obesity? Were these criteria equal for all patients, or did they depend on ethnicity?

      We would like to apologize to the Reviewer for any confusion. The criteria to determine hyperlipidemia and obesity are ICD-10-CM codes as recorded in the TriNetX platform. The ICD-10-CM codes for obesity are E65-E68 and the ICD-10-CM code for hyperlipidemia is E78.5. Please also see the Methods section of this manuscript where all the ICD-10-CM codes are described.

      (5) The section material and methods should provide information regarding quality assurance checks and any steps to eliminate data suspected to be unreliable or invalid, to process missing data, consisting of data or claim duplicates. If quality assurance of data hadn't been conducted, it should have been noticed in the study limitations.

      We thank the Reviewer for this suggestion. We have revised the Methods section to explicitly describe the data quality assurance procedures inherent to the TriNetX platform. Specifically, we clarified that TriNetX applies standardized data mapping to controlled clinical terminologies (ICD, CPT, RxNorm), performs automated quality checks and excludes records that do not meet platform-defined standards.

      (6) It's not clear why the authors didn't include in the analysis the information regarding taking painkillers or anti-inflammatory drugs by patients. Maybe there is no such data in EHR. However, if the patient has some chronic inflammatory or autoimmune disease, she should be prescribed medication. I recommend specifying this issue in the section Material and Methods and/or study limitations.

      We would like to thank the Reviewer for this important suggestion. We have now clarified this point in the limitations section of the discussion. Specifically, we added text explaining that over-the-counter analgesics and anti-inflammatory medications are not reliably captured by EHR or within the TriNetX platform and therefore could not be evaluated in our analysis.

      (7) The authors should provide the Table or complete Supplementary Tables 2 and 3 with the parameters of patients used for matching.

      We apologize to the Reviewer for any confusion. The parameters used for propensity score matching are described fully in the Methods section of the paper. Table 2 – Figure Supplement 5 and Table 3 – Figure supplement 6 display baseline characteristics for patients before and after the 1:1 propensity score matching using these parameters. We have now also added the propensity score matching parameters to the table descriptions to provide fluidity and further clarification.

      (8) The authors found out that women with PCOS and pain have higher RR for ovary cysts and liver diseases compared to women with PCOS who have higher RR for infertility, obesity, and T2D. Discussion includes thoughts regarding a higher risk of ovary cysts and liver disease in women with PCOS and pain, but there is not any suggestion as to why women with PCOS and without pain have a higher risk of infertility, obesity, and T2D. If there is no data explaining this phenomenon, I recommend noting the need for additional research.

      We would like to thank the Revier for this helpful feedback. The Discussion section now includes deeper insights into the pathophysiology behind the two distinct PCOS phenotypes (PCOS overall vs. PCOS and Pain) and their differing risk profiles for future health outcomes.  Specifically, we note that while women with PCOS overall may be more metabolically driven (higher risk of infertility, obesity, and T2D), women with PCOS and Pain show a higher risk of ovarian cysts and liver disease. We clarify that these findings are observational and hypothesis-generating and emphasize the need for future longitudinal and mechanistic studies.

      (9) The authors suggested that systematic contraceptives, metformin, or spironolactone reduce pain in PCOS women. The reduction is significant, but the number of patients with beneficial effects is low (2.5-7.5%). Is it enough to recommend prescribing this medication not only for PCOS treatment but against pain?

      We thank the Reviewer for this important comment. We agree that although the reduction in pain diagnoses following treatment with COCPs, metformin, or spironolactone was statistically significant, the absolute proportion of patients experiencing benefit was modest. Our intention was not to recommend prescribing these medications solely for pain management, but rather to highlight that standard PCOS therapies may have additional benefits in reducing pain symptoms. We have clarified this point in the Discussion to emphasize that these findings are observational and hypothesis-generating, and that prospective studies are needed before these medications can be considered specifically for pain management in PCOS.

      Reviewer #2 (Recommendations for the authors):

      (1) Including a subtype analysis of specific diseases on digestive, respiratory, and mental health diseases rather than generalizing the system will enhance the content.

      We would like to thank the Reviewer for this helpful suggestion. In the revised manuscript, instead of the generalized disease systems we previously reported on, we have included analysis for the top 5 most prevalent conditions. Specifically, we included analysis for anxiety disorder, depressive episodes, essential hypertension, Gastro-esophageal reflux disease (GERD), and acute pharyngitis. We observed that 17.1%, 11.5%, 10.5%, 10.0% of patients with PCOS and 20.1%, 13.7%, 13.5%, 13.3% of patients with PCOS and Pain were at risk of developing anxiety, depression, acute pharyngitis, and GERD respectively.

      (2) Including the prevalence of dysmenorrhea among healthy populations would allow readers to better compare its impact on the lives of individuals with PCOS.

      We would like to apologize to the Reviewer for any confusion. The prevalence of dysmenorrhea for cases and control cohorts can be found in Table 2 – Figure Supplement 5 and Table 3 – Figure Supplement 6 before and after propensity score matching.

      (3) Introducing an analysis of age subgroups will provide readers with a clearer understanding of the prevalence of pain and specific diseases across different age groups.

      We would like to thank the Reviewer for this helpful suggestion. For this revision, we did a sub-analysis to explore the prevalence of PCOS and PCOS and Pain stratified by 10-year age groups. A barplot of these results can be found in Figure 4 - Figure Supplement 7.

      Thank you again to the Reviewers for the positive and constructive feedback for this manuscript. We have made the appropriate edits and changes to the final revisions of the manuscript.

    1. eLife Assessment

      Rickert and colleagues demonstrate that the host peptidoglycan-binding protein PGLYRP1 has both beneficial and detrimental effects on Bordetella pertussis infection in mice. Using a solid array of techniques, the study provides useful insights into how the peptidoglycan fragment tracheal cytotoxin alters host immune responses, dampening inflammatory responses later in B. pertussis infection. These studies indicate that release of peptidoglycan fragments with particular structures can be used by bacteria to modulate NOD1 versus NOD2 responses to their advantage.

    2. Reviewer #1 (Public review):

      Summary:

      The authors aim to demonstrate that PGLYRP1 plays a dual role in host responses to B. pertussis infection. PGLYRP1 signaling is known to activate bactericidal responses due to recognition of peptidoglycan. Through NOD1 activation and TREM-1 engagement, it appears PGLYRP1 also has immunomodulator activities. The authors present mouse knockout studies and gene expression data to illustrate the role of PGGLYRL1 in relation to B. pertussis peptidogylcan. Mice lacking PGLYRP1 had slightly lower pathology scores. When TCT peptidoglycan was removed from the bacteria, surprisingly IL23A, IL6, IL1B and other pro-inflammatory genes encoding cytokines increased. The relationship to TCT and PGLYRP1 suggest the pathogen uses this strategy to decrease immune activation. The authors when on to show the relationship between PGLRP1 and TREM-1 as mediated with PGN using various versions of peptidoglycan. The study presents multiple angles of data to back up its findings and demonstrates an interesting strategy used by B. pertussis to down-regulate innate responses to its presence during infection.

      Strengths:

      Use of knockout mice of the key factor being considered paired with isogenic B. pertussis strains to reveal the mechanism of immune modulation to benefit the bacteria. The authors used in vivo gene expression paired with in vivo assays to establish each aspect of the mechanism.

      Weaknesses:

      The main focus was on innate responses, but some analysis of antigen specific antibody responses could improve the impact of the findings.

      Comments on revised version.

      I have no further input to add.

    3. Reviewer #2 (Public review):

      Since its original discovery, the mechanistic basis for TCT-mediated pathogenesis of Bordetella pertussis has been a moving target and difficult to uncouple from confounding variables. The current study provides some exciting data that suggest PGLYRP-1 modulates host responses upon 'activation' by TCT. While there are some strengths associated with the unbiased approaches and collective data to support the claims associated with TCT and PGLYRP-1's function in this system, caution should be used when interpreting and extrapolating some the information provided. While many of the initial concerns were addressed, one concern remains: using whole, intact PG sacculi from other species for comparative studies with a fragment of released PG (i.e., TCT).

      Comments on revised version.

      I have no further comments.

    4. Reviewer #3 (Public review):

      Summary:

      This study evaluates the contributions of the mammalian PG-binding protein PGLYRP1 to Bordetella infection. The authors find potential roles for PGLYRP1 in both bacterial killing (canonical) and regulation of inflammation (non-canonical). While these are interesting findings and the idea that PG fragment release has differential impacts on infection depending on fragment structure, the study is ultimately limited by the lack of connection between the in vivo and in vitro experiments and determining the precise mechanism of how PGLYRP1 regulates host responses and bacterial fitness during infection requires further study.

      Strengths:

      (1) The combination of scRNAseq with in vitro and in vivo assays provides complementary views of PGLYRP1 function during infection.

      (2) The use of TCT-deficient B. pertussis provides a useful control and perturbation in the in vitro assays.

      Weaknesses/Areas for future study:

      (1) The study does not ultimately resolve the initial early versus late phenotype divergence. While the in vitro assays suggest explanations for their in vivo observations, further mechanistic links are lacking and necessary for the author's conclusions throughout. To state one example, what is the early and late infection phenotype of TCT- Bp in mice lacking PGLYRP1? RNAseq data is reported from these mice but there are no burden or pathology studies. Furthermore, what are the neutrophil phenotypes (NOD-1/TREM-1 activation) in vivo? And are they dependent on PGLYRP1 and/or TCT? This will be an important topic of future study, as noted by the authors in their response.

      (2) It is unclear whether or how the NOD1 and TREM-1 pathways interact.

      (3) Many of the study's conclusions rely on the use of HEK293 reporter lines in the absence of bacterial infection, which may not be physiologically representative.

      Comments on revised version.

      The authors have responded adequately to my comments.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      The authors aim to demonstrate that PGLYRP1 plays a dual role in host responses to B. pertussis infection. PGLYRP1 signaling is known to activate bactericidal responses due to recognition of peptidoglycan. Through NOD1 activation and TREM-1 engagement, it appears PGLYRP1 also has immunomodulator activities. The authors present mouse knockout studies and gene expression data to illustrate the role of PGLYRP1 in relation to B. pertussis peptidoglycan. Mice lacking PGLYRP1 had slightly lower pathology scores. When TCT peptidoglycan was removed from the bacteria, surprisingly IL23A, IL6, IL1B, and other pro-inflammatory genes encoding cytokines increased. The relationship to TCT and PGLYRP1 suggests the pathogen uses this strategy to decrease immune activation. The authors went on to show the relationship between PGLRP1 and TREM-1 as mediated by PGN using various versions of peptidoglycan. The study presents multiple angles of data to back up its findings and demonstrates an interesting strategy used by B. pertussis to downregulate innate responses to its presence during infection.

      Strengths:

      Use of knockout mice of the key factor being considered, paired with isogenic B.

      pertussis strains, to reveal the mechanism of immune modulation to benefit the bacteria. The authors used in vivo gene expression paired with in vivo assays to establish each aspect of the mechanism.

      Weaknesses:

      The main focus was on innate responses, and some analysis of antigen-specific antibody responses could improve the impact of the findings.

      The authors thank the reviewer for their careful reading of the manuscript. We agree that understanding the impact of peptidoglycan recognition in adaptive immunity, including antibody responses, would be beneficial. This is particularly apparent due to the pressing need for novel vaccination strategies for pertussis. To this end, we have modified the discussion section to highlight this and are embarking on detailed studies of the adaptive response generated with B. pertussis strains releasing alternative peptidoglycan structures.

      Reviewer #1 (Recommendations for the authors):

      (1) This reviewer is of the opinion that describing the PGLYRP1 as a "bactericidal protein" seems misleading. "To determine whether PGLYRP1 has bactericidal activity against B. pertussis, we performed in vitro and ex vivo killing assays." Bactericidal activity was measured in normal or knockout neutrophils, but this seems to say the PGLYRP1 itself is an antimicrobial peptide. It clearly plays a role in the response,e but it is a regulator and not a killing agent.

      We agree that ‘bactericidal’ is not the most accurate description and have revised the manuscript accordingly to be more accurate throughout results section 1.

      (2) PGN can induce IgM production. Antibody production of any type was absent from this study. Would IgA/IgM/IgG levels to B. pertussis or its TCT change due to PGLYRP1? To this reviewer, it would be good to use the serum and perform some ELISA analysis. It is also likely that T cell responses could be impaired, but that may be out of the scope of this manuscript, but could be acceptable to consider for future studies.

      The authors thank the reviewers for this suggestion. We have added text to the discussion section to highlight the importance and potential of this suggestion.

      (3) Please include sources of mice (vendors) and strain numbers for transparency.

      The authors have added the relevant detail to the methods section to address this valid concern.

      (4) Were female or male mice or both used?

      For PGLYRP1 vs BALB/c comparisons both male and female mice were used. These are presented as combined data. No discernible differences were noted between male and female mice following infection. For single cell RNA sequencing studies only female mice were used, to be consistent with the published pertussis mouse model and avoid sex-based complications in analysis. We have clarified these details in the text.

      (5) It appears B. pertussis was cultured in SSM or BG. What condition was used for the bacteria used for the mouse challenge? SSM or BG?

      For mouse studies, bacteria were grown on BG agar supplemented with 10% defibrinated sheep blood for 48 hours and inoculum prepared by suspending in PBS in accordance with our established protocols. For in vitro studies liquid cultures were grown to mid-log in SSM. This has now been clarified in the methods section

      (6) Are the raw RNAseq and scRNAseq reads deposited in SRA?

      Raw data has now been deposited in the Gene Expression Omnibus (GEO) under number GSE324217

      (7) Is the scRNAseq data from one mouse or a pooled set of mice? If pooled, were the individual mice barcoded?

      scRNAseq data was obtained from barcoded individual samples and replicates were pooled and integrated during analysis, but the individual mouse each cell came from is still noted in the downstream analysis. This is now clarified in methods.

      (8) Why were some studies done by aerosol and others were done by intranasal delivery?

      The authors thank the reviewer for careful reading of the manuscript which erroneously listed aerosol infections. All infections in these studies were intranasal. This has now been rectified in the text.

      Reviewer #2 (Public review):

      Since its original discovery, the mechanistic basis for TCT-mediated pathogenesis of Bordetella pertussis has been a moving target and difficult to uncouple from confounding variables. The current study provides some exciting data that suggest PGLYRP-1 modulates host responses upon 'activation' by TCT. While there are some strengths associated with the unbiased approaches and collective data to support the claims associated with TCT and PGLYRP-1's function in this system, caution should be used when interpreting and extrapolating some of the information provided. For instance, the amount and purity of TCT used in the studies are unclear, and the in vitro activity of PGLYRP1 on B. pertussis is questionable. Different mouse backgrounds are used for various assays throughout, and it is known that the PRRs vary in these systems, so the confounding variables are difficult to uncouple. Additional concerns include the types of statistical tests being performed to support some of the claims and the relevance of using whole, intact PG sacculi from other species for comparative studies with a fragment of released PG (i.e., TCT).

      We thank the reviewer for their insightful suggestions to improve the standard of our manuscript and for highlighting several important considerations regarding our interpretation of TCT mediated host responses. We have addressed the points made in the revised manuscript. In particular, we have amended the Methods section to include a description of the purification and quantification of tracheal cytotoxin. These additions clarify the dosing of TCT used throughout the manuscript. We have revised the Results and Discussion sections to avoid overstating the bactericidal activity of PGLYRP1 against B. pertussis and to more carefully describe in vitro observations. Our revised interpretation emphasizes the role of PGLYRP1 in modulating host immune responses. Additionally, we have clarified experimental design and strain usage descriptions in the Methods section. The reviewer provided valuable and insightful comments on the solubility and structure of muropeptides studies. In response, we have revised the Results and Discussion sections to acknowledge these differences and the limitations they pose. Further, we have removed conclusions regarding the specific role of the 1,6anhydro bond. The statistical analyses have been reviewed and validated as well as clarified throughout the manuscript and Methods and figure legends updated.

      We appreciate the reviewer’s comments and believe the revisions have improved the clarity and rigor of the manuscript while maintaining the central conclusions about how peptidoglycan recognition influences host inflammatory responses during B. pertussis infection.

      Reviewer #2 (Recommendations for the authors):

      Major Points:

      (1) The concentration, purity, etc. of TCT seems like it is entirely unknown. Only a couple of experiments actually state the amount used, and it's unclear how the author determined the concentration because this is not trivial. Given the long-standing concerns with purity and co-purifying contaminants, this issue is paramount and needs to be properly addressed.

      TCT was purified by HPLC in the Goldman lab (UNC). Concentration was determined by comparing the peak area of each preparation to a purified TCT standard quantified by amino acid analysis. We have added these details to the Methods and now report concentrations throughout the manuscript.

      (2) Related to the effects of bacterial PG, studies performed are comparing TCT (a muropeptide) to commercially acquired, insoluble PG sacculi from B. subtilis and S. aureus. One cannot make these comparisons. There are flaws in terms of solubility (one goes into solution, the other does not), the amount used, the molar concentrations, etc. The authors also state that these are non-1,6 anhydro PG samples. That is not true. They contain plenty of 1,6 anhydroMurNAc, the moiety just exists in a different form. Finally, B. subtilis PG is not just mDAP, it's amidated, which is known to have effects on host response(s).

      We thank the reviewer for this important critique. We agree that differences in solubility and structural composition between TCT and PG sacculi limit direct comparisons. We have revised the Results and Discussion to remove statements implying direct equivalence and instead frame these experiments as highlighting how structural and physical properties of PGN fragments influence PGLYRP1-mediated activation of TREM-1. We have also removed statements regarding the 1,6-anhydro bond which were not adequately supported.

      (3) The claim that PGLYRP-1 is bactericidal in vitro is not supported by the data. Figure 1G shows that 24 hours after incubation, there is no difference. The comparison is being made to BSA, which is much higher (possibly because they're catabolizing it?) and thus entirely inappropriate. All other data in Figure 1 suggest no effect in vitro. In fact, it's this reviewer's position that none of the studies in Figures 1G, H, and I are convincing and should be entirely excluded.

      The authors agree that language describing the bactericidal assays is not optimal and have made revisions. The text in this results section has been modified to more carefully describe bacterial killing assays and accurately describe the effects the data suggest, primarily removing claims of bactericidal effects. BSA was chosen as a control protein (concentration matched with PGLYRP1), based on published controls for PGLYRP bactericidal assays (Lu et al 2006, JBC) similar results were obtained with PBS (volume matched with PGLYRP1). Descriptions of Fig1G,H,I have been updated. Data in 1H demonstrates that TCT release does not protect against effects of PGLYRP1, despite free PGN inhibiting PGLYRP1 bactericidal activity in published literature, while 1I suggests that extracellular polysaccharides contribute to protection against PGLYRP1 activity, preventing a more bactericidal phenotype which were not observed in the earlier assays when B. pertussis retained its capacity to produce bps polysaccharide.

      (4) Histology studies are unclear, and the data presented do not support the claims. Not only are the methods and results text describing the analysis contradictory, but nowhere are the actual statistical tests supporting the claims that they are different provided. This might be an oversight, but based on the variation, I would be surprised if they were statistically significantly different if proper tests are being used.

      Significance for pathology scores were initially determined using 2-way ANOVA as we had 4 groups (WT&KO at 4&7DPI) providing p-values of 0.01 for WT vs KO at 7DPI and 0.003 at 4DPI. Following reviewers’ suggestions, we have reanalyzed these data using a Mann Whitney U test, which is more appropriate for comparisons between two groups. This analysis yielded p-values of 0.013 (4DPI) and 0.00316 (7DPI) respectively confirming that the observed differences remain statistically significant. Statistical methods are now described in the methods and figure legends.

      (5) The NOD reporter studies are not well controlled and should include a) mouse vs human for both NOD1 and NOD2; b) defined details in terms of how spent culture media was treated, amount of material normalized, etc., c) concentrations of all materials used.

      We appreciate the reviewer’s comments regarding the NOD reporter assays. In response: (a) We have clarified and articulated the murine/human NOD reporter assays and included both human and mouse NOD1, along with controls. (b) We have supplemented descriptions of how conditioned (spent) culture media were collected, processed, and normalized in the ‘Bacterial strains and infections’ and ‘Reporter Cell Assays’ methods sections; (c) and the final concentrations of all agonists and test materials used in the reporter assays are now specified in the Methods and corresponding figure legends. Together, these additions address the requested controls and clarify the experimental conditions

      (6) The scRNA-seq studies are provocative and informative, but the data shown are selectively included for the purposes of the paper. This is justified in terms of 'telling a story', but it's a disservice to the community not to include all the raw data attained. These should be deposited in an open-source system.

      The complete dataset has now been deposited in GEO (GSE324217) enabling full access for the community. The analyses presented in the manuscript focus on the datasets most relevant to the central conclusions.

      Minor points:

      (1) The authors refer to arthropod PGRPs but call them PGLYRPs. It is best to stick with the established nomenclature and use the proper names to distinguish each. There are a few sentences in the abstract that don't make sense as they're written.

      The authors thank the reviewer for their careful reading of the manuscript and have altered the manuscript to use PGRP for arthropod peptidoglycan recognition proteins.

      (2) The reciprocal result of bacterial burden at different time points in the context of PGLYRP-1 production in mice could be simply explained - it is bactericidal early, and the accumulation of dead/dying bacteria releases large pieces of PG that are not released during growth (anhydro) but rather lysis. It is the latter that causes the inverse relationship later.

      The authors believe this is an interesting and plausible explanation for differences in responses at different stages of disease. Further, we believe that elucidating the mechanism by which ‘large pieces of PG not released during growth” are recognized differently than PG from lysed bacteria is worthwhile. We speculate that the release of TCT could be a mechanism by which B. pertussis takes advantage of host differences in PG recognition. We thank the reviewers for this thought and have included this possible interpretation in the text.

      (3) The results section references Figure 1G while discussing results presented in Figure 1H.

      This has now been corrected.

      Reviewer #3 (Public review):

      Summary:

      This study evaluates the contributions of the mammalian PG-binding protein PGLYRP1 to Bordetella infection. The authors find potential roles for PGLYRP1 in both bacterial killing (canonical) and regulation of inflammation (non-canonical). While these are interesting findings and the idea that PG fragment release has differential impacts on infection depending on fragment structure, the study is limited by the lack of connection between the in vivo and in vitro experiments, and determining the precise mechanism of how PGLYRP1 regulates host responses and bacterial fitness during infection requires further study.

      Strengths:

      (1) The combination of scRNAseq with in vitro and in vivo assays provides complementary views of PGLYRP1 function during infection.

      (2) The use of TCT-deficient B. pertussis provides a useful control and perturbation in the in vitro assays.

      Weaknesses:

      (1) The study does not ultimately resolve the initial early versus late phenotype divergence. While the in vitro assays suggest explanations for their in vivo observations, further mechanistic links are lacking and necessary for the author's conclusions throughout. To state one example, what is the early and late infection phenotype of TCT- Bp in mice lacking PGLYRP1? RNAseq data are reported from these mice, but there are no burden or pathology studies. Furthermore, what are the neutrophil phenotypes (NOD-1/TREM-1 activation) in vivo? And are they dependent on PGLYRP1 and/or TCT?

      (2) It is unclear whether or how the NOD1 and TREM-1 pathways interact.

      (3) Many of the study's conclusions rely on the use of HEK293 reporter lines in the absence of bacterial infection, which may not be physiologically representative.

      (4) The methods lack detail overall, and the experimental procedures should be described more concretely, especially for the scRNAseq datasets.

      We thank the reviewer for their comprehensive and fair assessment of our study and for highlighting both its strengths and areas where clarification could improve the manuscript. As noted in the review the possibility that peptidoglycan fragment structure impacts disease pathogenesis is interesting and the role of PGLYRP1 in regulating host and bacterial fitness during infection requires further study.

      We have addressed the points made by the reviewer in the revised manuscript. We edited the Methods section to provide additional experimental detail, particularly for the scRNA-seq analyses and reporter assays. We also clarified the experimental design and interpretation of the in vitro studies to avoid overstating mechanistic conclusions.

      Studies with TREM-1 and NOD are attempting to assess multiple aspects of PGN/PGLYRP mediated enhancement of inflammatory responses via NFkB/MAPKs. No attempts have been made to assess synergistic, overlapping or compensatory effects between these systems. Other work from our group highlights the role of peptidoglycan in driving inflammatory responses via NOD receptors (doi: https://doi.org/10.1101/2025.08.08.669383) and TREM-1 (doi: 10.1128/IAI.00126-21). Work in this paper assesses the contribution of these pathways to the observed immune modulation noted by PGLYRP1.

      We have clarified figure legends and analyses, including interpretation of neutrophil transcriptional programs identified in scRNAseq datasets and comparisons to known neutrophil phenotypes.

      We appreciate the reviewers feedback and the opportunity to improve the clarity of our manuscript and optimize the conclusions and central findings.

      Reviewer #3 (Recommendations for the authors):

      (1) Please clarify in Figure 1C what the axis means, since the text refers to both uninfected and infected cells. What data allow the conclusion that PGLYRP1 expression "expanded" to other cell subsets?

      We thank the reviewers for catching this oversight. We were relying on data which we had not best represented in Figure 1C, so we updated this figure and corresponding text so that this violin plot demonstrates increased PGLYRP1 expression levels and an increasing or expanding number of cell types following infection. This is now also reflected in the text. Expression of PGLYRP1 is apparent in more cell types and to a greater extent following infection (red) with B. pertussis compared to PBS challenge (black). Expression represents normalized and transformed unique molecular identifier counts per gene per cell.

      (2) Please revise the Figure 1 legend to match the Figure panels, and mention the time point of the mPGLYRP1 killing assay in 1H/I. Were these assays performed at 6 or 24 hours? This could affect the interpretation of the data.

      This has been revised to reflect timing of data.

      (3) The text at the end of the first Results section is overstated, as the data in Figure 1 do not relate to immune-mediated clearance apart from expression levels.

      This text has been revised and reference to immune mediated clearance removed

      (4) More detail is needed in the explanation of Figures 3E-G. Do the neutrophil subsets correspond to known subsets from the literature?

      When we overlaid established neutrophil signatures from the literature onto our dataset the NOD2+ neutrophils most closely resembled inflammatory or activated neutrophil programs described previously (Xie et al. 2020 Nat. Immuno., Veglia et al. 2021 J. Exp. Med)- specifically, high il1a, Ccl3 and Ptgs2 expression. In contrast, NOD1+ neutrophils showed greater overlap with resolving or regulatory neutrophil states- including genes associated with lipid mediator metabolism and NFkB dampening. Importantly, the clustering itself was not driven by NOD1 or NOD2 expression alone. NOD expression segregated within transcriptionally distinct neutrophil programs that are consistent with previously described inflammatory versus regulatory subsets. We included descriptions of these inflammatory neutrophils and related them to previously identified neutrophil populations, supporting our findings and improving the representation and articulation of the single cell neutrophil data analysis. We deeply thank the reviewers for their help in improving this section.

      (5) The Methods section describes qPCR, but this is not presented in the Results.

      This has now been removed. We thank the reviewer for their careful and complete review of the manuscript.

    1. eLife Assessment

      This study provides a fundamental finding regarding the context-dependent roles of the JAK-STAT pathway (JSP) across different cellular compartments within the breast cancer microenvironment, supported by convincing evidence. The comments of the reviewers were sufficiently addressed.

    2. Reviewer #1 (Public review):

      Summary:

      In their manuscript, Zhou and colleagues present a detailed look at how the JSP functions differently in the various cells of a breast tumor. The authors have effectively shown that the JSP acts as a double-edged sword, as it helps T cells fight cancer but also allows tumor cells to grow and avoid ferroptosis. These findings are important because they identify a useful biomarker to predict how TNBC patients might respond to PD-1 inhibitors.

      Strengths:

      This work is important because it provides a clear explanation for the conflicting roles of the JSP in the tumor environment. The evidence is solid, as it combines data from thousands of patients with single-cell analysis and lab experiments to confirm the role of STAT4 in cancer progression and immunity.

      Comments on revised version:

      The authors made a significant effort to improve the manuscript. My comments were sufficiently addressed.

    3. Reviewer #2 (Public review):

      Summary:

      The JAK-STAT pathway (JSP) exhibits cell-type-specific functional heterogeneity in breast cancer. This study investigates the JSP in breast cancer and its response to anti-PD‑1 immunotherapy. JSP displays distinct cell‑type heterogeneity: it promotes malignant phenotypes and immunosuppression in tumor cells, while enhancing cytotoxicity and reducing exhaustion in T cells. Elevated JSP expression correlates with improved immunotherapy responses, especially in triple‑negative breast cancer. These findings highlight the paradoxical roles of JSP, indicating that broad inhibition may compromise anti‑tumor immunity.

      Strengths:

      The major strengths of this study include the comprehensive characterization JSP heterogeneity across epithelial, tumor, and T cells in breast cancer. The identification of JSP and STAT4 as predictive biomarkers for immunotherapy response, particularly in triple‑negative breast cancer, provides clinically relevant insights for patient stratification.

      Weaknesses:

      The corresponding content has been revised.

    4. Reviewer #3 (Public review):

      Summary:

      This multi-omics study by Zhou et al elucidates the context-dependent roles of the Janus kinase-signal transducer and activator of transcription (JAK-STAT) pathway (JSP) across different cellular compartments in the breast cancer tumor microenvironment. While bulk JSP activity is associated with a favorable prognosis, single-cell analysis reveals a paradoxical landscape: high JSP in T cells drives anti-tumor cytotoxicity and reduces exhaustion, whereas high activity in tumor epithelial cells promotes malignancy and immunosuppression via the MIF-CD74 signaling axis. The JSP score (immune-related) serves as a robust predictive biomarker for response to anti-PD-1 immunotherapy, particularly in triple-negative breast cancer (TNBC). Furthermore, the study identifies the STAT4/SLC47A1 axis as a critical mechanism through which tumor cells resist ferroptosis, facilitating disease progression. These findings suggest that broad JAK-STAT inhibition may be counterproductive in cancer therapeutics; instead, therapeutic success depends on precise modulation and carefully timed interventions to preserve its T-cell-associated functions. This study may inspire future studies to explore specific factors that selectively modulate JAK-STAT activity in immune cells to achieve favorable therapeutic outcomes.

      Strengths:

      Significant therapeutics implications

      Weaknesses:

      Limited molecular mechanisms

      Comments on revised version:

      The authors have addressed my comments

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This multi-omics study provides a comprehensive characterization of the context-dependent roles of the JAK-STAT pathway (JSP) across different cellular compartments within the breast cancer microenvironment. The authors present convincing evidence that high JSP activity paradoxically drives anti-tumor cytotoxicity in T cells but promotes malignancy and immunosuppression in tumor epithelial cells, leading to the fundamental discovery that broad JAK-STAT inhibition could be therapeutically counterproductive. Ultimately, the identification of the immune-related JSP score and the STAT4 axis as predictive biomarkers for anti-PD-1 immunotherapy response, particularly in triple-negative breast cancer, offers critical insights for precise patient stratification and targeted therapeutic interventions.

      We greatly appreciate the editor’s insightful and comprehensive summary of our study.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In their manuscript, Zhou and colleagues present a detailed look at how the JSP functions differently in the various cells of a breast tumor. The authors have effectively shown that the JSP acts as a double-edged sword, as it helps T cells fight cancer but also allows tumor cells to grow and avoid ferroptosis. These findings are important because they identify a useful biomarker to predict how TNBC patients might respond to PD-1 inhibitors.

      We highly appreciate Reviewer #1’s generous comments and thorough understanding of our study.

      Strengths:

      This work is important because it provides a clear explanation for the conflicting roles of the JSP in the tumor environment. The evidence is solid, as it combines data from thousands of patients with single-cell analysis and lab experiments to confirm the role of STAT4 in cancer progression and immunity.

      Weaknesses:

      However, there are areas for improvement in the scope of the review, the depth of analysis, and the potential for broader clinical implications. The authors are encouraged to address these issues to enhance the scientific and clinical impact of the study.

      We greatly appreciate the positive recognition and insightful comments from the reviewer. We are grateful that you acknowledge our solid evidence and the significance of clarifying the dual roles of JSP and STAT4. We will fully address your suggestions to expand the research scope, deepen the analysis, and strengthen the clinical implications in the revised manuscript.

      Major Issues:

      (1) The authors demonstrate that STAT4 upregulates SLC47A1, but this is currently supported only by expression correlation and western blot data. To confirm a direct link, the authors are encouraged to perform ChIP-qPCR or luciferase reporter assays to show that STAT4 binds directly to the SLC47A1 promoter.

      We highly appreciate this insightful and important comment. Due to time constraints, the first author has left the laboratory for clinical practice, and this manuscript is critical for fulfilling his degree requirements at Sichuan University. We are making every effort to supplement additional mechanistic experiments where feasible. In the meantime, we have performed protein–nucleic acid docking analysis between STAT4 protein and the SLC47A1 promoter region, and the corresponding results have been added to the supplementary figures.

      (2) The conclusion that the MIF-CD74 axis drives immunosuppression is based on computational inference. To support this, the authors could consider mining publicly available breast cancer spatial transcriptomics data to show the co-localization of MIF and CD74. Alternatively, performing simple dual-color immunofluorescence staining on a few clinical sections would effectively demonstrate the physical proximity of these cells.

      We sincerely appreciate your careful review and valuable suggestions. We fully agree that the conclusion regarding the MIF-CD74 axis driving immunosuppression requires further spatial evidence. Although we plan to collect additional clinical specimens for direct co-localization validation, the related ethical approval is still ongoing and cannot be completed in a short time. Therefore, we have supplemented analyses on publicly available breast cancer spatial transcriptomics datasets, which now provide solid bioinformatic evidence to support the spatial co-localization and interaction of the MIF-CD74 axis in the tumor microenvironment in the revised manuscript.

      (3) TNBC is highly heterogeneous and includes subtypes like mesenchymal and immunomodulatory groups. The authors should analyze whether the JSP score or STAT4 levels vary significantly between these subtypes, as this could further refine the selection of patients for JAK1 inhibitors.

      Thank you for this insightful suggestion. We have supplemented the expression levels of JSP score and STAT4 in two independent TNBC cohorts to explore their heterogeneity across the four TNBC subtypes (Fig. S5B-C).

      (4) While the JSP score works well in the current datasets, the authors should consider validating its predictive accuracy in additional independent immunotherapy cohorts, such as the TONIC trial, to ensure the biomarker is robust across different treatment settings.

      We sincerely appreciate this valuable suggestion regarding the validation of the JSP score in independent cohorts. To address your concern about the robustness of our biomarker across different treatment settings, we would like to provide the following clarification and updates:

      Status of TONIC-trial Data Access:

      We fully recognize the significance of validating the JSP score in the TONIC-trial (Nat Med 2019; https://www.nature.com/articles/s41591-019-0432-4), a seminal study exploring immune induction strategies for PD-1 blockade in metastatic TNBC. We have made persistent efforts to obtain these data. However, our previous application to the Data Access Committee (DAC) of the European Genome-phenome Archive (EGA, Study ID: EGAS00001003535) was declined. The official reason provided was a restriction on data sharing imposed by the US Department of Justice, related to Executive Order 14117, which prohibits the transfer of bulk sensitive personal data to certain countries.

      Compensatory Validation in Available Anti-PD-1 cohorts:

      Despite the limitation on the TONIC-trial data, we have rigorously evaluated the predictive accuracy of the JSP score in two additional, independent, and publicly available anti-PD-1 treated breast cancer cohorts to thoroughly demonstrate its generalizability (Fig. S5A):

      GSE194040 (I-SPY2-990, Pembrolizumab, anti-PD-1): A cohort investigating anti-PD-1 therapy in metastatic breast cancer.

      GSE173839 (I-SPY2 trial, Durvalumab, anti-PD-L1): A cohort evaluating neoadjuvant anti-PD-L1 therapy in TNBC.

      We believe these additional validations adequately address your comment.

      Minor Issue:

      The manuscript mentions a U-shaped trajectory of JSP activity during tumor transition. A more detailed biological explanation of why the pathway activity initially drops and then rises would add depth to the discussion.

      We greatly appreciate this constructive comment. The JAK–STAT pathway (JSP) is essential for maintaining normal epithelial growth; its expression is higher in normal epithelium than in tumor tissues and increases during normal epithelial differentiation. In datasets containing both normal and tumor cell populations, JSP activity naturally declines during the transition from normal epithelium to early tumor lesions. In the subsequent tumor differentiation stage, JSP activity gradually rises, which may be driven by intrinsic tumor heterogeneity and pathway-dependency among different subtypes. This dynamic trend is consistent with JSP pathway activity score, which is independent of pseudotime cell trajectory analysis. We have added this explanation in the first paragraph of the Discussion.

      Reviewer #2 (Public review):

      Summary:

      The JAK-STAT pathway (JSP) exhibits cell-type-specific functional heterogeneity in breast cancer. This study investigates the JSP in breast cancer and its response to anti-PD‑1 immunotherapy. JSP displays distinct cell‑type heterogeneity: it promotes malignant phenotypes and immunosuppression in tumor cells, while enhancing cytotoxicity and reducing exhaustion in T cells. Elevated JSP expression correlates with improved immunotherapy responses, especially in triple‑negative breast cancer. These findings highlight the paradoxical roles of JSP, indicating that broad inhibition may compromise anti‑tumor immunity.

      Strengths:

      The major strengths of this study include the comprehensive characterization of JSP heterogeneity across epithelial, tumor, and T cells in breast cancer. The identification of JSP and STAT4 as predictive biomarkers for immunotherapy response, particularly in triple-negative breast cancer, provides clinically relevant insights for patient stratification.

      Weaknesses:

      The findings rely heavily on public dataset analyses.

      We sincerely appreciate the reviewer’s insightful recognition and comprehensive summary of our study, as well as the positive comments on our strengths.

      We fully agree that the current findings are mainly based on multi‑omics analyses of public datasets. In response to this comment, we have supplemented additional validation using independent cohorts (e.g., FUSCC‑TNBC and METABRIC) to reinforce the reproducibility of the cell‑type-specific heterogeneity of the JAK–STAT pathway and the predictive value of JSP/STAT4 for immunotherapy response in TNBC.

      Moreover, we have clearly discussed this limitation in the Discussion section and explicitly proposed further prospective experimental validation and clinical sample verification in our future work.

      We have carefully revised the manuscript in full accordance with all of your valuable suggestions to further improve the quality and rigor of our work.

      Reviewer #3 (Public review):

      Summary:

      This multi-omics study by Zhou et al elucidates the context-dependent roles of the Janus kinase-signal transducer and activator of transcription (JAK-STAT) pathway (JSP) across different cellular compartments in the breast cancer tumor microenvironment. While bulk JSP activity is associated with a favorable prognosis, single-cell analysis reveals a paradoxical landscape: high JSP in T cells drives anti-tumor cytotoxicity and reduces exhaustion, whereas high activity in tumor epithelial cells promotes malignancy and immunosuppression via the MIF-CD74 signaling axis. The JSP score (immune-related) serves as a robust predictive biomarker for response to anti-PD-1 immunotherapy, particularly in triple-negative breast cancer (TNBC). Furthermore, the study identifies the STAT4/SLC47A1 axis as a critical mechanism through which tumor cells resist ferroptosis, facilitating disease progression. These findings suggest that broad JAK-STAT inhibition may be counterproductive in cancer therapeutics; instead, therapeutic success depends on precise modulation and carefully timed interventions to preserve its T-cell-associated functions. This study may inspire future studies to explore specific factors that selectively modulate JAK-STAT activity in immune cells to achieve favorable therapeutic outcomes.

      Strengths:

      Significant therapeutic implications.

      Weaknesses:

      Limited molecular mechanisms.

      We sincerely appreciate the reviewer’s highly positive recognition and insightful summary of our work. Fully addressing your comment regarding limited molecular mechanisms, we have comprehensively supplemented and enriched the mechanistic elaborations in the revised manuscript—including detailed explanations of the dual cell-type-specific roles of the JSP pathway, the downstream MIF-CD74 axis, and the STAT4/SLC47A1-mediated ferroptosis resistance mechanism. All related revisions have been carefully incorporated into the text to strengthen the molecular depth and robustness of our findings.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) The Graphic Abstract in the current version fails to provide brief information about the submission.

      We appreciate your comment on the Graphic Abstract. We have redrawn a new, concise Graphic Abstract that clearly summarizes the key findings, workflow, and core message of our submission. The updated version now provides brief but complete information about the study.

      (2) Information regarding the epidemiology of breast cancer and TNBC is recommended to be included in the Introduction section.

      In response to your comment, we have supplemented up-to-date epidemiological data for both breast cancer and triple-negative breast cancer (TNBC) in the revised Introduction section.

      (3) Attention should be paid to the superscript, particularly for CD8+.

      We have revised the plus sign in CD4/8+ to the standard superscript format (CD8⁺) throughout the entire manuscript.

      (4) Typos are present, such as the error in "2.1" (please verify and correct accordingly).

      We have carefully checked and revised the entire manuscript, especially the section 2.1 Bioinformatical profiling. All typos, grammatical errors, and formatting inconsistencies pointed out in your comment have been fully corrected throughout the text.

      (5) Relevant information about MCF-10A cells in the cell culture protocol is missing.

      We sincerely apologize for the omission of MCF-10A cell culture details. We have supplemented the complete cell culture protocol for MCF-10A cells in 2.2.1 Cell culture.

      (6) For the Western blot experiments, information about the dilution ratios (of primary/secondary antibodies) is required.

      We have supplemented the detailed dilution ratios for all primary and secondary antibodies used in the Western blot experiments.

      (7) The Ethics Approval Number must be provided.

      We have supplemented the official ethics approval number for animal experiments in Section 2.2.6.

      (8) For the IHC staining experiments, information about the dilution ratios (of antibodies) is required.

      We have supplemented the detailed antibody dilution ratios for all primary antibodies used in the IHC staining experiments in Section 2.2.7 Immunohistochemistry (IHC).

      (9) Up-to-date citations are necessary, especially those published in 2026.

      We have thoroughly updated the reference list according to your suggestion in epidemiology of breast cancer.

      (10) Proofreading the language is recommended in order to enhance the fluency and readability of the manuscript.

      We have carefully polished the full manuscript with the help of a native English speaker to improve linguistic fluency, readability, and academic expression. All revisions have been completed strictly following your suggestions, and we deeply appreciate your efforts to help optimize this work.

      Reviewer #3 (Recommendations for the authors):

      Major points for the authors:

      (1) Please provide an overview figure of the datasets and approaches used in this study, as Figure 1.

      We sincerely appreciate your valuable suggestion. We have supplemented an overview figure (designated as Figure 1A) that systematically summarizes all datasets and experimental approaches used in this study, including the detailed workflow of bioinformatic profiling, pseudotime analysis, and functional validation.

      (2) The authors need to improve the organization of figure panels, as they appear cluttered in some regions, which impedes understanding of the figures.

      We sincerely appreciate your constructive comment. To address the cluttered figure panels that impeded understanding, we have redrawn Figures 2, 3, 5, and 6, and fine-tuned the image size, layout, and spacing of the panels.

      (3) The experimental section utilizes female mice for the MDA-MB-231 xenograft models. Given that a central finding of the paper is the pathway's role in T-cell-mediated anti-tumor immunity, the authors should discuss how the absence of a functional T-cell compartment in nude mice affects the interpretation of tumor growth data, or, ideally, provide data from immunocompetent syngeneic models.

      We thank the reviewer’s valuable comment. The MDA-MB-231 xenograft model in nude mice only supports our conclusion that STAT4 promotes tumor growth, given the deficient T-cell immune compartment in this model.

      We are currently constructing an orthotopic breast cancer model with stable STAT4 overexpression in 4T1 cells using immunocompetent mice, which possesses a complete immune microenvironment to further validate our immune-related findings. In addition, we plan to establish conditional STAT4 overexpression via the Cre/LoxP system in the MMTV-PyMT transgenic breast cancer mouse model. However, these elaborate in vivo validations cannot be completed within a short time frame due to experimental duration and technical limitations.

      This manuscript is critically important for the first author to complete their doctoral degree at Sichuan University. We sincerely appreciate the reviewer’s understanding and generous support for accepting our current data and future follow-up validation plans.

      (4) While the study links STAT4 to SLC47A1 upregulation, adding direct mechanistic evidence - such as ChIP-seq or luciferase reporter assays - would confirm that STAT4 directly binds the SLC47A1 promoter rather than acting through intermediary signaling.

      We highly appreciate this insightful and important comment. Due to time constraints, the first author has left the laboratory for clinical practice, and this manuscript is critical for fulfilling his degree requirements at Sichuan University. We are making every effort to supplement additional mechanistic experiments where feasible. In the meantime, we have performed protein–nucleic acid docking analysis between STAT4 protein and the SLC47A1 promoter region, and the corresponding results have been added to the supplementary figures.

      (5) Are there any potential upstream selective regulators of STAT4 in immune cells?

      IL‑12 acts as the upstream activator of STAT4 in immune cells. This cytokine binds to IL12R‑β1/β2, triggering Tyk2/Jak2 signaling to induce STAT4 phosphorylation, dimerization and nuclear translocation, thereby upregulating IFN‑γ transcription and enhancing T cell‑ and NK cell‑mediated antitumor immunity. We have added these details in the Discussion.

      (6) Recent studies have identified CD74+ lipid-associated macrophages (LA-MAMs) as a conserved niche in multi-organ metastasis of breast cancer. Linking the tumor-derived MIF-CD74 axis results to this broader metastatic framework could emphasize the clinical relevance of the findings.

      Recent study defines a conserved MIF-CD74 LA-MAM axis driving T-cell exhaustion and multi-organ metastasis in breast cancer, predicting poor patient survival. Our work further reveals that tumor-intrinsic JAK-STAT signaling reinforces this immunosuppressive cascade, while T-cell STAT4 activation reverses immune suppression. Combining MIF-CD74 blockade with precise STAT4-targeted strategies may synergize to remodel the metastatic niche and enhance immunotherapy efficacy in TNBC. We have supplemented the relevant mechanistic details and literature discussion in the revised Discussion section.

      Minor points for the authors:

      (1) The use of "spokesperson" to describe STAT4's role as a representative of the JAK-STAT pathway is somewhat informal for a scientific manuscript. Adopting more standard academic phrasing, such as "primary mediator" or "key transcriptional orchestrator," would enhance the professional tone.

      Thank you for your valuable comment. We have revised the manuscript accordingly by replacing the informal term "spokesperson" with the standard academic phrase "key transcriptional orchestrator".

      (2) The JSP score achieved a predictive AUC of 0.70-0.76. The authors could improve the work by testing whether combining the JSP score with existing clinical biomarkers, such as PD-L1 IHC or Tumor Mutational Burden (TMB), significantly enhances predictive accuracy.

      We have made every effort to collect publicly available breast cancer immunotherapy datasets for further validation. Unfortunately, none of these datasets provided immunohistochemistry (IHC) data for PD-L1/PD-1 expression. To address your valuable suggestion, we instead integrated mRNA expression levels of PD-L1/PD-1 with the JSP score to predict immunotherapy response.

      In cohorts GSE194040 and GSE173839 (Fig. S5A), this combined model exhibited improved predictive performance with an AUC exceeding 0.8, which is superior to using the JSP score alone. The corresponding results have been added and presented in the supplementary figures.

      (3) There is a potential contradiction in which bulk JSP scores correlate with better survival, whereas tumor-intrinsic JSP scores correlate with poor survival. A clearer discussion or a specific figure reconciling how the dominant immune signal overrides the pro-tumor signal in bulk analysis would be beneficial.

      In survival profiling, higher T-cells- and normal epithelial-specific JSP scores correlate with favorable patient survival, whereas elevated tumor-intrinsic JSP scores are associated with poor prognosis. This can be attributed to the predominant expression of JSP in T cells, which enhances T cell mediated anti-tumor immunity and counterbalances its pro tumor effects within cancer cells. We have added detailed clarification of this dual regulatory mechanism in the Discussion section.

      (4) The authors cite recent publications regarding the benefits of late-stage or intermittent JAK inhibition. Providing a more detailed proposed dosing schedule or "therapeutic window" based on their differentiation data could offer more actionable insights for clinical trial design.

      Based on the above clinical evidence and our findings, administering JAK–STAT inhibitors before or concurrently with immunotherapy may impair T‑cell cytotoxicity and disrupt normal epithelial differentiation in breast cancer patients. Instead, sequential delivery of JAK inhibition following immunotherapy represents a promising immune‑sensitizing strategy, particularly for the TNBC subtype. We have added corresponding descriptions in the third paragraph of the discussion section.

      (5) The authors note that they are unable to refine the analysis for TNBC subtypes, such as mesenchymal-like (MES), due to data limitations. If possible, using the METABRIC cohort (which was already accessed) to perform a secondary validation of JSP activity across these specific molecular subtypes would add significant depth.

      We appreciate this constructive suggestion. To address the subtype heterogeneity of JSP activity in TNBC, we have collected two TNBC datasets (FUSCC-TNBC and 2024_Nat.Comm.) and conducted further validation and analysis across different TNBC molecular subtypes in Fig. S5B-C.

      (6) The discussion evaluates both broad JAK inhibitors (Ruxolitinib) and STAT3-selective inhibitors (TTI-101). Explicitly comparing the potential biological impact of selective STAT3 inhibition versus selective STAT4 activation could clarify the most promising therapeutic direction.

      We greatly appreciate this valuable suggestion. We have supplemented the Discussion (in the penultimate paragraph) by proposing a translational strategy utilizing the specific cytokine IL‑12 to activate STAT4 for immune sensitization, while explicitly comparing the distinct biological effects and therapeutic directions between selective STAT3 inhibition and targeted STAT4 activation.

      In summary, we sincerely thank the editors and reviewers for their constructive comments and valuable suggestions. We have carefully addressed all the comments and revised the manuscript accordingly.

    1. eLife Assessment

      This study presents a valuable framework for the rational design of bacterial probiotics to protect against respiratory infections. The evidence supporting the central claim - that metabolic niche overlap predicts probiotic efficacy - is solid, combining innovative in vitro modeling with in vivo validation, though the model appears less effective for probiotics that rely on antimicrobial metabolite production.

    2. Reviewer #2 (Public review):

      Summary:

      This study aims to establish a rational framework for designing bacterial probiotics against respiratory infections. The central hypothesis is that in vitro antagonism, particularly through metabolic niche overlap with a pathogen, predicts in vivo efficacy.

      Strengths:

      (1) Systematic pipeline: The study integrates bacterial isolation, in vitro characterization, model development, and in vivo validation into a cohesive workflow.

      (2) Quantitative model: The introduction of the Niche Index (NI) and Niche Index Fraction (NIF) provides a novel, quantitative tool for predicting probiotic efficacy based on ecological principles.

      (3) Mechanistic insight: The work dissects different modes of action, clearly demonstrating that inhibition can be driven by specialized metabolite production (CP8) or carbon resource competition (e.g., CP7), with lactate utilization identified as a key factor.

      Weaknesses:

      (1) Limited model generalizability: The predictive power of the NI model is not universal. It fails to account for the in vivo inefficacy of CP8 (a metabolite-dependent inhibitor) and cannot explain the short-term protection conferred by some non-inhibitory CPs in vivo, suggesting unmodeled mechanisms like immune priming are at play.

      (2) Preliminary nature of key findings: The emphasis on lactate consumption as a critical predictor, while interesting, is not sufficiently explored to establish its general importance beyond the specific strains and conditions tested.

      Appraisal:

      The authors successfully achieve their aim of establishing a rational probiotic-design pipeline. The data robustly support the conclusion that metabolic niche overlap predicts efficacy for many strains, while also clearly delineating the model's limitations, as acknowledged by the authors.

      Impact:

      This work provides a valuable methodological framework for hypothesis-driven probiotic discovery. The quantitative Niche Index offers immediate utility to the field and, with further refinement, has the potential to become a fundamental tool for developing respiratory therapeutics.

      Comments on revised version.

      I thank the authors for their meticulous revisions.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      A summary of what the authors were trying to achieve:

      (1) Identify probiotic candidates based on the phylogenetic proximity and their presence in the lower respiratory tract based on phylogenetic analysis and on meta-analysis of 16S rRNA sequencing of mouse lung samples.

      (2) Predefine probiotic candidates with overlapping and competing metabolic profiles based on a simple and easy-to-applicable score, taking carbon source use into consideration.

      (3) Confirm the functionality of these candidate probiotics in vitro and define their mechanism of action (niche exclusion by either metabolic competition or active antibacterial strategies).

      (4) Confirm the probiotic action in vivo.

      Strengths:

      The authors attempt to go the whole 9 yards from rational choice of phylogenetic close lower respiratory tract probiotics, over in silico modelling of niche index based on use of similar carbon sources with in vitro confirmation, to in vivo competition experiments in mice.

      Weaknesses:

      (1) The use of a carbon source is defined as growth to OD600 two SD above the blank level. While allowing a clear cutoff, this procedure does not take into account larger differences in the preferences of carbon sources between the pathogen and the probiotic candidate. If the pathogen is much better at taking up and processing a carbon source, the competition by the probiotic might be biologically irrelevant.

      While the definition of carbon utilization in this work is a commonly used definition, we agree that there are numerous ways that one could define carbon utilization. We also agree that it is possible that inclusion of additional features of carbon consumption such as the order of prioritization of carbon sources by CP could improve the model. Our data in Figure 3H and 3I do suggest that certain carbon sources may be disproportionately important for predicting antagonistic phenotypes. However, given that the objective of this work was to develop a simple model to aid in the design of probiotic communities, we feel that the current definition of carbon utilization allows maximum accessibility and is suitable for our needs. Work is currently underway to identify additional features, such as carbon source processing efficiency, that may improve the model’s utility.

      (2) The authors do not take into account the growth of candidate probiotics in the presence of Bt. In monoculture, three of the four most potent candidate probiotics grow to comparable levels as Bt in LSM.

      Yes, our model only accounts for a one way interaction (effect of pathogen on CP). This is for two reasons (1) We are only interested in characterizing and modeling the antagonistic potential of the CP on the pathogen as this antagonism, we propose, is what gives a CP therapeutic potential (2) The degree to which co-culture with Bt impacts CP activity will be captured by the performed competition experiments and therefore any inhibition of the CP by Bt will be accounted for.

      While further investigation of the effect of Bt on the growth of the CP may not be necessary to achieve our objective, we agree that ecologically it would be interesting to understand this dynamic better. To explore this, we conducted co-culturing studies between each CP and Bt or media-only control and measured the amount of CP after 24 hours of co-culture. From the data it appears that only a small number of organisms (CP4, CP7 and CP19) are significantly inhibited by the pathogen at the 1:1 ratio tested. This result is perhaps unsurprising as these CPs have the highest niche index and therefore have a greater metabolic overlap with the pathogen.

      These data have been incorporated into Figure S1B and additional text has been added to line 157 and the methods at line 712.

      (3) Niche exclusion in vivo is not shown. Mortality of hosts after infection with Bt is not a measure for competition of CP with the pathogen. Only Bt titers would prove a competitive effect. For CP17, less than half of the mice were actually colonized, but still, there is 100% protection. Activation of the host immune system would explain this and has to be excluded as an alternative reason for improved host survival.

      We have revised the manuscript to address these issues as follows:

      (1) We include Bt titer data as suggested, displayed in a new figure (Figure 5F). The results indicate that CP8 fails to reduce Bt titers as compared to the no-CP control, whereas the other CPs tested (CP13, CP17, CP19, CP20, and CP26) do reduce Bt titers to statistically significant degrees (p-value < 0.05 by ANOVA/Tukey). These results support the idea that the CPs competitively exclude Bt in vivo (as they do in vitro), with the notable exception of CP8 (which competitively excludes in vitro but not in vivo, consistent with the mortality results). Further, additional spearman correlation analysis was performed to understand the relationship between the Niche Index value for a given CP and pathogen instantiation when pre-treatment with a given CP is performed. We found that there was a strong relationship between NI<sub>CP</sub> and pathogen load (r = -0.84, P<0.0001, 95% CI [-0.90 to -0.76], N = 77) such that prophylactic treatment with a CP with a high Niche Index value strongly correlated with lower pathogen load following Bt challenge. Text describing these findings has been added at line 471.

      (2) We include survival studies of mice prophylactically treated with non-viable CPs, displayed in a new figure (Figure S7). Viability is required for niche exclusion, so protection conferred by non-viable CPs must be due to other effects such as elicited immune responses. We found that non-viable CPs provide some protection when administered at 3 days prior to Bt challenge, though not to the same degree as viable CPs. Together, our data suggest that with the day 3 dosing schedule there are alternative mechanisms of protection (potentially including immune priming) that our current model does not capture. These results are described in further detail at line 460.

      Appraisal:

      (1) Based on phylogenetic comparison and published resources on lower respiratory tract colonizing bacteria, the authors find a reasonably good number of candidate probiotics that grow in LSM and successfully compete with the pathogenic target bacterium Bt in vitro.

      (2) In vivo, only host survival was tested, and a direct competition of CP with Bt by testing for Bt titers was not shown.

      Impact:

      Niche exclusion based on competition for environmentally provided metabolites is not a new concept and was experimentally tested, e.g. in the intestine. The authors show here that this concept could be translated into the resource-poor environment of the respiratory tract. It remains to be tested if the LSM growth-based competition data in vitro can be translated into niche exclusion in vivo.

      Reviewer #2 (Public review):

      Summary:

      This study aims to establish a rational framework for designing bacterial probiotics against respiratory infections. The central hypothesis is that in vitro antagonism, particularly through metabolic niche overlap with a pathogen, predicts in vivo efficacy.

      Strengths:

      (1) Systematic pipeline: The study integrates bacterial isolation, in vitro characterization, model development, and in vivo validation into a cohesive workflow.

      (2) Quantitative model: The introduction of the Niche Index (NI) and Niche Index Fraction (NIF) provides a novel, quantitative tool for predicting probiotic efficacy based on ecological principles.

      (3) Mechanistic insight: The work dissects different modes of action, clearly demonstrating that inhibition can be driven by specialized metabolite production (CP8) or carbon resource competition (e.g., CP7), with lactate utilization identified as a key factor.

      Weaknesses:

      (1) Limited model generalizability: The predictive power of the NI model is not universal. It fails to account for the in vivo inefficacy of CP8 (a metabolite-dependent inhibitor) and cannot explain the short-term protection conferred by some non-inhibitory CPs in vivo, suggesting unmodeled mechanisms like immune priming are at play.

      The NI model is not able to identify antagonism of metabolite-dependent inhibitors as their inhibitory activity is unrelated to the variables for which the model accounts. Based on the NI model, CP8 is predicted to have the least metabolic overlap with the pathogen which may explain its in vivo inefficacy. We do agree that short-term protection is only moderately related to NI (r = 0.48, P<0.0001, 95% CI [0.33 to 0.62], N = 115) and may represent an unmodeled alternative mechanism of protection as discussed at line 445, 466 and 523. We have added additional data in Figure S6 and corresponding text at line 444 which gives additional information about CP8 colonization in the context of infection.

      (2) Preliminary nature of key findings: The emphasis on lactate consumption as a critical predictor, while interesting, is not sufficiently explored to establish its general importance beyond the specific strains and conditions tested.

      Indeed, our model and assertions about critical predictors of antagonism only extend to the specific strains and conditions tested. While we cannot assert that lactate consumption is a critical predictor of antagonism universally, several other studies have indicated the importance of lactate in infection at other body sites [53-57].

      To further characterize the role of lactate utilization in the respiratory context, we performed an ex vivo experiment to measure lactate concentrations in respiratory tissue with or without treatment with a key isolate - CP19. After 24 hours of incubation, we found that lactate levels were significantly reduced in the CP19-containing homogenate compared to the PBS-only control (Figure S8A). Additionally, the pathogen was unable to grow in the CP19 conditioned homogenate but was able to grow in the untreated homogenate (Figure S8B). This indicates that CP19 can deplete the total lactate in lung tissue, and that this conditioning can inhibit pathogen growth in the lung tissue. These results are reported in a new supplementary figure (Figure S8) and summarized in corresponding text (line 485), with a description of the experimental procedure in the Methods section (line 924). While this does not prove our theory about the importance of lactate utilization universally, we believe that our work contributes to the growing body of evidence around lactate and its role in infection. Work is ongoing to expand the number of strains screened and determine the generalizability of particular carbon sources and their role in interbacterial antagonism.

      Appraisal:

      The authors successfully achieve their aim of establishing a rational probiotic-design pipeline. The data robustly support the conclusion that metabolic niche overlap predicts efficacy for many strains, while also clearly delineating the model's limitations, as acknowledged by the authors.

      Impact:

      This work provides a valuable methodological framework for hypothesis-driven probiotic discovery. The quantitative Niche Index offers immediate utility to the field and, with further refinement, has the potential to become a fundamental tool for developing respiratory therapeutics.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for improved or additional experiments, data or analyses.

      (1) CP titers at the end of the coculture experiment are missing in LSM.

      To quantify pathogen abundance after co-culture, cultures were plated on carbenicillin-100 to select for only colonies of the pathogen. As a result, no data about CP abundances were collected in the original experiments. However, we agree that ecologically it would be interesting to understand this dynamic better. We have added additional data about the impact of the pathogen on CP in co-culture to Figure S1B.

      (2) Bt titers in mice are essential to claim niche exclusion happens in vivo, and immune-mediated effects have to be excluded.

      Please see response to question 3 of the public review.

      (3) The definition of the use of carbon sources should be refined. Qualitative differences between the pathogen and the CP with regard to the usage of a given carbon source might have a substantial impact on the actual competitive effect.

      The definition of carbon utilization is stated at line 811. While we agree that there may be other carbon-consumption related variables (rate of growth on a particular carbon source, amount of biomass generation on that carbon source etc) that could be used in the model, for the purposes of this study a binary (can versus cannot grow on the carbon source) was sufficient. Work is currently ongoing to determine if metrics of growth on carbon sources such as those listed would improve the predictive capability of the model.

      Reviewer #2 (Recommendations for the authors):

      (1) Experimental & Analytical Suggestions:

      (a) To further validate the role of lactate, consider measuring lactate concentration in the airways of mice colonized by key CPs (e.g., CP7, CP19) versus controls. This would directly test if in vivo protection correlates with local lactate depletion.

      Unfortunately due to the funding for this project ending, we weren’t able to perform additional animal experiments. However, we were still able to test lactate utilization by CP19 in the respiratory context via an ex vivo experiment. We inoculated mouse lung homogenates with 10<sup>6</sup> CFU of CP19, or PBS as a negative control, and co-incubated for 24 hours. After 24 hours, we measured lactate levels and found that they were significantly reduced in the CP19-containing homogenate compared to the PBS-only control (A). Additionally, we measured the growth of the pathogen in CP19 conditioned (+CP19) and untreated (-CP19) homogenates and found that the pathogen was unable to grow in the CP19 conditioned tissues (B). This indicates that CP19 can deplete the total lactate in lung tissue, and that this conditioning can inhibit pathogen growth in the lung tissue. These results are reported in a new supplementary figure (Figure S8) and summarized in corresponding text (line 485), with a description of the experimental procedure in the Methods section (line 924).

      (b) The finding that CP8 provides no in vivo protection despite in vitro efficacy warrants further investigation. We suggest quantifying CP8 and Bt loads in co-colonized mice to determine if the probiotic fails to persist during infection or if the pathogen evades inhibition.

      Please see updated Figure S6 and accompanying text at line 444.

      (2) Quantitative Analysis:

      Please consider adding a brief justification in the manuscript explaining why the specific Niche Index formula (based on electron equivalents of shared carbon sources) was selected over alternative ecological metrics for quantifying niche overlap.

      Text was added starting at line 264 explaining our reasoning for choosing this model.

    1. eLife Assessment

      This study identifies apoptotic retinal ganglion cells as a potential source of ATP-mediated activation of PANX1 channels that initiate developmental retinal Ca²⁺ waves and coordinate microglial activation and vascular outgrowth with postnatal maturation. The work is important because it proposed an integrative framework linking programmed cell death, spontaneous neural activity, immune responses, and angiogenesis into a self-regulating developmental loop. The multimodal data are solid, but the mechanistic conclusions would be strengthened by complementary genetic approaches, such as PANX1 or BAX knockout models, to establish direct causality.

    2. Reviewer #1 (Public review):

      Summary:

      This study presents a potentially important integrative model linking spontaneous retinal waves, apoptosis, microglial activity, and vascular development during postnatal retinal maturation. Its significance lies in proposing a mechanistic framework that could reshape understanding of how neural activity and tissue remodeling are coordinated in the developing central nervous system. The evidence is strengthened by the use of multiple complementary techniques, including Ca++ imaging, high-throughput electrophysiology, transcriptomics, histology, and pharmacology.

      Strengths:

      (1) Multimodal Validation: The authors correlate large-scale functional imaging (calcium imaging and MEA) with high-resolution structural and molecular data (scRNA-seq and IHC), providing strong topographical evidence for the "centrifugal expansion" pattern.

      (2) The primary significance lies in identifying apoptotic Retinal Ganglion Cells (RGCs) as the physiological "pacemakers" for stage II retinal waves. By linking programmed cell death directly to neural activity and subsequent angiogenesis, the authors propose a self-regulating developmental loop.

      Weaknesses:

      (1) While the PANX1 pharmacological data provide compelling functional support, extending these conclusions to the broader CNS may be premature. Additional direct mechanistic validation would further strengthen the claim of causality.

      (2) While the manuscript beautifully illustrates the co-occurrence of events during retinal development, strengthening the distinction between correlation and direct causation would enhance the impact of the findings.

    3. Reviewer #2 (Public review):

      Summary:

      Savage et al. investigate the synchronization of retinal Ca2+ waves with developmental cell death, microglia activation, and vascular outgrowth. These developmental processes occur through a mechanism where apoptotic cells release ATP through Panx-1 channels to stimulate both Ca2+ retinal waves and microglia activation. Using scRNAseq, the authors classify autofluorescence cell clusters (ACCs) at the leading edge of vasculature outgrowth as Hmox-1+ microglia. From here, they show microglia engulfment of apoptotic RGCs, and the potential release of ATP may contribute to Ca2+ wave generation. The authors demonstrate these mechanisms through the use of two pharmacological agents to either block the ATP release from Panx-1 or block receptor binding to ATP. Furthermore, while previous studies have described the site of initiation of retinal Ca2+ waves as random, this study shows that the initiation of Ca2+ waves is biased to the leading edge of vascular growth in the developing retina. To do this, the authors use a combination of wide-field Ca2+ imaging and multi-electrode arrays to pinpoint the sites of Ca2+ wave initiation in the developing retina.

      Strengths:

      The authors use several techniques to interrogate these mechanisms, including single-cell RNAseq, wide-field Ca2+ imaging, and multi-electrode arrays. With these experiments, this manuscript proposes several novel ideas, such as ATP as the Ca2+ wave-initiating cue, and the localization of the Ca2+ wave initiation to the leading edge of vascular growth.

      Weaknesses:

      The main weakness of the manuscript is the overreliance on only two pharmacological agents to test the central hypotheses. These conclusions would be strengthened if, in addition to their pharmacological manipulations, they used genetic knockout models to perturb programmed cell death or ATP release (i.e., BAX-KO, Panx-1 KO).

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents a potentially important integrative model linking spontaneous retinal waves, apoptosis, microglial activity, and vascular development during postnatal retinal maturation. Its significance lies in proposing a mechanistic framework that could reshape understanding of how neural activity and tissue remodeling are coordinated in the developing central nervous system. The evidence is strengthened by the use of multiple complementary techniques, including Ca++ imaging, high-throughput electrophysiology, transcriptomics, histology, and pharmacology.

      Strengths:

      (1) Multimodal Validation: The authors correlate large-scale functional imaging (calcium imaging and MEA) with high-resolution structural and molecular data (scRNA-seq and IHC), providing strong topographical evidence for the "centrifugal expansion" pattern.

      (2) The primary significance lies in identifying apoptotic Retinal Ganglion Cells (RGCs) as the physiological "pacemakers" for stage II retinal waves. By linking programmed cell death directly to neural activity and subsequent angiogenesis, the authors propose a self-regulating developmental loop.

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Weaknesses:

      (1) While the PANX1 pharmacological data provide compelling functional support, extending these conclusions to the broader CNS may be premature. Additional direct mechanistic validation would further strengthen the claim of causality.

      We agree with the reviewer that the conclusions would be greatly solidified with more direct mechanistic validation. However, we are unable to conduct more experimentation as the grant is finished and the Sernagor lab is in the process of being shutdown, after the unexpected passing of the PI.

      In order to make clearer that this mechanism was found in retinal tissue, not CNS, we have moved any mention of the implications of our work to a broader CNS mechanism to the discussion section. We will add text into the discussion highlighting the need for more mechanistic investigation to uncover the full extent of the developmental processes described herein.

      (2) While the manuscript beautifully illustrates the co-occurrence of events during retinal development, strengthening the distinction between correlation and direct causation would enhance the impact of the findings.

      We have been clear to only present our findings as correlational as we were unable to fully explore the causational nature within the mechanisms presented. In the discussion, we have used published evidence and experimental papers to bolster our understanding of the causal aspects of this research. We will also include sections of text to address what experimentation is required to examine the causal interactions more directly.

      Reviewer #2 (Public review):

      Summary:

      Savage et al. investigate the synchronization of retinal Ca2+ waves with developmental cell death, microglia activation, and vascular outgrowth. These developmental processes occur through a mechanism where apoptotic cells release ATP through Panx-1 channels to stimulate both Ca2+ retinal waves and microglia activation. Using scRNAseq, the authors classify autofluorescence cell clusters (ACCs) at the leading edge of vasculature outgrowth as Hmox-1+ microglia. From here, they show microglia engulfment of apoptotic RGCs, and the potential release of ATP may contribute to Ca2+ wave generation. The authors demonstrate these mechanisms through the use of two pharmacological agents to either block the ATP release from Panx-1 or block receptor binding to ATP. Furthermore, while previous studies have described the site of initiation of retinal Ca2+ waves as random, this study shows that the initiation of Ca2+ waves is biased to the leading edge of vascular growth in the developing retina. To do this, the authors use a combination of wide-field Ca2+ imaging and multi-electrode arrays to pinpoint the sites of Ca2+ wave initiation in the developing retina.

      Strengths:

      The authors use several techniques to interrogate these mechanisms, including single-cell RNAseq, wide-field Ca2+ imaging, and multi-electrode arrays. With these experiments, this manuscript proposes several novel ideas, such as ATP as the Ca2+ wave-initiating cue, and the localization of the Ca2+ wave initiation to the leading edge of vascular growth.

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Weaknesses:

      The main weakness of the manuscript is the overreliance on only two pharmacological agents to test the central hypotheses. These conclusions would be strengthened if, in addition to their pharmacological manipulations, they used genetic knockout models to perturb programmed cell death or ATP release (i.e., BAX-KO, Panx-1 KO).

      We thank the reviewer for their insightful suggestions for further experimentation to bolster the research. Initially, we utilised pharmacological interventions as they provided acute and quick answering of the research question. At the outset of the research, we were not certain that purinergic release through PANX-1 channels was the mediator for the developmental mechanisms described. We tested a wide variety of specific agonists and blockers before seeing any profound effects on wave generation. These agonists and antagonists have been used before and are proven to deliver reliable results. In addition, since the ACCs had never been reported before we were unsure if a knockout animal would display the same anatomical phenotype. Furthermore, it is known that knockout mouse lines, especially connexin and hemichannel pores, do not lose function but rather have other isoforms or compensation mechanisms which can substitute the original function. For the retina, for example, it was shown that Cx36 can functionally replace Cx45 after Cx45 KO (Frank et al, 2010).

      We agree that while direct mechanistic validation would significantly reinforce the arguments, we are limited in conducting further experiments since the grant has been completed and the Sernagor lab is in the process of shutting down following her passing.

      In order to address the omission of mechanistic validation in the paper we will add text into the discussion highlighting the need deeper investigation in the causality of the developmental processes described herein.

      M. Frank et al., Neuronal connexin-36 can functionally replace connexin-45 in mouse retina but not in the developing heart, J. Cell Sci. 123, 3605 (2010).

    1. eLife Assessment

      This important study deepens our understanding of how populations of a given species may diverge in their molecular and physiological patterns as a result of adaptation to different thermal regimes. By approaching this question from multiple directions, the authors provide convincing evidence for adaptive changes in three strains of the diamondback moth after only three years of experimental evolution. This work will be of interest to anyone working on the response of pest species to environmental change and to workers on adaptive evolution in general.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Lei and co-workers aim to uncover the genetic underpinnings of thermal adaptation across three strains of the diamondback moth (Plutella xylostella) through experimental evolution over three years under three different thermal regimes. They identify systematic differences in trait responses (e.g., survival, fecundity), metabolic profiles, gene expression, and in the amino acid sequence of the PxSODC gene, among others. These results suggest that the diamondback moth has a strong potential for rapid physiological adaptation to different thermal regimes. Overall, this is a comprehensive and generally well-executed study that addresses an important question in the face of ongoing climate change.

      Strengths:

      The authors employ multiple approaches to identify signatures of thermal adaptation across the three strains, such as trait performance comparisons, metabolomics, transcriptomics, and amino acid sequence comparisons. All these different angles form a convincing picture of the underlying factors that underpin thermal adaptation in this experimental system. The manuscript is also generally well written and easy to understand.

    3. Reviewer #2 (Public review):

      Summary:

      In this paper, the authors set out to better understand the genetic mechanisms underlying thermal adaptation in insects. They experimentally evolved diamondback moth (Plutella xylostella) populations - a pest species with a wide distribution - under both hot (12h:12h 32{degree sign}C/27{degree sign}C) and cold (15{degree sign}C/10{degree sign}C) thermal conditions, and conducted phenotypic assays and metabolic and transcriptomic profiling to analyze how populations changed to deal with this thermal stress compared to the nonevolved ancestral population (constant 26{degree sign}C). Phenotypic assays showed that evolved hot populations had increased survival at high temperatures (42-43{degree sign}C) while evolved cold populations had lower freezing points compared to the ancestral population. When measured at the constant 26{degree sign}C conditions, metabolic and transcriptomic profiles of 3rd instar larvae from the evolved population were distinctive from the ancestral population, with a set of overlapping metabolic and transcriptomic pathways that were significantly differentially expressed in both hot and cold evolved populations compared to the ancestral. The authors narrowed down this set of candidate genes further by focusing on genes with high expression levels overall, whose expression profile was correlated with differentially expressed metabolites, and that contained mutants in both hot and cold strains. From this set, they chose the PxSODC gene for further functional validation, as it has previously been shown to be involved in the response of insects to abiotic stress with its antioxidative role in cellular defense. At the constant 26{degree sign}C, this gene showed lower expression across development in evolved strains compared to the ancestral population, while it showed similar expression patterns under thermal stress. Knockdown of PxSODC resulted in decreased survival rates at high temperatures and higher freezing points compared to the ancestral population. Based on this validation, the authors hypothesize that the non-synonymous mutation in the PxSODC gene that they found in the cold and hot evolved populations might alter the conformation of the PxSODC protein, increasing enzyme capacity. Their experimental evolution experiment furthermore indicates the capacity of the pest species, the diamondback moth, to adapt to a wide range of temperatures, providing insights into its capacity for global dispersal.

      Strengths:

      (1) The authors did a tremendous amount of work to characterize the mechanisms underlying thermal adaptation in the diamondback moth, artificially selecting populations for three years in the lab and characterizing how they evolved as a result at different biological levels: from phenotypes in different life stages, to larval metabolites and gene transcription, to functionally validating how one of the resulting gene candidates influences the capacity to deal with thermal stress.

      (2) The paper identifies and provides further evidence for candidate genetic mechanisms that might be particularly important for thermal adaptation in insects, including lipid metabolism, oxidoreductase activity, and DNA methylation. It is furthermore interesting that the authors found similar mechanisms to be involved in both the adaptation to cold and hot environments. Their functional validation of some of the genes involved in these mechanisms is very useful to understand how these genes might be causally involved in insect thermal adaptation.

      (3) The paper also has applied value: the diamondback moth is a pest species with a wide distribution, so understanding its adaptive capacity to different thermal environments is important for predicting the prevalence and potential further range expansion of this species under future climate change.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study deepens our understanding of how populations of a given species may diverge in their molecular and physiological patterns as a result of adaptation to different thermal regimes. By approaching this question from multiple directions, the authors provide solid evidence for adaptive changes in three strains of the diamondback moth after only three years of experimental evolution, and support the causal involvement of the PxSODC gene in thermal adaptation to both cold and hot temperatures. This work would benefit from more sophisticated phylogenetic analyses, better statistical support, and a more detailed discussion of the differences in the three strains at the pathway level.

      We sincerely thank the editors for this positive and constructive assessment. In the revised manuscript, we have addressed the highlighted points by: (1) re-inferring the phylogenetic tree of the PxSODC gene using a model-based Maximum Likelihood method (IQ-TREE) to ensure a robust evolutionary analysis; (2) substantially expanding the description of our statistical methods across all data types to ensure reproducibility and clarify multiple-testing corrections; and (3) adding a more detailed discussion of the pathway-level differences between the hot and cold strains, particularly integrating how their distinct transcriptomic responses align with their shared metabolic adjustments and phenotypic traits.

      Reviewer #1 (Public review):

      (1) The authors identify pathways that are enriched in different strain comparisons (Figure 3E), but do not provide a detailed interpretation of these results. It would be great if the authors could explain in more detail how the physiological processes of a cold-adapted strain of this species may differ from those of a warmer-adapted strain.

      We agree. We have addressed this by directly integrating our pathway enrichment results (Figure 3E) with the observed life-history phenotypes (concurrently addressing Reviewer 2's Comment 36a). We expanded the Discussion to explain that while both strains share convergent adjustments in core pathways (e.g., lipid metabolism for energy reallocation), their specific physiological strategies differ. The cold-adapted strain relies on broader transcriptional reprogramming to maintain homeostasis and support extended longevity/cold hardiness, whereas the hot-adapted strain utilizes broader metabolic rewiring to actively fuel its accelerated development and higher fecundity.

      (2) The authors reconstruct a phylogenetic tree of the PxSODC gene using the neighbor-joining algorithm. The limitations of this algorithm have been known for many years now, especially for sequences separated by long evolutionary distances. According to Wang et al. (2016), the last common ancestor of the species shown in Figure S4C occurred 392-350 million years ago. Given this, I would strongly recommend that the authors infer a phylogenetic tree using model-based methods, such as those implemented in RAxML-NG or IQ-TREE. Also, in the absence of a valid outgroup sequence, I would show the gene tree as unrooted or rooted based on the corresponding species tree.

      Agree. We have re-inferred the phylogenetic tree of the PxSODC gene using the model-based Maximum Likelihood (ML) method implemented in IQ-TREE. As recommended, in the absence of a valid outgroup sequence, the revised tree is now presented as unrooted. Supplemental Figure S4C (Figure 5-figure supplement 1C) and the corresponding text in the manuscript have been updated.

      (3) There is a key piece of the puzzle that is currently missing: the structural mechanism behind the mutational effects described in this study (e.g., Figure 5). The authors could leverage AlphaFold to generate structural models of different mutants and conduct molecular dynamics simulations to examine their conformational dynamics.

      We thank the reviewer for this excellent suggestion. We generated AlphaFold structural models of the wild-type (WT) and mutant (MU) PxSODC proteins and conducted 100 ns molecular dynamics (MD) simulations using GROMACS 2022.3 at three physiologically relevant temperatures: 15°C (cold stress), 26°C (favorable baseline), and 32°C (heat stress). Using 26°C as the physiological baseline, three key structural parameters support enhanced thermostability of the mutant protein (Figure 5–figure supplement 3). First, RMSD analysis revealed that under heat stress (32°C), the WT underwent severe conformational drift (RMSD increased from the 26°C baseline of 1.62 to 2.49, an increase of 0.87), while MU remained remarkably stable (from 1.59 to 1.66, an increase of only 0.07). Second, MU possessed a significantly more compact structure, with lower SASA values at 15°C (118.39 vs. 127.29 nm²) and 26°C (113.82 vs. 125.61 nm²), indicating optimized hydrophobic core packing. Third, the intramolecular hydrogen bond network of MU demonstrated dual stress resistance: under cold stress, MU actively increased hydrogen bonds from its baseline (113→119), whereas WT lost bonds (117→112); under heat stress, MU fully maintained its bond count (113→113). These results provide a direct structural mechanism for the enhanced catalytic efficiency of the mutant SOD at lower expression levels.

      Reviewer #1 (Recommendations for the authors):

      (4) The experimental evolution component of this study is described in the text as lasting for three years. It would help if the number of generations per strain were also reported.

      We have added the number of generations per strain. Over the three-year period, the hot strain completed ~75 generations and the cold strain ~15 generations. The ancestral strain was continuously maintained at 26°C throughout this period. The revised text has been updated in both the Introduction and Materials and Methods.

      (5) In Figure 3B: There is a typo in the word “Statistics”.

      Corrected. The typo in “Statistics” in Figure 3B has been fixed.

      (6) In Figure 3D: “CS” appears twice.

      Corrected. The duplicated “CS” label in Figure 3D has been replaced with the correct label.

      (7) Figure 4: This is not accessible to colorblind readers, who will clearly not be able to tell each color apart. As a non-colorblind person, I, too, have trouble figuring out which color label in panel B corresponds to which color in panel A. For example, I do not know off the top of my head how 'blue' differs from 'midnightblue', 'royalblue', or 'skyblue'. I recommend that the authors replace colors with identifiers, such as 'g1' for group 1 and so on.

      We appreciate this suggestion. We have replaced all color-based module labels with alphanumeric identifiers (M1, M2, M3, etc.) and added a corresponding legend. The main text and supplementary materials have been updated accordingly.

      (8) Lines 246-247: "Its secondary structure mainly consisted of strands, helices and coils." This sentence is redundant. These three are the only possible secondary structural elements, according to most bioinformatics tools such as PSIPRED, which the authors used. This sentence would be more useful if the authors could report the percentage breakdown of each secondary structural element.

      We have removed the redundant sentence and updated the text to report the specific percentage breakdown of the secondary structural elements based on our PSIPRED predictions (approximately 55.24% random coils, 16.19% alpha helices, and 28.57% extended strands). The revised text has been updated in the Results section.

      (9) Lines 260-261: "This suggests that the PxSODC gene can alter its expression pattern and function in response to environmental change...". I find this sentence a bit imprecise. Would it not be more precise to mention that the expression of this gene is regulated by temperature triggers?

      We agree that the original phrasing was imprecise. We have revised the sentence in the manuscript to state: “This suggests that the expression of the PxSODC gene is regulated by temperature triggers, and its altered function contributes to temperature-adaptive evolution in P. xylostella.”

      (10) The data points in Figures S1 and S7 are very small and hard to tell apart without zooming in a lot. Perhaps the authors could change the orientation of those pages to landscape and increase the size of the figures.

      Done. We have changed the orientation of Supplemental Figures S1 (Figure 1-figure supplement 1) and S7 (Figure 5-figure supplement 4) to landscape and increased the size of the figures and individual data points to improve visibility.

      (11) In Figure S2, the panel labeled as 'C' should be 'B' (based on the caption) and vice versa.

      Corrected. The panel labels ‘B’ and ‘C’ in Supplemental Figure S2 (Figure 2-figure supplement 1) have been swapped. The Supplementary Materials have been updated accordingly.

      Reviewer #2 (Public review):

      (1) The paper in its current form is hard to digest and would benefit from improved clarification of the storyline, as well as a tighter integration between the phenotypic, omics, and functional validation data. Currently, it is not always clear what the relevance is of all the reported results, nor why certain decisions were made, or how all the different methods the authors used fit together. For example, the authors functionally validated a second gene, PxDnmt1, but it is unclear why this particular gene was chosen, nor how it relates to their selection regimes when looking at the results obtained with the phenotyping and omics data collection. Seeing how much work the authors did, this makes the paper overwhelming and difficult to read.

      We sincerely appreciate this constructive feedback. In the revised manuscript, we have made significant structural revisions to improve the storyline and logical flow. We have streamlined the Results section (moving extensive descriptive data like life table curves and detailed metabolomics of mutant strains to the Appendix 1-3) to focus on the key findings. Furthermore, we have clarified the logical transitions between experiments. For instance, regarding the choice to validate PxDnmt1, we now explicitly explain in the Results that our untargeted metabolomic analysis of the PxSODC mutant strains revealed consistent alterations in 5-hydroxymethyluracil (involved in DNA demethylation) and 5'-deoxyadenosine (a precursor to the primary methyl donor S-adenosylmethionine) across all developmental stages. This specific metabolic signature provided a strong, data-driven hypothesis linking PxSODC function to epigenetic regulation via DNA methylation, prompting us to functionally validate PxDnmt1. By explicitly stating these rationales, the narrative is now much clearer and cohesive.

      (2) The authors at times stretch their results too far, as the ecological relevance of their study design and results is not clear, limiting the generalizability and value of the results for understanding species' adaptive potential under climate change. For example, the selection regimes used present the minimum and maximum known temperatures at which the species can survive and develop, but it is unclear how the temperatures relate to the natural environment of the source population, to what extent wild populations might experience these temperatures, and whether they would experience them at the extended duration used (12h at max/min temperature). Moreover, I wonder whether the comparisons made would identify the genes that matter under natural conditions, as unevolved populations were kept under constant conditions compared to 12h:12h temperature regimes for the evolved populations, and the metabolic and transcriptomic profiling was done under a constant favorable 26°C rather than under thermal stress in a, as far as I can tell, randomly chosen life stage (larval stage).

      We appreciate the reviewer raising these important points regarding ecological relevance and experimental design. In the revised manuscript, we have added context and acknowledged these limitations in the Methods and Discussion sections. First, regarding ecological relevance: The source population is from Fuzhou, a subtropical region where summer high temperatures frequently exceed 32°C and winter lows can drop below 10°C, making our selection temperatures ecologically relevant extremes for this population. The 12h:12h cycling temperatures were designed to simulate severe but natural diurnal fluctuations.

      Second, regarding constant control vs. cycling regimes: The constant 26°C represents the established optimal developmental temperature and standard laboratory condition for P. xylostella. We acknowledge that comparing cycling selection regimes against a constant control might conflate adaptation to absolute temperature extremes with adaptation to thermal fluctuation itself. We have added this as a caveat in the Discussion. Third, regarding omics profiling conditions: The transcriptomic and metabolomic profiling was conducted under common garden conditions (26°C) specifically to identify constitutive, genetically fixed adaptations resulting from evolutionary selection, rather than immediate physiological plasticity under stress. We have clarified these rationales in the text.

      (3) The paper in its current form does not adequately describe the statistical analyses underlying the results, nor do the authors share their code, making it very hard to judge whether the analyses used are appropriate and the results trustworthy. I have concerns about the inappropriate use of t-tests, the lack of correcting for confounding variables, and the need for multiple testing corrections.

      We sincerely appreciate this concern. In the revised manuscript, we have made substantial improvements to the description of statistical analyses throughout the Methods section:

      (1) Statistical methods for each data type are now described separately and in detail, specifying the tests used, the number and type of comparisons, and sample sizes.

      (2) For metabolomic data, we have clarified that FDR correction was applied alongside multi-criteria thresholds (|log<sub>2</sub>Fold Change| ≥ 1, VIP ≥ 1, FDR < 0.05). For transcriptomic data, FDR correction (Benjamini and Hochberg, 1995) was applied via DESeq2.

      (3) For WGCNA, we have specified the total number of correlation tests (29 modules × 30 metabolites = 870) and the stringent dual threshold (|r| > 0.8, P < 0.05) used to control for false positives, following standard practice.

      (4) For life table parameters, the paired bootstrap method with 100,000 replications was used for all pairwise comparisons among strains.

      (5) For all other experimental data (qRT-PCR, SOD activity, O<sub>2</sub><sup>-</sup> levels, survival rates, supercooling/freezing points, etc.), we have specified that t-tests were used only for two-group comparisons, while one-way ANOVA with Tukey's or Tamhane's T2 test was used for three or more groups, with non-parametric alternatives applied when normality assumptions were not met.

      (6) The raw data have been deposited in public repositories (see Data availability), and all statistical procedures are now described in sufficient detail to enable independent reproduction of the results.

      Reviewer #2 (Recommendations for the authors):

      Title

      (4) I don't feel the title adequately captures the work, I would instead of 'adaptive evolution' use 'experimental evolution' and I would not use the word 'underpins' but instead 'indicates', as it is not clear from your work whether the adaptations to the lab conditions you used would be ecologically relevant nor whether they are involved in thermal adaptation in wild populations.

      Accepted. The title has been revised to: “Experimental evolution to thermal stress indicates climate resilience in a cosmopolitan arthropod.”

      Abstract

      (5a) Please add the phenotype results to the abstract.

      We have added key phenotype results to the abstract. The revised text now reads: “The hot strain showed accelerated development, higher fecundity, and increased survival under extreme heat, while the cold strain exhibited lower supercooling and freezing points, indicating enhanced cold hardiness.”

      (6b) The Abstract doesn't really detail the answer to your research question yet: so what insights into the genetic mechanisms underlying thermal adaptation did you gain that are novel?

      We agree. We have revised the Abstract to explicitly highlight the novel genetic and molecular mechanisms we discovered. Specifically, we now detail that thermal adaptation is driven by a coordinated mutational, metabolic, and epigenetic (1) an energy-efficient genetic mechanism where non-synonymous mutations in PxSODC enhance superoxide scavenging efficiency, enabling effective oxidative stress management at lower gene expression levels; (2) convergent metabolic adjustments, notably a reduction in lipid metabolism to conserve energy; and (3) epigenetic regulation of thermal tolerance via DNA methylation. The revised text has been updated in the Abstract accordingly.

      (7c) Line 3: replace 'ectotherms' with 'arthropods' to match the title?

      Done. “Terrestrial ectotherms” has been replaced with “terrestrial arthropods” in the abstract.

      (8d) Line 9: replace 'demographic' with 'life history'?

      Done. “Demographic” has been replaced with “life history” in the abstract.

      Introduction

      (9a) The storyline is a bit unclear. Do you want to focus on the increased threat from insect pests under climate change or on the threat of climate change on insect persistence? Please pick one and adapt your storyline accordingly. I would suggest focusing on the first and talking more about the range extension of pest species under climate change (which would also require adaptation to cold extremes).

      We agree and have refocused the Introduction on the increased threat from insect pests under climate change, emphasizing that range expansion into new regions requires adaptation to both heat and cold extremes. Both the first and second paragraphs have been revised accordingly.

      (10b) Line 31-33: What do you mean by 'shows a positive relationship between the thermal tolerance range and the level of climatic variability'? Are they able to tolerate a larger range of temperatures?

      This sentence has been revised as part of the restructured Introduction, which now focuses on the range expansion of pest species under climate change. The revised text reads: “Such range expansion requires adaptation not only to warmer conditions in existing habitats but also to cold extremes encountered during colonization of higher latitudes or elevations (Harvey et al., 2020).”

      (11c) Line 33-35: Is this information relevant here?

      Agreed. This sentence has been removed as part of the restructured Introduction, which now focuses on the threat of pest range expansion under climate change.

      (12d) Line 55-56: What exactly do we not know yet about the mechanisms that enable thermal adaptation that you aim to fill in this paper? Please rephrase your knowledge gap to be more concrete (e.g., "but we do not yet know how...").

      We have rephrased the knowledge gap to be more concrete and aligned with the revised storyline. The revised text now reads: “...we do not yet know how long-term thermal selection drives coordinated changes across gene function, metabolic networks, and life history traits to enable thermal adaptation and range expansion in pest species.”

      (13e) Line 57: Also, here, the storyline is unclear. Why did you use the diamondback moth as your model species? You provide many different reasons, but it would help if you emphasized one reason that is in line with whichever storyline you want to focus on: is it because it is an insect pest that can tolerate a wide range of temperatures?

      We have streamlined this paragraph to focus on the primary rationale: P. xylostella is a globally distributed pest that thrives across a wide range of thermal environments, making it an ideal model for studying the genetic mechanisms of thermal adaptation. Supporting details on genomic resources are retained briefly as they enable the multi-omics approach used in this study.

      (14f) Line 65: Demonstrated how? Please give a short summary of the evidence for their genetic capacity to tolerate future climates.

      We have added a brief summary of the evidence. Specifically, genome-wide SNP analysis of field populations from 114 locations across diverse biogeographical zones revealed climate-adaptive genetic variability, indicating that P. xylostella can tolerate projected future climates in most regions (Chen et al., 2021).

      (15g) Line 72: What does 'Age-stage' mean? Should it read 'Aged-staged'?

      “Age-stage, two-sex life table” is an established demographic method developed by Chi (1988) that simultaneously accounts for both age and developmental stage in both sexes. This is a standard term in the field (Chi et al., 2020), so we have retained the original wording but added a brief clarification upon first use.

      (16h) Line 78-80: This needs a bit more explanation. Why does an increased ability to scavenge superoxide anions affect adaptability under extreme temperature environments?

      We have added a brief explanation. Extreme temperatures induce oxidative stress by elevating intracellular reactive oxygen species (ROS), including superoxide anions, which can damage cellular structures. Enhanced scavenging capacity thus helps maintain cellular homeostasis under thermal stress.

      (i) Line 82-86: Please be more precise. What novel insights did you gain about the genetic mechanisms underlying thermal adaptation?

      We have revised this sentence to more precisely summarize the novel insights, encompassing both the multi-omics findings and the functional validation of PxSODC.

      Results

      (18a) The results section is very long and presents an overload of information at the moment, overwhelming the reader. Consider moving some sections to the Supplements (for example, a large part of the phenotypic data that cannot be linked to the omics data and the metabolic profiling of the mutant strains) or leave them out of the paper altogether.

      We agree that the Results section was too dense. We have streamlined it by moving the following content to the Supplementary Materials:

      (1) Detailed age-stage survival and fecundity curve data for the ancestral, hot and cold strains (Supplementary Text S1).

      (2) Detailed life table analysis of the PxSODC mutant strains (Supplementary Text S2).

      (3) Detailed untargeted metabolomic profiling of the SODC-MU mutant strains across developmental stages (Supplementary Text S3).

      The main text now retains only the key life history comparisons, extreme temperature tolerance results, omics-based evidence linking transcriptomics and metabolomics, functional validation of PxSODC, and the DNA methylation findings, with brief summaries and cross-references to the Supplements for supporting details.

      (19b) Please also provide the effect sizes for the different effects you report, for example, how many degrees difference was there between ancestral and cold strains in the supercooling/freezing points, and what was the variation?

      We have added specific effect sizes (mean ± SEM and between-group differences) for all key comparisons throughout the Results section, including preadult duration, stage-specific survival rates under extreme heat, supercooling/freezing points, and SODC-MU mutant strain comparisons. For example, the supercooling points of CS pupae (-23.99 ± 0.18°C) were 0.90°C lower than AS (-23.09 ± 0.26°C), and the freezing points were 2.66°C lower (-14.24 ± 0.61°C vs. -11.58 ± 0.52°C). Please refer to the revised manuscript for all updated values.

      (20c) Line 93-94: "Intrinsic and finite rate of increase" of what?

      Clarified. These are population growth parameters. The revised text now specifies “intrinsic rate of increase (r) and finite rate of increase (λ) of the population.”

      (21d) Line 98-99: Please start the paragraph with this summary of the results and then further detail them.

      We have restructured this paragraph by moving the summary sentence to the beginning, followed by the supporting details.

      (22e) Line 100-109: Why did you look at daily survival and fecundity rates? Please add why this is relevant.

      As part of the overall streamlining of the Results section, this paragraph on detailed age-stage survival and fecundity curves has been moved to Supplementary Text S1. A brief justification for their relevance has been added there, noting that these curves capture stage-specific variation in survival and fecundity that summary life table parameters alone may obscure.

      (23f) Line 106: What do HS, AS, and CS stand for? And please provide the statistics for comparison of daily survival rates between the strains.

      We have defined the abbreviations (HS = hot strain, AS = ancestral strain, CS = cold strain) at their first appearance in the Results section. This paragraph on daily survival and fecundity has been moved to Supplementary Text S1, where the abbreviations are also defined. The survival rates reported are the maximum daily survival rates derived from the age-stage specific survival rate curves (s<sub>xj</sub>), and the statistical comparisons among strains are presented in Supplemental Table S1.

      (24g) Line 144-146: Why are these differential metabolites likely to play a crucial role?

      We agree this statement was speculative. It has been removed from the revised manuscript.

      (25h) Line 159-161: Why is a reduction of lipid metabolites evidence for adaptive evolution?

      We have revised this sentence to clarify the reasoning. The reduction in lipid metabolites in both independently evolved hot and cold strains suggests a convergent metabolic response, indicating that lipid metabolism adjustment is a shared adaptive strategy rather than a random change.

      (26i) Line 184-185: It is difficult to judge from Figure 3E the extent of overlap in KEGG pathways between the hot and cold strains. Can you adjust the figure to emphasize that overlap more?

      Agree. To intuitively emphasize the extent of overlap in KEGG pathways between the hot and cold strains, we have completely redesigned Figure 3E. Instead of presenting two separate panels with unaligned vertical axes, we have consolidated the data into a single back-to-back (mirrored) bar chart with a shared central y-axis.

      (27j) Line 211: Not only the red module, but also the blue and green module correlates with many of the shared differential metabolites.

      We agree. We have revised the text to acknowledge that the blue and green modules also showed strong correlations with shared differential metabolites, while noting that the red module had the highest number of significantly correlated metabolites and was therefore selected for further analysis.

      (28k) Line 215: I would rephrase this as genes being interesting candidates for being involved in thermal adaptation or 'seem to be important for the adaptation of...', as you don't know from these results whether these genes play a critical regulatory role.

      Agreed. We have toned down the language to reflect the correlative nature of these results.

      (29l) Line 233: Do you mean that you further analyzed 15 genes of the 79 identified candidate genes in the previous paragraph?

      Yes, exactly. From the 79 candidate genes, we selected 15 that were both annotated in the genome and had high expression levels (FPKM > 10) for further analysis. We have clarified this in the revised manuscript.

      (30m) Line 238: What does SOD stand for?

      We have spelled out the abbreviation upon first use in this section.

      (31n) Line 254-255: Please provide the stats for this result.

      We have added the specific allele frequencies for each strain. The Leu194-Met194 mutation frequency was determined by direct sequencing of 10 individuals per strain, and the frequencies are now reported in the revised text.

      (32o) Line 303-304: How did you test for enhanced stability to temperature fluctuations? And enhanced compared to what?

      This observation was based on the survival rate data in Figure 5C, where mutant pupae at 43°C showed no significant difference from the ancestral strain, whereas other life stages (eggs, larvae, adults) at 42°C showed significantly reduced survival in the mutant strains. We have revised the text to clarify the comparison.

      (33p) Line 324-326: Why do decreased expression levels demonstrate increased O₂⁻ scavenging capacity? And why is that beneficial for adaptation to thermal stress? Please explain.

      We have revised this sentence to clarify the logic. The non-synonymous mutations in the hot and cold strains likely alter the protein conformation of SOD enzymes, increasing their catalytic efficiency per molecule. This allows effective O<sub>2</sub><sup>-</sup> scavenging at lower expression levels, which is energetically favorable under thermal stress where energy conservation is critical for survival.

      (34q) Line 404-406: I'm confused. Is there a direct link between the gene you knocked out here and the results you presented up until now? How do the reduced levels of 5-methylcytosine relate to the metabolite results you present at the beginning of the paragraph, other than that both could be involved in DNA methylation?

      We have revised this paragraph to clarify the logical chain. Among the three metabolites consistently altered across all developmental stages in the SODC-MU strains, 5-hydroxymethyluracil is involved in dynamic DNA demethylation and 5'-deoxyadenosine is a precursor to S-adenosylmethionine (the methyl donor for DNA methylation). This suggested a link between PxSODC deletion and DNA methylation. To test this, we examined PxDnmt1 expression and activity in the thermally adapted strains and found both were significantly reduced. We then used RNAi to silence PxDnmt1 and confirmed that reduced DNA methylation (lower 5-mC levels) directly impaired thermal tolerance. Thus the connection is: PxSODC deletion → altered methylation-related metabolites → reduced DNA methyltransferase activity → decreased thermal tolerance.

      (35r) Line 410: Saying that your knockdown of a gene that did not directly pop up in any of your other analyses confirms that DNA methyltransferase is associated with the response to thermal selection is a stretch. Please rephrase.

      We agree this was overstated. We have toned down the language to reflect that the RNAi results provide preliminary evidence for a potential role of DNA methylation in thermal tolerance, rather than confirmation.

      Discussion

      (36a) The phenotype data are currently not discussed at all. Please add it to the discussion and try to integrate it more with the omics data you collected.

      We agree. To provide a cohesive narrative and avoid redundancy, we have addressed this comment in conjunction with our pathway interpretation (please see our response to Reviewer 1, Comment 1). In the revised Discussion, we explicitly integrated our specific phenotypic findings (e.g., accelerated development, increased fecundity, and heat survival in the hot strain; prolonged lifespan and lowered supercooling points in the cold strain) with the distinct transcriptomic and metabolomic profiles. This integration demonstrates how molecular and metabolic rewiring directly underpins the divergent life-history traits without engaging in unwarranted speculation.

      (37b) Line 433-434: I don't think this adequately represents the relevance of your particular study. I would suggest changing it to be more in line with the storyline of understanding the capacity for global dispersal in insect pests under climate change.

      We agree. We have revised this sentence to align with the storyline of pest range expansion under climate change.

      (38c) Line 476: This is a very odd statement; don't all species' genomes have genes encoding proteins involved in thermal adaptation? The reference also doesn't seem to be appropriate. I would suggest deleting this sentence.

      Agreed. This sentence has been removed.

      (39d) Line 483: Please write out SOD the first time you use it in a new section.

      Done. SOD has been spelled out at its first use in the Discussion.

      (40e) Line 544-548: This is a bit too specific to be the last sentence of the discussion. Try to formulate it more broadly in terms of what future research should focus on in general, not just your specific research.

      We agree. We have broadened the final sentence to address future research directions more generally.

      Figures

      (41a) Figure 1A: I don't think t-tests are appropriate here since you are not simply comparing two treatments, but testing for the effects of 5-6 different temperatures. And how did you correct for replicate populations in your analysis?

      Clarified. In Figure 1A, our comparisons are independent pairwise tests between exactly two strains (HS vs. AS) at each specific temperature and time point, making t-tests statistically appropriate. We were not testing for a continuous effect across temperatures. Regarding replicate populations, the individuals used in these assays were drawn from across the six replicate populations per treatment, with each biological replicate (n = 6, with 20 individuals per replicate) comprising individuals pooled from across the replicate populations to account for inter-population variation. We have clarified this in the revised figure legend.

      (42b) Figure 1B, Figure 5D, Figure 7: bar graphs are used for count data, so do the data represent the number of individuals with a certain trait value? If they are instead showing the mean of the population/treatment group, please use mean points ± standard errors instead.

      Accepted. The data in these figures represent continuous physiological traits (e.g., supercooling/freezing points) showing the mean of the populations, rather than count data. To align with current data visualization standards for continuous variables and to provide full transparency of the underlying data distribution, we have replaced the bar graphs in Figures 1B, 5D, and 7 with scatter plots. These revised figures now display the mean ± SEM overlaid with all individual biological replicate data points.

      (43c) Figure 3B: There is a typo in the graph, it reads 'Stattistics' instead of 'Statistics'.

      Corrected. The typo ‘Stattistics’ in Figure 3B has been fixed.

      (44d) Figure 3C: I don't understand what the colors of the graph mean here. Is it the average differential expression of each replicate compared to the ancestral?

      Clarified. We have updated the figure legend to explain that the colors represent the Pearson correlation coefficient (r) between pairs of biological replicates, indicating the degree of transcriptomic similarity among samples.

      Methods

      (45a) Please start each new methods paragraph with the purpose of the method/analysis, for example, "To investigate XX, we used method X to measure X". It is at the moment hard to understand why certain things were done.

      We agree. We have revised each Methods paragraph to begin with a clear statement of purpose, so that the rationale for each analysis is immediately apparent. All changes are shown in the revised manuscript.

      (46b) Line 575-578: Why were the selection regimes with cycling temperatures and the control with constant?

      The cycling temperatures in the hot (32°C/27°C) and cold (15°C/10°C) regimes were designed to simulate diurnal temperature fluctuations (12h light/12h dark) that more closely reflect natural thermal environments. The control was maintained at a constant 26°C, which is the established optimal developmental temperature for P. xylostella (Liu et al., 2002) and represents the standard laboratory rearing condition. We acknowledge this asymmetry and have added a justification in the revised manuscript.

      (47c) Line 581: How many generations was the ancestral population kept in the lab before the start of the selection experiment? And for how many generations were the populations selected?

      The ancestral population was maintained in the laboratory for approximately ~170 generations (from July 2012 to the start of the selection experiment) before the thermal selection began. The hot strain was selected for ~75 generations and the cold strain for ~15 generations over the three-year experiment. We have added this information to the revised manuscript.

      (48d) Line 585-586: I don't understand what you mean by randomly selecting six replicate populations per treatment for downstream experiments when you only had six replicate populations per treatment to begin with (as detailed in Line 574)?

      We apologize for the confusion. All six replicate populations per treatment were used for downstream experiments. We have corrected this sentence to remove the misleading “randomly selected” wording.

      (49e) Line 590: Were these 90 eggs also randomly selected, like for the individual life tables? And were these kept at the baseline temperature conditions?

      Yes, the 90 eggs were randomly selected and maintained under the baseline favorable temperature (26°C). We have clarified this in the revised manuscript.

      (50f) Line 606: Which life history and population fitness parameters were calculated?

      We have specified all parameters calculated in the revised manuscript.

      (51g) Line 609: Link to software doesn't work.

      We have updated the software link to the current working URL.

      (52h) Line 611: Please spell out what 'BT' stands for.

      Done. “BT” has been spelled out as “bootstrap” upon first use.

      (53i) Line 612-613: How many tests did you do? Did you correct for multiple testing? Using what method?

      The paired bootstrap method implemented in TWOSEX-MSChart inherently accounts for multiple pairwise comparisons through 100,000 bootstrap replications. We have clarified the scope of comparisons in the revised manuscript.

      (54j) Line 620-621: What does biological replicate mean here? Individual eggs / larvae / pupae / adults, or were all or some life stages pooled? Also, you now only detailed which samples were collected for metabolomic profiling, were the same samples used for transcriptomic profiling, or a subset?

      Each biological replicate consisted of pooled individuals at the same developmental stage. The same sample collection strategy was used for both metabolomic and transcriptomic profiling, but from independent biological replicates (six for metabolomics, three for transcriptomics). We have clarified this in the revised manuscript.

      (55k) Line 637: Also here, how many tests did you do? Were p-values corrected for multiple testing? Using what method?

      Differential metabolites were identified through pairwise comparisons using Student's t-test with FDR correction for multiple testing. A multi-criteria threshold of |log<sub>2</sub>Fold Change| ≥ 1, VIP ≥ 1, and FDR < 0.05 was applied. This approach was used for all metabolomic comparisons, including HS vs. AS, CS vs. AS, and SODC-MU vs. AS. We have clarified this in the revised manuscript.

      (56l) Line 662: And here: how many tests did you do? Did you correct for multiple testing? Using what method?

      In the WGCNA analysis, Pearson correlations were calculated between each module eigengene and each of the 30 common differential metabolites, resulting in a total of 29 × 30 = 870 correlation tests. Following standard WGCNA practice, rather than applying FDR correction, we used a stringent dual threshold of |correlation coefficient| > 0.8 and P < 0.05 to identify significant module-metabolite associations, which effectively controls for false positives (Langfelder and Horvath, 2008). We have clarified this in the revised manuscript.

      (57m) Line 663: How did you select these modules? The ones that significantly correlated with differential metabolites? Why did you not use the phenotype data here?

      Modules were selected based on significant correlations (|correlation coefficient| > 0.8, P < 0.05) with differential metabolites shared between the hot and cold strains. We chose metabolites rather than phenotype data as the trait input for WGCNA because metabolites serve as intermediate molecular phenotypes that bridge gene expression and organismal phenotypes, providing a more direct link to the underlying regulatory mechanisms. This approach allowed us to identify gene modules most closely associated with the metabolic changes driven by thermal adaptation, which could then be connected to the observed life history and fitness divergence.

      (58n) Line 666: move RNA extraction details to before RNAseq methods description.

      Done. The “RNA extraction and cDNA synthesis” section has been relocated to before the “Transcriptomic profiling” section for better logical flow.

      (59o) Line 836: This paragraph describing the statistics is very short, and it is unclear to what data the described analyses apply. As the different types of data are very different, I expect the analyses to differ as well. Please describe the statistical analyses for each data type in more detail, specifying what tests you used, which, and how many comparisons were performed.

      We agree. The statistical methods for life table analysis, metabolomics, and transcriptomics have been detailed in their respective method sections. We have expanded the Data analysis section to specify the statistical tests for the remaining experimental data.

      (60p) Line 837: Please include your SPSS scripts to ensure the reproducibility of your results.

      The statistical analyses in SPSS were performed using the graphical user interface. As all statistical tests, parameters, and comparison groups have been described in detail in the revised Methods section, and the raw data have been deposited in public repositories (see Data availability), we believe the analyses are fully reproducible. We are happy to provide additional details if needed.

    1. eLife Assessment

      This fundamental work demonstrates that ABHD6 regulates AMPAR gating kinetics in a TARP γ-2-dependent manner. The evidence in this study is compelling. This study will be of interest to readers in the field of synaptic transmission.

    2. Reviewer #1 (Public review):

      Summary:

      This research sheds light on the nuanced role of ABHD6 in regulating AMPARs, highlighting its interaction with TARP γ-2 as a critical factor in modulating receptor gating kinetics. It is crucial to understand that although ABHD6 alone does not alter AMPAR kinetics, its presence alongside TARP γ-2 accelerates AMPAR deactivation and desensitization, thereby affecting synaptic transmission dynamics.

      Strengths:

      Important findings in the research include:<br /> - ABHD6 does not affect the gating kinetics of GluA1 and GluA2(Q) homomeric receptors independently.<br /> - In the presence of TARP γ-2, ABHD6 accelerates deactivation and desensitization of these receptors, regardless of their splicing or editing isoforms.<br /> - The effect is consistent for both homomeric GluA1 and GluA2(Q) receptors and heteromeric GluA1i/GluA2(R)i-G receptors.<br /> - The recovery from desensitization of GluA1 with the flip splicing isoform is slowed by ABHD6 in the presence of TARP γ-2.

    3. Reviewer #2 (Public review):

      Summary:

      Cong et al. investigated the regulatory effects of ABHD6 on AMPARs. The authors performed adequate electrophysiology recordings to show the exact pattern of this regulation and covered major critical points.

      Strengths:

      The authors have performed high-quality ephys recordings and examined all potential regulatory aspects of ABHD6 on AMPARs. This is important to understand the AMPAR functions.

      Weaknesses:

      (1) The authors discussed CNIH-2 extensively from line 92-110 in the introduction, however, they did not perform related experiments. I suggest they move this part to the discussion where they also discussed the roles of CNIH.

      (2) The authors need to report the "n" for all the experiments they have presented in this manuscript. How many cells were recorded in each condition? How many batches? This information has to be in all of the figure legends, but it is missing except Fig. 4.

      (3) One question is what the physiological meanings of this regulatory effect are. The authors may consider adding some discussions.

      (4) About statistics. The authors need to add more details and make sure their statistics sound. For example, they also need to check the equality of variances. In their Table EVs, where the P values are reported, the authors need to report which statistics they have used, one-way ANOVA, K-W test, or others, and the exact post-hoc test type for each comparison. For one-way ANOVA, report the F values simultaneously with the P values in all figure legends.

      (5) Fig. 3J, the authors need to correct the label of the Y axis. It is shifted.

      Comments on revised version.

      In the revised manuscript, the authors have addressed all my concerns. The manuscript has been substantially strengthened by additional data and discussion.